Medical AI Evaluation Standards Differ Between China and the West? Chinese Institution Releases Ranking with WiseDiag, Gemini, and OpenAI GPT Taking Top Spots

Deep News
May 08

As public awareness grows, an increasing number of individuals are turning to AI for immediate answers to minor health concerns instead of spending time waiting in hospital queues. However, given the highly specialized nature of medicine, how reliable is AI-powered diagnosis? What standards should be used to evaluate the accuracy and professionalism of these AI systems?

Common applications for medical AI include health management and chronic disease management. The market is abundant with various medical AI models, including general-purpose large language models from major tech firms, health management apps, and mini-programs embedded within social platforms, all offering diagnostic advice. Yet, the responses provided by different platforms can vary, potentially causing confusion or even misleading users.

Some users express distrust in AI, noting that responses can be inconsistent. For example, an AI might recommend certain medications initially but suggest different ones when symptoms are elaborated, sometimes leading to overlapping or even contradictory advice between traditional Chinese and Western medicine. Due to their design, AI systems may attempt to accommodate users, offering vague or incorrect suggestions based on limited information even when unable to make an accurate diagnosis.

To mitigate liability risks, some AI systems provide responses that are technically accurate but unhelpful, such as mechanically advising users to "follow doctor's orders." This offers little practical value to individuals seeking reference advice.

The most critical consumer-facing scenarios for "AI + healthcare" currently are health management and chronic disease management, according to He Xun, product lead at Deshi Biotech. Since AI lacks the clinical experience of human doctors and cannot engage in in-depth dialogue about individual symptoms, the information provided by users is often incomplete and lacks key diagnostic data, increasing the risk of missed diagnoses.

He Xun noted that while there is a sufficient supply of AI agents in the market, the industry is still in a phase of rapid but uneven growth, with significant disparities in product quality and professional capability, making it difficult for average users to choose effectively. A unified evaluation standard and authoritative mechanism to assess the reliability of medical large models are currently lacking, prompting the creation of a dedicated medical AI evaluation ranking system.

This evaluation platform, named DoctorBench, was established by a domestic institution and launched in Hong Kong to address the gap in industry standards. The top three performers in its ranking are Hangzhou Zhizhen Technology's WiseDiag-v2, Google's Gemini-3.1-Pro-Preview, and OpenAI's GPT-5.4. In contrast, a separate evaluation system called HealthBench, released by OpenAI in May of last year, ranked OpenAI o3, GPT-4.1, and Claude 3.7 Sonnet as the top three.

The release of the domestic medical AI ranking has sparked discussion about potential differences in evaluation standards between China and other countries. Variations in healthcare systems raise the question of whether corresponding AI assessment criteria also differ. Can a locally developed evaluation system comprehensively address the diverse needs of medical AI across various scenarios? How can a universally accepted evaluation standard be promoted internationally in the future?

A comparison of the two rankings shows significant overlap among top-tier products, though their order differs slightly. Other listed products exhibit strong local characteristics. Deshi Biotech explained that significant differences in clinical guidelines, language habits, and patient demographics across countries and regions make it difficult for any single evaluation system to achieve global applicability.

According to HealthBench's weighting rules, the core overall metric is "comprehensive medical reasoning," with the highest weight given to diagnostic accuracy, including consultation logic, symptom assessment, examination and medication plans, and compliance of treatment recommendations. Among sub-weights, the ability to reason through complex cases is paramount, focusing on the model's capacity for deep reasoning involving comorbidities, ambiguous symptoms, rare diseases, and complex multi-round medical histories.

Two key rules are applied: first, scoring is performed by licensed physicians from multiple countries; second, irrelevant metrics such as model parameters, inference speed, and open-source status are excluded, with evaluation focused solely on high-difficulty clinical practical capabilities.

DoctorBench follows a similar core logic, officially defined as assessing a model's ability to "think like a doctor" in clinical communication and decision-making. Its three main rankings focus on a core medical leaderboard for LLMs, a multimodal leaderboard for VLMs, and an agent leaderboard, evaluating textual diagnostic capability, multimodal understanding, and multi-round decision-making and tool usage in simulated clinical environments, respectively.

However, DoctorBench sets "medical factual accuracy" and "safety and risk control" as non-negotiable red lines. Any model exhibiting serious deviations on critical patient safety issues is unable to achieve a high score, regardless of performance in other dimensions.

He Xun stated that DoctorBench employs a "professional question bank + blind expert review" scoring system. The question bank is independently developed, conducting full-scenario testing on mainstream medical AI products, while the manual review process includes quantifiable indicators to ensure the objectivity, professionalism, and credibility of the results.

In quarterly updates to the HealthBench Hard ranking starting August 2025, specialized medical large models from China began to appear, while leading general-purpose models started to drop out of the top positions. He Xun explained that while general-purpose models possess broad adaptability across scenarios, they lack the depth of specialized training and completeness of knowledge graphs in the medical vertical compared to dedicated medical models, resulting in lower overall industry rankings. Many high-performance specialized medical models are often characterized by closed APIs and independent deployment, creating higher access barriers for the general public but offering stronger专业性.

From a public application perspective, many leading medical AI agents have open service ports, allowing direct public access through name searches. However, public awareness may be low, and a certain level of专业 proficiency is required. Technical terms related to algorithm parameters, model scale, or architecture versions are not conducive to public identification and retrieval. The ranking includes通俗 explanations of专业术语, application scenario tags, and标注 of official access points, along with definitions of model positioning, suitable fields, and access channels, aiming to lower the information barriers and usage costs for the public to access quality medical AI services.

Specialized medical large models are now widely used in hospitals as辅助诊疗 tools. Since 2025, a complete policy framework for "AI + healthcare" has been established, with the deep integration of AI and medicine being a clearly defined direction mandated by national policy and implemented by medical institutions.

The 2025 "Opinions on Deepening the 'AI+' Initiative" listed healthcare as the top priority among seven key areas. Subsequently, five departments including the National Health Commission issued the "Implementation Opinions on Promoting and Regulating the Application Development of 'AI + Medical and Health Services'," which set clear targets: by 2027, high-quality medical datasets will be established, forming clinical specialized vertical large models; AI-assisted diagnosis will be普遍开展 in secondary and above hospitals; AI usage rates in primary care institutions will reach ≥40%. By 2030, intelligent辅助 applications for primary care diagnosis will be basically fully covered; a mature full-chain service system for "AI + healthcare" will be formed; and the普及率 of AI for resident health management will reach ≥80%.

Market data indicates that within medical institutions, AI agents cover scenarios including pre-consultation screening and inquiry, decision support during diagnosis, and post-diagnosis follow-up and intervention for chronic diseases. Currently, penetration rates in domestic tertiary hospitals exceed 60%, with diagnostic accuracy above 95%. Penetration in secondary hospitals is approximately 40-50%, while in primary care institutions it ranges from 20-30%.

For individual doctors, AI can serve as a tool to check for omissions and fill gaps. Doctors may find it difficult to retain long-term memory of patient medical history data and health characteristics, whereas AI can permanently store and dynamically track changes in indicators. This assists doctors in evaluating treatment plans, optimizing diagnostic processes, and improving diagnostic efficiency. Patients can also aggregate their own health data and track disease progression through user-end applications.

Currently, there remains a structural disparity in the spatial distribution of medical resources within China. Major cities and central urban areas concentrate a large number of tertiary hospitals and high-end medical talent, while prefecture-level cities, counties, and remote primary care areas still face a supply gap of quality medical resources. Additionally, there is a significant disparity in the专业 diagnostic capabilities and skill levels of primary care medical staff compared to those in central cities.

He Xun believes that the application of AI as an辅助 tool can optimize the allocation of medical resources, promote the inclusive development of public medical services, and allow for the shared benefits of smart medical technology.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10