DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models
A team of 11 AI expert agents built to explain how a classical Chinese herbal formula actually works
Researchers from Guangzhou University of Chinese Medicine and collaborators built DeepTCM1.0, a multi-agent system powered by the general-purpose large language model DeepSeek V3.2, in which 11 simulated experts collaborate to explain the mechanism of traditional Chinese medicine (TCM) herbal formulas. They tested it on Guizhi Decoction, a foundational classical prescription, and found it scored significantly higher than a single LLM in blind evaluation. Notably, the system works without any additional database construction or model fine-tuning, relying entirely on prompt engineering.
METAL MEDIA explanatory visual
A team of 11 AI expert agents built to explain how a classical Chinese herbal formula actually works
- 01Existing methods like data mining, network pharmacology, and molecular docking failed to deeply integrate classical TCM theory with modern science, while directly querying general-purpose LLMs suffered from poor adaptation to TCM theory and reasoning hallucinations (AI generating plausible but ungrounded content)
- 02The team designed a three-tier collaborative architecture with 11 agents total: a Principal Investigator and Recruiter for task decomposition, 7 fixed core experts plus 4 dynamically recruited specialists for multi-round reasoning, and a Critic and Report Integration Specialist for quality control and final output, following a three-round iterative discussion process (independent interpretation, integration, optimization)
- 03Applied to Guizhi Decoction, the framework linked the classical TCM concept of 'disharmony between ying and wei' (nutritive and defensive qi) to modern neuro-immune-metabolic regulatory networks and proposed three novel mechanistic hypotheses addressing longstanding controversies
- 04Four independent LLMs evaluated five anonymized reports across five rounds each (100 total scoring assessments), yielding intra-rater reliability (consistency when the same evaluator repeats scoring) of ICC 0.789-0.867 and inter-rater reliability of 0.868; one-way ANOVA showed highly significant group differences (F=66.969, P<0.001), and DeepTCM1.0 significantly outperformed a single general-purpose LLM group (U=191.00, P<0.001, effect size r=0.5403)
- 05On radar chart visualization, DeepTCM1.0 achieved near-perfect scores (4.8-5.0) across all evaluation dimensions with the highest scoring consistency, while offering lightweight deployment, zero fine-tuning requirement, and high reproducibility as practical advantages
What they did
- Existing methods like data mining, network pharmacology, and molecular docking failed to deeply integrate classical TCM theory with modern science, while directly querying general-purpose LLMs suffered from poor adaptation to TCM theory and reasoning hallucinations (AI generating plausible but ungrounded content)
- The team designed a three-tier collaborative architecture with 11 agents total: a Principal Investigator and Recruiter for task decomposition, 7 fixed core experts plus 4 dynamically recruited specialists for multi-round reasoning, and a Critic and Report Integration Specialist for quality control and final output, following a three-round iterative discussion process (independent interpretation, integration, optimization)
- Applied to Guizhi Decoction, the framework linked the classical TCM concept of 'disharmony between ying and wei' (nutritive and defensive qi) to modern neuro-immune-metabolic regulatory networks and proposed three novel mechanistic hypotheses addressing longstanding controversies
- Four independent LLMs evaluated five anonymized reports across five rounds each (100 total scoring assessments), yielding intra-rater reliability (consistency when the same evaluator repeats scoring) of ICC 0.789-0.867 and inter-rater reliability of 0.868; one-way ANOVA showed highly significant group differences (F=66.969, P<0.001), and DeepTCM1.0 significantly outperformed a single general-purpose LLM group (U=191.00, P<0.001, effect size r=0.5403)
- On radar chart visualization, DeepTCM1.0 achieved near-perfect scores (4.8-5.0) across all evaluation dimensions with the highest scoring consistency, while offering lightweight deployment, zero fine-tuning requirement, and high reproducibility as practical advantages
Why it matters
This shows that for a field like TCM, which requires handling both vast classical texts and modern biomedical data, orchestrating multiple expert roles on top of an off-the-shelf language model can produce reliable analysis without any domain-specific fine-tuning. The approach of reducing hallucination and integrating specialized knowledge purely through prompt engineering and role division offers a reusable, low-cost blueprint applicable to other complex traditional knowledge systems and interdisciplinary research.
Terms in this paper
- Large Language Model (LLM) · An AI model trained on vast text data to understand and generate natural language, e.g. GPT, DeepSeek
- Hallucination · When an AI generates plausible-sounding but factually ungrounded content
- Network Pharmacology · A research method that maps relationships among drug compounds, target proteins, and biological pathways to explain mechanisms of action
- Intraclass Correlation Coefficient (ICC) · A statistic measuring how consistent repeated scores from the same rater are
- Disharmony between Ying and Wei · A classical TCM pathological concept describing imbalance between the body's nutritive function (ying) and defensive function (wei)
Original abstract (English)
Background: Mechanistic elucidation of traditional Chinese medicine (TCM) compound formulas remains a central challenge in the modernization of TCM. Conventional approaches, including data mining and network pharmacology, are insufficient for achieving deep integration between classical TCM theory and modern scientific research. In addition, direct question-answering using general-purpose artificial intelligence large language models is limited by inadequate adaptation to TCM theoretical frameworks and susceptibility to reasoning hallucinations. Consequently, there is an urgent need to develop intelligent analytical methods aligned with the holistic principles of TCM. Objective: To establish a multi-expert intelligent agent framework integrating classical TCM theory with modern life sciences, thereby enabling systematic and interpretable mechanistic analysis of TCM compound formulas, with Guizhi Decoction serving as a representative validation case. Methods: The DeepTCM1.0 framework was constructed based on the general-purpose large language model DeepSeek V3.2. It adopts a three-tier collaborative architecture and a three-round iterative quality-control workflow, simulating the collaborative analytical process of 11 interdisciplinary intelligent agents. The framework was applied to the mechanistic interpretation of Guizhi Decoction from the dual perspectives of classical traditional Chinese medicine theory and modern scientific research. Framework performance was comprehensively evaluated through double-blind five-dimensional scoring, intraclass correlation coefficient (ICC) reliability testing, Mann-Whitney U tests, and effect size analysis. The evaluation employed four independent large language models as evaluators, each conducting five rounds of repeated scoring on five anonymized reports, resulting in a total of 100 independent scoring assessments.
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressA fix for AI models that get penalized by their teacher even when they're reasoning correctly
- PersonalBench: Measuring the Authorship Gap in LLM PersonalizationAI can be prompted to write 'like someone,' but its own voice never fully disappears