Skip to main content

Aug 24, 2026 Louis-Antoine Mullie MD | Chief AI Researcher

LLMs Alone Won’t Bring Us Medical Superintelligence

LLMs Alone Won’t Bring Us Medical Superintelligence  article hero image

LLMs Alone Won't Bring Us Medical Superintelligence

We believe increasingly capable AI systems will integrate more deeply into medicine. They will help clinicians synthesize evidence, reason across complex cases, identify patterns, and make medical knowledge available at unprecedented speed and scale.

Yet today's large language models, despite their growing usefulness, have fundamental limitations. They do not reliably learn from experience. They cannot consistently follow exact instructions. They can produce a correct answer once and fail on an essentially identical problem later. And because they generate probabilistically, even facts that are precisely knowable can become sources of uncertainty.

If AI systems are to become substantially more capable while remaining worthy of trust in medicine, these limitations have to shape how we build them.

We believe LLMs alone do not possess the abilities required for the emergence of medical superintelligence.

A superintelligent medical AI would learn from its failures, discover methods that improve its future performance, and eventually improve the process by which those methods themselves are discovered. Those improvements would have to persist without erasing established knowledge or introducing new failures into capabilities that already work.

In medicine, this process must also remain grounded in clinical reality. Continuous oversight and physician governance are necessary to determine which failures matter, which improvements are valid, and which knowledge deserves to persist.

For this reason, we envision medical superintelligence as more than a model. It will be a living system for accumulating, verifying, and transferring an ever-expanding body of medical knowledge among clinicians, researchers, and machines.

Learning from experience

An AI system can receive feedback and correct an answer. A more capable system can carry that lesson forward and perform better the next time.

Self-improvement begins when the lesson produces a persistent change in the system itself. So-called "recursive self-improvement," or RSI, begins when the system becomes better at producing those changes.

For medical AI, the important question is what happens after a failure. Can the system identify the broader class of errors that produced it? Can it discover a general solution? Can it preserve that solution and apply it to future problems? And can solving that problem improve its ability to identify and correct the next class of failures?

Consider the difference between discovering a new treatment and discovering a better method for finding treatments. The first can advance the care of one disease. The second can advance the care of many diseases.

Recursive self-improvement requires the latter: second-order learning. Each solved problem leaves behind more than an answer. It leaves behind a better method for solving future problems.

That creates the possibility of compounding progress. Better methods produce better discoveries. Those discoveries reveal new ways to improve the methods themselves. Over many iterations, the system's capacity to learn can advance alongside the knowledge it accumulates.

But compounding works in both directions: if improvements persist, errors can persist too.

Preserving what is known

Medicine contains profound uncertainty. Evidence conflicts. Patients differ. Data are incomplete. Many clinical decisions require judgment under conditions where several answers may be defensible.

Generative models are remarkably useful in precisely these settings. They interpret language, synthesize evidence, connect incomplete information, generate hypotheses, and reason through problems without a single predetermined answer.

Medicine also contains knowledge that can be established exactly.

A unit conversion has an answer. An infusion rate can be calculated. A weight-based dose follows from defined inputs. A clinical score can be computed from explicit criteria. A drug formulation, laboratory reference range, or validated guideline recommendation can often be retrieved from an authoritative source.

Once an answer is available through an exact mechanism, generating it probabilistically introduces unnecessary uncertainty.

This distinction becomes increasingly important as AI systems grow more capable. Medical superintelligence will require generative and normative components working together.

The generative component interprets the clinical problem, reasons across incomplete information, retrieves and synthesizes evidence, chooses appropriate tools, and communicates conclusions.

The normative component executes validated functions, retrieves governed facts, enforces units and bounds, applies explicit rules, and verifies claims against the evidence from which they follow.

This architecture gives the system a body of knowledge and operations whose validity can survive changes to the generative intelligence around them.

A neural model can determine what needs to be known. A deterministic system can establish what follows from known inputs and rules.

The boundary between those domains will evolve. New evidence can turn uncertainty into knowledge. Better measurement can turn judgment into calculation. New discoveries can establish facts that previous generations could only estimate.

Medical intelligence therefore requires the ability to recognize when a question has crossed that boundary.

Exactitude gives the system a way to preserve what medicine has already learned while continuing to reason about everything medicine has yet to resolve.

The problem of recursive error

This becomes especially important when a system begins improving itself.

Today's model errors are often transient. A hallucinated fact may disappear with the end of an interaction. A poor reasoning strategy may vanish when a new model replaces the old one.

A self-improving system changes the consequences of error.

Today's outputs can influence tomorrow's tools, rules, memories, evaluations, data, and learning processes. An error can survive the interaction in which it first appeared. If the system generalizes from that error, the mistake can become embedded in the machinery used to solve future problems.

Repeated over many cycles, a small inaccuracy can become part of the substrate of the system, inherited by every generation that follows.

The ability to improve therefore creates a corresponding need for invariants: knowledge and operations that remain stable as the system changes.

A verified calculation can remain fixed across model updates. A governed fact can retain its provenance. A validated safety constraint can persist across generations. A clinical rule can remain inspectable even as the intelligence deciding when to invoke it becomes dramatically more capable.

Recursive improvement can then accumulate around a growing body of knowledge that has already earned a stronger epistemic status. The system can change without requiring everything it knows to change with it.

Physician governance

A self-improving medical system needs a definition of improvement.

Benchmarks provide part of that definition. They can measure factual accuracy, retrieval performance, calibration, adherence to instructions, and outcomes on standardized clinical tasks. But governance, on the whole, is a problem that computation alone cannot solve.

Clinical quality extends beyond any individual benchmark.

An answer can score higher while omitting a clinically essential qualification. A recommendation can improve average outcomes while becoming unsafe for a particular population. A system can become more efficient while introducing a rare failure with catastrophic consequences. It can optimize successfully for an objective that turns out to be an incomplete representation of good medicine.

Clinical expertise determines which distinctions matter.

Physicians can identify consequential error classes, establish acceptable evidence, define clinically meaningful outcomes, surface edge cases, and determine whether a technical improvement translates into better care.

This principle is at the heart of Doximity's PeerCheck™ program, which has enlisted more than 12,000 U.S. healthcare providers to help improve the quality of Doximity Ask.

As AI systems begin learning continuously, this role becomes even more important.

Physicians move inside the learning process itself.

Their judgment shapes the evaluations from which the system learns, the evidence it treats as authoritative, the failures it prioritizes, the constraints under which improvements are retained, and ultimately the definition of progress the system attempts to optimize.

A medical AI capable of recursive self-improvement therefore becomes inseparable from the community governing its evolution.

Its intelligence grows through an ongoing interaction between machine learning and human clinical knowledge.

A working hypothesis

We believe these observations point toward three properties of medical superintelligence.

First, it must genuinely self-improve. The system must learn from grounded failures, discover generalizable solutions, preserve successful improvements, and eventually improve the mechanisms responsible for future learning.

Second, it must preserve established knowledge. Information that can be retrieved, computed, or verified deterministically should pass through mechanisms capable of retaining its provenance and exactness as the broader system evolves.

Third, its definition of progress must remain clinically governed. Physicians must help determine which failures matter, which evidence is acceptable, which outcomes are meaningful, and whether increasing capability produces better medicine.

Together, these properties describe a system capable of expanding the frontier of its knowledge while protecting the knowledge already established behind it.

It would learn from realistic clinical experience. It would generalize from its failures. Successful solutions would become durable capabilities. Better learning methods would accelerate the discovery of subsequent improvements. Verified knowledge would accumulate as the intelligence surrounding it continued to change.

And physicians would continuously shape the direction in which that intelligence develops.

The result would be less like a static model and more like a living medical knowledge system: a community in which discoveries made by one clinician, researcher, or AI process can be evaluated, verified, generalized, and made available to every participant that follows.

Conclusion

Medicine has spent centuries constructing systems for cumulative intelligence.

Observations became reproducible. Quantities became measurable. Hypotheses became experiments. Experiments became evidence. Evidence became guidelines, calculations, classifications, therapies, and standards of care.

Each generation inherited a larger body of established knowledge and used it as the foundation for the next set of discoveries.

Medical superintelligence extends this process to intelligence itself.

For the first time, we can begin to imagine systems that improve the methods by which medical knowledge is discovered and applied, while preserving the discoveries that those methods produce.

Such a system could allow improvements to compound across millions of clinical questions and scientific problems. A failure encountered in one context could produce a lesson available everywhere. A better method discovered while solving one problem could become part of the machinery used to solve thousands more.

The ultimate promise is therefore larger than an AI that knows more medicine. It is a medical system that becomes progressively better at learning medicine.