AI is in the exam room and the question for doctors and patients is to understand not what the diagnosis is but WHY it reached its conclusion and- whether that conclusion should be trusted.
Doctors treat the person with the illness, not the illness. That’s a big difference. They are not just looking at an image or lab values. The diagnosis is often uncertain since available information may be incomplete. The consequences of being wrong can be dire.
As AI is being used more by patients to answer their questions. In fact, almost a third of the US population now uses a chatbot to answer health related questions and many are even uploading their health records. But AI arrives at the statistically most likely answer and still fails at the more important task of deciding what that answer means for that particular patient. There is an important example of how this happens.
In March 2016, AlphaGo played Lee Sedol, one of the greatest Go players in the world. During the second game, AlphaGo made what became known as Move 37. To a human expert, the move looked almost absurd. It was so unexpected that some observers initially wondered whether the machine had made a mistake.
READ: Sreedhar Potarazu | AI is the mirror. The Sutras are the questions. Socrates is still asking. (September 21, 2026)
AlphaGo ultimately won the game and went on to defeat Lee four games to one. What made Move 37 so remarkable was not simply that the machine had found a move no human expected. It was that AlphaGo combined pattern recognition with a mechanism for exploring the consequences of different possibilities.
AlphaGo did not just look at the board and choose the move that appeared most likely based on patterns it had seen before. Its neural networks generated promising possibilities, but its search process explored what could happen after those possibilities were played. It maintained a representation of the game, considered alternative futures and evaluated the consequences of different moves.
But the practice of medicine is more complicated. In Go, the board is visible. The rules are known. The possible actions are defined. The consequences unfold according to a constrained system.
In the doctor’s office today, a patient often arrives with an incomplete history, a multitude of symptoms and emotions, missing information, previous treatments, medications, allergies, genetic factors, social circumstances, and a diagnosis from ChatGPT. This is often a different body of information not available to the LLM, even when there is an AI scribe.
Large language models are good at producing plausible answers. They have learned patterns from enormous quantities of human language and can synthesize those patterns into explanations, differential diagnoses, treatment options and recommendations. Open Evidence is remarkably good at investigating references much faster than humanly possible
But the practice of medicine is not just about an answer. It requires a justification for how we reached that conclusion.
An experienced ophthalmologist can look at an optic nerve and recognize a pattern suggesting glaucoma. AI is also powerful at this kind of pattern recognition. And this may be one of the greatest strengths of modern AI.
AI has been designed to produce longer chains of reasoning. Rather than immediately generating an answer, newer models can break problems into intermediate steps, perform calculations, reconsider earlier conclusions and use those intermediate outputs to arrive at a final response.
Clinical reasoning is different because a physician does not simply generate a sequence of plausible statements about a patient. When confronted with diagnostic uncertainty, the physician develops competing hypotheses, evaluates the evidence supporting and contradicting each hypothesis, identifies missing information and determines which additional observation, test or question would most effectively reduce the uncertainty.
As new evidence arrives, the physician changes the relative probability of those hypotheses and may eventually abandon the original diagnosis altogether. The reasoning process is therefore dynamic because the physician is maintaining and updating a model of the patient’s clinical state rather than simply generating an increasingly elaborate explanation.
Showing the math
Current large language models generally do not provide a clinical ledger of reasoning inside the model that allows us to examine which facts it considers established, which findings it regards as uncertain, which competing diagnoses it is maintaining, what evidence supports each diagnosis, what evidence contradicts them and what information would cause it to change its conclusion.
The system may therefore produce a highly sophisticated answer without giving the physician a reliable way to determine whether the model actually evaluated competing hypotheses or simply generated the language most strongly associated with the clinical circumstances presented to it.
This is important in medicine because the model may tell us that a patient most likely has condition A, but it may not be clear whether conditions B and C were actually considered and rejected because of contradictory evidence or whether they were simply less statistically probable within the patterns learned by the model. It may also be unclear whether the system recognized that an important piece of information was missing, whether it understood that the patient falls outside the population represented in the data used to develop the model, or whether it knew what additional evidence would cause it to change its conclusion.
This creates a particularly difficult problem for medicine because an explanation generated after an answer is not necessarily the same thing as an auditable record of how the answer was reached.
Physicians can certainly make mistakes, and clinical reasoning is hardly perfect, but a doctor is expected to be able to explain the basis for a diagnosis, identify the alternatives that were considered, describe the evidence that influenced the decision and reconsider the diagnosis when new information contradicts the original hypothesis.
With many current AI models, we can examine the language of the explanation without necessarily having access to an equivalent underlying representation of the system’s beliefs and the changes those beliefs underwent during the reasoning process.
READ: Sreedhar Potarazu | AI is burning through free cash flow (July 25, 2026)
The result is a new form of clinical risk in which an AI system can be highly persuasive without being equally accountable. An obviously incorrect answer is relatively easy to recognize and challenge, whereas a medically plausible answer that is supported by a convincing explanation may be accepted precisely because it appears to demonstrate a depth of reasoning that the underlying system may not actually possess.
The diagnosis is therefore a working hypothesis that should become stronger or weaker as evidence accumulates, rather than a fixed conclusion generated from the first set of observations.
Treatment has options
The limitations of current AI become even more consequential when the system moves from diagnosis to treatment recommendations because treatment is not simply a matter of identifying the therapy most commonly associated with a particular disease.
Treatment is a decision under uncertainty in which the potential benefits and risks of several alternatives must be considered in relation to the characteristics, circumstances and preferences of an individual patient.
A clinical guideline may establish that treatment A produces better outcomes than treatment B in a particular population, but a physician does not treat an abstract population. The physician treats an individual who may have impaired kidney function, another medication that creates an interaction, a history of an adverse reaction, financial limitations, difficulty accessing follow-up care or preferences that materially alter the balance between risk and benefit.
The fact that AI has access to all of these facts does not necessarily mean that it has constructed an explicit clinical state in which those facts are factored in the treatment decision.
The problem, therefore, is not that AI will always give the wrong treatment recommendation. In many circumstances it may provide an excellent recommendation. The problem is that physicians may not know when the recommendation reflects genuine integration of the patient’s circumstances and when it represents a statistically plausible continuation of medical knowledge.
Probability and prediction is enormously useful, but it cannot substitute for individual clinical judgment when the characteristics of the patient materially alter the decision.
Anchored to AI
There is another risk that has do with human psychology. Instead of beginning with the question, “What do I think is happening?” the physician may unconsciously begin with the question, “Does the AI’s recommendation make sense?” Those questions sound similar, but they represent very different cognitive processes. The first requires the physician to generate and evaluate hypotheses independently, while the second begins with an algorithmically generated hypothesis and asks the physician to validate it.
Once that shift occurs, the physician can become anchored to the AI’s interpretation, particularly when the system communicates its recommendation with confidence and provides a coherent explanation. Contradictory evidence may receive less attention because the machine has already provided a narrative that organizes the available information into a seemingly logical conclusion.
Patients may face an even greater challenge because conversational AI removes much of the friction associated with obtaining medical information. A patient can describe symptoms in the middle of the night, upload a photograph, paste a pathology report into a chatbot or ask whether a physician’s recommendation is consistent with current medical literature, and receive an immediate explanation written in language that may be considerably easier to understand than a medical journal or clinical guideline.
That access can be enormously empowering, but it also creates a new form of health literacy. Patients will need to understand that an AI-generated explanation is not necessarily evidence and that a probable diagnosis is not equivalent to an established diagnosis.
A patient asking whether a particular finding could indicate glaucoma, for example, may receive a response explaining that the findings and risk factors are consistent with glaucoma. The patient may interpret that statement as meaning that the AI has diagnosed glaucoma, even though the system may not have examined the optic nerve, measured intraocular pressure, evaluated the visual field or considered the full clinical history. The patient has received an answer, but the patient has not necessarily received a diagnosis.
This creates a paradox in which AI can make doctors patients dramatically more informed while simultaneously making them more confident in conclusions that remain uncertain.
AI for a reason
The lesson of AlphaGo was not just that machines can imitate human intuition. What made AlphaGo remarkable was the combination of pattern recognition with a mechanism that could represent a position, explore possible futures, evaluate consequences and use that process to select an action. Medicine requires an even more sophisticated version of that idea because the clinical “board” is never fully visible, the patient’s state changes over time, the available information is incomplete and the consequences of an action are uncertain.
The next generation of medical AI should therefore move beyond systems that are primarily optimized to produce the most plausible answer and toward models capable of maintaining explicit representations of clinical knowledge, uncertainty and competing hypotheses.
The ultimate test of medical AI should not be whether it can sound like a doctor or even whether it can achieve a higher diagnostic accuracy than an individual physician under controlled conditions.
AI should participate in a reasoning process that is sufficiently transparent and evidence-based for a physician to challenge its assumptions, understand its uncertainty and recognize when the machine may be wrong.
AI may produce the right answer often enough, explain that answer persuasively enough and integrates into clinical workflows deeply enough that physicians and patients gradually stop asking how the system reached its conclusion, what evidence would change it and whether the patient in front of them is sufficiently similar to the population from which the AI learned its patterns.
In medicine, where every diagnosis belongs to a particular person and every treatment carries consequences that cannot be reversed simply by generating a better answer afterward, the difference between prediction and reasoning is not academic
It may ultimately determine whether artificial intelligence becomes a tool that genuinely improves clinical judgment or merely a more sophisticated way of making plausible decisions without fully understanding why they are right.


