Can an AI Actually Diagnose You? Here’s Where the Technology Stands
Type a set of symptoms into a chatbot and you’ll get an answer in seconds, delivered with total confidence. Whether that answer deserves the confidence is a harder question than most headlines admit. In a randomized trial published in Nature Health in February 2026, GPT-4o working alone scored 82.9% on six written diagnostic cases, higher than the doctors who were allowed to use it. Those were cases on paper, though, and the distance between paper and patient is where the real story sits.
What the Best Evidence Actually Shows
The Nature Health trial recruited 60 licensed physicians in Pakistan, and 58 of them finished. Every one first completed a 20-hour course on what large language models can and can’t do. Half then worked through up to six clinical cases with GPT-4o alongside their usual resources, such as medical databases and search. The other half used only the usual resources.
The gap was large. Doctors with the AI averaged 71.4% on a rubric that graded the whole reasoning process, from differential diagnoses to proposed next steps. Doctors without it averaged 42.6%, a difference of 27.5 percentage points. Time per case was about the same in both groups, so the AI didn’t cost anyone speed.
Then the researchers ran the AI by itself, three times per case. It averaged 82.9%.
That’s the number that travels. It scored higher than both groups of doctors. The fine print travels less well, and the authors put it in their own limitations section: the cases were written vignettes adapted from real patients rather than real patients, only six were used, and only one model was tested. They also note that the doctors without AI scored far below the 74% reported in an earlier US trial, so a boost of this size may say as much about the setting as about the software.
One detail complicates the tidy version of the story. The AI alone outscored doctors using it, yet in 31.4% of cases the physician-plus-AI pairing beat the AI’s median performance. The authors read that as a sign the two can complement each other: the model faltered on cases that needed nuanced clinical judgment or subtle red-flag context, and trained physicians could cover those gaps.
The Catch Behind the Good News
The same research group ran a second randomized trial built around a different question: what happens when the AI is wrong? In that study, posted as a preprint in August 2025, 44 AI-trained physicians worked through six cases with ChatGPT-4o recommendations available on request. Half received clean recommendations. The other half received recommendations with deliberate, clinically meaningful mistakes planted in three of the six cases.
Doctors who got clean advice scored 84.9% on diagnostic reasoning. Doctors who got the flawed advice scored 73.3%, an adjusted gap of 14.0 percentage points. On the single most likely diagnosis, accuracy fell from 90.5% to 76.1%.
Nobody forced them to look. The doctors chose whether to consult the AI, and roughly two-thirds did in both groups.
That changes how the first trial reads. The AI’s advantage is real, but it comes with a condition. When the advice is right, the people using it do better. When it’s wrong and sounds just as sure, they inherit the mistake.
Experience didn’t work as a shield either. In the subgroup breakdowns, doctors with 10 or more years in practice lost 16.6 points when the advice was flawed, against 9.1 for their less experienced colleagues. With only 44 doctors split into smaller groups, those figures are rough estimates, but they point away from the comfortable assumption that seniority protects against a confident machine. The first trial had a mirror-image pattern on the upside: doctors at or below the median of 8.5 years in practice gained 30.0 points from the AI, versus 24.5 for the more experienced.
One caution runs the other way. The authors of the second trial flag that mistakes in real clinical use may be subtler and harder to spot than the ones their team planted on purpose, so the study can’t say how large the effect would be outside the lab.
What This Means If You’re the Patient
Neither trial tested a patient reading an AI’s answer about their own symptoms. Both tested trained physicians, working through written cases, with a chatbot. That’s a long way from typing symptoms into an app at midnight, and the studies don’t say what happens across the distance.
What they do show is where the risk lives. In the second trial it wasn’t the model’s raw ability that pulled scores down, it was deference. These physicians had 20 hours of training in how such tools fail, and a wrong recommendation still moved their scores. It’s fair to assume a layperson without that training has less to check an answer against, not more.
The first trial adds a clue about preparation. Its authors point to an earlier trial in which doctors handed GPT-4 with no AI-literacy training saw no improvement in diagnostic reasoning. Set beside their own result, the pattern suggests that how people are prepared to use the tool matters about as much as the tool itself.
That points to a sensible way to use these tools, and it’s an interpretation rather than a finding: treat an AI answer as a list of questions to bring to a clinician, not as a verdict. For anything worrying, the clinician comes first.
Sites Worth Bookmarking for the Next Chapter
If a study like this leaves you wanting more plain-English explanation of the technology behind it, a few sites are worth bookmarking. The Decoder, founded in 2022 and part of heise medien since 2024, reports on AI without the hype. Metamandrill, active for 4 years, explains the technology changing how people live, work, and interact, such as its AI Doctor coverage. Freethink, a solutions-focused media company, tells daily stories about the people and technology building the future.
So can an AI diagnose you? On paper, it can outscore a room of trained doctors. The open question, and the one neither trial answers, is who checks the answer when the person reading it was never trained to doubt it.