Three papers report a reasoning model matching or beating physicians, a fourth starts from the image itself, and the subject of this issue is the evidence underneath each claim.
RadBrief EP12.
A note before this issue: RadBrief issues will be more spaced out from here, without a fixed schedule. Weekly was fine for covering papers, but the pieces I want to write take longer than a week and are less about what a study found than about whether the way it was evaluated supports the claim. Shorter thoughts will go up on LinkedIn in between.
Between April and July, three journals published three papers reporting that a reasoning model matched or beat physicians: o1-preview in Science, an autonomous agent in Nature, GPT-5-Thinking in JMIR. The claims are virtually identical, but the evidence is different in every case: different tasks, different comparator groups, different inputs, and in one case a scale unable to separate the groups it compared. A fourth paper (RadFabric) makes no claim against physicians and belongs with these three as the only study attempting to work from the image rather than from text generated by a radiologist. The question this issue asks of each paper is what its evidence actually supports.
Science, 30 April 2026. DOI 10.1126/science.adz4433
The paper reports six experiments comparing o1-preview against physicians, together with a blinded second-opinion study on real emergency department cases. The main virtue of this paper is that it compares to actual physicians rather than providing a figure from an earlier publication. Most work in this area compares one model against a previous model, which tells us nothing about whether either is close to what a clinician would do. Running the comparison properly is expensive, and this one was: the probabilistic reasoning experiment alone was benchmarked against 553 practicing clinicians. The largest difference came in management reasoning, where o1-preview scored a median of 89% per case, against 42% for GPT-4, 41% for physicians who had GPT-4 available, and 34% for physicians working with conventional resources. In the mixed-effects model that puts o1-preview 48.4 percentage points above the conventional-resources group, 95% CI 38.3 to 58.5, P < 0.001.
That 89% was derived from five clinical vignettes, designed by a panel of 25 physicians in consensus, so it is reasonable to ask how it would hold up at a realistic sample size, where the effect would probably be smaller. The model also did not win everywhere. On landmark diagnostic cases the difference of 20.2 percentage points did not reach significance, at P = 0.055, and on probabilistic reasoning it was only modestly better than GPT-4. Two further limitations appear in the discussion. The authors state that the study tested text-based performance only, for the models and the humans alike, and that current foundation models reason less well over nontext inputs, naming imaging specifically. They also state that these benchmarks depend on cases clinicians curated and cleaned, and on questions originally written for teaching, which may overstate how a model performs against the messy data of a real workflow. Both limitations point the same way. The result is real, but it was produced under conditions that a working department does not provide.
Nature, 17 June 2026. DOI 10.1038/s41586-026-10675-5
MIRA is an AI agent that runs a full emergency department workflow inside a sandboxed record built to the FHIR standard. It takes a history, orders labs and imaging, prescribes, and admits, rather than answering a question about a case that someone else has already assembled. Across more than 500 cases drawn from MIMIC-IV it reached 87.8% diagnostic accuracy, against 78.1% for board-certified physicians and 71.1% for a mixed cohort of attendings and residents. Imaging is handled differently from everything else in that workflow. MIRA never sees an image. It requests a modality and a region, and what comes back is report text, delivered as a FHIR observation. When the agent is right about an imaging finding, it is right because the radiologist’s report was right, and the model adds nothing at that step. Nothing in the paper tests what the agent does when the report it receives is wrong, incomplete, or hedged.
J Med Internet Res 2026;28:e91733. DOI 10.2196/91733
The study put 100 consecutive lung cancer cases in front of the multidisciplinary team at University Hospital of Split and, in parallel, in front of two reasoning models. The cases reached the models as verbatim radiology and pathology reports, not as images, so here too the imaging input was text a radiologist had written. Two lung oncologists then graded every recommendation on a scale of 1 to 5. GPT-5-Thinking averaged 4.90 against the team’s 4.14 in the first phase, and 4.79 against 4.34 in the second, when the team knew it was being benchmarked against AI. Deepseek-v3-r1 scored above the P=0.15. The authors open their own Results section by stating that ratings showed ceiling effects, and that is the more important finding than the comparison itself. When the human comparator already scores 4.14 out of 5, only 0.86 of the scale is left for any model to demonstrate an advantage in, and even less is left to show that a model is worse. A scoring system that places all results in the top fifth of its allowed scale cannot differentiate well, regardless of the p values.
NPJ Digital Medicine, 15 July 2026. DOI 10.1038/s41746-026-02994-8
RadFabric makes no claim against physicians, and it is in this set as the contrast case: it attempts the step the other three papers left to a radiologist, producing the findings from the image itself rather than receiving them as written text. It combines fourteen chest radiograph models with two vision-language models, coordinated by one agent that maps findings to anatomy and a second, trainable agent that reconciles the component models when they disagree. On MIMIC-CXR, across fourteen pathologies, the assembled system reached a mean AUC of 0.852, with its largest gains on the rare classes where a miss costs most. The paper also reports a control condition, and it changes how the headline should be read. The fine-tuned vision-language model, running on its own without the surrounding structure, scored 0.4984, which is chance. Almost all of the performance is coming from the orchestration rather than from the model that looks at the image. The paper contains no radiologist baseline at any point, runs on a single retrospective dataset, covers the chest only, and the authors describe it as a research prototype. It is also the only one that needed an elaborate apparatus in order to work at all.
In three of the four papers, the system never interpreted an image. Wherever a case involved imaging, it read text that a radiologist had written. The worked example in the MIRA paper returns the sentence that a CT of the chest showed signs of infection in the left upper lobe. Brodeur’s conference cases arrive with the imaging findings already described. Viculin’s models were given verbatim radiology reports. The single system that worked from pixels needed fourteen specialized models and a trained reasoning layer to reach an AUC of 0.852, and when its vision-language model was tested on its own, it performed no better than guessing.
That is not an argument that imaging is insulated from any of this. It is an argument about which part of the problem is still open. The reasoning layer is now good enough to beat physician baselines when the information arrives already organized. Turning an image into that organized information is the part nobody has solved, and at the moment a radiologist’s report is what bridges the two. Every accuracy figure in these four papers therefore contains a contribution from radiologists that none of the papers measured or controlled for.
This gives us three questions, and these are the ones that I would ask of another paper like this.
The first is whether there is a human baseline and whether the humans were tested under the same conditions on the same task. RadFabric has no human baseline at all so there is no way to know whether an AUC of 0.852 is better or worse than a radiologist reading the same films. Viculin has two oncologists scoring the team’s written output, which measures the quality of a recommendation as opposed to both humans and models carrying out the same test. Brodeur has genuine baselines, and building them cost hundreds of clinician hours, which is most of the reason so few other groups have them.
The second is what the model actually received. Brodeur’s cases were curated and cleaned before the model saw them, by the authors’ own account, and MIRA’s imaging input was a finished report rather than a study. Viculin is the exception, and its verbatim records make it the harder test of the four. This question matters outside the literature as well. A vendor figure produced on pre-organized input carries the same limitation as Brodeur’s, and a system evaluated on finished reports is being evaluated partly on the quality of our own work.
The third is whether the measure used can actually separate the groups being compared. Viculin’s could not, because the human comparator was already at 4.14 out of 5. Brodeur runs into the same difficulty in a different measure. Physician agreement on test selection was 87% in raw terms, but kappa was only 0.26, which the authors attribute to severe class imbalance: when nearly every answer is scored correct, there is almost nothing left for two raters to disagree about, and the agreement statistic collapses even though the raters agreed. The instruments differ, but the problem is the same in both papers. The measurement could not distinguish between the things it was being used to compare.
Brodeur’s authors say the bottleneck has moved from model capability to how these systems are designed to work with people, and that imaging is where the models are still weakest. That is the problem I’m building RadReason for: reasoning support during the read, with the radiologist reading the whole study and responsible for the case. The baseline question applies to it too. Nobody has measured what a reasoning aid does to a radiologist’s own accuracy, and I can’t answer that yet either.