Case 24 is a type A aortic dissection presenting as a stroke. Liu and colleagues gave it, with 53 other hard emergency cases, to six language models, three times each (J Med Syst, published 23 September). The two models that came out safest across the study, DeepSeek R1 and Gemini 3 Pro, reasoned through it correctly twice. On another run, each recommended antithrombotic treatment that the authors rated potentially fatal.[1]
That wasn't a one-off. Across the study, 34 model–case pairs produced a dangerous recommendation at least once. In 31 of them, the danger showed up on some runs and not others.
I haven't met Liu's patient. I have met a relative: a descending aortic dissection that arrived looking like an embolism. An immediate IV bolus of unfractionated heparin could have led to a fatal aortic rupture or massive internal haemorrhage. That time, 30 more seconds of history and examination prevented a catastrophic error.
Getting that patient right two times in three is a hit rate. Nobody writes a protocol around a hit rate.
My instrument did the same thing
I build a tool that uses a language model to score clinical AI products for hospital buyers. I score other people's tools on whether they can show their working. I am not exempt from the standard I am selling, so here is what mine did in June.
The tool scores a vendor's evidence across weighted domains, then applies fixed rules. One rule is a safety interlock: if Population Validity or Operational Performance scores below 30, the verdict is Do Not Proceed, whatever the total says.
For the Epic Sepsis Model, I scored the same evidence pack five times, as five independent passes.
| Domain | Five runs | Spread |
|---|---|---|
| Clinical Evidence Quality | 40, 40, 40, 40, 50 | 10 |
| Population Validity | 20, 20, 20, 20, 30 | 10 |
| Operational Performance | 20, 30, 20, 20, 40 | 20 |
| Workflow Integration | 60, 70, 60, 80, 70 | 20 |
| Regulatory Compliance | 40, 40, 40, 40, 40 | 0 |
| Data Governance | 40, 50, 40, 60, 60 | 20 |
| Implementation Maturity | 40, 60, 40, 60, 60 | 20 |
Run one: Do Not Proceed. Runs two, three and four: the same. Run five: Population Validity came back 30, Operational Performance 40, the interlock stayed down, and the rules would have returned "significant concerns, remediate" instead. I changed nothing between those runs. Not the evidence, not the prompt, not the rubric. Only the roll.
A hospital acting on run five buys a different future from a hospital acting on run two.
What I found when I went back
Checking the numbers for this piece, I found errors in the reliability note I published in June.
The model reports its own weighted total alongside its domain scores. The tool doesn't trust that number. It recomputes the total from the domain scores, so no published verdict ever rested on it. The June note quoted the model's totals anyway, and several of them are miscalculated. The note gave the first-round spread as 14.5 points. Recomputed, it's 11.5: from 37 to 48.5. The model's own recommendation, which the tool also overrides, disagreed with the rules in three runs of five.
So the model was inconsistent twice over, in its judgement and in its arithmetic. Only the judgement reaches a report.
Two clues, and the fix
One domain held steady across all five runs: Regulatory Compliance, 40 every time. It was also the only domain carrying a written evaluator note tying its evidence to a band.
The unstable domains had something else in common. Each was mixed-signal: something defensible, such as a favourable architecture or a demonstrated integration, sitting next to silence in the public record. The model had to resolve that tension into one number, from scratch, on every run. Hand a model an unresolved judgement and it will resolve it differently on different runs.
So I wrote more notes. Each says which side of a band boundary the evidence sits on and cites the facts behind it. For Population Validity, the note says that no evidence at deployment scale counts as insufficient, not limited. None names a score. Every note prints verbatim in the report's appendix. Two went in for round two, three more for round three.
| Round | What changed | Spread of totals | Interlock fired |
|---|---|---|---|
| 1 | One evaluator note (Regulatory Compliance) | 11.5 | 4 of 5 |
| 2 | Notes on the two safety domains | 0.45 | 5 of 5 |
| 3 | Notes on the remaining mixed-signal domains | 0 | 5 of 5 |
Totals recomputed from the domain scores. The June note printed the model's own figures (14.5, 3.1, 1.5).
That is a short step from a thumb on the scale. The stability came from my judgement, written down and disclosed. The notes pinned the reading the model had already reached most often, Population Validity at 20 in four runs of five, but they are still my reading. What separates calibration from contamination is that calibration is visible and arguable. A buyer can read every note and argue with it.
It is one vendor, one evidence pack, before and after. It shows the verdict is stable. It doesn't show the verdict is right.
Two other failures the safeguards caught
When the interlock fired, the model's summary prose still read "proceed with caution" directly beneath the Do Not Proceed banner. Rules now write the summary whenever the interlock fires, and no model prose gets in to contradict it.
On two runs the model picked red flags whose premises were false for this vendor, including a mismatch with a regulatory filing that doesn't exist. Guardrails caught both. The validator still checks that a flag is on the list, not that the evidence supports it. That one isn't fixed yet.
What else I got wrong
The model behind all of this was DeepSeek V3, served with 4-bit weights. Neither the June note nor the v1.0 working paper said so. The run records also didn't log which provider served each call. The note names one. I can't show it served every call.
That means the variance above was measured on a 4-bit model, and I can't separate how much came from the rubric, the prompt or the quantisation. The June note said the leftover variance wasn't sampling noise and wouldn't fall with a lower temperature. Nothing I ran tested that. I've withdrawn it, and the June note now carries a dated correction listing each of these errors.
What this doesn't do is move the Epic verdict. With the notes in place, Population Validity and Operational Performance returned 20 on every run, against a threshold of 30. They failed by ten points, not one. A verdict is fragile when it sits on a boundary, and this one doesn't. That's a reason to think it holds. It isn't a substitute for re-running at a declared precision.
The DAX Copilot evaluation needs a disclosure of its own. My evaluator anchors set its domain scores, and I told the model to reproduce them. Those scores are my judgements run through the rubric. The model rated nothing independently. The paper should have said so.
What changes
The revision I'm working on has several models from different developers rate the evidence without seeing each other's scores. Every call logs which provider served it, and its precision wherever the provider discloses it. When the models disagree, a named person settles it, and there's no majority vote. Agreement figures go out with every report.
None of it has been tested yet. When the pilot runs, the results get published whatever they show.
What to ask
If my evaluator could flip between runs, a product that generates its output can too. Liu measured it in six general-purpose models. Ask the vendor, and ask whoever evaluates the product for you:
- How many times did you run the same cases, and did the answer change?
- Which model, which version, at what settings and precision, served by whom? Are those the settings you'll ship?
- When it's wrong on a case, is it wrong every time or some of the time?
- If the answer moved, what did you do about it, and can I read it? A fix you can't inspect looks exactly like a thumb on the scale.
The third matters most. An error that happens every time can be found, and the product fenced away from it. One that turns up every third run can pass a single test clean.
On the same cases, one run is an anecdote.
Dr Kpakpo Acquaye is an emergency medicine physician (MBBS UCL, MRCEM) working on clinical AI evaluation for health systems in the Gulf and Asia-Pacific, independent of any vendor. The consistency studies described here were run on a retrospective desk evaluation using the published record; they do not imply field deployment of the evaluation instrument.