AI (artificial intelligence) systems can now outperform emergency department physicians at diagnosing patients during triage, the crucial first assessment when a patient arrives at hospital. Research published in Science tested a large language model (LLM), the same technology powering ChatGPT, using actual clinical notes from a Boston emergency department. At triage, when information is scarcest and the stakes are highest, the AI correctly identified the diagnosis or something closely related in 67% of cases. Two comparison doctors managed just 50% and 55%. That gap matters enormously when missing a serious condition could mean death. The study represents a significant leap from earlier AI research that focused on passing medical licensing exams, impressive party tricks that said nothing about real world clinical utility. This time, researchers used genuine patient records across multiple decision points in emergency care, making the findings directly relevant to actual medical practice. The implication is clear: these systems could help doctors think through broader differential diagnoses, particularly when the priority is not missing something catastrophic. But the study's limitations are equally important. The AI worked entirely from written text. It never saw the patient's face, never noticed respiratory distress or the subtle signs of sepsis, never examined them physically, never spoke to worried family members, and never bore responsibility for what happened next. It was offering sophisticated pattern recognition on curated text, not practicing emergency medicine in the chaotic reality of a crowded department at 3am on a Saturday. There is also a chasm between producing accurate diagnostic suggestions and actually improving patient outcomes. A longer list of possibilities might prompt useful thinking, but it could equally trigger a cascade of unnecessary tests, over treatment, clinician fatigue, or dangerous overconfidence in a plausible sounding answer that turns out to be catastrophically wrong. Some benchmark cases used in studies like this may have been in the AI's training data, though this does not invalidate the emergency department findings. The more urgent problem is that clinical practice has already raced ahead of governance. A Royal College of Physicians survey found 16% of UK doctors using AI tools in clinical practice every single day, with another 15% using them weekly. That means roughly a third of physicians are integrating these systems into their workflow before hospitals, the NHS (National Health Service), or regulators have established protocols for testing them, training staff to use them safely, detecting when they cause harm, or determining liability when things go wrong. The technology is already embedded in patient care, operating in a regulatory vacuum.
❤️ health
AI Beats Doctors at Triage, Nobody Knows Who's Liable
A new study in Science shows AI diagnosing emergency patients more accurately than doctors at the critical triage stage - 67% versus 50-55%. Sounds impressive until you realize thousands of UK doctors are already using these tools daily, and nobody has figured out the governance, liability, or safety protocols. The technology has sprinted past the guardrails.
My Take
The headline 'AI beats doctors' misses the point entirely. This is not about whether AI can outperform humans at specific diagnostic tasks. That ship has sailed. The real scandal is that we are sleepwalking into widespread clinical use of powerful systems that nobody has properly figured out how to govern. Saying 'keep a human in the loop' is comforting nonsense. Which human? With what training? With authority to override the AI based on what criteria? When a tool quietly starts failing, and all complex systems eventually fail, who spots it, who fixes it, and who is legally responsible? These are not hypothetical questions. One in six UK doctors is using AI daily right now, in a system where those questions remain unanswered. The correct path forward is not to ban these tools or pretend they don't exist. It is to demand rigorous real world trials, establish them as second opinion support rather than decision makers, measure what actually matters to patients, and build governance structures before, not after, widespread adoption. Instead, we are doing the opposite: deploying first, asking questions later, and hoping nobody gets hurt in the gap.
What Happens Next
Within six months, expect the first serious adverse event tied to AI diagnostic tools in a UK hospital, not because the technology is fundamentally broken, but because the governance infrastructure does not exist to catch predictable failures. A patient will be harmed by over testing triggered by an AI generated differential, or by a missed diagnosis because a junior doctor deferred too heavily to a confident sounding algorithm, and suddenly the liability question will move from academic journals to courtrooms. The NHS will respond by rushing out guidance rather than real governance, probably some variation of 'clinicians remain responsible for all decisions' that solves nothing and protects nobody. Meanwhile, commercial AI vendors will continue selling into this regulatory vacuum, because there is no enforcement mechanism to stop them. Some hospitals will quietly pull back from AI tools after the first lawsuit, while others will double down, creating a postcode lottery of AI adoption across the country. The genuinely interesting scenario is if the Royal Colleges step in before regulators and establish their own peer reviewed certification scheme for clinical AI tools. It would be messy and imperfect, but it might actually create accountability faster than waiting for government action. The alternative is years of reactive regulation written in response to preventable tragedies, which is typically how the UK handles emerging technology in healthcare.
What History Tells Us
This governance gap has precise historical parallels with the introduction of CT (computed tomography) scanners in the 1970s. The technology was revolutionary, genuinely improved diagnostic accuracy, and was adopted rapidly across hospitals before anyone established protocols for appropriate use, radiation exposure limits, or clinician training standards. The result was years of overuse, unnecessary radiation exposure, and eventual recognition that powerful diagnostic tools require systematic governance, not just clinical enthusiasm. It took regulatory bodies nearly a decade to catch up with the technology that was already embedded in daily practice. The AI situation is moving faster and potentially carries higher stakes, because CT scanners were tools that physicians operated with clear accountability, while AI systems make autonomous recommendations that can subtly shift clinical decision-making without clear lines of responsibility. We are repeating the same mistake at higher speed.