n o ren
AI & Technology

When AI Claims Expertise, It Tries to Replace It

IBM trained Watson for Oncology on one elite cancer center's judgment, then sold it worldwide as expertise.

An expert system's confidence and its coverage are set by different things, and only one of them is visible to the user. Coverage comes from what the model was trained on; confidence comes from the interface. A ranked list of treatment options with a highlighted first choice looks identical whether it rests on a thousand institutions or on one, and the reader has no way to tell which. That gap is where the failure lives: the system does not overstate its accuracy so much as understate its provincialism, presenting a single institution's convention as settled practice.

IBM began training Watson for Oncology with Memorial Sloan Kettering in 2012, and by the mid-2010s was selling it to hospitals in Asia and beyond. STAT News reported in September 2017 that the system had largely been taught on synthetic cases authored by a small group of Sloan Kettering physicians rather than mined from real-world outcomes. What Watson had learned was how one leading American cancer center preferred to treat, encoded as though it were oncology itself. Doctors elsewhere found its recommendations diverging from their own guidelines, with no visible reason why. MD Anderson, which had spent roughly $62 million on a related Watson project, shelved it in early 2017 without reaching clinical use, according to a University of Texas System audit. IBM sold the Watson Health assets to Francisco Partners in 2022.

No regulator shut Watson down, and no patient injury was ever established — which is what makes the case useful rather than lurid. The product failed because its answers could not explain their own jurisdiction. Any system that returns a single best answer inherits this problem, and the remedy is not more accuracy. It is making the boundary of the training set as visible as the recommendation sitting on top of it.

Coverage comes from training data, confidence comes from the interface, and users only ever see the second.
Watson for Oncology encoded one American cancer center's treatment preferences and shipped them abroad as general expertise.
The remedy is exposing the training boundary, not raising the accuracy number.

A system whose training boundary is invisible gets trusted outside that boundary by default, which is exactly where it is weakest.

Buyers evaluating AI tools test accuracy on familiar cases — the one condition under which a provincial model looks universal.

1
Take the AI tool your team relies on most and write down in one sentence whose data or judgment it was trained on; if nobody can answer within five minutes, put that question to the vendor in writing this week.
2
Pull the last ten recommendations that tool produced and count how many carried any indication of what it was uncertain about; a count of zero means the interface is asserting a confidence the model never expressed.

The underlying issue is distributional rather than technical. A model trained at a single high-volume referral center sees a patient mix skewed toward complex and rare presentations, and it learns decision rules tuned to that mix. Deployed at a community hospital with a different population and different drug availability, the same rules can be defensible medicine in one place and poor advice in another, with nothing marking the transition. Medicine calls this external validity, and it quietly defeats systems that never claim to be wrong about anything.

The commercial dynamics amplified the problem. Watson's fame came from winning Jeopardy! in 2011, a general-knowledge feat with little bearing on clinical reasoning, and that reputation did sales work the oncology system underneath could not support. Buyers were purchasing a brand's implied breadth rather than a documented evidence base. When the gap finally surfaced, it surfaced through journalists — not through any evaluation the deployments themselves had required.