AI & Technology
The Watson Hangover
IBM's Watson beat human champions at Jeopardy in 2011; six years later its flagship cancer project was shelved.
2026-08-211 min read
Watson's 2011 Jeopardy win was a real result on a hard benchmark: parse a pun-laden clue, search a large corpus, rank candidate answers, and buzz with calibrated confidence. The inference drawn from it was that the hard part of medicine was the same hard part — retrieve the right knowledge, rank the options — and that a system good at one would transfer to the other. MD Anderson Cancer Center began a Watson oncology project in 2013. It was shelved in 2017, after tens of millions of dollars, without reaching routine use in patient care.
The failure was not that the model was inaccurate on its own terms. Reporting in 2018 on internal IBM documents described physicians calling some of Watson for Oncology's treatment recommendations unsafe and incorrect, and the reason sat upstream of the model. The system had been trained largely on hypothetical cases constructed by a small group of physicians at one institution, so it had learned that group's practice patterns rather than outcomes across a population. A University of Texas audit found the project had been built alongside, not inside, the hospital's electronic records rollout, which left getting real patient data into it a manual exercise. IBM sold off the Watson Health assets in 2022.
A benchmark measures a model in its native environment, where inputs arrive clean, the question is well-posed, and being wrong costs a point. A clinical workflow supplies none of those conditions. The cost of retrofitting is not compute; it is provenance, auditability, and a defined answer for what happens when the system is wrong — work done mostly by people who are not model engineers, and work that does not get cheaper as the model gets better.
Key insights
Benchmark accuracy is measured where inputs are clean and errors are free; production is neither.
Training data drawn from one institution's constructed cases encodes a house style, and it generalizes about as well as a house style does.
Why it matters
Reading a benchmark result as a deployment forecast makes the schedule wrong by years, because the work that dominates the timeline never appears in the benchmark.
Credibility gets spent announcing the model before the integration work has been scoped, so the team starts the real project already behind.
Use this tomorrow
1Take your newest model-backed feature and write down every system it must read from or write to; count how many of those you can produce a schema and a named owner for today.
2Pick three outputs your model produced last week and reconstruct, from stored data alone, why it produced each one; count how many you can finish without re-running the model.
Go deeper
Machine learning has a name for part of this — distribution shift — but the version that ended the project was organizational as much as statistical. A benchmark fixes the input distribution and the scoring function in advance; a hospital fixes neither, and revises both on its own schedule. Gartner's hype cycle labels the aftermath the trough of disillusionment, though that label names a mood rather than the mechanism, which is that integration was always the majority of the work and was never on the plan.
There is a staffing consequence that surfaces about a year in. The work that decides whether a model reaches production — data contracts, audit trails, consent and retention rules, fallback behavior — is not model work, and engineers recruited on the strength of the model rarely want it. Teams that fail to name that work as the main deliverable at kickoff end up assigning it to whoever is left over, usually the person least equipped to argue with a compliance officer about liability.