AI & Technology
Where AI Quietly Makes Experts Worse
In a 2023 Harvard–BCG study, GPT-4 lifted consultants’ ideation work, yet those given a subtler business case got it wrong more often.
2026-09-201 min read
Generative AI does not have a clean border of competence. It is superb at some tasks and confidently wrong at others that look almost identical, and nothing on the screen tells you which side of that line you are standing on. Researchers at Harvard Business School and Boston Consulting Group called this the jagged technological frontier, and in 2023 they measured what happens when professionals cross it without noticing.
Nearly 800 BCG consultants were randomly assigned to work with or without GPT-4. On creative product-development tasks inside the model’s range, the AI users finished more work, faster, and produced output that graders rated markedly better. A separate task was built to sit just outside that range: recommending where a company should invest by combining spreadsheet figures with interview notes, where the right answer hinged on a detail in the interviews that the numbers hid. There, consultants using GPT-4 were about 19 percentage points less likely to reach the correct recommendation than those working alone. The model’s answer was fluent and plausible, and many of them adopted it.
The lesson is not that the model is bad; it is that a good model lowers your guard. Fabrizio Dell’Acqua, the study’s lead author, had earlier found the same pattern with recruiters: people handed a more accurate AI screened résumés less carefully than people handed a weaker one, and made worse calls. Your judgment matters most exactly where the tool’s help feels most convincing, so decide in advance which kinds of work you will verify line by line.
Key insights
AI’s competence border is jagged: tasks that look alike can fall on opposite sides of it.
In the Harvard–BCG study, the same tool that raised quality inside its range cut accuracy on a task just outside it.
Better tools can make checking lazier, so verification has to be a rule you set in advance, not a feeling you wait for.
Why it matters
The errors AI introduces cluster in subtle, judgment-heavy calls, which are exactly the ones clients and bosses pay experts to get right.
Because quality rises on routine work, team-level results can show AI helping even while the rare, costly mistakes quietly increase.
Use this tomorrow
1List the five task types you used AI for this week and mark each “inside” or “outside” based on whether you have ever caught it making a substantive error there.
2Before accepting the next AI answer on a judgment call, write your own one-line conclusion first, then compare; count how many of your next three comparisons disagree.
Go deeper
The study, “Navigating the Jagged Technological Frontier,” was released as a Harvard Business School working paper in September 2023, with Fabrizio Dell’Acqua, Ethan Mollick and Karim Lakhani among its authors. Mollick and his co-authors described two patterns among consultants who used the tool well: “centaurs,” who divided the work cleanly between themselves and the AI, and “cyborgs,” who worked with it continuously, checking as they went. Neither pattern involved simply passing the model’s answer along. The useful point is that this is a choice about workflow, not about which model you buy.
The effect has an older name in aviation and medicine: automation bias, the tendency to favor a machine’s suggestion over contradicting evidence in front of you. Linda Skitka and Kathleen Mosier documented it in the 1990s, finding that people monitoring automated aids missed problems the system did not flag and followed its advice when it was wrong. Generative AI widens the exposure because its mistakes arrive in the same confident prose as its correct answers. The old remedies still apply: explicit cross-checks, and making one person accountable for the final verdict rather than for the draft.