When a language model spits out a paragraph faster than a human can read, the immediate reaction is to trust the output and move on. The hidden cost, however, is that speed creates a false sense of completion, prompting users to over‑edit, re‑prompt, and double‑check every sentence. This “latency‑induced polishing” stems from a cognitive bias: the brain treats rapid, algorithmic output as low‑effort, so it feels obligated to add human effort to restore perceived value. The result is a feedback loop where the original time‑saving disappears, and the workflow becomes slower than a manual draft.
Consider a data‑science team of eight that adopted an internal LLM for generating quarterly narrative reports. The model produced a first draft in under a second, but the team’s analyst, Maya, spent the next forty‑five minutes re‑phrasing each bullet, checking citations, and inserting domain‑specific jargon. By the time the report was signed off, the total effort was roughly twice the time it would have taken to write the narrative from scratch. The paradox is that the faster the model, the more work it spawns, because users feel compelled to “human‑proof” every instant result.
The underlying mechanism is simple: rapid output reduces the perceived cost of the initial step, inflating the perceived cost of the subsequent step. When the marginal effort of polishing feels higher than the marginal benefit of the raw draft, teams waste time that could have been allocated to higher‑impact analysis. The paradox intensifies as models improve, because each speed gain amplifies the polishing impulse.
The cure is not to slow the model but to redesign the workflow so that speed is paired with explicit “accept‑as‑is” checkpoints, preventing unnecessary over‑editing.