In Goethe’s poem The Sorcerer’s Apprentice, the young apprentice sets the broom to carry water while the master is away. It works. It works splendidly. The water pours in, the work runs by itself, and the apprentice stands in the middle of his own success and realises too late that he can neither stop the broom nor grasp what he has set in motion. The story is not about a faulty broom. It is about a sense of magical control that does not match the actual control.
The same shift happens every time a language model delivers a draft in thirty seconds. The task leaves behind a clear experience: that went quickly. The experience is no doubt real. But it does not always match reality.
The research does not find the time employees feel they save
Two studies underline the gap. In a randomised trial run by METR, 16 experienced developers completed 246 real tasks in their own codebases, with and without AI tools. They were 19 % slower when allowed to use AI. Afterwards they estimated that they had been roughly 20 % faster. The distance between experience and measurement was almost 40 percentage points – in the same direction, among the same people, on the same task. METR has since run a follow-up trial pointing the other way, but with confidence intervals so wide and selection problems so clear that the researchers themselves call it weak evidence. The gap between experience and measurement itself stands unchallenged. [1]
The Danish picture points the same way, just at a larger scale. Anders Humlum and Emilie Vestergaard have linked survey data from 25,000 employees at 7,000 Danish workplaces with register data. Users save an average of 2.8 % of their working hours. That is real, but far from the 15 to 50 % found by controlled experiments on single tasks [2]. And the effect on pay and working hours is a precise zero. [3]
I do not read those numbers as an argument against AI. I read them as an argument against using the feeling as a metric – and against being blinded by the promises of efficiency gains. We measure something that feels like progress and call it productivity.
The feeling of flow is a poor indicator of value
Why do we misjudge our own work? One way of looking at it is that we confuse friction with effort. What a language model removes first and foremost is the blank page, the waiting and the uncomfortable starting point. It feels as if the task got smaller. But what got smaller was the resistance at the beginning – not necessarily the work as a whole. And since the friction that drains us and the friction we grow from are not the same, this shift is not innocent either.
Three things move at once, and they pull in different directions:
- The work shifts from producing to judging. A draft that arrives in thirty seconds still has to be read, checked and cut to size. That time sits elsewhere in the day, often with another person, and it is rarely recorded as part of the task.
- Quality becomes harder to assess. A well-phrased answer signals competence, even when the substance is thin. The researchers behind the BCG experiment with 758 consultants call it a jagged frontier: inside AI’s capability, participants completed 12.2 % more tasks, 25.1 % faster and at more than 40 % higher quality – but on a task just outside the frontier, they were 19 percentage points less likely to answer correctly. [4] The help and the error feel the same along the way.
- And the work can move to a colleague. In 2025 Harvard Business Review gave the phenomenon a name with the term workslop: AI-generated material that looks finished but is empty. 41 % of respondents had received it within a month, and each instance cost around two hours of clean-up. [5] The sender saved time. The organisation did not.
Clarity is amplified by AI. But so, it turns out, is a lack of clarity.
The gap is not a fault in employees, but a condition of new technology
I want to hold on to the fact that nobody is lying in this calculation. The developers in the METR trial were experienced professionals answering honestly about their own day. It is not their attitude that fails, but the instrument: people are poor at measuring their own use of time, and best of all at remembering the moments when something came easily.
The economists Erik Brynjolfsson, Daniel Rock and Chad Syverson have described the pattern as a J-curve. General purpose technologies demand large complementary investments – new processes, new skills, new ways of organising work – and those investments are invisible in the accounts while they are being made. Productivity therefore first appears to fall, and later to rise more than it really does. The gap between perceived and measured effect is, in other words, not a sign that something has gone wrong. It is the shape of the transition itself. [6]
But – and this is where the responsibility sits with us – a condition is not the same as an excuse. The J-curve explains why the gain is delayed. It does not promise that it will arrive on its own.
The leadership task is to make the gain a choice rather than a feeling
If the feeling cannot carry the decision, something else has to. Three moves go a long way, and none of them requires new systems or a measurement apparatus nobody can be bothered to use.
- The first is a baseline on one concrete task. Not an organisation-wide measurement of AI maturity, but a single, well-defined piece of work: how long does it take today, how many steps are there, how often does it need correcting. Without a starting point, any later claim about effect becomes a matter of temperament.
- The second is to move verification back to the sender. Whoever passes on an AI-assisted draft is responsible for having read it, checked the sources and being willing to stand behind it. It sounds banal, but it is precisely the rule that decides whether the time saving stays in the organisation or simply moves to the next person in the chain.
- The third is to decide what the freed-up time should be used for. Time saved that nobody has taken a position on disappears into the everyday. In the Danish data, 80 % of users move the saving to other work tasks, while fewer than 10 % spend it on breaks [3] – which is entirely reasonable, but it also means that the gain never becomes visible as anything other than a feeling of getting more done. Whether the time is taken as a saving or moved to tasks the organisation otherwise did not get to is a leadership choice. It has to be made, said out loud and followed up.
None of the three moves is about the technology. They are about judgement, about what we want to use it for, and about daring to check whether it holds true. The same applies to motivation and AI competencies: it is people, not tools, who determine the return.
What we can feel, and what we can answer for
The sorcerer’s apprentice ends up calling for the master. The point is not that he should have left the broom alone. The point is that he set it going without being able to judge what was happening while it happened.
That is the phase we are in. We have brought a tool that feels like speed into organisations that rarely measure where they started.
Three questions are worth taking into the leadership team:
- Which concrete task can we measure before and after – and who owns that measurement?
- Where in our workflows is time saved in one place and spent in another, without anyone seeing it?
- And what would we use the time for if we actually got it – have we decided, or are we assuming it will show itself?
Organisations that ask those questions get further than those that simply buy more licences. Not because measuring creates value in itself – but because it forces us to say out loud what we wanted to achieve. That is also what separates leading AI in the organisation from implementing a tool.
Notes and sources
[1] METR: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, July 2025. Randomised trial with 16 experienced contributors to large open source projects, 246 real tasks. With AI tools the tasks took 19 % longer. Participants expected a 24 % speed-up beforehand and believed afterwards that they had been 20 % faster. METR followed up with a new trial starting in August 2025 with 57 developers and more than 800 tasks across 143 codebases. It shows a speed-up: 18 % less time among returning participants (confidence interval from 38 % faster to 9 % slower) and 4 % among new ones (from 15 % faster to 9 % slower). METR itself stresses that the result is weak evidence: participants dropped out rather than work without AI, and they mainly submitted tasks where AI could help. The figure for speed itself is therefore open in both directions, while the difference between perceived and measured time from the first trial has not been withdrawn. The first trial · the follow-up, February 2026
[2] Noy and Zhang: Experimental evidence on the productivity effects of generative artificial intelligence, Science, 2023. In a writing experiment, time spent fell by 40 % and the average grade rose by 18 %. Brynjolfsson, Li and Raymond find 14 % more cases resolved per hour in customer service. The 15 to 50 % range is the comparison Humlum and Vestergaard themselves use for controlled experiments. CEPR’s review · Brynjolfsson, Li and Raymond
[3] Humlum and Vestergaard: Large Language Models, Small Labor Market Effects, NBER working paper 33777. Around 25,000 employees at 7,000 Danish workplaces across 11 occupations, linked to register data, surveys from 2023 and 2024. Users report an average saving of 2.8 % of working hours, from 0.6 % among teachers without encouragement from their employer to 6.8 % among marketing professionals with it. Effects on pay and working hours larger than 1 % can be ruled out. 80 % move the saved time to other work tasks, fewer than 10 % spend it on breaks, and 25 % spend more time on the very tasks they saved time on. NBER
[4] Dell’Acqua et al.: Navigating the Jagged Technological Frontier, Harvard Business School and BCG, 2023. 758 consultants, about 7 % of BCG’s individual contributor-level consultants. Inside AI’s capability: 12.2 % more tasks, 25.1 % faster and more than 40 % higher quality. Outside it: 19 percentage points lower probability of a correct answer. The weakest participants improved by 43 %, the strongest by 17 %. SSRN
[5] Niederhoffer et al.: AI-Generated “Workslop” Is Destroying Productivity, Harvard Business Review, September 2025, based on a survey by BetterUp Labs and Stanford Social Media Lab among 1,150 US desk workers. 41 % had received workslop within a month, and each instance cost around two hours, equal to 186 dollars per employee per month. HBR
[6] Brynjolfsson, Rock and Syverson: The Productivity J-Curve: How Intangibles Complement General Purpose Technologies, American Economic Journal: Macroeconomics, 2021. General purpose technologies require complementary, intangible investments that are not measured while they are being made. Productivity growth is therefore first understated and later overstated. AEA · free NBER version
