Listen to this article
Jump to a section
In this article
Executive Summary
In the past eighteen months, controlled studies in schools, workplaces, clinics and laboratories have measured the same thing from different angles: people who let AI do the cognitive work perform worse when the AI is taken away. Turkish high-school students who practised with an unrestricted chatbot scored 17 percent lower on the unassisted exam. Developers who learned a new library with an AI assistant scored 50 percent on a comprehension quiz against 67 percent for those who coded by hand. Experienced endoscopists who worked with AI detection for three months found fewer adenomas once it was switched off. Ten minutes of GPT-5 help on fraction problems was enough to cut adults’ unassisted solve rate from 73 to 57 percent and nearly double the rate at which they gave up. The evidence is early, the samples are small, and one of the most widely reported studies of the year was retracted in August. The direction is consistent all the same. The same studies also show what protects the skill: tools that give hints instead of answers, users who ask the model to explain instead of produce, and regular practice without the tool at all.
17%
lower unassisted exam scores for high-school students who practised with an unrestricted GPT-4 tutor, against no loss for students whose tutor gave hints only, Bastani et al., PNAS, 2025
50 vs 67
quiz scores for developers who learned a new Python library with AI assistance against those who coded by hand, largest gap on debugging, Anthropic, January 2026
28 to 22%
adenoma detection rate in unassisted colonoscopies by 19 experienced endoscopists, before and after three months of routine AI-assisted work, Budzyń et al., The Lancet Gastroenterology & Hepatology, 2025
73 to 57%
unassisted solve rate on fraction problems after about ten minutes of GPT-5 assistance, with the skip rate nearly doubling, Liu et al., CMU, Oxford, MIT and UCLA, 2026
Core conclusions
- The loss is real and it is fast. Every controlled study that removed the AI after a period of use found a drop in unassisted performance, in some cases within a single ten-minute session. The mechanism is old and well understood: a skill that is not practised is not kept, and the tool removes the practice.
- The evidence is thinner than the headlines. The largest samples are self-report surveys, the EEG study had 54 participants and is a preprint, and the April 2026 study that put “AI did most of the thinking” into news stories around the world was retracted four months later. Anyone making policy on this should read the methods sections.
- The design of the tool and the habits of the user decide the outcome. Hints instead of answers removed the exam loss entirely. Developers who asked the model to explain retained what they built. The fix is a change in how the tools are used, in schools, at work and at home, and it has to start with the people who set the rules.
The argument, in ten slides
Save it, or send it to whoever sets the AI rules at your school or your company.










What was observed in schools
The cleanest study is a field experiment in a large Turkish high school, published in PNAS in June 2025 by Hamsa Bastani and colleagues at Wharton and Penn. Nearly a thousand students in grades nine to eleven were given one of two GPT-4 tutors for their maths practice sessions, or none. The first tutor behaved like ordinary ChatGPT. The second was prompted to give teacher-designed hints and to withhold the answer. During practice, both groups improved sharply: 48 percent better than the control for the plain chatbot, 127 percent better for the hints-only tutor. Then the tools were removed for the exam. Students who had practised with the plain chatbot scored 17 percent lower than students who had never used one. Students who had practised with the hints-only tutor scored the same as the control.
That pair of results is the whole subject in miniature. The tool that did the work for the student made the practice look better and the learning worse. The tool that made the student do the work kept the learning intact.
A second study, posted in April 2026 by researchers at Carnegie Mellon, Oxford, MIT and UCLA, measured how fast the effect sets in. Across three experiments with 354, 667 and 201 adult participants, people worked through fraction arithmetic or SAT-style reading comprehension for about ten minutes with GPT-5 available, then continued without it. In the first experiment the unassisted solve rate fell from 73 to 57 percent and the share of problems participants skipped rose from 11 to 20 percent. In the third experiment the skip rate went from 1 percent to 8 percent. The effect was worse when the AI had been giving full solutions than when it gave hints. Ten minutes was enough to reduce both the skill and the willingness to try.
Cross-sectional data points the same way. Michael Gerlich’s 2025 survey of 666 people in the United Kingdom, with follow-up interviews, found a significant negative relationship between how heavily people used AI tools and how they scored on a critical-thinking assessment, with cognitive offloading as the mediator. The youngest group, aged 17 to 25, was the most dependent on the tools and scored lowest. Higher education was protective regardless of how much AI a person used. That study is self-report and cannot show cause, and I grade it accordingly below, but its shape matches the experiments.
What was observed at work
Anthropic ran a randomised trial on its own kind of user, published in January 2026. Fifty-two developers, mostly junior, with at least a year of weekly Python, were asked to learn Trio, an asynchronous programming library most of them had never used, either with an AI assistant or by hand. The AI group finished about two minutes faster, which was not statistically significant. On a quiz about the concepts they had just used, the AI group averaged 50 percent and the hand-coding group 67 percent. The widest gap was on debugging questions. The developers who scored well with AI had used it differently: they generated code and then asked the model to explain it, or asked only conceptual questions and fixed the errors themselves. The ones who delegated the whole task learned the least.
Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 real tasks they had done with generative AI, and presented the results at CHI in 2025. The finding that matters for managers is about confidence. The more a worker trusted the AI for a given task, the less critical thinking they reported doing on it. The more they trusted their own ability on that task, the more critical thinking they did. The work itself shifted from producing to checking: verifying facts, integrating the output, steering the task. Whether that checking happens depends on whether the person believes they could have done the work themselves.
A third workplace study is worth mentioning because of what happened to it. In April 2026 the American Psychological Association’s journal Technology, Mind, and Behavior published a paper reporting that 58 percent of 1,923 adults who used commercial AI on ten simulated work tasks agreed that “AI did most of the thinking”, with lower confidence in their own reasoning and less ownership of their ideas. It was reported everywhere, including in trade press for laboratory and office managers. On 31 August 2026 the journal retracted it. A reader had flagged mathematical inconsistencies in the statistics, the methodology described after publication did not match the paper, two of the four primary metrics were inconsistently defined, and the author declined to provide the data. A university listed as an affiliation stated it had never approved or supervised the research. The direction of that study agreed with everything else in this post. It was still wrong to rely on it, and the press releases that announced it in April still carry no correction.
What was observed in clinical practice
The study I would show to any board is the one from four colonoscopy centres in Poland, published in The Lancet Gastroenterology & Hepatology in August 2025. Nineteen experienced endoscopists, each with more than two thousand procedures behind them, began using an AI polyp-detection system routinely at the end of 2021. The researchers compared the 1,443 colonoscopies these physicians performed without AI, before and after the AI was introduced. The adenoma detection rate in the unassisted procedures fell from about 28 percent to about 22 percent, a six-point absolute drop, over roughly three months of AI exposure. These were experts at the top of their skill, and they lost part of it while the tool was on and did not have it back when the tool was off.
The clinical literature had a name for this before generative AI existed: deskilling. It was documented in aviation when autopilots became standard and in navigation when GPS did. The International AI Safety Report published in February 2026, written by an international panel of experts for participating governments, cites the endoscopy study as its lead example under risks to human autonomy and describes the broader research base on AI, cognitive offloading and critical thinking as “nascent”. Both things are true. The clinical result is solid and the field around it is young.
What was observed in everyday use
The MIT Media Lab study that gave the subject its vocabulary was posted in June 2025 by Nataliya Kosmyna and colleagues in Pattie Maes’s group. Fifty-four people wrote essays across three sessions using ChatGPT, a search engine, or nothing but their own heads, while wearing EEG caps. The group using ChatGPT showed the weakest and least distributed brain connectivity of the three. They struggled to quote a sentence from the essay they had just written. They reported the lowest sense of ownership of their work. In a fourth session, eighteen participants swapped conditions. Those who moved from ChatGPT to unassisted writing carried the under-engagement with them. Those who moved from unassisted writing to ChatGPT showed stronger recall and engagement than people who had used ChatGPT from the start. The authors called the pattern cognitive debt: the convenience is taken now and the cost arrives later. It is a preprint with a small sample and the authors say so. It is also the only study to date that looked at the brain rather than the score.
The same group published a follow-up in June 2026 on a different task. Sixty-seven people spent four weeks judging whether news items were real or fake, with an AI chatbot to consult. With the chatbot they were 21 percent more accurate. When it was withdrawn they were no better than before. Four weeks of practice with the tool had built no lasting skill, because the tool had been doing the judging.
How the decline happens
None of this needs a new theory. Psychologists have studied cognitive offloading, the habit of putting a mental task onto an external aid, since long before language models. A 2011 study in Science showed that when people expect a fact to remain available on a computer, they remember where to find it rather than the fact itself. Evan Risko and Sam Gilbert’s 2016 review in Trends in Cognitive Sciences set out the general pattern: offloading is efficient, people do more of it as the aid becomes more reliable, and the skill that is offloaded gets less practice.
Generative AI changes two things. It can take on the whole task rather than one step of it, so the practice removed is the thinking itself and not just the arithmetic or the recall. And it answers fluently and confidently, which makes checking feel unnecessary. Put those together and you get a loop that the studies above keep finding.

The tool produces the answer. The person accepts it, so the reasoning step is not practised. The unpractised skill fades, which the person experiences as the task feeling harder without the tool. Confidence moves from the person to the tool, which the Microsoft data shows is exactly the condition under which critical thinking stops. The next task is delegated a little more completely. In the endoscopy suite this loop took three months. In the fraction experiment it took ten minutes.
The part that should worry anyone responsible for a school or a workforce is that the loop feels like productivity from the inside. The Turkish students practising with the plain chatbot were 48 percent better than their peers, right up to the exam. The developers were two minutes faster. The endoscopists’ detection rate with the AI on was presumably fine. Every measure that an ordinary dashboard tracks improved. The loss was only visible when someone took the tool away and measured again, and almost nobody does that.
How strong the evidence is
I want to be precise about this, because the subject attracts more certainty than the data supports in both directions.
The strong results are the controlled ones where the tool was removed and unassisted performance was measured: the Turkish school experiment, the three CMU and Oxford experiments, the Anthropic trial, the Polish endoscopy study and the two MIT studies. Every one of them found a loss. Their weakness is scale and duration. The largest has about a thousand participants, the smallest 52, and the longest follow-up is four months.
The weak results are the surveys. Gerlich’s 666-person study is the most cited number in the field and it is cross-sectional self-report, which means it cannot tell whether AI use lowers critical thinking or whether people with weaker critical thinking use AI more. It also required a correction in September 2025 for a duplicated table, which did not change its conclusions. The retracted Technology, Mind, and Behavior study was also a survey of self-reported reasoning.
There is also evidence pointing the other way, and it is larger than anything above. A meta-analysis in Nature Human Behaviour in 2025 by Jared Benge and Michael Scullin pooled 57 studies and 411,430 adults with an average age of 69 and found that people who used digital technology had 58 percent lower odds of cognitive impairment and a slower rate of decline, after adjusting for education, income and health. That is about computers, phones and the internet in older adults, not about chatbots in students, and the direction of cause is not settled. But it is a reminder that “technology rots the brain” has been claimed about every tool since writing, and has usually been wrong. What the generative AI studies show is narrower and more defensible: a skill that a tool performs for you is a skill you stop practising, and it fades.
| Setting | Study | Sample | What was removed | Result without the tool |
|---|---|---|---|---|
| High-school maths | Bastani et al., PNAS, 2025 | About 1,000 students, Turkey | GPT-4 tutor after a term of practice | 17% lower exam scores with an answer-giving tutor; no loss with a hints-only tutor |
| Adult arithmetic and reading | Liu et al., CMU, Oxford, MIT, UCLA, 2026 | 354, 667 and 201 adults | GPT-5 after about ten minutes | Solve rate 73% to 57%; skip rate 11% to 20% |
| Software development | Anthropic, 2026 | 52 developers | AI assistant during learning of a new library | Quiz 50% with AI against 67% by hand; widest gap on debugging |
| Knowledge work | Lee et al., Microsoft and CMU, CHI 2025 | 319 workers, 936 tasks | Nothing removed; self-report | More trust in AI, less critical thinking; more self-confidence, more critical thinking |
| Colonoscopy | Budzyń et al., The Lancet Gastroenterology & Hepatology, 2025 | 19 expert endoscopists, 1,443 unassisted procedures | AI detection after three months of routine use | Adenoma detection about 28% to about 22% |
| Essay writing | Kosmyna et al., MIT Media Lab, 2025 preprint | 54 adults, four months | ChatGPT after three sessions | Weakest EEG connectivity; could not quote own essay; lowest ownership |
| Judging news | Rani et al., MIT Media Lab, 2026 | 67 adults, four weeks | AI chatbot after four weeks | 21% more accurate with the tool; no lasting gain without it |
| General population | Gerlich, Societies, 2025 | 666 adults, UK, survey | Nothing removed; self-report | Heavier use, lower critical-thinking scores; youngest group worst |
Only the studies that removed the tool and measured again can speak to cause. The two self-report surveys show association. The retracted April 2026 study is deliberately left out.
What protects the skill
The encouraging part is that the same studies that measured the loss also measured what prevented it, and the answers agree with each other.
Hints instead of answers. This is the strongest single finding. The Turkish students whose tutor withheld the answer improved more during practice than the students with the plain chatbot and lost nothing on the exam. The CMU and Oxford experiments found the loss was significantly smaller when the AI gave hints than when it gave solutions. A tool designed to make the person do the last step keeps the person’s skill.
Explain instead of produce. In the Anthropic trial, developers who used AI and still scored well had asked it to explain the code it generated, or had asked only conceptual questions and written the code themselves. The MIT misinformation study’s authors proposed the same thing: a Socratic mode in which the chatbot asks the user guiding questions rather than delivering the verdict. The tool is the same. The prompt is different.
Unassisted practice, on purpose. The endoscopy study is the model. The physicians’ unassisted detection rate was measurable, it was measured, and the drop was caught. The response the gastroenterology literature has settled on is ongoing monitoring of unassisted performance, feedback and structured training, so that core detection skill is preserved alongside the tool. The International AI Safety Report describes the same idea for organisations as periodic testing or reliance drills. If you cannot measure how your people perform without the tool, you will not know they have lost anything until it matters.
Self-confidence built before the tool arrives. The Microsoft study found that workers confident in their own ability on a task kept thinking critically about the AI’s output on that task. Gerlich found that education protected critical thinking regardless of AI use. The Singapore Ministry of Education’s position, stated in Parliament on 6 May 2026, follows the same logic: pupils in Primary 1 to 3 learn about AI but are not given AI tools to use, and from Primary 4 they use only supervised, education-specific tools that give feedback rather than answers, once the basics have been mastered. The Ministry has also funded a longitudinal study from 2027 to track how children’s AI use affects learning. That is the right order: skill first, tool second, measurement throughout.

Free tool
Human-AI Interaction & Decision Quality Dashboard
Benchmark your team’s automation bias and decision-acceptance rates, the two numbers that rise as the offloading loop above takes hold.
What to do about it
I do not think the answer is less AI. The productivity gains are real, the tools are not going away, and a child or an employee who is kept from them will be at a disadvantage against one who has learned to use them well. The answer is to change what “using them well” means, and that has to be set by the people who write the rules: ministries, school heads, chief executives and parents.
| Where | What to do | What it rests on |
|---|---|---|
| Schools | Teach the underlying skill before introducing the tool for that skill. Primary years learn about AI and do not use it for schoolwork. | Bastani et al.; Gerlich’s education finding; Singapore MOE’s Four Learns sequence |
| Schools | Procure or configure tools that give hints and feedback and withhold answers. Reject tools that cannot be configured that way. | Bastani et al.; Liu et al. |
| Schools | Keep regular unassisted assessment, and report the unassisted score alongside the assisted one. | Every removal study in this post |
| Employers | Measure unassisted performance for the roles where it matters, on a schedule. Treat a falling unassisted score as a finding. | Budzyń et al.; International AI Safety Report reliance drills |
| Employers | Build skill before granting the tool for that skill. Junior staff get AI for tasks they have already demonstrated they can do by hand. | Anthropic; Lee et al. confidence finding |
| Employers | Set the prompt discipline: explain, then produce. Write it into the AI use policy and the onboarding. | Anthropic usage-pattern finding |
| Employers | Keep a named person accountable for every decision the tool touched, so that checking is the job and not a courtesy. | International AI Safety Report mitigations for automation bias |
| Individuals | Think first. Write the answer, the plan or the diagnosis in a sentence before asking the tool. Then compare. | Kosmyna et al. crossover finding; Lee et al. |
| Individuals | Ask the tool why, not just what. If you cannot explain the output to someone else, you have not learned it. | Anthropic |
| Individuals | Do some of the work without the tool every week, and notice whether it has got harder. | Budzyń et al.; Liu et al. |
Each row traces to a specific study above. Where two rows rest on the same study, the study measured both the loss and the protective condition in the same sample.
For the generation now in primary school the stakes are different from ours. An adult who offloads a skill they already have can usually get it back with practice. A child who offloads a skill before acquiring it has nothing to get back. That is why the school rows in the table come first and why the sequence matters more than the tool. The Turkish students who practised with hints did better than everyone. The ones who practised with answers did worse than students with no AI at all. The difference between those two groups was a paragraph in a system prompt, decided by an adult.
Free tool
Board AI Oversight Checklist
Check whether your organisation’s AI policy names who is accountable for decisions the tool touched and whether unassisted competence is measured anywhere.
Evidence & Methodology
This post leans on controlled studies where the tool was removed and unassisted performance was measured, and I have graded those higher than the surveys. One study widely reported in April 2026 was retracted in August and is excluded. My own claims about what organisations should do are practitioner judgment built on those studies.
| Claim | Source | Grade |
|---|---|---|
| Students with an unrestricted GPT-4 tutor improved 48% in practice and scored 17% lower on the unassisted exam; hints-only tutor improved 127% with no exam loss | Bastani, Bastani, Sungu, Ge, Kabakcı and Mariman, PNAS, June 2025, field experiment, about 1,000 students, one Turkish high school; minor correction August 2025 | Measured, randomised |
| About ten minutes of GPT-5 assistance cut unassisted solve rate from 73% to 57% and raised skipping from 11% to 20%; effect worse with solutions than hints | Liu, Christian, Dumbalska, Bakker and Dubey, CMU, Oxford, MIT, UCLA, arXiv April 2026, three experiments, 354, 667 and 201 adults | Measured, preprint |
| Developers learning with AI scored 50% against 67% by hand, largest gap on debugging, two minutes faster and not significant | Anthropic, January 2026, randomised, 52 mostly junior developers, published by the vendor of the tool used | Measured, vendor-run |
| Higher confidence in AI predicts less critical thinking; higher self-confidence predicts more | Lee et al., Microsoft Research and CMU, CHI 2025, 319 knowledge workers, 936 self-reported tasks | Measured, self-report |
| Expert endoscopists’ unassisted adenoma detection fell from about 28% to about 22% after three months of routine AI use | Budzyń et al., The Lancet Gastroenterology & Hepatology, August 2025, observational, 19 endoscopists, four Polish centres, 1,443 unassisted procedures | Measured, observational |
| ChatGPT essay writers showed the weakest EEG connectivity, could not quote their own essays, reported lowest ownership; effect carried into unassisted writing | Kosmyna et al., MIT Media Lab, arXiv June 2025, 54 participants over four months, 18 in the crossover session; not peer reviewed | Measured, preprint, small sample |
| Four weeks of AI help raised fake-news accuracy 21% and built no lasting skill | Rani, Danry, Liang, Lippman and Maes, MIT Media Lab, June 2026, 67 participants | Reported, from the lab’s own summary |
| Heavier AI use associated with lower critical-thinking scores, mediated by offloading; 17 to 25 age group most dependent; education protective | Gerlich, Societies, January 2025, 666 UK participants, cross-sectional survey and interviews; table correction September 2025 | Measured, self-report, no causal claim |
| Digital technology use in older adults associated with 58% lower odds of cognitive impairment | Benge and Scullin, Nature Human Behaviour, 2025, meta-analysis of 57 studies, 411,430 adults, mean age 69; observational | Measured, different population |
| The research base on AI, offloading and critical thinking is “nascent”; reliance drills, accountability and AI literacy are the proposed mitigations | International AI Safety Report 2026, section 2.3.2 | Measured, expert synthesis |
| 58% of 1,923 workers said AI did most of the thinking | Baldeo, Technology, Mind, and Behavior, April 2026; retracted 31 August 2026 for statistical inconsistencies and unavailable data | Retracted, not used |
| Singapore pupils in Primary 1 to 3 learn about AI without using AI tools; supervised education-specific tools from Primary 4; longitudinal study from 2027 | Ministry of Education, parliamentary reply, 6 May 2026 | Measured, government text |
| Skill first, tool second, unassisted measurement throughout is the right sequence for schools and employers | My own read, from AI governance work with organisations rolling these tools out | My call |
Sources
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26).
- Liu, G., Christian, B., Dumbalska, T., Bakker, M. A., & Dubey, R. (2026). AI assistance reduces persistence and hurts independent performance. arXiv preprint.
- Anthropic. (2026). How AI assistance impacts the formation of coding skills.
- Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. (2025). The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. CHI 2025.
- Budzyń, K., Romańczyk, M., Kitala, D., et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: A multicentre, observational study. The Lancet Gastroenterology & Hepatology, 10.
- Kosmyna, N., Hauptmann, E., Yuan, Y. T., Situ, J., Liao, X.-H., Beresnitzky, A. V., Braunstein, I., & Maes, P. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task. arXiv preprint.
- MIT Media Lab. (2026). AI’s impact on cognitive ability: MIT study reveals more troubling data.
- Gerlich, M. (2025). AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies, 15(1), 6.
- Benge, J. F., & Scullin, M. K. (2025). A meta-analysis of technology use and cognitive aging. Nature Human Behaviour.
- Bengio, Y., et al. (2026). International AI Safety Report 2026.
- Technology, Mind, and Behavior. (2026). Retraction of “Generative artificial intelligence reliance and executive function attenuation: Behavioral evidence of cognitive offload in high-use adults” by Baldeo (2026).
- Ministry of Education, Singapore. (2026). AI usage in schools. Parliamentary reply, 6 May 2026.
- Sparrow, B., Liu, J., & Wegner, D. M. (2011). Google effects on memory: Cognitive consequences of having information at our fingertips. Science, 333(6043).
- Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9).
Where to start this quarter
Pick one skill that matters to your organisation and that your people now do with AI: writing a credit memo, reading a contract, triaging an incident, debugging a service. Take a sample of the people who do it and measure how well they do it without the tool, the way the Polish endoscopists were measured. Most organisations have never taken that measurement, and the first one is the baseline you will wish you had a year from now.
Then change two things. Configure the tool, or the prompt your people are told to use, so that it explains and hints before it produces, and make “explain it back” part of how the output is reviewed. Repeat the unassisted measurement in six months. If the score has held, the tool is doing what it should. If it has dropped, you have found it while it is still cheap to fix, which is more than the endoscopists’ patients could have said.
For a school, the same thing with one subject and one cohort, and the tool withheld until the basics are in. For a parent, the same thing at the kitchen table. Think first, then ask, then check.

The Levelling of Cognitive Assets: What Remains Scarce When Articulation Becomes Free
The market-side view of the same shift: when fluent output costs nothing, judgment is what is left to pay for.

AI Can Triage. It Can't Be Accountable. Most Enterprises Have the Order Backwards.
Why accountability is the workplace version of unassisted practice, and what happens when it is removed.

What CHROs Are Worried About With AI and How to Respond
The workforce questions this post raises, from the perspective of the person who owns skills and hiring.

AI Could Do the Job Tomorrow. The Math Still Caps How Much You Can Automate.
Oversight capacity depends on reviewers who could still do the work themselves; this post is about how that capacity erodes.
If you are deciding how AI tools are introduced across a workforce or a school system, the unassisted baseline above is where I would start the conversation. My consulting work covers the use policy, the measurement and the governance around it.
Was this useful?






