Is AI making us dumber?
Nobody knows yet: we could not find a single study that tested whether AI lowers intelligence. What experiments have tested is narrower and harder to shrug off: when AI does the thinking for us, we often learn less, remember less, and follow it even when it is wrong.
Short answer
- Smarter or dumber? Untested. We found no study that gave people a standard IQ test before and after they started using AI, so the claim is unproven, not disproven.
- Learning is where the warning lights are brightest. In a randomized trial with nearly 1,000 high-school math students in Türkiye, unrestricted ChatGPT raised practice scores, then left those students 17% lower on a closed-book exam than students who had no AI.1
- We often follow it, even when it is wrong. In preprint experiments, when people chose to consult an AI on trick puzzles, they took its answer on about four in five of those trials even when it had been secretly set to be wrong.2
- Authorship blurs. One week later, people who had mixed their own ideas with a chatbot’s often could not tell which ideas were theirs.3
- How you use it changes the outcome. The same ChatGPT model, set up to give hints instead of answers, largely avoided the exam loss.1
- Not shown: brain damage, or decline over years. Most experiments here test people minutes to weeks after the AI help.4,5,6,7
The question sounds like clickbait. The research behind it is not. Economists, psychologists, doctors and the AI companies themselves have now studied what happens when people let ChatGPT, Claude and other tools do the work.1,2,8,9,10 Their findings don’t add up to a verdict on intelligence. They add up to a map of where AI help turns into AI dependence, and what seems to stop it. Here it is, study by study, weak spots marked.
Is AI lowering our IQ? No study we found has measured it
Start with the headline question. We found no study in which people took a standard IQ test before and after they started using AI. So the claim that AI is lowering intelligence is untested, not disproven. Anyone who says the science proves AI makes us dumber, or proves it doesn’t, is getting ahead of the evidence.
What researchers have measured is narrower: exam scores after AI-assisted practice, problems solved once help is removed, recall of one’s own work, skill quizzes and self-reported thinking.1,5,11,12,13 Those are skills and habits, not general intelligence. They are what the rest of this article is about.
A huge international dataset shows a link, but only a correlation. In PISA 2025, which tested 15-year-olds in 91 countries and economies, students who said they never used AI chatbots for specific schoolwork tasks (summarising readings, preliminary research, drafting writing) scored higher in science, on average, than students who did.14 Students who used AI “to help me learn” about weekly scored about the same as non-users once socio-economic status was taken into account.14 The OECD says these associations do not show that AI caused the lower scores.14
A widely shared survey by Microsoft Research and Carnegie Mellon found that among 319 knowledge workers, higher confidence in AI was associated with less self-reported critical thinking.13 It measured what workers reported, not their ability, and cannot show cause.13
Public worry is already widespread. In a June 2025 Pew survey of 5,023 U.S. adults, 53% said the increased use of AI will make people worse at thinking creatively, and 40% said worse at making difficult decisions.15 Those are expectations about people in general, not measurements of anyone.15
About half of Americans now use AI chatbots. How much of our thinking are we handing over?
Use is up, and fast. In Gallup’s surveys of U.S. employees, the share using AI at work at least a few times a year rose from 21% in May 2023 to 52% in May 2026, and daily use rose from 4% to 15%.16
Among U.S. adults, 49% said in February 2026 that they use AI chatbots, up from 33% in 2024 (when the question was asked only of adults aware of chatbots), and 24% use them daily.17 In the UK, the share of full-time undergraduates using generative AI to help with assessed work went from 53% in 2024 to 94% in the survey fielded in December 2025.18
AI use has become ordinary
* The 2024 item was asked only of adults aware of chatbots.
† The 2025 figure is restated as 89% in the 2026 report; the pollster changed from UCAS to Savanta after 2024.
Using a chatbot is not the same as handing it your thinking. What matters is what we ask for. Economists at OpenAI, Duke and Harvard estimated that, averaged over the period they studied, 49% of consumer ChatGPT messages were “Asking” (for information or advice), 40% were “Doing” (asking ChatGPT to perform a task, such as drafting text or code) and 11% were “Expressing”.8 Among work-related messages, Doing made up about 56%.8 That study is a working paper, not peer reviewed, and several of its authors work at OpenAI.8
Anthropic, the company behind Claude, tracks what it calls “directive” use: users hand over a whole task with minimal back-and-forth.10 In its company reports (not peer reviewed), the directive share of sampled Claude.ai conversations rose from 27% in January 2025 to 39% in August 2025, then eased to 32% in November 2025.10
In another company report, Anthropic’s analysis of about 575,000 university-student conversations with Claude found nearly half were “direct” requests for answers or content with minimal engagement.19
A big share of AI use is “do it for me”
A large share of AI use asks the machine to do the work rather than help us do it.8,10 The question is what that costs.
Haven’t we panicked about this before?
Yes, many times. An early recorded version is in Plato’s Phaedrus, where Socrates recounts the Egyptian king Thamus warning that writing “will create forgetfulness in the learners’ souls, because they will not use their memories”.20 It is often credited to Plato himself, but the words belong to Thamus, a king in a story Socrates tells.20
In 2008, Nicholas Carr asked in The Atlantic, “Is Google Making Us Stupid?”21 He wrote that the Net seemed to be “chipping away my capacity for concentration and contemplation”, while admitting that such fears return with new tools.21
Calculators got the same treatment, and here we have data. A 1986 meta-analysis of 79 studies found that, except in grade four, calculators used alongside traditional teaching improved average students’ paper-and-pencil basic skills, while sustained calculator use by average fourth-graders appeared counterproductive for basic skills.22 The tool did not decide the result. How and when it was used did.22
Automation research named the deeper problem early. Lisanne Bainbridge’s 1983 paper “Ironies of Automation” argued that automation can erode the skills operators need when they must take over, because “physical skills deteriorate when they are not used” and process knowledge “develops only through use and feedback about its effectiveness”.23
So what is actually new?
A calculator does one step that you set up. A chatbot can hand back the whole thing: the essay, the code, the plan, the verdict. The “Doing” and “directive” numbers above are rough measures of how often we ask it to.8,10 That is our argument from usage data, not a finding of any single study.
Better practice, worse exam: do we learn less when AI does the work?
Often, yes, if the AI does the work while we are supposed to be learning. This is where the evidence is strongest.
Practice up, exam down
The landmark study ran in Türkiye.1 In a preregistered (planned and published in advance) field experiment with nearly 1,000 high-school math students, published in PNAS, one group practised with unrestricted ChatGPT, one with a ChatGPT “tutor” built to give hints instead of answers, and one with no AI at all.1
Compared with the no-AI group, practice scores rose 48% with unrestricted ChatGPT.1 Then the AI was taken away for a closed-book exam. The unrestricted ChatGPT group scored 17% lower than the no-AI group, while the tutor group’s score was not significantly different from it.1
Read that twice. The students with unrestricted ChatGPT beat the no-AI group in practice and lost to it once the AI was gone.1 And what avoided the loss was not banning AI. It was changing what the AI was allowed to do.1
With AI: better practice. Without it: a worse exam, unless the AI only gave hints.
Take the AI away, and people give up more often
A 2026 study, peer reviewed and accepted at the Conference on Language Modeling (COLM), ran three online randomized experiments with U.S.-based adults; 1,060 were analysed.5 Participants in the AI group got about 10 to 15 minutes of on-demand help from a ChatGPT-based assistant on fraction or SAT-style reading problems, and then the assistant was removed without warning.5
In the first experiment, people who had had help solved 57% of the final problems, against 73% for people who never had it. In the largest experiment, which also fixed an uneven dropout between groups in the first, the gap was smaller but still there (71% vs 77%). In two of the three experiments, people who had had help also skipped more problems.5
Remembering less, hours to weeks later
In a randomized trial with 120 undergraduates at a Brazilian university, students who prepared with ChatGPT scored 57.5% on a surprise test 45 days later, against 68.5% for students who studied the traditional way; 85 of the 120 took the test.7
In a lab experiment at an Australian university, posted as a preprint, first-year computer science students using ChatGPT scored higher on C programming tasks than students using web search (89% vs 69%).11 But they recalled less about those tasks right afterwards (41% vs 53%) and 48 hours later (39% vs 52%).11
Once the AI was gone
Three more studies point the same way:
- Among 405 secondary-school students in England, taking notes, alone or alongside an AI chatbot, beat using the chatbot alone on comprehension and retention three days later, even though more students preferred the chatbot.6
- Across seven experiments with more than 10,000 people, those who learned from LLM summaries rather than web links reported shallower knowledge and wrote sparser, less original advice.24
- In a randomized trial at a large U.S. public university (a working paper), offering students a course-integrated AI tutor lowered final grades by 0.37 standard deviations within the same courses, though the full-sample drop was smaller and only marginally significant.25
The OECD’s 2026 review of the research put it plainly: “When students depend too heavily on GenAI, metacognitive engagement... drops. This results in a misalignment between task performance and genuine learning.”26
Do we follow AI when it’s wrong? Often, and experts too
Cognitive surrender
In three experiments posted as a preprint in 2026, Wharton researchers Steven Shaw and Gideon Nave gave people trick reasoning puzzles and an AI assistant whose accuracy they secretly controlled.2 When people chose to consult it, they accepted its answer on about 93% of those trials when it was right, and on about four in five when it had been set to give a wrong answer.2
Access to the AI also raised their confidence (by 11.7 percentage points in the first experiment), even though about half its answers were wrong.2 The authors’ name for the pattern is “cognitive surrender”.2
Radiologists, clinicians, consultants
In a controlled reading experiment with a purported AI system, inexperienced radiologists rated about 80% of mammograms correctly when the AI suggested the right category, but only about 20% when it suggested a wrong one; for very experienced radiologists the figures were 82% and 46%.9 Experience helped. It did not make them immune.9
When the AI’s suggestion was wrong, most ratings were wrong too
Human-factors researchers call this automation bias. A widely cited 2010 review concluded that it occurs in novices and experts alike and that, in the studies it reviewed, training or instructions alone did not prevent it.27
In a randomized vignette study of 457 U.S. hospital clinicians, published in JAMA, systematically biased AI predictions lowered diagnostic accuracy by 11.3 percentage points from a 73.0% baseline, and image-based explanations of the AI did not significantly reduce that drop.28 When the AI was not biased, it improved their accuracy.28
In a small randomized trial posted as a preprint, 44 physicians in Pakistan, all with 20 hours of AI-literacy training, were randomly given ChatGPT recommendations that were either error-free or contained deliberate errors in 3 of 6 cases. Those given the flawed recommendations scored 14 percentage points lower on diagnostic reasoning (after adjustment).29
In a randomized field experiment with 758 Boston Consulting Group consultants, ChatGPT made them faster and better on 18 tasks inside AI’s capability “frontier”: 12.2% more tasks completed, 25.1% faster.30 But on a task chosen to lie outside that frontier, consultants using AI were 19% less likely to reach the correct solution than those working without it.30
When it reaches a courtroom, or the office
In Mata v. Avianca (2023), a U.S. federal judge imposed a $5,000 penalty on two lawyers and their firm for submitting non-existent judicial opinions generated by ChatGPT and then standing by them.31 The judge noted that existing rules give attorneys “a gatekeeping role” over the accuracy of their filings.31
At work in general, checking lags behind use. In a 47-country survey of 48,340 people, 42% of employees who use AI at work said they sometimes to very often relied on AI output without evaluating it, and only 56% said they verify its accuracy most of the time or always.32 These are self-reports.32
A week later, whose idea was it?
Here is a stranger cost. In a preregistered experiment published at CHI 2026, a peer-reviewed computing conference, 184 adults came up with ideas and wrote them out, sometimes alone and sometimes with an AI chatbot built on a ChatGPT model.3 One week later, many could not say which ideas were their own and which came from the AI.3 The researchers call it “the AI Memory Gap”.3
Confusion was worst when they had mixed their own work with the AI’s. People who did both steps alone named an idea’s source correctly about 92% of the time (a model estimate). When the AI supplied the idea and they wrote it out, that fell to about 38%; when they supplied the idea and the AI wrote it out, about 64%.3
The programming students in the Australian preprint point in a related direction: those using ChatGPT credited themselves with only 45% of their code, against 81% for students using web search.11 That measures how much of the work they felt was theirs, not whether they later misremembered who wrote what.
A week later, whose idea was it?
“Your Brain on ChatGPT”: what the MIT study did, and did not, find
This is the study behind many alarming ChatGPT headlines, so precision matters. In a small MIT Media Lab experiment, still an unreviewed preprint, 54 people wrote essays with ChatGPT, with a search engine, or with no tools, while EEG measured their brain connectivity.4
People writing with ChatGPT showed the weakest connectivity of the three groups, a measure of how strongly signals from different brain regions were linked during the task, not of brain health or intelligence: up to 55% lower total connectivity than people writing without tools, in low-frequency networks.4 In the first session, 15 of the 18 people in the ChatGPT group could not correctly quote from the essay they had written minutes earlier, against 2 of 18 in each of the other groups.4
Now the caveats, which are large. The study is still a preprint, not peer reviewed, and it has drawn methodological criticism.4 It measured connectivity during one writing task, which is not brain damage and not a measure of intelligence.4 Its famous phrase, “cognitive debt”, comes from the paper’s title; it is the authors’ metaphor, not something the study measured.4
Is AI deskilling the experts? The doctors’ data is split
Maybe. Here the evidence is genuinely mixed.
A much-discussed case comes from medicine. At four Polish endoscopy centres, experienced doctors’ adenoma detection rate in colonoscopies done without AI fell from 28.4% to 22.4% in the three months after AI assistance was introduced.33 It was an observational before-and-after comparison, so it is a signal of possible deskilling, not proof that AI caused the drop.33
A 2026 prospective trial did not find the same thing. Across 13 endoscopists and 5,013 colonoscopies, it found no deskilling or upskilling in non-AI colonoscopies after a period of AI-assisted detection, and AI raised detection for inexperienced endoscopists only while it was switched on.34 The two studies used different detection measures, so they are not a direct head-to-head.34
Aviation offers an older parallel, from cockpit automation rather than AI. In a simulator study of 16 airline pilots, hand-flying skills were largely retained, but cognitive skills such as tracking position and recognizing instrument failures showed frequent problems, and pilots spent more time on unrelated thoughts while automation flew (20.0% under autoflight vs 6.9% when hand-flying from raw data).35
The pattern is consistent with Bainbridge’s 1983 warning: the skills at risk are the ones you stop practising while the machine does them.23,35 This section describes research on professional skill. It is not guidance for clinical practice or for patients.
Why AI can feel faster than it is
Partly because in the moment it often does help, and partly because our sense of how much it helps can be unreliable.
In a study of 1,237 U.S. adults doing short everyday tasks, posted as a preprint, people expected AI help to cut completion time by about 68 seconds.36 AI-assisted completion was not faster overall: it was faster only on the harder tasks, by about 26 seconds, and significantly so on just 3 of 24 tasks.36 People still reported lower subjective effort when they had AI.36
In a 2025 randomized trial, also a preprint, 16 experienced open-source developers took 19% longer on tasks where early-2025 AI tools were allowed, yet afterwards they believed AI had cut their time by about 20%.37 One small study of early-2025 tools does not show that AI slows programmers down.37 The striking part is the gap between feeling and measurement.
Feeling faster is not being faster
A real assessment showed a related pattern: faster, not better. In a German study of first-year economics students on a timed reasoning assessment, the 38 who chose to use AI chatbots finished significantly faster than matched non-users but did not produce significantly better answers.38
Ask people about their own lives and far more say chatbots help than hurt. In February 2026, more U.S. adults said AI chatbots help rather than hurt their productivity (30% vs 5%), how informed they are (28% vs 5%) and their creativity (21% vs 11%); these are shares of all adults, including non-users.17
That feeling is real. But feeling helped is not the same as keeping a skill: in the Türkiye trial, practice scores rose while closed-book exam scores ended up 17% below the no-AI group’s.1
In a three-wave survey of 589 Chinese students and early-career workers, handing core thinking to AI and using it as a scaffold came with comparable immediate benefits, but the people who handed over their core thinking reported less deep processing and independent judgment at the final wave.39 The authors suggest this kind of reliance may be hard to notice from the inside, though they did not test that.39
What the makers of ChatGPT and Claude found in their own research
The companies that sell these tools have published some of the most useful evidence about their risks. Treat research on your own product with extra caution; their findings are mixed, like everyone else’s.
Anthropic’s coding experiment. In a randomized experiment by Anthropic researchers, posted as a preprint, 52 developers learned an unfamiliar Python library, half of them with an AI assistant.12 The AI group scored 4.15 points lower on a 27-point skills quiz and was not significantly faster.12
But how they used it mattered. Developers who used the AI to understand the code, rather than to hand the work over, kept most of their learning: their three usage patterns averaged 65% to 86% on the quiz, against 24% to 39% for the three delegating patterns.12 Each pattern covered only 2 to 7 people, and the patterns were described after the fact, not randomly assigned.12
What Claude users do. Anthropic’s analysis of 9,830 Claude.ai conversations from one week in January 2026, a company report that is not peer reviewed, found that when Claude produced an artifact such as code or a document, users were less likely to identify missing context (−5.2 percentage points), check facts (−3.7) or question the model’s reasoning (−3.1).40
In a June 2026 company report based on a survey linked to usage, people who delegated more of their work reported learning more with AI at about the same rate as everyone else.41 Anthropic added its own warning: “these are self-assessments, and skills can erode even as they become more valuable and as someone reports learning more, so the data do not rule out skill erosion.”41
OpenAI’s research. The ChatGPT usage study above was co-written by OpenAI economists.8 A separate OpenAI and Bocconi University working paper reports that in a randomized trial with 1,053 first-year business students, a short causal-reasoning training made students’ solutions more mechanism-based, more falsifiable and more diverse, but did not raise expert rubric scores, while ChatGPT access did.42 The AI improved the graded output; the training changed how students reasoned.42
What the research does not say
A page like this can tip into doom. Here is the counterweight, at full strength.
- It does not show brain damage. The EEG study measured connectivity during a task, which is not brain damage or intelligence.4
- Offloading is not bad in itself. A meta-analysis in Memory & Cognition found that offloading (reminders, note-taking) improves memory-task performance while the aid is available; and, going by the authors’ coding data, none of its studies involved AI.43 The documented cost shows up in what people keep once the aid is gone.5,44
- Direct answers don’t always hurt. In online experiments posted as a preprint, people who practised rewriting a cover letter with an AI writing tool later wrote better cover letters without AI than people who practised alone or did not practise, and the advantage held a day later.45 Simply viewing one AI-revised example helped as much (the writing was scored mainly by AI).45
- AI can help people learn. In a randomized experiment with 211 undergraduates (a working paper), AI access raised knowledge-test scores by 0.27 standard deviations, and about three-quarters of that gain persisted a week later.46 In a randomized crossover study of 194 Harvard physics students, a carefully designed AI tutor that followed teaching best practices produced more than double the median learning gains of an in-class active-learning lesson.47 At one firm that rolled out an AI assistant to customer-support agents in stages (not a randomized trial), issues resolved per hour rose 15% on average, with the largest gains for less experienced workers and evidence that it helped them learn.48
- But big average benefits shrink under scrutiny. A 2026 meta-analysis of 49 controlled STEM studies found that the apparent overall learning benefit of generative AI largely disappears once publication bias is corrected for.49
- Most evidence is short-term. Most experiments here test people minutes, days or weeks after the AI is taken away.5,6,7,46 None of the AI studies in this article followed people for years.
- None of these studies tested Nodalist, which publishes this article, or any canvas or node-based thinking tool like it.
The pattern underneath: it’s how you use it
Line the studies up and one thread runs through them. When AI does the thinking in place of us, learning and judgment often suffer. When we keep doing the thinking and use AI to check, explain or extend it, the costs shrink, and sometimes disappear.
- Same model, different job. ChatGPT as an answer machine left students 17% lower on a closed-book exam than students without AI. The same ChatGPT model as a hint-giving tutor largely avoided that loss.1
- Judge the AI’s work; being judged by it didn’t help. In a randomized experiment with 744 Chinese undergraduates, students assigned to check and revise AI-written drafts later wrote better reports without AI than students who never used AI, while students whose own drafts were critiqued by AI did no better than the no-AI group.50 The benefit faded among students with a strong habit of handing judgment to AI.50
- Answers vs analysis. In a randomized experiment with 130 people making legal-judgment predictions, an AI that gave direct recommendations produced the highest decision accuracy (68.8%) but less skill improvement than having no AI, while AIs that offered analytical support or evaluative feedback instead were reported to foster skill gains.51
- Augment vs substitute. In the STEM meta-analysis, whether AI augmented students’ own thinking or substituted for it was one of the factors that shaped results.49 The OECD’s review reached a similar view: tools built on explicit teaching models show more promise than general-purpose chatbots.26
None of this proves that “how” beats “whether” in every setting. It does mean the tool alone doesn’t decide the outcome.
What seems to help
None of these habits depends on a particular AI tool. None is a guarantee, and the first is openly contested.
- Try it yourself first, but don’t treat that as a law. In a 10-week randomized classroom experiment at one Turkish university, students who did each reading task on their own before they could consult ChatGPT outscored both students who used ChatGPT while reading and students who never used it.52 In a peer-reviewed logic-puzzle study, people who spent more time reasoning alone before asking a simulated AI for help gained more skill (an association within the study, not a controlled comparison).53 But in a Berlin lab experiment (a working paper), forcing students to read alone for 10 minutes before a ChatGPT-based tutor unlocked did not significantly beat studying the textbook alone, and students with continuous access did marginally better than the forced-delay group.54 Treat “think first” as a good default, not a proven rule.
- Commit to your answer before you see the AI’s. In a 199-person experiment, making people commit to their own answer first, ask for the AI’s suggestion only on demand, or wait before seeing it reduced how often they followed a wrong AI suggestion.55 It used a simulated AI, before chatbots, and it did not improve overall accuracy.55 People liked these designs least, which is worth knowing before you try it.55
- Take your own notes. Notes beat chatbot-only study three days later.6
- Notice how you use it. In a preprint experiment with 704 UK adults, feedback showing people how they had been using an AI assistant roughly halved the odds they asked it for complete answers and modestly raised their odds of solving later problems alone; both effects were significant only under one-sided tests.56
- Pick the harder path when you are learning. Sixth-graders who found and fixed errors in worked examples gained more on a test a week later than peers who solved and explained problems themselves, even though they liked the lesson less (the delayed benefit was clear at only one of the two schools); the tutor was a 2012 web tool, not generative AI.57 What feels pleasant is not always what sticks.57 The report-writing study above fits the same logic: judging the AI’s draft helped, letting the AI judge yours did not.50
The key studies and their limits
This table covers the 28 studies that carry this article’s main findings and its main counterweights. Most surveys and reviews, and all usage reports and historical sources, are in the References only. “Status” was last checked on October 2–3, 2026.
| Study | Design | Sample | Status | What it found | What it does not show |
|---|---|---|---|---|---|
| Bastani et al. 20251 | Preregistered randomized field experiment | Nearly 1,000 high-school math students, Türkiye | Peer-reviewed (PNAS) | Unrestricted ChatGPT: practice +48%, closed-book exam −17% vs no AI; hint-only tutor: no significant exam difference | Effects beyond one school, one subject and the unrestricted arm |
| Liu, Christian et al. 20265 | Three online randomized experiments | 1,060 U.S.-based adults analysed | Peer-reviewed (COLM 2026) | After 10–15 min of ChatGPT help, fewer problems solved once it was removed (57% vs 73%, Exp. 1; 71% vs 77%, Exp. 2, the largest); more skipping in 2 of 3 experiments | How long the effect lasts |
| Barcaui 20257 | Randomized trial | 120 undergraduates, Brazil (85 tested) | Peer-reviewed | Surprise test 45 days later: 57.5% (ChatGPT) vs 68.5% (traditional study) | How much each student relied on ChatGPT |
| Bergh et al. 202611 | Lab experiment, ChatGPT vs web search | 55 first-year CS students, Australia | Preprint | Higher task scores (89% vs 69%) but lower recall (41% vs 53%; 39% vs 52% at 48 h); less code credited to self (45% vs 81%) | A comparison with no help at all |
| Kreijkes et al. 20266 | Preregistered randomized experiment | 405 secondary students, England | Peer-reviewed | Notes, alone or with an LLM, beat LLM-only on comprehension and retention 3 days later; more students preferred the LLM | That LLMs damage memory |
| Melumad & Yun 202524 | Seven experiments | More than 10,000 participants | Peer-reviewed (PNAS Nexus) | LLM summaries led to shallower self-reported knowledge and sparser, less original advice than web links | Results on a standard knowledge test |
| Liu, Sweet et al. 202625 | Cluster-randomized trial, 34 course sections | 2,379 students, U.S. university | Working paper (EdWorkingPaper 26-1598) | Offering a course AI tutor lowered final grades 0.37 SD within the same courses; full-sample effect marginal | The effect of actually using the tutor |
| Shen & Tamkin 202612 | Preregistered randomized experiment | 52 developers | Preprint (Anthropic) | 4.15 of 27 points lower on a skills quiz (d = 0.74), not significantly faster; understanding-oriented users kept most of their learning (tiny subgroups, not randomized) | Longer-term skill |
| Shaw & Nave 20262 | Three preregistered experiments, trick reasoning puzzles | Adult participants | Preprint | When consulted, AI answers accepted on ~93% of trials (AI right) and ~4 in 5 (AI secretly wrong); confidence up | Real work decisions or lasting skill loss |
| Dratsch et al. 20239 | Controlled reading experiment | Radiologists reading mammograms | Peer-reviewed (Radiology) | ~80% correct with a right AI suggestion vs ~20% with a wrong one (inexperienced); 82% vs 46% (very experienced) | A drop from reading without AI |
| Qazi et al. 202529 | Small randomized trial | 44 physicians, Pakistan | Preprint | Randomly given flawed vs error-free ChatGPT advice: the flawed-advice group scored 14 points lower on diagnostic reasoning (adjusted) | A comparison of AI vs no AI |
| Dell’Acqua et al. 202630 | Preregistered randomized field experiment | 758 BCG consultants | Peer-reviewed (Organization Science) | ChatGPT: 12.2% more tasks, 25.1% faster inside the frontier; 19% less likely to be correct outside it | That today’s models behave the same way |
| Zindulka et al. 20263 | Preregistered online experiment | 184 adults, US and UK | Peer-reviewed (CHI ’26) | A week later, ideas’ source named correctly ~92% after doing both steps alone vs ~38% (AI supplied the idea) and ~64% (AI wrote it out); model estimates | Learning or content memory |
| Kosmyna et al. 20254 | Lab EEG experiment | 54 participants | Preprint | Weakest connectivity in the ChatGPT group (up to 55% lower than the no-tools group in low-frequency networks); in session 1, 15 of 18 in the ChatGPT group could not correctly quote their essay | Brain damage or intelligence |
| Budzyń et al. 202533 | Observational before/after | 4 endoscopy centres, Poland | Peer-reviewed (Lancet Gastro Hep) | Non-AI detection rate 28.4% → 22.4% after AI arrived | That AI caused the drop |
| Pedersen et al. 202634 | Prospective multicentre trial | 13 endoscopists, 5,013 colonoscopies | Peer-reviewed (Endoscopy) | No deskilling after AI; gains for novices only while AI was on | Proof that deskilling never happens |
| Casner et al. 201435 | Simulator experiment | 16 airline pilots | Peer-reviewed (Human Factors) | Hand-flying largely retained; frequent problems with cognitive skills; mind-wandering 20.0% vs 6.9% | A measured decline in the same pilots over time |
| Yu et al. 202636 | Preregistered online experiment | 1,237 U.S. adults | Preprint | Expected ~68 s saved; not faster overall; lower felt effort | Effects on thinking quality |
| Becker et al. (METR) 202537 | Randomized trial at task level | 16 experienced developers | Preprint | 19% longer with early-2025 AI tools; developers believed AI had cut their time by ~20% | That AI slows programmers in general |
| OECD PISA 202514 | Cross-sectional international assessment | 15-year-olds, 91 countries and economies | OECD report | Non-users of chatbots for some schoolwork tasks scored higher in science on average | That AI caused lower scores |
| Lee et al. 202513 | Cross-sectional survey | 319 knowledge workers | Peer-reviewed (CHI ’25) | More confidence in AI associated with less self-reported critical thinking | Measured ability, or cause |
| Fan & Chang 202650 | Randomized field experiment | 744 Chinese undergraduates | Peer-reviewed (Behavioral Sciences) | Judging AI drafts beat no AI (d = 0.35); AI critiquing your draft did not | Long-term retention |
| Lira et al. 202645 | Preregistered online experiments | Online participants | Preprint | Practising with an AI writing tool improved later unaided cover letters | Skills beyond cover-letter writing |
| Al-khresheh et al. 202652 | Randomized classroom experiment, 10 weeks | 90 EFL students, one Turkish university | Peer-reviewed (Journal of Intelligence) | Own work first, then ChatGPT, beat ChatGPT-while-reading and no AI | Retention after the course |
| Fischer, Rau & Rilke 202554 | Preregistered lab experiment | 334 students, TU Berlin | Working paper | A forced 10-minute read-first delay did not significantly beat textbook-only; continuous tutor access scored highest (only marginally above the delay group) | Long-term learning |
| Kestin et al. 202547 | Randomized crossover study | 194 Harvard physics students | Peer-reviewed (Scientific Reports) | Designed AI tutor: more than double the median learning gains of an active-learning class | That general-purpose chatbots help like this |
| Contractor & Reyes 202646 | Randomized experiment | 211 undergraduates | Working paper | AI access raised test scores 0.27 SD; about three-quarters persisted a week later | Effects beyond one college and one week |
| Boolzen et al. 202649 | Meta-analysis, 49 controlled studies | STEM education | Peer-reviewed (Artificial Intelligence Review) | Overall benefit largely disappears after publication-bias correction; augment vs substitute mattered | That AI is harmful on average |
Nodalist publishes this article, and none of the studies cited here tested Nodalist or any product of ours. The working method in the Our view box is the idea behind the product; we explain how we built Nodalist around this idea.
Frequently asked questions
Is ChatGPT making us stupid?
No study we found has tested whether it lowers intelligence. But in randomized studies, students given ChatGPT, either unrestricted for practice or to prepare with, did worse later than comparison groups: 17% lower on a closed-book exam in one trial, and lower on a surprise test 45 days later in another (57.5% vs 68.5%).1,7 Set up to give hints instead of answers, the same ChatGPT model largely avoided the exam loss.1
Is AI making us smarter?
In some settings it helped people learn. In one randomized Harvard physics study, a carefully designed AI tutor produced more than double the median learning gains of an in-class active-learning lesson, and in a working paper, AI access raised test scores by 0.27 standard deviations.46,47 But a 2026 meta-analysis of STEM studies found that the apparent overall benefit largely disappears after correcting for publication bias.49
Does AI make you lazy?
It makes work feel easier. In a preprint study of 1,237 U.S. adults, people reported lower effort with AI even though they were not faster overall.36 In a peer-reviewed set of three experiments, people who had had AI help skipped more problems once it was removed, in two of the three.5 “Lazy” is a judgment; what the studies show is less felt effort and, sometimes, less persistence.
Does using AI reduce critical thinking?
A survey found an association: workers more confident in AI reported less critical thinking.13 Experiments show people often go along with wrong AI answers, from radiologists reading mammograms to adults who chose to consult an AI on trick puzzles in a preprint study.2,9 But students assigned to judge AI-written drafts later wrote better without AI than students who never used it.50 In that study, what mattered was who did the judging.
What is automation bias?
The tendency to follow an automated suggestion even when it is wrong. Inexperienced radiologists rated about 80% of mammograms correctly with a right AI suggestion and about 20% with a wrong one; a widely cited 2010 review found the bias in novices and experts alike.9,27
Does AI make you a worse programmer?
The evidence is thin and specific. In Anthropic’s preprint experiment, developers learning an unfamiliar library with an AI assistant scored 4.15 of 27 points lower on an immediate skills quiz, but those who used it to understand the code, rather than to hand it over, kept most of their learning.12 That measures learning something new, not losing skills you already have.
What is cognitive offloading?
It is “the use of physical action to alter the information processing requirements of a task so as to reduce cognitive demand”, such as setting a smartphone reminder.58 It improves performance on memory tasks while the aid is available.43 The cost, in lab experiments, is that people remembered less of what they offloaded.44
What is cognitive debt?
It is a phrase from the title of the 2025 MIT preprint “Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task”.4 That study measured brain connectivity during essay writing and whether people could quote their essays afterwards; “cognitive debt” is the authors’ metaphor, not something it measured, and the study is still a preprint.4
What is cognitive surrender?
It is the name Wharton researchers Steven Shaw and Gideon Nave gave, in a 2026 preprint, to the pattern they found: when people chose to consult an AI on trick puzzles, they took its answer on about four in five of those trials even when it had been secretly set to be wrong.2
What did the MIT “Your Brain on ChatGPT” study find?
It measured how strongly signals from different brain regions were linked while 54 people wrote essays. The ChatGPT group showed the weakest connectivity, and in the first session most of them could not correctly quote their own essay minutes later. It is a preprint, it measured one task, and it is not evidence of brain damage or lower intelligence. This article covers skills, learning and judgment, not health.4
Does AI lower IQ?
No study we found has measured it: none tested people’s IQ before and after they started using AI. Untested is not the same as disproven.
Methodology and corrections
Scope. This article reviews research on skills, learning and judgment when people work with AI. It is not medical or psychological advice and does not assess anyone’s health.
What we included. Peer-reviewed studies come first. Preprints, working papers and company reports appear where they are the best evidence available, and each one is labelled as such in the text and in the table.
How we checked. Every figure and quote was taken from the original paper or report (for a few paywalled papers, from the published abstract), not from press coverage, and checked against that source in two separate passes. We used AI research agents for the checking and reviewed the results ourselves. Verbs follow design: “caused” or “lowered” only for randomized experiments, “associated with” for surveys and observational data. Where a study used a particular ChatGPT model, we simply say ChatGPT; the exact model is named in each source. This article cites 58 sources, listed below. Study status (peer review, preprint versions, corrections, retractions) was last checked on October 2–3, 2026, and is re-checked on publication day and at every update.
Conflict of interest. The author founded Nodalist, which publishes this article. No study cited here tested Nodalist.
Charts. Every chart was redrawn from reported values; no figure was copied from a paper. Charts are free to reuse under CC BY 4.0. Please credit: “Source: Nodalist, <chart title>, https://nodalist.ai/is-ai-making-us-dumber/”.
Corrections. If you find an error, tell us through our contact page. Every correction is logged below with its date.
Update log. October 3, 2026: first version.
References
- 1. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. 2025. “Generative AI without guardrails can harm learning: Evidence from high school mathematics.” PNAS 122(26): e2422633122. Peer-reviewed (correction 10.1073/pnas.2518204122 fixes an affiliation only). https://doi.org/10.1073/pnas.2422633122
- 2. Shaw, S. D., & Nave, G. 2026. “Thinking—Fast, Slow, and Artificial: How AI is Reshaping Human Reasoning and the Rise of Cognitive Surrender.” PsyArXiv yk25n_v1 (also SSRN 6097646). Preprint. https://doi.org/10.31234/osf.io/yk25n_v1
- 3. Zindulka, T., Goller, S., Fernandes, D., Welsch, R., & Buschek, D. 2026. “The AI Memory Gap: Users Misremember What They Created With AI or Without.” Proceedings of CHI ’26, pp. 1–22. Peer-reviewed conference paper. https://doi.org/10.1145/3772318.3791494
- 4. Kosmyna, N., Hauptmann, E., Yuan, Y. T., Situ, J., Liao, X.-H., Beresnitzky, A. V., Braunstein, I., & Maes, P. 2025. “Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task.” arXiv:2506.08872 (v2, 31 Dec 2025). Preprint. https://doi.org/10.48550/arXiv.2506.08872
- 5. Liu, G., Christian, B., Dumbalska, T., Bakker, M. A., & Dubey, R. 2026. “AI Assistance Reduces Persistence and Hurts Independent Performance.” COLM 2026 (Third Conference on Language Modeling). Peer-reviewed conference paper. https://openreview.net/forum?id=u9VvZd6CiY · arXiv: https://doi.org/10.48550/arXiv.2604.04721
- 6. Kreijkes, P., Kewenig, V., Kuvalja, M., Lee, M., Hofman, J. M., Vitello, S., Sellen, A., Rintel, S., Goldstein, D. G., Rothschild, D., Tankelevitch, L., & Oates, T. 2026. “Effects of LLM use and note-taking on reading comprehension and memory: A randomised experiment in secondary schools.” Computers & Education 243: 105514. Peer-reviewed. https://doi.org/10.1016/j.compedu.2025.105514
- 7. Barcaui, A. 2025. “ChatGPT as a cognitive crutch: Evidence from a randomized controlled trial on knowledge retention.” Social Sciences & Humanities Open 12: 102287. Peer-reviewed. https://doi.org/10.1016/j.ssaho.2025.102287
- 8. Chatterji, A., Cunningham, T., Deming, D. J., Hitzig, Z., Ong, C., Shan, C. Y., & Wadman, K. 2025. “How People Use ChatGPT.” NBER Working Paper No. 34255. Working paper (several authors are OpenAI employees). https://doi.org/10.3386/w34255
- 9. Dratsch, T., Chen, X., Rezazade Mehrizi, M., Kloeckner, R., Mähringer-Kunz, A., Püsken, M., Baeßler, B., Sauer, S., Maintz, D., & Pinto dos Santos, D. 2023. “Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance.” Radiology 307(4): e222176. Peer-reviewed. https://doi.org/10.1148/radiol.222176
- 10. Anthropic Economic Research. 2025–2026. “Anthropic Economic Index report: Economic primitives” (January 2026) and “Uneven geographic and enterprise AI adoption” (September 2025). Anthropic. Company reports, not peer reviewed. https://www.anthropic.com/research/anthropic-economic-index-january-2026-report; and Handa, K., et al. 2025, Anthropic Economic Index report (February 2025), arXiv:2503.04761.
- 11. Bergh, C., Tag, B., Vassar, A., & Renzella, J. 2026. “Your Programming Students’ Cognition with ChatGPT: Higher Performance, Lower Retention, and Reduced Ownership.” arXiv:2609.21194. Preprint. https://doi.org/10.48550/arXiv.2609.21194
- 12. Shen, J. H., & Tamkin, A. 2026. “How AI Impacts Skill Formation” (published by Anthropic as “How AI assistance impacts the formation of coding skills”). arXiv:2601.20245. Preprint. https://doi.org/10.48550/arXiv.2601.20245
- 13. Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. 2025. “The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers.” Proceedings of CHI ’25, pp. 1–22. Peer-reviewed conference paper. https://doi.org/10.1145/3706598.3713778
- 14. OECD. 2026. PISA 2025 Results (Volume I): Future-Ready Students, with the companion PISA 2025: Insights and Interpretations. OECD Publishing, Paris. Report (official publication, not journal peer review). https://doi.org/10.1787/73451bc5-en
- 15. Kennedy, B., Yam, E., Kikuchi, E., Pula, I., & Fuentes, J. 2025. “How Americans View AI and Its Impact on People and Society.” Pew Research Center. Survey report. https://www.pewresearch.org/science/2025/09/17/views-of-ais-impact-on-society-and-human-abilities/
- 16. Gallup (Kemp, A.). 2026. “Global Indicator: Artificial Intelligence” (workplace AI adoption, U.S. employees). Gallup. Survey report. https://www.gallup.com/699797/indicator-artificial-intelligence.aspx
- 17. Gottfried, J., Bishop, W., Anderson, M., Faverio, M., Park, E., & McClain, C. 2026. “Americans and AI 2026: Chatbots, Smart Devices and Views on Impact.” Pew Research Center. Survey report. https://www.pewresearch.org/internet/2026/06/17/americans-and-ai-2026-chatbots-smart-devices-and-views-on-impact/
- 18. Stephenson, R., & Armstrong, C. 2026. Student Generative AI Survey 2026 (HEPI Report 199). Higher Education Policy Institute, with Kortext. Survey report; with Freeman, J. 2025, Student Generative AI Survey 2025 (HEPI Policy Note 61). https://www.hepi.ac.uk/reports/student-generative-ai-survey-2026/
- 19. Handa, K., Bent, D., Tamkin, A., McCain, M., Durmus, E., Stern, M., et al. 2025. “Anthropic Education Report: How University Students Use Claude.” Anthropic. Company report. https://www.anthropic.com/news/anthropic-education-report-how-university-students-use-claude
- 20. Plato (trans. B. Jowett). c. 370 BCE. Phaedrus. Project Gutenberg eBook #1636. Book. https://www.gutenberg.org/cache/epub/1636/pg1636.txt
- 21. Carr, N. 2008. “Is Google Making Us Stupid? What the Internet is doing to our brains.” The Atlantic, July/August 2008. Magazine essay. https://www.theatlantic.com/magazine/archive/2008/07/is-google-making-us-stupid/306868/
- 22. Hembree, R., & Dessart, D. J. 1986. “Effects of Hand-Held Calculators in Precollege Mathematics Education: A Meta-Analysis.” Journal for Research in Mathematics Education 17(2): 83–99. Peer-reviewed. https://doi.org/10.2307/749255
- 23. Bainbridge, L. 1983. “Ironies of Automation.” Automatica 19(6): 775–779. Peer-reviewed. https://doi.org/10.1016/0005-1098(83)90046-8
- 24. Melumad, S., & Yun, J. H. 2025. “Experimental evidence of the effects of large language models versus web search on depth of learning.” PNAS Nexus 4(10): pgaf316. Peer-reviewed. https://doi.org/10.1093/pnasnexus/pgaf316
- 25. Liu, J., Sweet, T., Chen, M. H., Engelberg, J., Masters, M. C., Clark, M., Persaud, A., Lancaster, A., Hollingsworth, J. K., & Rice, J. K. 2026. “The Effects of Course-Integrated AI Tutoring on Student Performance and Engagement: A Randomized University Trial.” EdWorkingPapers No. 26-1598, Annenberg Institute, Brown University. Working paper. https://doi.org/10.26300/y3f8-vh05
- 26. OECD. 2026. OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education. OECD Publishing, Paris. Report (intergovernmental evidence review). https://doi.org/10.1787/062a7394-en
- 27. Parasuraman, R., & Manzey, D. H. 2010. “Complacency and Bias in Human Use of Automation: An Attentional Integration.” Human Factors 52(3): 381–410. Peer-reviewed (review). https://doi.org/10.1177/0018720810376055
- 28. Jabbour, S., Fouhey, D., Shepard, S., Valley, T. S., Kazerooni, E. A., Banovic, N., Wiens, J., & Sjoding, M. W. 2023. “Measuring the Impact of AI in the Diagnosis of Hospitalized Patients: A Randomized Clinical Vignette Survey Study.” JAMA 330(23): 2275–2284. Peer-reviewed. https://doi.org/10.1001/jama.2023.22295
- 29. Qazi, I. A., Ali, A., Khawaja, A. U., Akhtar, M. J., Sheikh, A. Z., & Alizai, M. H. 2025. “Automation Bias in Large Language Model Assisted Diagnostic Reasoning Among AI-Trained Physicians.” medRxiv. Preprint. https://doi.org/10.1101/2025.08.23.25334280
- 30. Dell’Acqua, F., McFowland, E., Mollick, E., Lifshitz, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. 2026. “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality.” Organization Science 37(2): 403–423. Peer-reviewed. https://doi.org/10.1287/orsc.2025.21838
- 31. Castel, P. K. (U.S. District Judge). 2023. Mata v. Avianca, Inc., Opinion and Order on Sanctions, Case 1:22-cv-01461 (PKC), S.D.N.Y., 22 June 2023. Legal record. https://storage.courtlistener.com/recap/gov.uscourts.nysd.575368/gov.uscourts.nysd.575368.54.0_3.pdf
- 32. Gillespie, N., Lockey, S., Ward, T., Macdade, A., & Hassed, G. 2025. Trust, attitudes and use of artificial intelligence: A global study 2025. University of Melbourne and KPMG International. Survey report. https://doi.org/10.26188/28822919
- 33. Budzyń, K., Romańczyk, M., Kitala, D., Kołodziej, P., Bugajski, M., Adami, H. O., Blom, J., … Mori, Y. 2025. “Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study.” The Lancet Gastroenterology & Hepatology 10(10): 896–903. Peer-reviewed (erratum: doi 10.1016/S2468-1253(25)00294-8). https://doi.org/10.1016/S2468-1253(25)00133-5
- 34. Pedersen, T. A., Mori, Y., Botteri, E., Engjom, T., Seip, B., Dimcevski, G. G., & Havre, R. F. 2026. “Learning and deskilling effects of artificial intelligence in colonoscopy among endoscopists with different levels of experience: a pragmatic, prospective trial.” Endoscopy 58(9): 1003–1014. Peer-reviewed. https://doi.org/10.1055/a-2858-7084
- 35. Casner, S. M., Geven, R. W., Recker, M. P., & Schooler, J. W. 2014. “The retention of manual flying skills in the automated cockpit.” Human Factors 56(8): 1506–1516. Peer-reviewed. https://doi.org/10.1177/0018720814535628
- 36. Yu, S., Cheng, M., Jabbar, A., Sucholutsky, I., Collins, K. M., Jurafsky, D., & Hawkins, R. D. 2026. “Cognitive offloading and the speedup illusion in human-AI interaction.” arXiv:2605.23177. Preprint. https://arxiv.org/abs/2605.23177
- 37. Becker, J., Rush, N., Barnes, E., & Rein, D. 2025. “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity.” METR, arXiv:2507.09089. Preprint. https://doi.org/10.48550/arXiv.2507.09089
- 38. Molerov, D., Federiakin, D. A., Zlatkin-Troitschanskaia, O., Shenavai, K., Trierweiler, L., & Nagel, M.-T. 2026. “The relationship between AI-chatbots use, student assessment performance and learning outcomes in higher education.” Unterrichtswissenschaft 54(3): 325–356. Peer-reviewed. https://doi.org/10.1007/s42010-026-00242-2
- 39. Zhu, Q., Li, X., Dong, Y., Chang, P., & Fan, M. 2026. “Not all cognitive offloading is equal: distinguishing dependent and autonomous offloading to generative AI.” Frontiers in Psychology 17. Peer-reviewed. https://doi.org/10.3389/fpsyg.2026.1878629
- 40. Anthropic (Education research team). 2026. “Anthropic Education Report: The AI Fluency Index.” Anthropic. Company report, not peer reviewed. https://www.anthropic.com/research/AI-fluency-index
- 41. Massenkoff, M., Lyubich, E., Sacher, S., Hitzig, Z., Zhang, S., Heller, R., & McCrory, P. 2026. “Anthropic Economic Index report: Cadences.” Anthropic. Company survey report. https://www.anthropic.com/research/economic-index-june-2026-report
- 42. Asirvatham, H., Betti, C., Brown, R., Camuffo, A., Chatterji, A., Fumagalli, C., Gambardella, A., Mariani, M., Pandey, A., Ramos, S., Salvucci, G., & Simic, V. 2026. “Training novices to think, or giving them LLMs? Evidence from an RCT.” OpenAI and Bocconi University (also CEPR Discussion Paper 21882). Working paper (OpenAI co-authors). https://cdn.openai.com/pdf/novices-and-llm-august-2026.pdf
- 43. Burnett, L. K., & Richmond, L. L. 2026. “Meta-analytic investigations of the effect of cognitive offloading on memory-based task performance and interindividual variability.” Memory & Cognition 54(1): 144–168. Peer-reviewed (meta-analysis). https://doi.org/10.3758/s13421-025-01743-8
- 44. Grinschgl, S., Papenmeier, F., & Meyerhoff, H. S. 2021. “Consequences of cognitive offloading: Boosting performance but diminishing memory.” Quarterly Journal of Experimental Psychology 74(9): 1477–1496. Peer-reviewed. https://doi.org/10.1177/17470218211008060
- 45. Lira, B., Rogers, T., Goldstein, D. G., Ungar, L., & Duckworth, A. L. 2026. “Coach not crutch: Evidence that AI can improve writing skill despite reducing effort.” arXiv:2502.02880 (v4). Preprint. https://doi.org/10.48550/arXiv.2502.02880
- 46. Contractor, Z., & Reyes, G. 2026. “Experimental Evidence on the Learning Impact of Generative AI.” arXiv:2607.08849. Working paper. https://doi.org/10.48550/arXiv.2607.08849
- 47. Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. 2025. “AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting.” Scientific Reports 15: 17458. Peer-reviewed. https://doi.org/10.1038/s41598-025-97652-6
- 48. Brynjolfsson, E., Li, D., & Raymond, L. 2025. “Generative AI at Work.” The Quarterly Journal of Economics 140(2): 889–942. Peer-reviewed. https://doi.org/10.1093/qje/qjae044
- 49. Boolzen, C., Kuhn, J., Flegr, S., Rott, E.-M., Stausberg, N., & Küchemann, S. 2026. “Evidence of impact and interpretational limits of generative AI in STEM education: a systematic review and meta-analysis on cognitive learning outcomes.” Artificial Intelligence Review 59(10): 223. Peer-reviewed (meta-analysis). https://doi.org/10.1007/s10462-026-11665-9
- 50. Fan, M., & Chang, P. 2026. “Telling Students to Evaluate Does Not Make It Happen: Task Stage, Offloading Tendency, and Error Detection in AI-Assisted Student Writing.” Behavioral Sciences 16(9): 1671. Peer-reviewed. https://doi.org/10.3390/bs16091671
- 51. Ding, H., Shen, Y., Chen, J., & Wang, P. 2026. “More vs. less cognitive offloading from AI assistants: impacts on novices’ collaborative performance and skill development.” Information Processing & Management 64(1): 105046. Peer-reviewed. https://doi.org/10.1016/j.ipm.2026.105046
- 52. Al-khresheh, M. H., Demirkol Orak, S., Alharbi, W., & Almayez, M. 2026. “The Effects of Immediate Versus Delayed Timing of AI Support on EFL Learners’ Reading Comprehension Performance: An Experimental Study.” Journal of Intelligence 14(9): 199. Peer-reviewed. https://doi.org/10.3390/jintelligence14090199
- 53. Wu, S., Belem, C. G., Fu, S., Steyvers, M., & Smyth, P. 2026. “How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles.” Proceedings of HCOMP 2026, pp. 88–96. Peer-reviewed conference paper. https://doi.org/10.1145/3834580.3838741
- 54. Fischer, M., Rau, H. A., & Rilke, R. M. 2025. “AI Tutoring Enhances Student Learning Without Crowding Out Reading Effort.” IZA Discussion Paper No. 18338. Working paper. https://docs.iza.org/dp18338.pdf · SSRN copy: https://doi.org/10.2139/ssrn.5992341
- 55. Buçinca, Z., Malaya, M. B., & Gajos, K. Z. 2021. “To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making.” Proceedings of the ACM on Human-Computer Interaction 5(CSCW1), Article 188. Peer-reviewed. https://doi.org/10.1145/3449287
- 56. Maier, S., Schwabe, K., Schneider, M., & Feuerriegel, S. 2026. “Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants.” arXiv:2609.20143. Preprint. https://doi.org/10.48550/arXiv.2609.20143
- 57. McLaren, B. M., Adams, D. M., & Mayer, R. E. 2015. “Delayed Learning Effects with Erroneous Examples: a Study of Learning Decimals with a Web-Based Tutor.” International Journal of Artificial Intelligence in Education 25(4): 520–542. Peer-reviewed. https://doi.org/10.1007/s40593-015-0064-x
- 58. Risko, E. F., & Gilbert, S. J. 2016. “Cognitive Offloading.” Trends in Cognitive Sciences 20(9): 676–688. Peer-reviewed (review). https://doi.org/10.1016/j.tics.2016.07.002