“Does using artificial intelligence assistance accelerate skill decay and hinder skill development without performers’ awareness?”
Macnamara et al. (2024)
Consider the case of an endoscopist with a decade of training who spends months working alongside an AI tool that flags suspicious lesions before her eye reaches them. Then, for a routine competency check, the tool leaves the room. Her unassisted detection rate falls (Budzyń et al., 2025).
That is no hypothetical but a finding from a multicenter observational study published in The Lancet Gastroenterology & Hepatology, serious enough that the International AI Safety Report (Bengio et al., 2026) cites it as emerging evidence of AI-driven skill erosion. The result matters beyond medicine. Any review workflow where AI increasingly does the first-pass catching, code review, underwriting, editing, diagnostics, is a candidate for the same erosion, quietly, without anyone noticing until the tool goes offline and the human is not ready.
My hypothesis: when AI systems take over detection work in review-based professions, the humans overseeing them lose unassisted detection ability over time. The fix is not watching the AI's accuracy score more closely. It is periodically testing whether the human can still do the job without it.
The Research
The endoscopy finding is not an isolated curiosity. Budzyń et al. (2025) found that adenoma detection rates dropped after clinicians had routine exposure to AI-assisted colonoscopy, a decline visible only once the AI was removed (Emerging). What makes this hard to dismiss as a one-off is a second, independent mechanism study: Macnamara et al. (2024) argue that AI assistance can accelerate skill decay and hinder skill development while masking both from the performer (Emerging). That unawareness is the whole problem. If people could feel their judgment eroding, they would compensate. They cannot, so organizations have to build the check for them.
The same pattern shows up in a different diagnostic domain. Dratsch et al. (2023) found that AI-generated BI-RADS suggestions in mammography shaped reader performance, evidence that automation-first detection changes human judgment across specialties, not just in one procedure (Emerging). This is where a Trust Architecture becomes the operative framework rather than a nice phrase. Trust Architecture is the deliberate design of transparency, feedback loops, and calibration mechanisms that keep human reliance on AI appropriate rather than automatic. The mechanism behind the erosion is well established: Parasuraman and Manzey (2010) laid out how automation breeds complacency and bias through reduced attentional engagement, a foundational account the International AI Safety Report (Bengio et al., 2026) still cites as the base explanation for why oversight degrades (Established). Skitka, Mosier, and Burdick (2000) showed that accountability structures change how much people defer to automated output (Established), which means the organizational lever exists. Nobody has to accept deskilling as inevitable. They have to design against it.
What This Means in Practice
In a code review pipeline where an AI linter or static-analysis tool catches most bugs before a human ever looks, the relevant question is when senior engineers last reviewed a diff with the tool switched off. In underwriting where a model flags risk before the underwriter sees the file, the same question applies: when did someone last underwrite blind and compare notes against the model's flags? The audit most organizations run, checking whether the AI is accurate, answers a different question than the one that matters for workforce capability. It confirms that the tool works. It reveals nothing about whether the workforce still can, without it.
This is not an argument for switching AI off. Ray (2025), responding directly to the endoscopy finding in Nature Reviews Gastroenterology & Hepatology, argues against simply withdrawing AI assistance and instead calls for structured practice that preserves unassisted skill alongside the tool (Emerging). The design question is how often, and under what conditions, humans verify independently rather than defer. Lyell and Coiera's (2017) systematic review ties automation bias to verification complexity and cognitive load, meaning the fix is procedural, not motivational (Established). Exhortation does not solve this. Scheduling the moments when independent verification is the only option available does.
Three Things to Take Away
Schedule unassisted-performance tests, separate from AI audits
Run periodic checks where reviewers work without the AI tool and measure their detection or judgment against a baseline, independent of any audit of the tool's own accuracy. Macnamara et al. (2024) argue that performers cannot be expected to detect their own skill decay, which is exactly why this cannot be left to self-report (Emerging). The test has to be scheduled, not volunteered.
Treat every AI-first-pass review workflow as deskilling risk, not just clinical ones
The mechanism generalizes. Macnamara et al. (2024) frame AI-assisted skill decay as a general performance phenomenon (Emerging), and Dratsch et al. (2023) show the same dynamic in a second diagnostic field, radiology, distinct from the original endoscopy finding (Emerging). Code review, underwriting, and editorial oversight sit on the same curve.
Build rotation into workflows so staff periodically work without AI assistance
Ray (2025) argues for structured unassisted practice rather than withdrawing AI tools altogether (Emerging), and Lyell and Coiera (2017) tie automation bias to the complexity of verifying the machine's output (Established). Rotation is the practical version of both findings: keep the muscle in use on a fixed schedule, not an emergency one.
My Two Cents
What unsettles me about this evidence is that the erosion is invisible to the person experiencing it, which means the standard organizational response, asking reviewers whether they still feel sharp, is worthless. Kosmyna et al. (2025) found measurably reduced neural engagement during AI-assisted writing tasks (Emerging), which tells me the offloading effect is not a medicine-specific artifact of one procedure but what happens whenever a tool absorbs the cognitive work a human used to do to stay good at doing it. Acemoglu, Kong, and Restrepo (2025) model how automating tasks reshapes the comparative advantage between humans and machines over time (Established), and deskilling is the mechanism by which that reshaping happens quietly, one unaudited quarter at a time. Organizations that maintain an AI accuracy dashboard but no equivalent dashboard for human capability are measuring the wrong half of the system.
Read to Learn More
Gerlich (2025), published in Societies, examines AI tool use and its relationship to cognitive offloading and critical thinking capacity, offering a broader empirical base for the mechanism discussed here.
Dunham (2026), writing for Human Resources Director, translates the emerging skill-decay evidence into workforce and training-design questions for HR and people leaders managing AI adoption.
References
Acemoglu, D., Kong, F., & Restrepo, P. (2025). Tasks at work: Comparative advantage, technology and labor demand. In Handbook of Labor Economics (Vol. 6, pp. 1-114). Elsevier. https://doi.org/10.1016/bs.heslab.2025.08.003
Bengio, Y. (Chair). (2026). International AI safety report 2026 (2nd ed.). International AI Safety Report Secretariat. https://www.internationalaisafetyreport.org
Budzyń, K., Romańczyk, M., Kitala, D., Kołodziej, P., Bugajski, M., Adami, H. O., Blom, J., Buszkiewicz, M., Halvorsen, N., Hassan, C., Romańczyk, T., Holme, Ø., Jarus, K., Fielding, S., Kunar, M., Pellise, M., Pilonis, N., & Mori, Y. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: A multicentre, observational study. Lancet Gastroenterology & Hepatology, 10(10), 896-903. https://www.thelancet.com/journals/langas/article/PIIS2468-1253(25)00288-2/abstract
Dratsch, T., Chen, X., Rezazade Mehrizi, M., Kloeckner, R., Mähringer-Kunz, A., Püsken, M., Baeßler, B., Sauer, S., Maintz, D., & Pinto Dos Santos, D. (2023). Automation bias in mammography: The impact of artificial intelligence BI-RADS suggestions on reader performance. Radiology, 307(4), e222176. https://doi.org/10.1148/radiol.222176
Dunham, J. (2026). Skill decay: Is AI eroding your workforce's ability to think? Human Resources Director. https://www.hcamag.com/ca/news/general/skill-decay-is-ai-eroding-your-workforces-ability-to-think/580678
Gerlich, M. (2025). AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies, 15(1), 6. https://doi.org/10.3390/soc15010006
Kosmyna, N., Hauptmann, E., Yuan, Y. T., Situ, J., Liao, X.-H., Beresnitzky, A. V., Braunstein, I., & Maes, P. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task (arXiv:2506.08872). arXiv. https://arxiv.org/abs/2506.08872
Lyell, D., & Coiera, E. (2017). Automation bias and verification complexity: A systematic review. Journal of the American Medical Informatics Association, 24(2), 423-431. https://doi.org/10.1093/jamia/ocw105
Macnamara, B. N., Berber, I., Çavuşoğlu, M. C., Krupinski, E. A., Nallapareddy, N., Nelson, N. E., Smith, P. J., Wilson-Delfosse, A. L., & Ray, S. (2024). Does using artificial intelligence assistance accelerate skill decay and hinder skill development without performers' awareness? Cognitive Research: Principles and Implications, 9, 46. https://doi.org/10.1186/s41235-024-00572-8
Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381-410. https://doi.org/10.1177/0018720810376055
Ray, K. (2025). AI-assisted colonoscopy and risk of endoscopist deskilling. Nature Reviews Gastroenterology & Hepatology, 22, 672. https://www.nature.com/articles/s41575-025-01122-3
Skitka, L. J., Mosier, K., & Burdick, M. D. (2000). Accountability and automation bias. International Journal of Human-Computer Studies, 52(4), 701-717. https://doi.org/10.1006/ijhc.1999.0349