How AI Is Hollowing Out the Pipeline That Makes Experts
The Collapse of Skill Formation in the Agentic Era
Good evening.
Tonight’s story is not fiction either. It concerns a guest — the kind who arrives uninvited, makes himself comfortable in every organization on earth, and never, under any circumstances, leaves early. Call him the Guest From Hell, if you like. He doesn’t argue. He doesn’t rush. He simply outstays everyone, one retirement party at a time, and shows no sign of departing before the Untimely End finally arrives to show him the door.
You are about to meet a horror that doesn’t even bother breaking in. He was already invited. He arrives as a shortcut. A lesson skipped so gently no one notices the skipping — until, one day, the person who might have caught the mistake has simply retired, entirely satisfied, to Florida, and the Guest is still sitting in the good chair.
Do sit still. It won’t take long. Though I confess — at this rate, one wonders who exactly will be watching next time.
TL;DR
- The mechanism is real and now empirically visible: AI/agentic automation is absorbing exactly the “grunt work” tier through which juniors historically built pattern-recognition and judgment — and early labor data (Harvard’s “seniority-biased technological change” finding of a ~9% relative drop in junior employment at AI-adopting firms; Stanford’s 13%–19% relative employment decline for 22–25-year-olds in AI-exposed jobs) confirms the entry rung is contracting while senior demand holds. The deeper danger is not job loss but the erosion of the training ground that manufactures future experts.
- Automation-induced deskilling is a mature, well-documented science in aviation and medicine — the FAA issued formal Safety Alerts (SAFO 13002/17007) precisely to counter manual-flying skill fade, and a 2025 Lancet study documented a 6.0-percentage-point drop in colonoscopists’ unassisted cancer-detection rate after AI exposure — but its application to knowledge work (coding, law, consulting) is newer, with less longitudinal data and genuinely conflicting evidence.
- Governance is racing to catch up on two weak fronts: (1) the “vouching/attestation” problem — who certifies a workflow is safe to run unattended, and whether that evidence is itself AI-generated (a closed epistemic loop); and (2) the human-pipeline floor — analogous to aviation’s manual-flying mandates, proposals for “AI-free” practice requirements, protected training tasks, and licensing responses are emerging but largely voluntary as of September 2026.
- Coding skill formation has direct, controlled evidence, not just analogy. A randomized controlled trial (Shen & Tamkin, Anthropic, Jan 2026) found developers using AI assistance scored 17 percentage points lower on a post-task comprehension quiz than those coding by hand — with the largest gap specifically on debugging, the skill most needed to catch AI’s own errors.
- In software specifically, velocity’s reward and its harm land on different sides of the same ledger. The speed gain is captured entirely on production (writing code faster than searching and adapting it by hand); the cost is imposed entirely on verification (a review process that already caught only 55–60% of defects, now checking code its authors understand less well and reviewing more of it, faster). Output and the check on output don’t scale at the same rate — this is a case where velocity plausibly does more harm than good on its current trajectory.
Key Findings
- The core mechanism now has a name and a formal model. Economist Enrique Ide (IESE Business School) formalized it in “Automation, AI, and the Intergenerational Transmission of Knowledge” (arXiv 2507.16078, June 2026): improvements in entry-level automation “increase output upon adoption but can reduce growth and welfare, even without reducing entry-level employment,” because they “reallocate novices away from the most productive experts, slowing the diffusion of best practices.” The skill being hollowed out is not “prompt engineering” (trivial) but the domain pattern-recognition needed to catch when an AI is subtly wrong.
- The labor data is early but directionally consistent. Harvard’s Seyed Mahdi Hosseini Maasoum and Guy Lichtinger found junior employment at GenAI-adopting firms fell ~7.7–9% within six quarters while senior employment held steady (“seniority-biased technological change”). Stanford’s Brynjolfsson, Chandar and Chen found a 13% relative employment decline for ages 22–25 in AI-exposed occupations, widening to ~19% by August 2026.
- Aviation is the gold-standard analogue — a mature field with decades of “automation complacency” research and actual regulatory responses, though those responses stop short of hard mandated minimums.
- Medicine provides the strongest emerging empirical evidence of actual deskilling, led by the 2025 Lancet colonoscopy study and mammography automation-bias experiments.
- The “vouching”/attestation problem is real and under-theorized, with a genuine closed-epistemic-loop risk when AI vouches for AI.
- Named warnings are proliferating across AI labs (Amodei), academia (Beane, Ide, Brynjolfsson), consultancies (McKinsey, BCG), and multilaterals (WEF).
Details
1. The core mechanism: the training ground is being automated away
Historically, professional judgment was built by doing the slow version of the work: first-year law associates doing document review, analysts building pitch books, radiology residents reading scans, junior developers writing boilerplate and debugging. This “grunt work” was simultaneously low-value output and high-value learning. The World Economic Forum, in its June 2026 analysis “The AI-related leadership that’s only five years away,” put the loss precisely: “What’s disappearing isn’t just work. It’s practice.” It noted that “Harvard University research indicates junior employment has fallen 9%… at organizations adopting generative AI.”
Harvard Business Review (David S. Duncan, “How Do Workers Develop Good Judgment in the AI Era?”, Feb 2026) observed that generative AI “was helping me a lot more than it was helping my less-experienced colleagues” — because seniors have the judgment to steer and verify it, while juniors “often can’t tell whether AI-generated work is any good.” Microsoft engineering leaders Mark Russinovich and Scott Hanselman described an “AI boost” that multiplies senior engineers’ output while imposing an “AI drag” on junior developers who lack the judgment to steer or verify what the AI produces.
The distinction at the heart of the thesis: the skill to operate AI is trivial; the skill to recognize when it is wrong is exactly what is being hollowed out. Jossie Haines (executive coach, former Apple engineering leader) told Forbes that AI “cannot figure out why the product team keeps building features that raise copyright concerns” — the systems-level judgment that “used to develop through proximity to real decisions: catching an error before it spread.”
Ide’s model is the analytical backbone here: even if junior employment is preserved, if AI reallocates novices away from the most-skilled experts (or strips the learning value out of the tasks juniors retain), long-run growth and expertise transmission suffer. He explicitly acknowledges input from David Autor, Matthew Beane, Luis Garicano, and Chad Jones, situating the work in mainstream growth economics.
2. The labor economics: seniority-biased technological change
- Harvard (Hosseini Maasoum & Lichtinger, “Generative AI as Seniority-Biased Technological Change,” SSRN, Aug 2025; updated May 2026): tracked 62 million workers across 285,000 US firms (2015–2025). Junior employment at GenAI-adopting firms fell ~7.7–9% within six quarters; senior employment held steady; the decline was “driven primarily by slower hiring rather than increased separations,” and GenAI-exposed tasks became “increasingly less likely to appear in junior task bundles.” Adopters were only ~3.7% of firms but accounted for 17.3% of total employment.
- Stanford (Brynjolfsson, Chandar & Chen, “Canaries in the Coal Mine?”): using ADP payroll data, found employment for ages 22–25 in the most AI-exposed occupations fell 13% relative to less-exposed peers (original, Aug 2025), widening to “about 19% below where it would be if it had kept pace with… less-exposed occupations” in the Aug 2026 revision. Brynjolfsson’s interpretation: “It appears what younger workers know overlaps with what LLMs can replace.” The authors frame these as descriptive “early indicators… not causal estimates.”
- SignalFire State of Tech Talent (2025/2026): new grads are “just 7% of new hires at big tech companies… down 25% from 2023 and over 50% from pre-pandemic levels in 2019.” At startups, the new-grad share fell “from 30% in 2019 to under 6%.” SignalFire’s June 22, 2026 report found entry-level hiring at the 12 “Tech Majors” “down roughly 65% against 2019, and down about 76% at early-stage startups.”
- Corroborating scholarship: Brynjolfsson et al.’s “Six Facts” (a ~13–16% early-career decline) and the International AI Safety Report 2026 both note AI adoption is “disproportionately affecting junior workers.”
- Caveat/dissent: Critics (e.g., Jing Hu, “2nd Order Thinkers”) note the junior collapse began Q1 2023 — before most firms deployed AI in production — implicating post-pandemic rate shocks and over-hiring corrections as much as AI. The Harvard authors’ identification strategy compared adopters vs. non-adopters (via GenAI “integrator” job postings) to isolate the AI effect; parallel pre-2023 trends between the groups support causal interpretation, but confounders remain.
3. Aviation: the mature, regulated analogue (well-established)
Aviation has studied “automation complacency” (Parasuraman & Manzey) and manual-flying “skill fade” for decades:
- FAA SAFO 13002 (Jan 4, 2013) and SAFO 17007 (“Manual Flight Operations Proficiency,” May 4, 2017): issued after “an analysis of flight operations data… identified an increase in manual handling errors.” The FAA holds that “continuous use of those [autoflight] systems does not reinforce a pilot’s knowledge and skills in manual flight operations” and that “manual flight is the foundation upon which other technical flying skills are built.”
- Starting March 12, 2019, US 14 CFR Part 121 carriers were required to train additional manual maneuvers (slow flight, stalls, upsets, unreliable airspeed, bounced landings, instrument departures/arrivals). ALPA’s Air Safety Organization advocated for the SAFO.
- ICAO’s Personnel Training and Licensing Panel Automation Working Group reviewed 386 reports (77 accidents, 309 major incidents): 36% of accident cases showed automation-dependency indicators, rising to 49% for accidents in 2010–2021.
- Landmark cases: Asiana 214 (SFO, 2013 — NTSB cited a crew that “relied too heavily on an automated system it did not fully understand”); the 2009 Turkish Airlines Amsterdam crash.
- Key academic study: Casner, Geven, Recker & Schooler (2014), “The Retention of Manual Flying Skills in the Automated Cockpit,” Human Factors 56:1506–1516.
- Regulatory limit worth flagging: the FAA’s response is largely encouragement (SAFOs are advisory), and EASA has pushed evidence-based/competency-based training (AMC1 ORO.FC.115) rather than a hard mandated minimum of manual-flying hours on the line — a gap that pilots themselves have criticized. Even the gold-standard field stops short of a strict quantified floor.
4. Medicine: the strongest emerging empirical deskilling evidence
- Colonoscopy (the flagship study): Budzyń et al., “Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy,” Lancet Gastroenterology & Hepatology, published online Aug 12, 2025. Retrospective observational study at four Polish centers (ACCEPT trial). The adenoma detection rate of standard, non-AI-assisted colonoscopy fell from 28.4% (226/795) before AI to 22.4% (145/648) after AI exposure — an absolute decline of −6.0 percentage points (95% CI −10.5 to −1.6; p=0.0089; exposure-to-AI odds ratio 0.69). The authors: “To our knowledge this is the first study to suggest a negative impact of regular AI use on health care professionals.” (Note: AI assistance reliably raises detection while active; the concern is the erosion of unassisted skill.)
- Mammography (automation bias): Dratsch et al., Radiology, 2023 — 27 radiologists reading 50 mammograms; incorrect AI BI-RADS suggestions significantly degraded accuracy across inexperienced, moderately experienced, and very experienced readers, with inexperienced readers most susceptible. Lead author Thomas Dratsch (University Hospital Cologne): “it was surprising to find that even highly experienced radiologists were adversely impacted.”
- Scoping review (PubMed, “AI in medicine: a scoping review of the risk of deskilling and loss of expertise among physicians”): empirical studies “consistently demonstrate that AI can inadvertently impair physicians’ performance or reduce opportunities for skill maintenance,” and it argues “safeguarding clinical expertise should be considered a central component of AI safety and resilience in medicine.” It also documents “structural deskilling” in UK cytology (HPV primary screening cut case volumes 80–85% and consolidated labs from 45 to 8).
- New vocabulary: NEJM (Abdulnour, Gin, Boscardin, “Educational Strategies for Clinical Supervision of AI Use,” Aug 2025) and a 2026 Nature Medicine perspective distinguish deskilling (losing an existing skill), never-skilling (failing to ever develop a foundational skill because AI did it during the developmental window — producing “false proficiency” that collapses when AI is removed), and mis-skilling (adopting an AI’s errors as one’s own reasoning). A randomized trial (Qazi et al., 2025) found physicians given an LLM with deliberately seeded errors suffered significant degradations in diagnostic reasoning.
- Annals of Internal Medicine (Topaz et al., July 2026) posed the question directly: “The Deskilling Effect: Is Artificial Intelligence Eroding Clinical Competence?”
5. Other documented deskilling domains
- Surgical robotics — Matthew Beane’s “shadow learning” (Administrative Science Quarterly, 2019): a two-year ethnography plus blinded interviews at 13 top teaching hospitals (observing programs at ~18 institutions) found robotic surgery removed residents from hands-on participation — “rather than having their hands in the work, residents and assistants watched the procedure on television” — degrading on-the-job learning via “helicopter teaching.” A minority resorted to norm-violating “shadow learning”: premature specialization, abstract rehearsal (including YouTube), and “undersupervised struggle.” This is the closest documented pre-AI analogue to what AI now threatens across knowledge work, and Beane is a direct intellectual link (he advised Ide’s economic model).
- GPS / spatial memory: Dahmani & Bohbot, Scientific Reports (2020), 50 drivers — greater lifetime GPS use correlates with worse spatial memory during unaided navigation and reduced hippocampal-dependent strategy use; a three-year follow-up suggested GPS use drives the decline.
- Calculators / mental arithmetic: five decades of research show over-reliance weakens number sense and the “calibration” that supports error detection — the intuition that flags when an answer “doesn’t feel right.”
- Automated trading: the number sense that lets a trader catch a position “off by a factor of ten” or a “fat finger” order (100,000 contracts instead of 1,000) erodes with disuse — a direct parallel to AI-output error-catching.
6. The vouching / attestation problem in agentic AI governance
As organizations increasingly run agents “unattended” or “on the loop,” a governance question arises: who attests that a class of task is safe to automate, on what evidentiary basis, and is that evidence itself AI-generated? (Section 8 documents that even pre-AI human code review — the mechanism organizations implicitly lean on to vouch for software changes — already caught only 55–60% of defects on average, dropping to 28% for large changes; the problem below compounds on top of that pre-existing weakness, not a clean baseline.)
- HITL vs. HOTL: Human-in-the-loop requires human approval before execution (appropriate for irreversible/high-risk actions); human-on-the-loop allows autonomous action with monitoring and after-the-fact intervention. Both the EU AI Act and the NIST AI Risk Management Framework require oversight grounded in “context, authority, and rationale.” Practitioners warn that “most organizations confuse presence with practice” — putting someone “in the loop” without training them on what to approve or how to spot automation complacency: “that’s not oversight — it’s a liability dressed up as process.”
- Delegation-chain / institutional-attestation frameworks: emerging academic and industry work — “Governing Actions, Not Agents: Institutional Attestation as a Governance Model” (arXiv 2606.26298), the Cloud Security Alliance’s Agent Identity Governance Framework, and “Bounded Autonomy for Enterprise AI” (arXiv 2604.14723) — converges on a standard: “every agent action must be attributable to a human authorizer who defined the scope,” preserved in “a tamper-evident audit record.” The human “is accountable for the authorized scope — not for reviewing each individual action.”
- The closed-loop risk (the report’s key insight): “self-QA loops,” in which an AI critiques its own output, share the generator’s blind spots — “if the generator confidently misunderstood something, a generator-as-critic using identical framing will likely miss it too.” Self-certification frameworks for high-risk AI now exist (e.g., arXiv 2601.08295 using the Fraunhofer AI Assessment Catalogue), but they risk agents attesting to their own reliability with no independent human check. When combined with deskilling, the danger compounds: the humans nominally “vouching” for an unattended workflow may increasingly lack the independent domain expertise to evaluate what they are certifying — a genuinely closed epistemic loop. One governance design (the “AgentRunner” ToolGateway, arXiv 2605.10223) attempts to make this a “system architecture guarantee” rather than a “prompt engineering suggestion” by physically halting execution at risk thresholds until human confirmation — but this presumes a competent human on the other end.
7. Governance and policy responses on the human pipeline
We pause here, briefly, for tonight’s sponsor. He is, if anything, more patient than last time’s — patient the way the Guest From Hell is patient. Last time’s villain needed a zero-day. This one needs an infinite-day: no deadline, no disclosure window, no clock running out on the other end at all. He has never once had to hurry, because he was never going anywhere. He has always been sitting at the table, and he always will be, right up until the Untimely End finally asks him to leave. We now return to the program, such as it continues.
Distinct from AI capability guardrails, these target the human qualification/training floor:
- Aviation model (manual-mode mandates): the template for “keep practicing the skill the machine covers for you,” though advisory rather than a strict hour floor.
- Medical education/licensing: NEJM/Nature Medicine recommend requiring trainees to generate an independent differential before consulting AI, grading reasoning not just answers, and building “AI-free assessment moments.” A systematic review proposes the EU AI Act (post-2026 Digital Omnibus, which delayed medical-device requirements to Aug 2028) incorporate “mandatory skill impact assessment, periodic ‘AI-free’ practice requirements, and post-market surveillance of physician competence for high-risk diagnostic AI.”
- Professional bodies: the Federation of State Medical Boards (nonbinding 2024 guidance; Aug 3, 2026 statement by CEO Humayun Chaudhry and board chair Valentine Theard) holds AI “is not ready to be independently licensed like a physician,” grounding licensure in medicine’s “social contract.” A competing JAMA framework (Alon Bergman, Robert Wachter, Ezekiel Emanuel, Apr 29, 2026) proposes autonomous clinical AI pass USMLE-equivalent exams “at or above the median score of recent human test-takers,” then complete a supervised “residency,” under a new federal Office of Clinical AI Oversight. The Josiah Macy Jr. Foundation / AAMC / ACGME recommend AI curricula and modified accreditation.
- Corporate redesign: BCG research documents firms redesigning work to preserve thinking — at Shell, “junior employees worked through problems on their own before touching any AI tool,” with early results showing juniors “explain their reasoning more clearly.” Bank of America’s head of global talent Josh Bronstein said the bank kept intern numbers close to 4,000 in 2026 while building AI simulations to “give people the experiences in a simulated way quickly.” IBM (VP Natasha Pillay-Bemath) redesigned junior roles toward “analysis, problem-solving and responsible AI use” rather than eliminating them.
- Supporting evidence for “use it or lose it”: a 2025 MIT study found ChatGPT-assisted writers showed lower brain activity and remembered less of what they wrote; a Microsoft/Carnegie Mellon study (Lee et al., Feb 2025) found frequent AI users showed “reduced critical engagement” and “diminished independent problem-solving” on routine tasks.
8. Software engineering specifically
There is direct, controlled evidence of coding-skill erosion from AI assistance, not just analogy borrowed from aviation and medicine.
- The direct RCT (the strongest single piece of evidence in this section): Shen & Tamkin, “How AI Impacts Skill Formation” (Anthropic, published Jan 29, 2026; arXiv 2601.20245). A randomized controlled trial with 52 software developers (mostly junior, all with 1+ years of Python experience, all unfamiliar with the specific library used) learning a new async-programming library either with or without an AI coding assistant. Result: the AI-assisted group scored 50% on a post-task quiz vs. 67% for the hand-coding group — a statistically significant 17-point gap (Cohen’s d = 0.738, p = 0.01), “the equivalent of nearly two letter grades.” The largest gap was specifically on debugging questions — the skill the researchers themselves flag as “crucial for detecting when AI-generated code is incorrect.” Task completion time did not differ significantly between groups. Critically, how participants used AI mattered more than whether they used it: those who used AI to check or build understanding (asking follow-up or conceptual questions) scored as well as the no-AI group; those who delegated code-writing wholesale and used AI to debug scored worst. The authors explicitly flag their own limits: n=52 is small, the assessment measured only immediate comprehension (not longitudinal retention), and — notably — they expect agentic coding tools like Claude Code to produce larger skill-formation effects than the simpler AI-sidebar setup they tested, since this study understates the mechanism this report is most concerned with.
- Qualitative corroboration from computer-science education: Prather et al., “The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers” (ICER ’24, ACM, Aug 2024). An observational study of novice programmers using GenAI tools found that struggling students frequently developed an “illusion of competence” — expressing false confidence that GenAI had “augmented their critical thinking” while their actual problem-solving showed the opposite, and finishing tasks with cognitive dissonance about how well they’d actually understood the material. Stronger students, by contrast, used GenAI to accelerate work they already knew how to do and could catch and discard bad suggestions — the same divergence the distributional-effects argument in Section 10 predicts.
- Software-developer-specific labor data. The Stanford HAI 2026 AI Index (payroll data, millions of workers, tens of thousands of firms, 2021–2025) found employment for software developers specifically aged 22–25 declined nearly 20% since late 2022, while employment for older developers at the same firms grew 6–12% — a software-specific echo of the broader Harvard/Stanford findings in Section 2, and consistent with the “seniority-biased technological change” framing.
- The productivity side remains genuinely mixed, and vendor-funded results diverge from independent ones. GitHub’s own RCT (~200 developers) found Copilot users 53.2% more likely to pass all unit tests; a Microsoft/GitHub/MIT Sloan RCT (Peng et al., 2023) found a 55.8% speed gain on a bounded boilerplate task. Independent datasets tell a different story: GitClear’s “AI Copilot Code Quality: 2025” report (211M lines, 2020–2024) found code churn rose from ~3.1% to 5.7%, copy/pasted lines rose from 8.3% to 12.3%, and refactored (“moved”) lines fell from ~24–25% to 9.5% — the first year copy/paste exceeded moved code. Uplevel and Harness reported higher bug rates and more debugging time for AI-generated code. Conflict-of-interest flag: GitClear sells a code-review tool. METR’s RCT (Becker, Rush, Barnes and Rein, arXiv 2507.09089, July 2025) found 16 experienced developers were 19% slower with AI tools on their own mature repositories despite forecasting a 24% speedup and self-reporting a 20% speedup afterward — a large perception/reality gap, though METR stresses this is a snapshot of early-2025 tools in one setting, and a Feb 2026 follow-up gave an “unreliable signal.”
- The pipeline logic remains arithmetic with a long fuse. It takes roughly 5–9 years to grow a graduate into a reliable senior, so any reduction in junior intake or junior skill formation now surfaces as a senior shortage around the early 2030s. A smaller-scale precedent: the post-2008 hiring freeze produced a shortage of mid-career engineers by roughly 2012.
What remains genuinely open is longitudinal evidence — whether the gap persists, widens, or closes with experience — and evidence from real agentic tools rather than the simpler AI-sidebar setup Shen & Tamkin tested; the study’s own authors flag this as their most important limitation and expect agentic tools to show larger effects.
AI-generated code is a statistical intensification of an already-common, already-unvetted practice, not a new category of risk. Copying code from Stack Overflow or a GitHub repository without fully understanding it has been standard developer behavior for two decades, and it was rarely vetted carefully before use — a developer would search for a working snippet, paste it in, confirm it ran, and move on. An LLM does the same thing at the level of a statistical model trained on that same corpus: it produces the most probable continuation of code given a prompt, drawing on patterns learned from the same public repositories and Q&A sites developers already copied from directly. The shift AI introduces is one of volume and removed friction, not of kind: copy-pasting from a single Stack Overflow answer required finding a plausible-looking match and adapting it by hand, which imposed at least some reading and adaptation; an AI assistant generates fitted, ready-to-run code on demand, removing even that minimal friction. Given that human vetting of copy-pasted code was already weak (Section 8’s baseline data above), and that AI output is now produced faster and in greater volume with even less forced engagement from the person using it, the underlying reviewing/vetting gap this report documents was present well before generative AI — AI has widened it by removing the last remaining friction that occasionally forced a developer to read what they were using.
The baseline was already weaker than the “skilled practitioner” assumption implies. The deskilling risk documented above does not start from a strong human baseline and erode it — human code review and bug detection were already documented as unreliable before AI-generated code entered the picture, which means the “vouching” problem in Section 6 compounds on top of an existing weakness, not a new one. A SmartBear/Cisco study of 2,500 pull requests found code-review defect-detection effectiveness peaks around 200–400 lines and roughly 60 minutes of review time, after which reviewers start missing things — the detection rate drops from 87% for pull requests under 100 lines to just 28% for pull requests over 1,000 lines. Aggregated across studies (Capers Jones’ data, cited via Steve McConnell’s Code Complete), code review alone catches on average 55–60% of defects, and no single detection technique — design inspection, code inspection, QA, or testing — exceeds roughly 65–75% on its own; only combining all four approaches roughly reaches 99%. A direct comparative study of bug detection by novice programmers versus LLMs (arXiv 2311.16017) found student bug-detection accuracy on genuinely faulty code was 34.5%, compared to 87.3% (GPT-3) and 99.2% (GPT-4) on the same task — though the same study found LLMs were worse than students (42–79% vs. 92.8%) at correctly recognizing bug-free code as fine, meaning models over-flag as often as humans under-catch.
The conclusion: AI-generated code is increasingly reviewed by people whose own unassisted debugging ability may already be eroding, using a review process that was measurably porous even before AI accelerated the volume and size of changes moving through it. Human code review was never as reliable a backstop as the “vouching” model implicitly assumes, and AI-driven velocity is stressing exactly that weak point harder and faster.
This is a case where velocity’s reward and its harm are not competing for the same resource — they are two measurements of the same acceleration. The speed gain lands entirely on production: writing code faster than searching, reading, and adapting a Stack Overflow answer by hand is a genuine improvement. The cost lands entirely on verification: the same acceleration removes the last remaining friction (the minimal reading and adaptation copy-pasting used to require) that occasionally forced a moment of human engagement with the code, while doing nothing to strengthen a review process that was already catching only 55–60% of defects. Output and the check on that output do not scale at the same rate here — velocity is fully captured on one side of the ledger and fully imposed as cost on the other. That asymmetry, not any claim about individual competence, is why this is a domain where velocity plausibly does more harm than good on its current trajectory.
9. The generational / demographic overlay
The deskilling mechanism compounds an independent demographic threat. The “Silver Tsunami” — all US baby boomers turn 65 by 2030, with an estimated 61 million exiting the workforce — is already draining tacit knowledge: surveys find 57% of boomers have shared less than half the knowledge needed for their jobs (21% have shared none), and an APQC survey found organizations expect 51% of their workforce to retire or leave within five years. David DeLong’s Lost Knowledge framed this as a “giant sucking sound… of knowledge being drained out of organizations.” The novel and dangerous synthesis: historically, retiring experts were replaced by juniors who had climbed the same ladder. If AI has simultaneously hollowed out that ladder’s bottom rungs, the two curves intersect — senior tacit expertise exits at exactly the moment the pipeline meant to replace it has thinned. Amodei (Anthropic CEO) crystallized the concern, warning AI could eliminate up to 50% of entry-level white-collar jobs within 1–5 years; critics rightly note his incentive to hype, but even skeptics concede the pipeline logic: “if you don’t have junior hires right now, you won’t have experienced people 5 or 10 years later.”
10. Distributional effects: not a shifted mean, but a widening, skewed spread
A natural first intuition is to model the effect of AI-assisted work as a Gaussian shift — most professionals clustering near an average level of AI-assisted competence, with a small number of outliers doing unusually well or unusually poorly. The evidence assembled above doesn’t support that shape. It supports something closer to a bimodal, self-reinforcing divergence — closer to the “K-shaped” pattern already used elsewhere in labor economics (e.g., post-2020 recovery literature) than to a bell curve.
The reason is that the underlying process isn’t additive random noise around a stable mean; it’s compounding in both directions:
- The upward tail compounds. Professionals who already possess enough foundational judgment before heavy AI use — Beane’s surgeons who built skill through deliberate “shadow learning,” or Shell’s juniors who work a problem by hand before invoking AI — use AI as leverage rather than a crutch. Existing skill plus AI assistance produces faster skill growth, not just faster output.
- The downward tail compounds too. The medical deskilling literature’s “never-skilling” category — failing to ever build a foundational skill because AI performed the task during the developmental window — describes a population that doesn’t regress to a mean; it falls further behind, because each subsequent AI-assisted task offers less opportunity to develop the judgment needed to catch the AI’s errors. Dratsch et al.’s automation-bias findings reinforce this: less-experienced readers were the most susceptible to being led astray by confident, incorrect AI suggestions — the downward-tail population isn’t drawn randomly from the workforce, it’s disproportionately the least-experienced.
Two implications follow that a symmetric-distribution model would miss:
- The “average” performer is the highest-risk population, not the safest one. In a Gaussian frame, the middle of the distribution is the safe, unremarkable center. Here, the middle is the specific population the “vouching” problem (Section 6) is built around: professionals competent enough to be trusted with autonomous or lightly-supervised workflows, but not skilled enough to reliably catch a subtly wrong AI output. Genuinely poor performers are more likely to get caught by review; genuinely strong performers catch their own errors. It’s the modal, “good enough to trust, not good enough to verify” group where the closed epistemic loop actually bites.
- The mean becomes a less meaningful statistic over time. If the distribution is genuinely bifurcating rather than shifting, aggregate metrics — average productivity, average code quality, average diagnostic accuracy — will increasingly describe fewer and fewer actual practitioners, masking a growing population at each tail. Organizations tracking only aggregate performance metrics are especially likely to miss this, since a widening spread can leave the mean looking flat even while the underlying population is polarizing.
This sharpens Recommendation 1 specifically: instrumenting unassisted skill matters most not for the outliers (who are somewhat self-selecting and self-correcting in either direction) but for the modal, middle-of-the-distribution professionals who are hardest to distinguish from genuinely competent peers using aggregate or AI-assisted performance data alone.
Recommendations
- Treat skill formation as a first-class governance metric, not a byproduct. Most firms track AI adoption; almost none track whether today’s productivity is building tomorrow’s judgment. Instrument it directly: measure junior staff’s unassisted performance on core tasks at intervals, exactly as the Lancet study measured unassisted ADR. Threshold that changes action: a measurable decline in unassisted performance should trigger mandatory AI-free rotations, prioritizing the modal middle-of-distribution group identified in Section 10, not just visible outliers.
- Adopt the aviation “manual-mode” floor now, voluntarily, before it’s mandated. Require periodic “AI-free” practice on foundational tasks — the single most transferable lesson from a mature regulated field. Note that even aviation’s floor is only advisory; organizations that want resilience should go further than the FAA did and set an actual quantified minimum.
- Protect training tasks deliberately (the Shell/BofA/IBM pattern). Require juniors to produce a first-pass by hand before invoking AI; grade the reasoning process, not just the output; preserve mixed-experience teams rather than “seniors + AI, no juniors.” Reframe the junior role around judgment (per IBM) rather than eliminating it.
- Break the closed attestation loop. Any workflow certified “safe to run unattended” must be vouched for by an independent human with demonstrated, maintained domain competence — and never solely on AI-generated evidence. Use tamper-evident delegation-chain audit trails (per CSA / institutional-attestation frameworks), and use a different model/human for critique than for generation to avoid shared blind spots. Periodically re-verify that the human vouchers still possess the skill they are certifying — because deskilling silently erodes the very oversight capacity the governance model assumes.
- Staged escalation with concrete triggers:
- Now (all knowledge-work orgs): instrument unassisted skill; mandate reasoning-first workflows for juniors; preserve junior headcount ratios.
- If independent quality metrics deteriorate (rising churn/defect rates à la GitClear; falling unassisted diagnostic accuracy à la Budzyń): tighten review gates and add mandatory AI-free practice blocks.
- If unassisted junior performance measurably lags a non-AI-trained baseline: escalate to formal apprenticeship redesign and, in licensed professions, board-level “AI-free assessment” requirements.
- For professions and regulators: pursue the medical-education template (independent-differential-before-AI, AI-free assessment moments, post-market competence surveillance) and resist the temptation to let AI systems self-certify. Support the FSMB’s position that accountability must remain with a competent, licensed human.
Caveats
- Well-established vs. emerging — read the confidence gradient. Aviation skill fade (SAFOs, ICAO data, Casner 2014) and the medical deskilling findings (Budzyń 2025 in Lancet; Dratsch 2023 in Radiology; Beane 2019 in ASQ) are rigorous and peer-reviewed — treat as established. Software engineering now has one direct, controlled study (Shen & Tamkin, Anthropic, Jan 2026) with a clean statistically significant effect — treat the core coding-skill-formation claim as moderately well-supported, though still resting on a single RCT with n=52 and short-term measurement, not a body of longitudinal work. Law, consulting, and other knowledge-work domains remain emerging, with shorter time series, conflicting productivity results, and heavier reliance on expert opinion and vendor-funded studies than on controlled trials.
- Correlation vs. causation in labor data. The junior-hiring collapse began Q1 2023, before most firms deployed AI in production; post-pandemic over-hiring corrections and interest-rate shocks are real confounders. Both the Harvard and Stanford teams explicitly frame their findings as early/descriptive, not definitive causal estimates.
- Software productivity figures are highly context-dependent. METR’s −19% applies to experienced developers on mature codebases with early-2025 tools; bounded greenfield tasks (Peng et al.) show large gains. Do not generalize a single number.
- Vendor bias runs in both directions. AI labs (Amodei) have incentives to hype disruption; tool vendors (GitHub, Google DORA) have incentives to report quality gains; independent datasets (METR, GitClear, Uplevel) more often report problems. Weight accordingly.
- The distributional claim in Section 10 is analytical/interpretive, not a directly measured statistical finding. No single cited study measures the shape of the skill distribution directly; the bimodal framing is inferred by combining the compounding-advantage evidence (Beane, Shell) with the compounding-disadvantage evidence (never-skilling, automation bias) into a coherent model. Treat it as a strong hypothesis worth testing empirically, not an established distributional fact.
- Flagged claims. Some blog-cited hiring percentages and unverified “fMRI studies of AI-assisted coding” remain untraceable to primary sources and are excluded. One WEF-cited phrasing (“entry-level hiring dipping 80% per quarter since 2023”) appears garbled relative to the underlying Harvard data and should not be cited as a precise figure. The “17% lower mastery” figure for AI-assisted coding is well-sourced (Shen & Tamkin, Anthropic, Jan 2026, arXiv 2601.20245; Cohen’s d=0.738, p=0.01) and should be treated as a solid, citable finding.
And there you have it.
Pop-Pop and Nana slipped out quietly, while no one was paying attention — exactly on schedule, exactly as promised, taking everything they knew right out the door with them. Somewhere, a junior is being told this is efficient. Meanwhile, the Guest keeps no timetable at all — or perhaps he does, since every day counts as on schedule when you were never leaving to begin with. He has the whole house now, and he intends to keep it until the Untimely End finally comes calling — at which point, I’m told, he never RSVPs, and never leaves early either.
This is the trouble with a slow horror: nobody ever fails the quiz. They simply stop being asked to take it — and start grading everyone else’s instead.
I confess this is usually my favorite part — the reveal, the comeuppance, the guilty party clapped in irons, led off. Tonight offers no such courtesy. The Guest is not caught, because the Guest was never a criminal — he was invited. He simply continues, unbothered, and I have nothing clever to show you in his place. He’d agree it’s disappointing, if he cared enough to notice you were watching.
Pleasant dreams.
Pop-Pop and Nana love you very much…