🔬

AI Replaceability in
Corpus Linguistics

A Multi-Dimensional Benchmark Study

Background · Karpathy (2025)

"Taxonomy of AI Replaceability"

Andrej Karpathy's 2025 essay proposes a systematic framework for scoring occupational exposure to AI. He mapped 342 occupations across a single 0–10 replaceability axis, sparking a broader conversation about methodology.

342
Occupations Mapped
0–10
Replaceability Scale
1
Dimension Only
Key Limitation: A single replaceability score collapses task complexity into a single number — omitting the why (what dimensions make a task hard?) and the when (current vs. projected AI capability).
Background · Lancaster CASS Series

"How Not to Be Replaced by AI"

Lancaster University's Centre for Corpus Approaches to Social Science (CASS) began an open-access series in 2024 exploring how corpus linguists should position themselves relative to AI. The series offers rich qualitative analysis but lacks a numeric scoring rubric.

📖 Scope: Covers 15+ sub-fields of corpus linguistics from lexicography to forensic linguistics, examining where AI tools currently assist vs. threaten human roles.
🔍 Analytical framing: Proposes that tasks requiring interpretive judgment, theoretical framing, and ethical deliberation remain most resistant to AI substitution.
⚠️ Core limitation: Arguments remain qualitative and essay-form — no quantified, task-level, multi-dimensional replaceability scores that could be compared or replicated.
Background · arXiv 2026

Agent-Driven Corpus Linguistics

A 2026 arXiv preprint proposes an autonomous LLM agent loop for end-to-end corpus research — from corpus selection to result synthesis. It demonstrates that certain pipeline stages can be delegated to AI with minimal human intervention.

🔄 Four-Step Agent Loop
1 · Corpus selection & sampling
2 · Query design & retrieval
3 · Statistical analysis & annotation
4 · Result interpretation & synthesis
What It Covers
Demonstrates AI capability on structured, deterministic corpus operations. Steps 1–3 are near-fully automated in the prototype.
Key Gap
No scored, multi-dimensional index of which tasks are most/least replaceable across the full corpus linguistics research spectrum.
Research Gap

What Is Missing

Karpathy covers occupations broadly — not task-level corpus linguistics, and uses only a single replaceability dimension.
Lancaster CASS provides rich analysis but no quantified, comparable scores that can be tracked across models or time.
The arXiv agent paper shows automation works for structured tasks but does not map which dimensions make tasks tractable to AI.
Research Question

"To what extent can frontier AI models replace human researchers across the full spectrum of corpus linguistics tasks, and which task-level dimensions best predict that replaceability?"

Section A

Original Scoring Dimensions

Five core dimensions · derived from the Karpathy framework · extended for corpus linguistics

1
Human Feasibility
How feasible is it for a human to perform this task?
2
Research Significance
How significant is this task to the overall research?
3
Current AI Replaceability
To what extent can today's AI replace a human on this task?
4
AI Replaceability in 10 Years
Projected replaceability given anticipated progress over the next decade.
5
Theoretical Maximum AI Replaceability
The upper-bound replaceability assuming no practical constraints.
Section B

Additional Indicators to Score

Eight sub-dimensions capturing technical and epistemic replaceability factors

6
Specification Extractability
Can the LM parse published methods into a machine-executable spec?
7
Parameter Explicitness
Does the literature report the parameters AI would need to extract?
8
Execution Determinism
Identical inputs → identical outputs?
9
Tool-chain Maturity
Can established libraries execute the task end-to-end?
10
Interpretive Judgement
How much qualitative, contextual human judgement is needed?
11
Ground-Truth Availability
Can AI replication success be objectively measured?
12
Robustness Tractability
Can AI vary parameters to map sensitivity of findings?
13
Verification Cost
Human time needed to audit AI's output vs. manual replication.
Methodology · Experiment Design

End-to-End Evaluation Pipeline

STEP 1 Task Definition 50 corpus-ling. tasks + synthesized task dataset STEP 2 LM Agent Execution Orchestrator: Nemotron 3 Super Receives task + dataset → completes the task STEP 3 Benchmarking LLM Panel Receives: task + reasoning + agent result Scores all 13 dimensions OUTPUT Scores ×13 dimensions per task Corpus research tasks Nemotron 3 Super (orchestrator) 8 frontier models in panel 4,909 total obs.
Each of the 50 tasks is run through this pipeline once per benchmarking model — yielding 50 × 8 × 13 = 4,909 scored observations in total.
Methodology · Model Selection

Models Used in This Study

🤖 LM Agent — Orchestrator
Nemotron 3 Super
Performs every corpus-linguistics task on the synthesized dataset. Its reasoning trace and output are passed to the benchmarking panel for scoring.
⚖️ Benchmarking Panel — 8 Judges
Claude Opus 4.7 Anthropic
Gemini 3.1 Pro Google DeepMind
GPT-5.4 OpenAI
GPT-5.5 OpenAI
Kimi K2.6 Moonshot AI
Nemotron 3 Super NVIDIA
Sonar-Pro Perplexity AI
Model Council Claude 4.7 + GPT-5.5 + Gemini 3.1
Empirical Findings

Scoring 50 Tasks
Across 8 Frontier Models

13 dimensions · 4,909 observations · 29 sub-fields

Findings · Overview

Dataset at a Glance

8 frontier models · 50 corpus-linguistics tasks · 13 scoring dimensions

0
Total Observations
0
Frontier Models
0
Corpus Tasks
0
Sub-fields
7.35 / 10
Mean current AI replaceability
8.94 / 10
Mean projected in 10 years
9.60 / 10
Theoretical maximum mean
Findings · Dimensions

Mean Scores Across All 13 Dimensions

Averaged across 50 tasks & 8 models · Scale 0–10 · Green = AI-friendly · Red = human-residual

Findings · Human Residual

Two Dimensions Resist Automation

Both strongly negatively correlated with current AI replaceability (r ≈ −0.95)

Dimension 10
Interpretive Judgement Load
4.34 / 10

Qualitative, theoretical, contextual judgement irreducible to rules. Corpus-based metaphor studies peak at 8.57/10 on this dimension — the highest in the dataset.

Dimension 13
Verification Cost
4.46 / 10

Human expert time required to audit AI output relative to manual replication. Forensic & authorship attribution tasks score 7.8/10 — auditing an AI forensic report takes near-equivalent time to doing it manually.

Together, D10 + D13 form the "irreducible human core" — tasks with high scores on both dimensions are the least economically viable to automate even at high capability levels.
Findings · Correlation

Interpretive Judgement vs. Current AI Replaceability

Each point = one corpus task · r = −0.95 · The more judgement a task demands, the less replaceable it is today

0 2 5 7 10 0 2 5 8 10 D10 Interpretive Judgement Load → D3 Current AI Replaceability → Wordlist / TTR Metaphor / CDA
Point Colour = Sub-field
Lexicography / TTR
Diachronic / learner
Pragmatic / social
Forensic
Metaphor / CDA
r = −0.95
Pearson correlation across 50 tasks
Findings · Task Rankings

Most vs. Least Replaceable Tasks

Current AI replaceability score (D3) · mean across 8 models · scale 0–10

🟢 Highest Replaceability
🔴 Lowest Replaceability
Findings · Sub-field Analysis

AI Exposure by Sub-field

Mean current replaceability per research area · quantitative → hermeneutic gradient

Findings · 10-Year Dynamics

Biggest Projected Catch-Up: Now → 10 Years

Tasks with the largest expected AI capability gain by 2035 · score delta on 0–10 scale

All top catch-up tasks are from interpretive sub-fields — metaphor, discourse, cognitive, forensic. Mechanical ceiling near-reached; interpretive frontier closes fastest over 2025–2035.
Findings · Model Agreement

Inter-Model Agreement on Replaceability

Model profiles · 5 key dimensions (0–10 scale)
Replaceability Spec. Extract. Exec. Det. Toolchain Verif. Cost
Sonar Pro (highest) Nemotron-3 (lowest)
Sonar Pro
Most AI-optimistic · current mean 7.89 · lowest verification cost 3.04
Nemotron-3
Most conservative · current mean 6.74 · highest judgement scores
Model spread: σ = 0.33–1.42
Low σ=0.33 on wordlist extraction · High σ=1.42 on metaphor interpretation — uncertainty highest where it matters most
Strong consensus on mechanics
All 8 models agree D1 Human Feasibility = high (mean 8.52); diverge on D10 Interpretive Load
Conclusion

What the Benchmark Tells Us

Corpus linguistics is already highly automatable — field-average current replaceability 7.35/10 · mechanical & lexicogrammatical tasks reach 9/10 today.
🧠 The irreducible human core — Interpretive Judgement Load (4.34/10) + Verification Cost (4.46/10) — concentrated in metaphor, discourse, cognitive & forensic corpus work.
📈 10-year catch-up fastest in interpretive tasks (+2.62 pts for MIPVU metaphor identification). Mechanical ceiling near-reached; interpretive frontier closes faster than expected.
🤝 Models agree on structured tasks, diverge on interpretation — σ=0.33 on wordlist extraction vs. σ=1.42 on metaphor. Epistemic uncertainty is highest where it matters most.