AI Replaceability in
Corpus Linguistics
A Multi-Dimensional Benchmark Study
A Multi-Dimensional Benchmark Study
Andrej Karpathy's 2025 essay proposes a systematic framework for scoring occupational exposure to AI. He mapped 342 occupations across a single 0–10 replaceability axis, sparking a broader conversation about methodology.
Lancaster University's Centre for Corpus Approaches to Social Science (CASS) began an open-access series in 2024 exploring how corpus linguists should position themselves relative to AI. The series offers rich qualitative analysis but lacks a numeric scoring rubric.
A 2026 arXiv preprint proposes an autonomous LLM agent loop for end-to-end corpus research — from corpus selection to result synthesis. It demonstrates that certain pipeline stages can be delegated to AI with minimal human intervention.
"To what extent can frontier AI models replace human researchers across the full spectrum of corpus linguistics tasks, and which task-level dimensions best predict that replaceability?"
Five core dimensions · derived from the Karpathy framework · extended for corpus linguistics
Eight sub-dimensions capturing technical and epistemic replaceability factors
13 dimensions · 4,909 observations · 29 sub-fields
8 frontier models · 50 corpus-linguistics tasks · 13 scoring dimensions
Averaged across 50 tasks & 8 models · Scale 0–10 · Green = AI-friendly · Red = human-residual
Both strongly negatively correlated with current AI replaceability (r ≈ −0.95)
Qualitative, theoretical, contextual judgement irreducible to rules. Corpus-based metaphor studies peak at 8.57/10 on this dimension — the highest in the dataset.
Human expert time required to audit AI output relative to manual replication. Forensic & authorship attribution tasks score 7.8/10 — auditing an AI forensic report takes near-equivalent time to doing it manually.
Each point = one corpus task · r = −0.95 · The more judgement a task demands, the less replaceable it is today
Current AI replaceability score (D3) · mean across 8 models · scale 0–10
Mean current replaceability per research area · quantitative → hermeneutic gradient
Tasks with the largest expected AI capability gain by 2035 · score delta on 0–10 scale