Weihang
Huang
I study how language models encode authorial identity — and how that encoding can be used for authorship attribution, forensic linguistics, and understanding what it means to write.
Scholar Metrics
Advisors:
Jack Grieve & Akira Murakami
Updates
Recent News
Paper published — “Attributing authorship via the perplexity of authorial language models” in PLOS ONE 20(7), e0327081.
Conference talk — “Estimation of Relative Frequency via Large Language Model” at the International Corpus Linguistics conference.
New preprint — “ChatGPT-generated texts show authorship traits that identify them as non-human” (Dentella, Huang, Mansi, Grieve & Leivada). arXiv:2508.16385.
New preprint — “Metaphor identification using large language models: A comparison of RAG, prompt engineering, and fine-tuning.” arXiv:2509.24866.
Paper published — “The sociolinguistic foundations of language modeling” in Frontiers in Artificial Intelligence, 7, 1472411.
Workshop paper — “Authorial language models for AI authorship verification” at CLEF 2024 (PAN). CEUR Workshop Proceedings 3740.
Focus Areas
Research Interests
My work sits at the intersection of computational linguistics, natural language processing, and forensic stylistics — asking how language models can be used as tools to understand human authorship.
Authorship Attribution
Using authorial language models (ALMs) to attribute texts to their authors — with applications in forensic linguistics, AI detection, and stylometry.
Large Language Models
Investigating how LLMs encode stylistic and sociolinguistic signals, and leveraging model perplexity as a linguistic measurement tool.
Corpus Linguistics
Building and analysing large-scale linguistic corpora to study language variation, frequency estimation, and sociolinguistic patterns in text.
Metaphor Identification
Comparing RAG, prompt engineering, and fine-tuning for automated metaphor identification in natural language using LLMs.
Forensic Linguistics
Applying computational stylistics to real-world forensic problems: authorship verification, AI-generated text detection, and legal linguistic analysis.
Translation & Bilingualism
Studying code-switching in bilingual children and phylogenetic approaches to tracing the origins of language varieties.
Scholarship
Selected Publications
A selection of recent journal articles and workshop papers. View the full list →
The sociolinguistic foundations of language modeling
Frontiers in Artificial Intelligence, 7, 1472411 38 citationsChatGPT-generated texts show authorship traits that identify them as non-human
arXiv:2508.16385Metaphor identification using large language models: A comparison of RAG, prompt engineering, and fine-tuning
arXiv:2509.24866Authorial language models for AI authorship verification
CEUR Workshop Proceedings 3740 — CLEF 2024 (PAN) 7 citationsBackground
Education & Positions
Education
Doctor of Philosophy — Applied Linguistics
University of Birmingham
Sep 2022 – Jul 2026 (expected)
Style in the Eye of Machine: Authorial Attribution via Large Language Models
Supervisors: Jack Grieve, Akira Murakami
Master of Arts — Linguistics
The Chinese University of Hong Kong
Sep 2019 – Jul 2020 · GPA: 3.73/4.00
Bachelor of Arts — English (Interpreting)
Beijing Sport University
Sep 2015 – Jul 2019 · GPA: 4.17/5.00
Academic Positions
Research Associate
Dept. of Linguistics & Communication, University of Birmingham
Aug 2023 – present
Research Assistant & Software Developer
Institute of Microbiology, Chinese Academy of Sciences
May 2021 – Feb 2022
Research Assistant
Language Acquisition Lab, CUHK
Aug 2020 – Jun 2021
Research Assistant
School of Sports Engineering, Beijing Sport University
Sep 2018 – Jun 2019
Competencies
Skills
Languages
Programming
ML & Modelling
Text Mining & NLP
Translation Certs.
Get in Touch
Happy to hear from collaborators, researchers, or anyone curious about language and computation.