Tairan Yang
Adv. Artif. Intell. Mach. Learn., - (-):-
1. Tairan Yang: Mulgrave School
DOI: 10.54364/AAIML.2026.65340
Article History: Received on: 05-May-26, Accepted on: 30-Aug-26, Published on: 05-Sep-26
Corresponding Author: Tairan Yang
Email: melody.yang0825@gmail.com
Citation: Tairan Yang. Content-Level Preferences in Large Language Model Resume Screening: A Counterbalanced Paired-Comparison Audit. Advances in Artificial Intelligence and Machine Learning. 2026. (Ahead of Print) https://dx.doi.org/10.54364/AAIML.2026.65340
Audits of large language model (LLM) hiring bias have focused overwhelmingly on demographic signals such as names, pronouns, and group labels. Whether LLMs also hold systematic preferences about the content of candidates’ experience (the kind of work they do or the social background they signal) remains underexamined, even as employment AI is increasingly subject to regulation. We audited three LLMs (GPT-4o, Qwen 2.5-7B, and Llama 3.1-8B) across 62,208 evaluations of 96 counterbalanced resume-pair templates. Each pair varied on exactly one of three factors: project-type framing (applied and metrics-driven versus speculative and creative), community-engagement framing (working-class-coded service versus aesthetic-elite cultural), or gender. Pairs were presented under three prompt conditions with full order counterbalancing, and content effects were separated from position artefacts using generalised estimating equations. All three models preferred applied-framed experience (selection rates 57–88%) and working-class-coded service engagement (52–96%), although the latter manipulation bundles class coding with activity type. Gender preferences were small and, for GPT-4o, non-significant after adjustment for position bias. On trials where all three models agreed, 94–98% of agreements pointed in the same direction, and no prompt condition reversed any preference. These findings indicate that LLM-assisted screening can exhibit systematic, cross-model content-level preferences that demographic-only audits would miss, that switching providers may not remove, and that the prompt-based mitigations tested here did not eliminate.