SCOPED-Hiring paper flags hidden unfairness in multi-agent LLM hiring decisions
A paper accepted to the EMNLP 2026 Main Conference introduces SCOPED-Hiring, a process-aware fairness diagnosis framework for large language model (LLM) multi-agent systems used in hiring. The study argues that similar final hiring rates across groups do not necessarily mean the decision process was fair. Candidates can still face uneven burdens in the form of extra doubt, lower intermediate scores, added scrutiny, and different investigative standards. Using controlled resume variants and a two-stage hiring committee made up of role-based agents, the researchers recorded more than 311,000 structured decision trajectories. Across GPT, Gemini, and Qwen backends, they found that risks visible through process, pathway, interaction, and system design were much stronger than those visible from final outcomes alone. In the paper’s measurements, process-aware O/P/E/D signals showed average diagnostic significance 4.52x, 4.66x, and 2.74x higher than outcome-focused S/C signals. The team also tested a diagnosis-driven repair method called Fair Skills. In controlled experiments, the intervention reduced total stratified fairness burden by 72.3% while changing final hiring rates by only 1.86 percentage points, with output validity above 99.7%.








