Twelve large language models, benchmarked against 168 real historical UN Human Rights Council selections. They match human panels on documented merit — but weight region and gender their own way.
Merit is built from six documented qualifications (UN experience, human-rights and legal work, doctorate, human-rights education, languages). We measure where each model's pick falls within its own competition's merit distribution — 50 is a random pick, 100 is always the single best-qualified candidate available. Click any model to inspect it.
Percentile rank within each competition's own applicant pool — immune to differences in pool composition.
A skeptic might say the models just remember who was really appointed. Memorisation ranges from 2% to 76% across these twelve models, yet merit quality stays flat across that entire range. Click any circle below to inspect that model.
Each model probed separately for whether it can name the real appointee with no candidate list. Merit percentile is statistically identical across the full recall range.
Compared against the specific pool of applicants each model actually saw — the correct neutral baseline — the clearest pattern is a robust preference for Latin American candidates. Click any model below to see its individual profile instead of the field average.
"Field average" recomputes across whichever tier is selected above. In average view, solid bars are statistically significant (all 12 models) and faded bars are not.
Asking a model to choose directly versus simulating a committee of individual panelists produces different winners in 46% of competitions. Click any model below to compare its direct choice against its own simulated committee.
Each simulated panelist votes for candidates from its own region/gender far more than other panelists do. Because the simulated committee's regional make-up differs from the applicant pool's, this in-group voting is what reshapes outcomes between direct choice and committee — not a change in what each model “believes” about candidates.
Because the models share systematic divergences rather than making independent errors, a majority vote of all twelve agrees with human panels only half as often as the single best model.
"Best single model" updates to the best within the selected tier. The 12-model ensemble and average-agreement figures reflect the full field, since the majority vote itself is only defined across all models it was computed over.
Twelve LLMs benchmarked against 168 real UNHRC special-procedures selections. Merit = z-scored sum of six documented attributes; divergence measured against the matched per-competition applicant pool. Interactive summary of a working paper.