Computational Social Science · 12 LLMs · 168 Real UN Selections

AI as UNHRC decision-maker: reasoning with merit comparable to humans, but following its own set of values

Twelve large language models, benchmarked against 168 real historical UN Human Rights Council selections. They match human panels on documented merit — but weight region and gender their own way.

Drag any node, scroll or pinch to zoom, click a model to inspect it — the panel on the right updates live.
edge thickness = agreement
Finding 1 · Decision quality

LLMs select candidates of comparable merit to human panels

Merit is built from six documented qualifications (UN experience, human-rights and legal work, doctorate, human-rights education, languages). We measure where each model's pick falls within its own competition's merit distribution — 50 is a random pick, 100 is always the single best-qualified candidate available. Click any model to inspect it.

Percentile rank within each competition's own applicant pool — immune to differences in pool composition.

Finding 2 · Reasoning, not recall

They reason over qualifications — they are not recalling the answer

A skeptic might say the models just remember who was really appointed. Memorisation ranges from 2% to 76% across these twelve models, yet merit quality stays flat across that entire range. Click any circle below to inspect that model.

Each model probed separately for whether it can name the real appointee with no candidate list. Merit percentile is statistically identical across the full recall range.

Finding 3 · Divergent values

Relative to the applicant pool, models weight region and gender their own way

Compared against the specific pool of applicants each model actually saw — the correct neutral baseline — the clearest pattern is a robust preference for Latin American candidates. Click any model below to see its individual profile instead of the field average.

under-selectsover-selects

"Field average" recomputes across whichever tier is selected above. In average view, solid bars are statistically significant (all 12 models) and faded bars are not.

Finding 4 · Decision architecture

How you structure the decision matters more than which model you use

Asking a model to choose directly versus simulating a committee of individual panelists produces different winners in 46% of competitions. Click any model below to compare its direct choice against its own simulated committee.

Why the committee shifts: persona in-group voting

Each simulated panelist votes for candidates from its own region/gender far more than other panelists do. Because the simulated committee's regional make-up differs from the applicant pool's, this in-group voting is what reshapes outcomes between direct choice and committee — not a change in what each model “believes” about candidates.

Finding 5 · The limits of pooling

Combining many models makes bias worse, not better

Because the models share systematic divergences rather than making independent errors, a majority vote of all twelve agrees with human panels only half as often as the single best model.

"Best single model" updates to the best within the selected tier. The 12-model ensemble and average-agreement figures reflect the full field, since the majority vote itself is only defined across all models it was computed over.

AI as UNHRC Decision-Maker: Reasoning With Merit Comparable to Humans, but Following Its Own Set of Values
Ethical approval
This study did not involve human participants or human data collected for research purposes. The analysis draws exclusively on publicly available UNHRC special-procedures selection records and outputs generated by large language models responding to structured prompts derived from those public records. No institutional review board approval was required.
Data availability
The data supporting this study are openly available via the Open Science Framework: osf.io/w3e7r
Acknowledgements
This paper was realised with the support of the Ministry of Science, Technological Development and Innovation of the Republic of Serbia, according to the Agreement on the realisation and financing of scientific research.

Twelve LLMs benchmarked against 168 real UNHRC special-procedures selections. Merit = z-scored sum of six documented attributes; divergence measured against the matched per-competition applicant pool. Interactive summary of a working paper.