Complexity Science · Interactive benchmark

Can you out-predict the AI?

We asked 27 language models to guess how 1,511 real people would react to social media posts. Now it is your turn. Step into a real person's shoes, react to their feed, and see whether you read them better than the machines did.

1,511
real participants
27
language models
120K+
AI personas
6.85M
predictions
Play · Read the room

Become one of our participants for ten posts

Each persona is built from the real survey traits we fed our AI agents: age, region, politics, trust in institutions, and taste. Pick one, react to ten posts the way you think that person would, then see your accuracy against how people of that profile actually reacted, and against the agents. The names and faces are invented, and no real participant is identified.

Choose a persona to play · names and faces are invented
Post 1 of 10
React the way  this person  would
You: 0/0
0/10
You
0%
Text classifier
69%
Best LLM agent
67%
Why the AI struggles too
01 · The experiment

One survey, then a simulated crowd of millions

Every participant answered a survey and reacted to 56 posts. Their answers became persona prompts at three levels of detail, each fed to 27 models. The agents then faced the same posts, producing a behavioral benchmark at a scale no lab could run with real users.

1,511
Serbian participants, each reacting to 56 real posts
27
language models from nine families, one prompt structure
6.85M
individual reaction predictions generated in total
02 · Which model reads people best?

The model you pick matters more than anything else

Accuracy swung 13 points from the best model to the worst, a bigger gap than richer personas, matched content, or any other design choice produced. Model choice, not persona craft, was the dominant lever.

Accuracy predicting human reactions · Study 1
Estimated marginal means across all five reaction types. Hover a bar for the model family.
03 · The twist

A plain text classifier beat every "intelligent" agent

Here is the finding that reframes the whole study. On the hardest test, predicting whether a specific person likes a specific unseen post, a simple TF-IDF classifier outscored the best LLM agent. The apparent understanding was mostly the words carrying the signal.

Text classifier (TF-IDF + logistic regression)
0.00
Best LLM agent (GPT-5.2)
0.00
Structured baseline (demographics only)
0.00
Matthews correlation coefficient (0 = no signal). The text model leads the best agent by 0.064 MCC and was also far better calibrated.

The agents were not useless. They clearly beat the structured baselines that had no access to the text. But once a conventional model could read the same persona and post text, the agent's edge vanished. The signal lived in the words, not in any capacity for agentic simulation.

04 · A lopsided ear

The agents hear applause better than dissent

Prediction accuracy was consistently higher for positive posts than negative ones, a systematic positivity bias. For platform governance that is the wrong ear to favor, since backlash and dissent are exactly what simulations most need to anticipate.

67.9%
accuracy on positively framed posts
54.8%
accuracy on negatively framed posts
a persistent up to 13-point gap, stable across models and prompts
05 · The scorecard

Six hypotheses, six honest verdicts

Not everything we expected held up. Richer personas barely helped, content domain flipped between models, and the agents never overtook the text classifiers.

06 · What it means

Genuine signal, modest reach

Persona-prompted agents do capture something real about how people react, enough to flag who might push back on a post 54% better than chance. The signal is semantic rather than agentic, and it favors the positive. Deploying swarms of them to sway a feed is a real risk, while trusting them to model one specific person is not yet warranted.