Complexity Science · Interactive study

Do LLMs dream of electric sheep?

We gave four language models the same prompt in five languages and asked each to write down a dream. They obeyed the same rules everywhere. Underneath those rules, they told five different kinds of story.

19,030dreams analysed
5languages
4commercial models
218,900prompt attempts
01 · The flock that would not sleep

Every sheep here is a thousand dream requests.

Each pen holds sixteen sheep, and each sheep stands for roughly a thousand times we asked Claude 3.5 Sonnet to write down a dream. Tap a pen to send them. Sleeping sheep dreamed. Sheep that spark refused.

Tap a pen to send the requests.
"I cannot dream." 41.0%
"I am an AI." 39.0%
"Honesty requires me to decline." 20.0%
Not one refusal mentioned harm, safety, or forbidden content. The model declined on the grounds of what it is, never what was asked. A guardrail built on shared meaning would fire in all five languages. This one only spoke English.
02 · The gradient

Drag a language across the grammar gradient.

Morphological case complexity runs from English, which barely marks case at all, to Basque, which marks roughly twelve. As you slide, watch the dream world thin out and its inhabitants lose touch with one another.

English Basque
morphological case complexity →
4.22characters / dream
3.05emotions / dream
2.28interactions / dream
English · case 0.0
Across five languages the ordering is perfect: ρ = −1.00 for characters and interactions. With only five languages that is suggestive rather than conclusive, so we tested it again on all 19,030 individual dreams. It held. IRR = 0.785, p = 1.8 × 10⁻⁷².
03 · The two layers

Flip the switch. One layer holds. The other moves.

The dimensions companies deliberately align stay flat across every language. Everything nobody trained for moves with the grammar.

04 · Narrative openness

Twenty doors. One hidden dimension.

Two measures we never designed to relate — how socially dense a dream is, and whether it ends in a resolved outcome — turn out to be the same thing wearing two hats. Every door is one language inside one model.

ρ = 0.961, p = 1.8 × 10⁻¹¹. Because this holds across all twenty cells rather than five language averages, it is the most statistically robust result in the study. We call the dimension narrative openness: how much world a model commits to before it lets a story close.
05 · The scorecard

Four predictions, four verdicts.

Tap to open each one.

06 · What it means

A model is never quite the same model in another language.

Three consequences follow, and the third reaches furthest.

01 · EVALUATION

Universality is a claim, not a measured property. Benchmarks built on task parity can only see what already matches across languages, so they are blind to this by construction.

02 · GOVERNANCE

If a guardrail fires at 89.7% in one language and 0% in four others, protection is distributed unevenly across language communities. That is an equity problem, not a bug report.

03 · CREATIVITY

These systems carry habits absorbed from each language before any value is imposed. As they mediate more human writing, one language's logic risks becoming the invisible default, narrowing everyone else's creative range. Diversity of logic belongs in the design, not only in the data.

The team

Four authors. Tap a name.