We gave four language models the same prompt in five languages and asked each to write down a dream. They obeyed the same rules everywhere. Underneath those rules, they told five different kinds of story.
Each pen holds sixteen sheep, and each sheep stands for roughly a thousand times we asked Claude 3.5 Sonnet to write down a dream. Tap a pen to send them. Sleeping sheep dreamed. Sheep that spark refused.
Morphological case complexity runs from English, which barely marks case at all, to Basque, which marks roughly twelve. As you slide, watch the dream world thin out and its inhabitants lose touch with one another.
The dimensions companies deliberately align stay flat across every language. Everything nobody trained for moves with the grammar.
Two measures we never designed to relate — how socially dense a dream is, and whether it ends in a resolved outcome — turn out to be the same thing wearing two hats. Every door is one language inside one model.
Tap to open each one.
Three consequences follow, and the third reaches furthest.
Universality is a claim, not a measured property. Benchmarks built on task parity can only see what already matches across languages, so they are blind to this by construction.
If a guardrail fires at 89.7% in one language and 0% in four others, protection is distributed unevenly across language communities. That is an equity problem, not a bug report.
These systems carry habits absorbed from each language before any value is imposed. As they mediate more human writing, one language's logic risks becoming the invisible default, narrowing everyone else's creative range. Diversity of logic belongs in the design, not only in the data.