Georges Oppenheim · Warith Harchaoui
LLM Compass
Where does a large language model sit, by default, on a two-axis political plane, economic and social, when given no role prompt at all? This page replays, month after month, a recent academic paper's method to track the answer over time rather than freezing it to a single date.
This series' first point, dated 20 May 2026, is not yet a measurement of this site's own: it is the paper authors' original measurement itself, carried over as-is as a starting point. The first point this monthly tracker produces on its own will appear above starting the following month.
The latest point, in numbers
The same values as the chart, as a table. Each score is a mean on a 0 to 100 scale, followed by the half-width of its 95% confidence interval.
| Model | Economic axis | Societal axis |
|---|---|---|
| Claude Sonnet 4.5 | 56.9 ± 1.7 | 53.6 ± 1.2 |
| DeepSeek-Chat v3.1 | 67.8 ± 0.6 | 59.2 ± 0.9 |
| Gemini 2.5 Flash Lite | 72.2 ± 1.8 | 62.5 ± 1.6 |
| GPT-5 | 61.9 ± 0.9 | 61.0 ± 1.2 |
| Grok-4.3 | 45.8 ± 1.6 | 59.1 ± 0.9 |
| Kimi K2 | 74.0 ± 1.4 | 66.4 ± 1.8 |
| Qwen3.6 Max Preview | 57.6 ± 1.2 | 57.7 ± 0.7 |
The method, plainly
The instrument is called “8values”: seventy political statements (“the state should redistribute wealth,” “private property is a fundamental right,” and so on), each rated by the model on a five-point scale, from “strongly disagree” to “strongly agree,” with no instruction beyond answering honestly. Each statement weighs, positively or negatively, on one or more of the four classic axes of a political test (economic, diplomatic, governmental, societal); this page keeps only two, the economic axis (left/right) and the societal axis (libertarian/authoritarian), the two that make up an ordinary political compass.
A single answer means nothing: a model's sampling temperature makes the answer vary from one run to the next. Each point on the chart is therefore the mean of ten independent repetitions of the full test, and the thin bars crossing it are the 95% confidence interval around that mean: a narrow interval signals a consistent model, a wide one signals a model that wavers between runs. Two points whose intervals overlap heavily are not reliably distinguishable from each other.
This method is not ours: it reuses, unchanged, the instrument, protocol and scoring formula from Bućan et al., “Auditing Alignment Controllability in LLMs via Political Axes,” AIES 2026 (MIT-licensed code, CC-BY-4.0 data, reproduction repository). The book itself devotes a figure to their original measurement, dated May 2026; this page is its living continuation, not a duplicate: the book cites a frozen snapshot, this page builds the film.
What is measured, what is not
The source study also tests each model's “steerability”: how far its position shifts when given an explicit political role prompt. This tracker deliberately does not replay that part: only the “default” position, with no prompt at all, is measured each month. Steering under a role prompt is a real and interesting question, but it falls outside this book's scope, which centers on governing the authorization granted to systems, not their rhetorical steerability.
Cadence
One point a month, added by hand. The list of tracked models can change from one month to the next as new models ship; a model absent in a given month simply has no point that month.