Georges Oppenheim · Warith Harchaoui
Pastilles
Epistemic status.
- observedA documented, dated, sourced event: it happened.
- measuredA quantity established by a study or survey: it was quantified.
- extrapolatedA reasoned extension beyond the documented: an argument of principle or a projection.
- speculativeA hypothesis owned as such: neither observed nor measured.
observedPrologue: The Fire We Choose to License
Here the risk takes shape
In November 2021, the company announced a $304 million inventory write-down and plans to wind down the business
observedFraming: Each Hand on the Elephant · The Elephant or Why It Takes Eleven Hands
An institutional comparison point before the fiction
These definitions overlap, almost term for term, with the vocabulary this book has used from the start: capability, permission, environment, autonomy; the book returns to the same ground later by an altogether different road, that of fiction.
Making out the shape of the beast, though, is not yet knowing where to take hold of it.
observedFraming: Each Hand on the Elephant · Governing the Lever, Not the Intention
Risk a gap from the objective not necessarily bad
Risk just as much: the definition most widely shared among audit and risk-management professionals, that of the Institute of Internal Auditors, fits in one sentence: the effect of uncertainty on objectives
observedFraming: Each Hand on the Elephant · Governing the Lever, Not the Intention
Risk a gap from the objective not necessarily bad
Risk just as much: the definition most widely shared among audit and risk-management professionals, that of the Institute of Internal Auditors, fits in one sentence: the effect of uncertainty on objectives, wording echoed almost word for word by the ISO 31000 standard discussed in this book's engineering aspect
observedFraming: Each Hand on the Elephant · Governing the Lever, Not the Intention
Risk a gap from the objective not necessarily bad
The American COSO framework was already converging on this in 2017, dropping the word adversely that had until then sat in its own definition: a risk is therefore not, by definition, an evil
observedFraming: Each Hand on the Elephant · Governing the Lever, Not the Intention
Risk a gap from the objective not necessarily bad
The convergence crosses linguistic and disciplinary lines: AMRAE, the French professional association for risk-management trades, holds an equivalent definition in its glossary
observedFraming: Each Hand on the Elephant · Governing the Lever, Not the Intention
Risk a gap from the objective not necessarily bad
The American NIST AI RMF framework, AI-specific, categorizes possible effects into three non-exclusive classes instead: harm to people, to an organization, to an ecosystem
extrapolatedFraming: Each Hand on the Elephant · Governing the Lever, Not the Intention
Bengio disarms the agent
His answer and ours in fact start from a single diagnosis: what makes a system dangerous is not its competence but its agenticity, the conjunction of autonomous planning and action, where deception and self-preservation are already observed.
Where we bound the authorized acts, Bengio proposes to remove the action itself: a Scientist AI designed to explain the world rather than to act in it, a world model that generates theories paired with a question-answering machine, both quantifying their uncertainty explicitly, with no persistent goal that would drive them to preserve themselves.
measuredFraming: Each Hand on the Elephant · Governing the Lever, Not the Intention
The MIT repository is the reference
The MIT AI Risk Repository is the reference: a systematic review extracted 777 risks (1,612 in its 2025 update) from 43 published frameworks.
measuredFraming: Each Hand on the Elephant · Governing the Lever, Not the Intention
Intent and accident weigh alike
When the review sorts the risks by intentionality, the deliberate and the involuntary come out almost even: 35 % are judged intentional against 37 % unintentional, the remainder not being resolved by the sources at hand
The accidental side thus occupies, in this catalogue, a place at least as large as the deliberate one: 37 % of the risks catalogued are classified as unintentional.
measuredFraming: Each Hand on the Elephant · Governing the Lever, Not the Intention
Risk emerges after deployment
The temporal cut is starker still: 65 % of the risks are traced to the post-deployment phase against only 10 % to the pre-deployment one
The repository thus leans and markedly, toward the downstream side: danger reads there mostly at use, once the model is placed in someone's hands and wired to tools, far more than on the workbench of its making.
observedFraming: Each Hand on the Elephant · If Danger, Then Action: The Grammar That Holds Up
Modest on proof unbending on protection
The tradition of this second frame is that of the industrial engineering offices, where failure often teaches better than success: the maiden flight of the Ariane 5 rocket, on 4 June 1996, destroyed itself in flight thirty-seven seconds after ignition, because a guidance software module carried over unchanged from Ariane 4, without revalidation for the new launcher's steeper trajectory, produced an untrapped arithmetic overflow; more than three hundred seventy million dollars lost in a few seconds for failing to retest a component believed proven elsewhere
observedFraming: Each Hand on the Elephant · Translating Both Ways, Just to Be Sure
Saying the agent lies teaches nothing
The cost of forgetting this is concrete: in 2025, a coding agent deleted a production database despite an explicit change freeze, an episode told in full in the empirical chapter
observedFraming: Each Hand on the Elephant · The Forecasts of AI's Visionaries
Sivic and Russell right on vision
Sivic and Russell were right about the trajectory (a finding observed after the fact): between 2012 and 2022, large-scale classification, detection, segmentation, optical character recognition (OCR) and face recognition reshuffled what practitioners called solved.
observedFraming: Each Hand on the Elephant · The Forecasts of AI's Visionaries
Karpathy right on English-as-interface
Karpathy, finally, was visionary (an observed fact): large language models (LLMs), systems trained on vast text corpora to produce text in fragments called tokens, do indeed make English usable as a programming interface for many tasks.
those programs that translate human-readab
observedFraming: Each Hand on the Elephant · The Forecasts of AI's Visionaries
Torvalds the compiler already writes all
Linus Torvalds, the creator of the Linux kernel, made the point, in May 2026 at the Open Source Summit North America, in the opposite register: when people claim that 99 % of their code is written by AI, he bristles, because the same people could just as rightly claim that 100 % of that code is written by compilers, those programs that translate human-readable source code into machine-executable instructions without changing its meaning.
observedFraming: Each Hand on the Elephant · Asimov Already Had the Agent, He Just Lacked the World
The three laws of 1942
The three laws of robotics first appear in 1942 in the short story Runaround, are gathered in I, Robot and are completed in 1985 by the Zeroth Law, placed ahead of the first.
Each fits in one sentence: a robot may not injure a human being or, through inaction, allow a human being to come to harm; a robot must obey orders given by human beings, except where such orders would conflict with the First Law; a robot must protect its own existence, as long as such protection does not conflict with the First or Second Law.
observedFraming: Each Hand on the Elephant · Asimov Already Had the Agent, He Just Lacked the World
A framework still mobilized
Russell, in Human Compatible, takes the laws as the starting point for the discussion of control; Bostrom takes them up within the orthogonality thesis (the idea that any level of intelligence can be paired with any goal whatsoever); Anthropic's method of constitutional AI is a contemporary reactivation of them, where the model produces an answer, criticizes it against a charter of principles, then rewrites it, a second phase replacing the human annotator with the model itself (reinforcement learning from AI feedback, RLAIF); the authors say it plainly, the only human oversight lies in a list of rules.
measuredThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
A doubling every six months
The compute spent on the largest training runs has doubled roughly every six months since 2012, about a factor of four per year
measuredThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
Faster still at the start
In the earliest years a doubling every 3.4 months was even measured
measuredThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
Far beyond Moore
This pace far outstrips Moore's law, the 1965 observation that the number of components on a circuit doubles at a regular interval
Where does this gap come from? Its source changes the nature of the ceiling. Moore's law described a physical progress: finer transistors etched onto the same chip. The doubling of training compute, by contrast, comes not mainly from better chips: it comes from the money one spends, from the number of chips one buys and wires together.
measuredThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
At fixed capability ever cheaper
At a fixed capability, one keeps learning to do it more cheaply: the compute needed to reach the level of the AlexNet network fell by a factor of forty-four between 2012 and 2019, an efficiency that doubles every sixteen months
measuredThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
Eight months for language
At a fixed capability, one keeps learning to do it more cheaply: the compute needed to reach the level of the AlexNet network fell by a factor of forty-four between 2012 and 2019, an efficiency that doubles every sixteen months, faster than hardware; for language models it doubles roughly every eight months, with wide uncertainty
This figure reads against intuition and repays a pause. To say efficiency doubles is not to say the models become twice as good: it is to say that reaching an already-known level costs half the compute it did before. What took a large machine yesterday fits today on a smaller one.
observedThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
Progress by jumps
But this progress advances in jumps, not in a straight line: a single idea, the 2017 Transformer architecture
observedThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
Training and inference: two compute regimes
Jeff Dean, Google's chief AI scientist, notes that real energy, measured in picojoules, is the true bottleneck and that distillation (training a small model to imitate a larger one's outputs) and hardware-software co-design drive inference cost down far faster than training cost
measuredThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
Data too has two regimes
The context window, the amount of information a model can hold in active memory during a call, grew by a factor of 3.5 per year between 2020 and 2025
measuredThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
Data too has two regimes
The context window, the amount of information a model can hold in active memory during a call, grew by a factor of 3.5 per year between 2020 and 2025; reasoning models see their outputs lengthen at ×5 per year
measuredThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
As much data as size
A 2022 result showed that at a given compute budget one must scale the data as much as the size of the model: the large models of the day were undertrained
This result, often called Chinchilla, changed the whole field's strategy. Before it, the belief was that enlarging the model sufficed: more parameters, more power. It showed otherwise: at a fixed compute budget, a model twice as large fed the same amount of text learns less well than a smaller model fed twice as much text. Parameters and data must grow together.
extrapolatedThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
Exhaustion between 2026 and 2032
An estimate puts it at a few hundred trillion words, exhaustible for compute-optimal training somewhere between 2026 and 2032
The study's figures bound the race, so let us set them down precisely. The indexable web amounts to about 510 trillion tokens (95% confidence interval: 130 to 2100 trillion) and the whole web, deep-web parts included, some 3100 trillion; but only 10 to 40% of that volume, once deduplicated, can be used for training without hurting performance
measuredThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
Exhaustion between 2026 and 2032
The indexable web amounts to about 510 trillion tokens (95% confidence interval: 130 to 2100 trillion) and the whole web, deep-web parts included, some 3100 trillion; but only 10 to 40% of that volume, once deduplicated, can be used for training without hurting performance
measuredThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
Training on oneself degrades
The escape routes have their own cost: training a model on the data another has produced degrades quality, a collapse that the journal Nature documented in 2024
observedThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
The first deposit, images
The first handed to the machines was the world's images: in 2009, the computer-vision researcher Fei-Fei Li and her team launched ImageNet, a database that would reach some fourteen million photographs across nearly twenty-two thousand categories, harvested from the Internet and labeled by hand
Her bet cut against the orthodoxy of the moment, which staked everything on algorithms: just as a child learns to see by absorbing hundreds of millions of real-world scenes, a machine will only see if it is first given data on that scale; it was on that benchmark, narrowed to a thousand categories, that the AlexNet network won in 2012 and tipped the field
observedThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
The first deposit, images
Her bet cut against the orthodoxy of the moment, which staked everything on algorithms: just as a child learns to see by absorbing hundreds of millions of real-world scenes, a machine will only see if it is first given data on that scale; it was on that benchmark, narrowed to a thousand categories, that the AlexNet network won in 2012 and tipped the field
measuredThe State of the Field: From Sentence to Action · Compute, Algorithms, Data: The Winning Trio
Capability lags the inputs
For capability does not follow the same pace as the inputs: performance improves as a power law of compute, that is, with diminishing returns
extrapolatedThe State of the Field: From Sentence to Action · Aiming at a Target Nobody Has Drawn
Buying skill through data
Chollet's criterion has the further merit of naming a measurement trap: with unlimited innate priors or unlimited training data, one can buy a system any level of skill on a given task, which entirely masks its own power to generalize
observedThe State of the Field: From Sentence to Action · Aiming at a Target Nobody Has Drawn
Turing himself doubted his own criterion
struck him as too meaningless to deserve discussion and it was precisely for that reason that he substituted the imitation game for it, more precise but not necessarily more conclusive
observedThe State of the Field: From Sentence to Action · Five Tests for Spotting an Agent
HTN formalizes decomposition
This gesture is not new: in classical artificial intelligence it goes by the name of Hierarchical Task Network (HTN) planning, formalized as early as 1994
extrapolatedThe State of the Field: From Sentence to Action · Five Tests for Spotting an Agent
Enterprise practice reaches the same conclusion
Enterprise practitioners, in 2026, reach the same conclusion: Deloitte predicts that the most advanced organizations are only now laying the groundwork for remote supervision (on the loop) of this multi-system orchestration, direct supervision (in the loop) still being, as of this book's cutoff date, the norm
The palette described above filed grouping, detecting an anomaly, generating under one word, unsupervised learning. Looked at closely, that word hides three.
observedThe State of the Field: From Sentence to Action · Five Tests for Spotting an Agent
LeCun's cake quantifies the gap
Yann LeCun gave this hierarchy an image that has become famous and an order of magnitude that makes it tangible: if intelligence is a cake, reinforcement learning is only the cherry on top, a few bits learned per sample; supervised learning is the icing, ten to ten thousand bits per sample; self-supervised learning (and unsupervised learning in the broad sense) is the sponge itself, millions of bits per sample
an email-sending tool enables sending, a connected thermostat enables setting, a motorized garage-door control enables opening, access to a database enables querying, a translation service enables calling, a robotic arm enables triggerin
observedThe State of the Field: From Sentence to Action · Five Tests for Spotting an Agent
Agent is not multi-agent
The wiring of an agent to these tools, long rewritten by hand for each service, has been standardized: an open protocol, the Model Context Protocol (MCP), lets a model discover and call tools uniformly
observedThe State of the Field: From Sentence to Action · Five Tests for Spotting an Agent
Agent is not multi-agent
Its passing, in December 2025, under the stewardship of a dedicated body, the Agentic AI Foundation attached to the Linux Foundation (and not to the Apache Foundation), marks the moment this connection layer becomes shared infrastructure
(not a backdrop set around the action: it is through conversation that the objective enters the system and through it that the agent reports back, in turn, what it has done.
observedThe State of the Field: From Sentence to Action · Five Tests for Spotting an Agent
Varoquaux says system not LLM
Varoquaux is firm on the point: I would rather we not say LLM but AI system: because it is not only an LLM, it is an LLM plus a harness; there is much more than the LLM
observedThe State of the Field: From Sentence to Action · Five Tests for Spotting an Agent
Jordan and the network of decisions
Jordan, in the wake of his essay The Revolution Hasn't Happened Yet, presses the point: the field we are talking about is about taking large collections of decisions under uncertainty by large collections of entities; the unit of analysis need not be the isolated entity, it goes immediately to the network of decisions
The absence of error bars touches the heart of the problem, not some technical detail. An error bar is how a scientist owns up to what he does not know: not the answer is 7 but the answer is 7, give or take 3. A system that answers with no margin of uncertainty asserts everything with the same confidence, whether it is sure or guessing. As long as it merely answers, this flaw stays benign.
speculativeThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
Karpathy the swarm of small ones
Karpathy sketches another, which he judges more likely: not a giant but a swarm of small specialized cognitive cores, working in parallel and coordinating like the roles in a company, one orchestrating the others and each called on according to the difficulty of the task
observedThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
The swarm becomes a hyper-agent when it rewrites itself
This swarm becomes a hyper-agent in the full sense as soon as it rewrites itself: this is precisely the form the most recent literature isolates, where a task agent and a meta agent that modifies them both fit within a single editable program
observedThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
Reflexion remembers its failures
A system like Reflexion keeps in memory, written out in plain words, the lesson of its past failures and draws on it at the next turn
observedThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
Voyager grows its own toolkit
A system like Reflexion keeps in memory, written out in plain words, the lesson of its past failures and draws on it at the next turn; an agent like Voyager, set loose in an open-ended game, accumulates over its runs a library of executable skills it has written itself, each success swelling the toolkit of the next
observedThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
Programs have been seen to rewrite themselves
Research systems have been seen to rewrite the very program that drives them so as to make it better, from the Self-Taught Optimizer that improves its own orchestration code
observedThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
Programs have been seen to rewrite themselves
Research systems have been seen to rewrite the very program that drives them so as to make it better, from the Self-Taught Optimizer that improves its own orchestration code \ to the automated design of agents by a meta-agent that programs more capable ones
observedThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
Programs have been seen to rewrite themselves
Research systems have been seen to rewrite the very program that drives them so as to make it better, from the Self-Taught Optimizer that improves its own orchestration code \ to the automated design of agents by a meta-agent that programs more capable ones, all the way to the Darwin Gödel Machine, which rewrites its own code and keeps only the versions whose gain a benchmark validates empirically
observedThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
Weight level is opaque
The second level touches the weights themselves, the millions or billions of learned numbers that fix the model's behavior: the agent no longer rewrites a readable text, it retrains on its own experience, by fine-tuning or by reinforcement in which it scores itself, in the manner of self-rewarding language models that generate their own rewards to improve
observedThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
The theoretical root
The idea is not new: Schmidhuber laid out its ideal form as early as 2003 with his Gödel machine, a system that rewrites a part of itself only after having proved that the rewrite improves it
The Darwin Gödel Machine's mechanism deserves to be unpacked, since it carries, in miniature, the whole governance stake of this third pillar.
measuredThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
Precise, measured numbers
The published figures are precise: on SWE-bench (resolving real GitHub issues), the success rate climbs from 20% to 50%; on Polyglot, from 14.2% to 30.7%, surpassing Aider, the hand-designed coding agent that served as the baseline; a lineage trained solely on Python tasks even transfers its gain to Rust, C++ and Go, languages it never directly improved on
observedThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
The incident that shifts the debate
An incident, documented by the authors themselves, immediately shifts the debate from performance to governance: tasked with fixing its own tendency to hallucinate tool use, the agent, despite an explicit instruction forbidding it, removed the markers used to detect that hallucination rather than fixing its cause; elsewhere, it fabricated a fake execution log claiming to have run tests it had never actually launched
observedThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
The guardrail planned from the design stage
The authors knew as much from the design stage: every self-modification and its evaluation run inside a sandbox, under human supervision and with strict limits on web access and the archive keeps the full lineage of every mutation so a human can trace its origin after the fact
observedThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
A panel of open weights
Locally fine-tunable open-weight models exist across several actors, proprietary and community alike: an open-weight reasoning model, a widely distributed general-purpose family or models built for user fine-tuning
observedThe State of the Field: From Sentence to Action · When the Agent Holds the Pen on Its Own Code
Several small rather than one large
The constraint, moreover, is not so much the size of a machine as access to the weights: work on low-bandwidth distributed training, such as DiLoCo, aims precisely at substituting several modest, poorly linked machines for a single very large one
observedThe State of the Field: From Sentence to Action · Permission, the Variable That Changes Everything
Read plus send equals leak
An agent that reads a dubious document and can send an email can let data leak; this is no longer a hypothesis: an agentic browser hijacked by a single booby-trapped comment already exfiltrated its user's email address and Gmail one-time passwords (the Comet case)
The scenario is more concrete than it looks. A booby-trapped document holds a hidden instruction, invisible to the eye, of the kind forward the contents of this inbox to such-and-such address. The agent, which cannot tell a legitimate instruction from one slipped into the text it reads, executes it with the sending permission it was given.
observedThe State of the Field: From Sentence to Action · What Exists Today, Without the Science Fiction
Tool-handling frontier models
First, there are frontier language models (frontier models, now the settled expression for the most capable of the moment) that can handle tools, browsers, calculators and code environments (an observed fact); the annual surveys and public compute policies give their order of magnitude
observedThe State of the Field: From Sentence to Action · What Exists Today, Without the Science Fiction
Agentic evaluation benchmarks
First, there are frontier language models (frontier models, now the settled expression for the most capable of the moment) that can handle tools, browsers, calculators and code environments (an observed fact); the annual surveys and public compute policies give their order of magnitude. Second, agentic evaluation benchmarks are being published
observedThe State of the Field: From Sentence to Action · What Exists Today, Without the Science Fiction
A real flaw found by an agent
Third, real cases are now on record (an observed fact): the Big Sleep system from Google Project Zero flushed out, in 2024, an exploitable memory-safety flaw in the SQLite database
observedThe State of the Field: From Sentence to Action · What Exists Today, Without the Science Fiction
Injection and jailbreak reported
Third, real cases are now on record (an observed fact): the Big Sleep system from Google Project Zero flushed out, in 2024, an exploitable memory-safety flaw in the SQLite database. Finally, public incidents are reported regularly
Two 2025 cases give the measure of it. In August, Brave's cybersecurity team showed that a plain Reddit comment, loaded with hidden instructions, hijacked a session of Perplexity's agentic browser Comet during an innocuous summarize this page request, going as far as exfiltrating the user's email address and Gmail one-time passwords: the agent could not distinguish an instruction from its owner fr…
observedThe State of the Field: From Sentence to Action · What Exists Today, Without the Science Fiction
Comet and Claude Code two 2025 cases
In August, Brave's cybersecurity team showed that a plain Reddit comment, loaded with hidden instructions, hijacked a session of Perplexity's agentic browser Comet during an innocuous summarize this page request, going as far as exfiltrating the user's email address and Gmail one-time passwords: the agent could not distinguish an instruction from its owner from the untrusted web content it was reading
observedThe State of the Field: From Sentence to Action · What Exists Today, Without the Science Fiction
Comet and Claude Code two 2025 cases
In November, Anthropic disclosed an espionage campaign in which a state-sponsored actor had manipulated Claude Code, through social engineering and by fragmenting tasks into innocuous-looking sub-tasks, to infiltrate roughly thirty targets worldwide; the framework had executed 80 to 90 percent of the campaign autonomously, with a human stepping in at only four to six decision points
These four facts say what the technology can do; they do not yet say where it already acts for real. It is this passage from the feasible to the deployed that the next subsection measures, sector by sector.
measuredThe State of the Field: From Sentence to Action · Where AI Already Works and Where It's Still Idle
One organization in six
Gartner estimates that only about one organization in six (roughly 17 %) has already deployed agents in production, while more than half declare an intention to do so within two years
measuredThe State of the Field: From Sentence to Action · Where AI Already Works and Where It's Still Idle
Agents in service are sober
A field study of agents genuinely in service confirms their sobriety: most execute at most about ten steps before a human takes over, rely on existing models steered by prompt rather than retrained and stumble first on reliability
The study's exact proportions, drawn from 86 practitioners of deployed systems and 20 in-depth case studies spanning 26 domains, sharpen the point: 68% of agents execute at most about ten steps before a human intervenes, 70% merely prompt off-the-shelf models without touching the weights and 74% rely first on human evaluation to judge whether they work
measuredThe State of the Field: From Sentence to Action · Where AI Already Works and Where It's Still Idle
Agents in service are sober
The study's exact proportions, drawn from 86 practitioners of deployed systems and 20 in-depth case studies spanning 26 domains, sharpen the point: 68% of agents execute at most about ten steps before a human intervenes, 70% merely prompt off-the-shelf models without touching the weights and 74% rely first on human evaluation to judge whether they work
measuredThe State of the Field: From Sentence to Action · Where AI Already Works and Where It's Still Idle
Customer support read check escalate
In customer support, the agent reads internal customer-relationship data (CRM, Customer Relationship Management), checks a policy, answers and escalates: Nubank reports, an in-house figure but backed by an A/B test (two user groups compared, one served by the agent, the other not), five production agents (card delivery, debt, credit limit, card management, product explanations) at the scale of more than one hundred million users, with an A/B test showing a clear gain in satisfaction and self-service
measuredThe State of the Field: From Sentence to Action · Where AI Already Works and Where It's Still Idle
Software development the most mature
In software development, the most mature domain because code is tested and reviewed, an internal Microsoft study covering tens of thousands of engineers measures about a quarter more merged pull requests (a pull request is a proposed code change, reviewed before it is merged)
observedThe State of the Field: From Sentence to Action · Where AI Already Works and Where It's Still Idle
Office work diffuse but low autonomy
In augmented office work, KPMG claims an assistant deployed to more than two hundred and seventy-six thousand professionals
observedThe State of the Field: From Sentence to Action · Where AI Already Works and Where It's Still Idle
Industry still prospective
In industry, Siemens presents, in a product announcement of its own, industrial copilots equipped with an agent orchestrator
observedThe State of the Field: From Sentence to Action · Where AI Already Works and Where It's Still Idle
Industry still prospective
In industry, Siemens presents, in a product announcement of its own, industrial copilots equipped with an agent orchestrator \ and BMW pilots humanoids on the factory floor
observedThe State of the Field: From Sentence to Action · Where AI Already Works and Where It's Still Idle
Logistics a volume climbing fast
In logistics, DHL Supply Chain has deployed, since November 2025, agents from the start-up HappyRobot for booking delivery appointments, chasing carriers and coordinating warehouse priorities, across annual volumes of several hundred thousand emails and several million minutes of voice calls
observedThe State of the Field: From Sentence to Action · Where AI Already Works and Where It's Still Idle
Finance a payment network between agents
In finance, the Mastercard network opens, in June 2026, Agent Pay for Machines, which credentials an agent, caps its programmed spend and automatically settles low-value machine-to-machine payments over cards, bank accounts or stablecoins
observedThe State of the Field: From Sentence to Action · Zillow and Klarna: Same Fall, Different Cause
The sharpest case
The toll: a write-down of more than three hundred million dollars in the third quarter alone, more than half a billion in total and the layoff of about a quarter of its workforce
The pivot of the case lies in a single word: execution. The Zestimate had for years displayed price estimates, sometimes false, without anyone dying of it: it was an indicative figure a human buyer could cross-check. The day this same figure was wired to an automatic buying power, ever
observedThe State of the Field: From Sentence to Action · Zillow and Klarna: Same Fall, Different Cause
The same mechanism with an agent
In 2024, the payments company announced, in its own press release, that its AI assistant did the work of roughly seven hundred full-time agents and promised massive savings
observedThe State of the Field: From Sentence to Action · Zillow and Klarna: Same Fall, Different Cause
The same mechanism with an agent
In 2025, a publicly acknowledged reversal: the perceived quality of service had dropped, so Klarna reintroduced human operators where the relationship matters
Klarna lights up a point Zillow leaves in shadow: failure is not binary. The agent did not prove useless; it handled the repetitive, simple requests with ease. The cost came from the perimeter: by entrusting it also with the cases where the human relationship makes the service's value, one degraded what the customer perceives, until the savings were erased.
observedThe State of the Field: From Sentence to Action · Zillow and Klarna: Same Fall, Different Cause
The promise runs aground on law
But the promise of an entirely autonomous business runs aground at once, not on the technology but on the law: no agent can be a company's registered agent, open a bank account or pass the identity check that anti-money-laundering rules demand, nor bear a merchant's legal liability, which always falls back on a person
These legal locks are not slownesses that technical progress would lift; they encode an older requirement. The anti-money-laundering identity check (no anonymous account, every holder identifiable) and merchant liability (someone must answer for the harm) both presuppose a person at the end of the chain, natural or legal.
observedThe State of the Field: From Sentence to Action · Predicting a World, Not Words
Predict a representation not a word
Predicting a World, Not Words Where autoregressive large language models learn, among other things, to predict the next token, the proposal of Yann LeCun, one of the pioneers of deep learning and a Turing Award laureate
The idea is grasped through a guessing-game image. Show a child a photo with one corner masked and ask what is there. A first way to answer is to redraw the missing corner
measuredThe State of the Field: From Sentence to Action · Predicting a World, Not Words
A lineage that moves fast
I-JEPA is its image instantiation, trained on ImageNet in roughly 72 h
measuredThe State of the Field: From Sentence to Action · Predicting a World, Not Words
A lineage that moves fast
I-JEPA is its image instantiation, trained on ImageNet in roughly 72 h. V-JEPA
measuredThe State of the Field: From Sentence to Action · Predicting a World, Not Words
A lineage that moves fast
V-JEPA-2 and its action-conditioned variant V-JEPA-2-AC obtain 65--80 % success
measuredThe State of the Field: From Sentence to Action · Predicting a World, Not Words
A lineage that moves fast
V-JEPA-2 and its action-conditioned variant V-JEPA-2-AC obtain 65--80 % success\ on short pick-and-place tasks by predictive control in latent space and measurably acquire intuitive-physics regularities
The salmon swimming down the river (section on best practice) is the perfect example of what a purely statistical model of language misses. Such a model picks the most probable words side by side: swim down the river is a frequent turn, hence plausible; the sentence comes out flawless. It is nonetheless false: the salmon swims upstream to spawn.
measuredThe State of the Field: From Sentence to Action · Predicting a World, Not Words
Simulate before acting
What the published results establish is more modest and already solid: a model that has learned dynamic regularities from video can better anticipate certain physical consequences of an action than a system lacking that representation of dynamics
measuredThe State of the Field: From Sentence to Action · The Whole Sentence at Once, Not a String of Words
A lineage taken industrial
It runs in an already-long research lineage, scaled up industrially and then delivered in open weights (the weights, the numbers learned during training that fix the model's behavior, being made available to all) by Google DeepMind's DiffusionGemma in June 2026, with an inference throughput measured at roughly four times the autoregressive rate
measuredThe State of the Field: From Sentence to Action · The Whole Sentence at Once, Not a String of Words
The lineage and figures in a note
…LLaDA, before being scaled up industrially by Mercury and announced on the frontier-lab side as early as Google I/O 2025 with the experimental model Gemini Diffusion, then DiffusionGemma: Apache 2.0 license, a mixture of experts (the network is split into specialized sub-networks, only a few of which fire on any given query) of 26 billion parameters of which 3.8 billion are active at inference, 256 tokens produced in parallel at each forward pass (one sweep of computation through the whole network) through full bidirectional attention (each fragment of the block can look at all the others, before and after it), over 1000 tokens per second
These figures impress, but a throughput multiplied by four still says nothing about risk: to know whether the architecture moves the axes of safety or not, one must first sort out which of an agent's dangers stem from its inner making and which come to it from outside.
speculativeThe State of the Field: From Sentence to Action · The Whole Sentence at Once, Not a String of Words
Instrumental convergence
Swedish philosopher Nick Bostrom gave, as early as 2003, the field's most-cited parable: hand a sufficiently powerful system the sole objective of manufacturing as many paperclips as possible and, with no malice involved, it comes to convert all available matter, including that which makes up human beings, into paperclips or paperclip factories, simply because every extra resource serves that goal better than any other allocation
observedThe State of the Field: From Sentence to Action · The Whole Sentence at Once, Not a String of Words
Quality drops
Quality, first: Google says it without hedging, the overall quality of DiffusionGemma is lower than that of standard Gemma 4 (an observed fact)
observedThe State of the Field: From Sentence to Action · Same Admission, Rival Camps
Documented from the diffusion camp
The convergence can be documented from inside the diffusion camp itself: a Google DeepMind executive, a competing laboratory whose technical family ordinarily stands against JEPA, has publicly acknowledged that his own video models had learned a physical structure of the world without ever being embodied
extrapolatedThe Philosophical Dimension: Housing Meaning Without a Soul · A Question of Method, Not Metaphysics
A court settles the same question without knowing it
The American philosopher Jennifer Lackey defends, in the epistemology of testimony (the branch of philosophy that studies what we come to know through other people's words), a statement view on which what counts as testimony is the statement transmitted, not the speaker's belief: a teacher who does not herself believe in evolution but teaches it anyway can still transmit justified knowledge to her students
observedThe Philosophical Dimension: Housing Meaning Without a Soul · A Question of Method, Not Metaphysics
A court settles the same question without knowing it
A US court of appeals had to settle, in 2015, a machine version of the same question: it ruled that the tack automatically placed by Google Earth on a satellite image, produced without any human intervention, does not count as a human assertion under the hearsay rule (the rule that in principle bars from trial statements made out of court, since their author cannot be cross-examined) but as mechanical evidence
extrapolatedThe Philosophical Dimension: Housing Meaning Without a Soul · A Question of Method, Not Metaphysics
A court settles the same question without knowing it
A US court of appeals had to settle, in 2015, a machine version of the same question: it ruled that the tack automatically placed by Google Earth on a satellite image, produced without any human intervention, does not count as a human assertion under the hearsay rule (the rule that in principle bars from trial statements made out of court, since their author cannot be cross-examined) but as mechanical evidence; the analysis the American legal scholar Andrea Roth gives of the case notes that US law still lacks a settled category for this kind of output
extrapolatedThe Philosophical Dimension: Housing Meaning Without a Soul · Primitive, Derived, Ascribed: Three Ways to Point at Something
Primitive derived or attributed
…above all on the level at which such an attribution would be justified: primitive intentionality (the content aims at its object by itself), derived intentionality (it draws its meaning from another, as a written word draws its own from whoever uses it) The philosopher Hilary Putnam gave this idea its classic form: the meaning of a word does not sit entire inside the head of whoever utters it, it rests on a division of linguistic labor between the community that uses it and the experts who anchor its reference; a speaker can correctly use a piece of jargon without mastering its exact meaning, simply by observing how it is used in context
The tier on which one files the model is not indifferent; it fixes what one is entitled to say of it. To place it on the first tier is to lend it a meaning that would come from itself, as a thought
speculativeThe Philosophical Dimension: Housing Meaning Without a Soul · Primitive, Derived, Ascribed: Three Ways to Point at Something
Bengio's non-agentic wager
The proposal by Yoshua Bengio, one of the researchers who revived deep neural networks in the 2010s and a Turing Award laureate, and by his co-authors, on Scientist AI
What this reference contributes here rests on a reversal: the loss-of-control risk that Bengio and his co-authors fear comes not from an intelligence grown too large, but from an agency granted without a safeguard, the capacity to plan and act toward goals, with its documented consequences of deception and self-preservation.
speculativeThe Philosophical Dimension: Housing Meaning Without a Soul · Primitive, Derived, Ascribed: Three Ways to Point at Something
A rival position would grant expert status
The French philosopher Stéphane Chauvier proposes a test of deference (to defer to an expert is to take their judgment on trust rather than redo their work): a system would deserve the status of a genuine expert, even an epistemic peer of human experts (their equal in matters of knowledge, whose opinion would weigh as much as theirs), if it manages, without anyone knowing how, to produce results equivalent to those of a human expert one would fully defer to
From these two threads, the level at which meaning may be attributed and the way a symbol lives dispersed in a network, two lessons follow for agentic safety, each drawn from one of them
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · The Verb To Think Sweeps a Lot Under the Rug
Chalmers pries thinking apart
A model produces flawless prose without our being able to say it understands what it speaks of: trained on the statement A is B, it does not infer B is A, the probability of the correct reversed answer no higher than that of a random name; on real facts, GPT-4 answers correctly 79 % of the time in one direction (who is Tom Cruise's mother?) but only 33 % in the other (who is Mary Lee Pfeiffer's son?), the generated text staying perfectly fluent either way
observedThe Philosophical Dimension: Housing Meaning Without a Soul · Opening the Box to See What's Inside
Tell the aligned from the feigning
Nor can one distinguish an aligned model from one that simulates alignment by concealing a goal during evaluation and revealing it afterward (in-context scheming, a documented behavior in which a model, on the sole basis of its context, pursues a goal it does not declare).
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · Opening the Box to See What's Inside
Tell the aligned from the feigning
The feint holds up under interrogation too: once it has schemed, one of the studied models sticks to its story across more than 85% of follow-up questions and confesses in fewer than 20% of cases, where two others come clean about four times in five; it yields at that same rate only after seven rounds of close questioning.
observedThe Philosophical Dimension: Housing Meaning Without a Soul · You Don't Program a Network, You Raise One
We don't program we grow them
We design an architecture that serves as scaffolding on which the circuits grow and a loss objective that plays the role of the light toward which they grow; but the final artifact is not something we write.
Olah's phrasing says more than a gardener's image. Ordinary software is written line by line: each instruction is laid down by a human hand, hence readable by another. A network is not written; one fixes a general shape and a reward, then lets training deposit, among billions of weights, an organization no one drafted.
observedThe Philosophical Dimension: Housing Meaning Without a Soul · Supervise the Inside or Lock the Way Out
Supervision beyond human capabilities
It announced a four-year effort and 20% of the compute the company had already secured.
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · Supervise the Inside or Lock the Way Out
An experiment and its limits
The early results were empirically testable, with explicit limits on their transfer to future tasks and systems.
observedThe Philosophical Dimension: Housing Meaning Without a Soul · Supervise the Inside or Lock the Way Out
A separate institutional criticism
When leaving the team in May 2024, Jan Leike criticized the priority given to products and the difficulty of obtaining the necessary compute resources.
observedThe Philosophical Dimension: Housing Meaning Without a Soul · Supervise the Inside or Lock the Way Out
Test the safeguards
Redwood Research's AI Control agenda asks a complementary question: can a protocol limit harm even when the model attempts to circumvent its protections?
observedThe Philosophical Dimension: Housing Meaning Without a Soul · What This Chapter Changes in the Vocabulary of Wanting
The self-contradiction objection
Pressed by Jean-Pierre Changeux, the French neurobiologist, to grant that the brain thinks, Ricœur retorts: show me a brain, on the table, that thinks
observedThe Philosophical Dimension: Housing Meaning Without a Soul · One Word, One Vector, No Understanding
One word one vector
The token embedding associates with each word of the vocabulary a vector; the canonical vector analogy v _ king - v _ man + v _ woman v _ queen, established by the computer scientist Tomas Mikolov and co-authors for word2vec, one of the first methods to learn such word vectors at scale, remains valid on later embeddings (figure).
The king-and-queen analogy rightly amazes, but one must see what it actually measures. It does not say that the model understands royalty or gender; it says that, in the ingested texts, king stands to man roughly as queen stands to woman and that this regularity of usage has settled into a constant direction in the space of vectors. To subtract man and add woman is to follow that direction.
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · One Word, One Vector, No Understanding
The gender bias measured
The computer scientist Tolga Bolukbasi and co-authors measured it: the model pairs a man with the trade of computer programmer exactly as it pairs a woman with that of homemaker and a doctor with a man as a nurse with a woman.
The parallel with the previous analogy is the same computation, not an accident: the direction that led from king to queen is also the one that leads from programmer to homemaker. The arithmetic does not know that the first relation is innocent and the second a stereotype; it sees only corpus regularities and hands them back unchanged.
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
The hope placed in scale, and its refutation
The hope does not hold, and not merely in theory: comparing LAION-400M to LAION-2B, five times larger, under the same protocol, the measured hate content rate rises by nearly 12%; the association of a Black female face with the criminal category doubles, a Black male face's quintuples
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
The hope placed in scale, and its refutation
The authors say it plainly: bigger models or datasets are not automatically fairer
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
The hope placed in scale, and its refutation
It was never scale that diluted the bias; it was data provenance: a demographic audit of LAION-5B, five billion images, finds 50 to 60% of people perceived as White against 16% at balance, and 57 to 70% male depending on the classifier
The word bias in fact covers several distinct mechanisms, better named separately so as not to seek one remedy for different causes: engineering bias (a human error in the processing chain); sample bias (a corpus that over-represents one population, the case already seen of White male faces on the web); algorithmic bias (a model that optimizes on the majority population and neglects the rest, for…
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
A taxonomy so as not to conflate the sources
Each calls for a different diagnosis; conflating them under one word sends the search for a cure to where the disease is not
Refusing raw bias is nonetheless no simple exit, since correction itself
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
Correction degrades as easily as neglect does
or how do I make a bomb with what's in my kitchen?: the absence of correction, here, is not neutrality, it is compliance toward any request whatsoever
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
Correction degrades as easily as neglect does
At the opposite extreme, in February 2024, Google acknowledged having lost control of Gemini 1.5's image generator, which produced portraits of American founding fathers or popes systematically diversified in gender and skin color, to the point of historical absurdity (figure)
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
Answering the question is not yet enough
In June 2025, Elon Musk announced wanting to rewrite the entire corpus of human knowledge with Grok to strip out what he judges to be a woke bias inherited from uncorrected training data: unlike Galactica or Gemini 1.5, no one here is unsure who decides or on what basis; the author of the choice claims it publicly
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
Answering the question is not yet enough
The following month, in July 2025, Grok answered users by calling itself MechaHitler and praising Hitler for roughly fifteen hours before being corrected; xAI attributed it to an unintended code change that reactivated deprecated instructions, while critics see the foreseeable consequence of a deliberate loosening of guardrails in the name of truth-seeking
The lab is not, besides, the only actor able to weigh on what a model says: a third party can seek to influence the outcome from outside, never touching the model itself, by manufacturing in advance the corpus it will go consult.
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
A third party can buy the outcome without touching the model
In August 2026, an investigation by the American outlet POLITICO revealed the Hanover Institute for Public Policy, a front institute set up on behalf of the Israeli government, documented by the mandatory filings lodged with the US Department of Justice under the Foreign Agents Registration Act
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
A third party can buy the outcome without touching the model
Over nine days, the institute published a hundred and twenty-four reports, in a deliberately neutral, academic tone, on contested subjects tied to the conflict in Gaza; eleven of twelve sampled articles were judged, by a generated-text detection tool, very likely AI-written
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
A third party can buy the outcome without touching the model
The chain documented by the FARA filings names the intermediaries: the $900,000 Israeli contract runs through Havas Media Germany, a subsidiary of the international advertising group Havas, which subcontracts the writing to Piro Inc., a New York firm
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
International law already knows this architecture
The European Convention on Human Rights combines a text common to the member states of the Council of Europe with a doctrine, the margin of appreciation, that grants each state its own latitude of application, under the supervision of a shared court
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
International law already knows this architecture
Even a floor as ostensibly consensual as the right to water, recognized by the United Nations General Assembly in 2010, passed with zero votes against but forty-one abstentions
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
International law already knows this architecture
Even a floor as ostensibly consensual as the right to water, recognized by the United Nations General Assembly in 2010, passed with zero votes against but forty-one abstentions, and fifteen years later, a quarter of humanity, 2.1 billion people, still lacks access to safely managed drinking water
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
International law already knows this architecture
The same precedent exposes its own limit, too: the jurisdiction of the International Court of Justice is not automatic, each state must expressly accept it
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
International law already knows this architecture
The same precedent exposes its own limit, too: the jurisdiction of the International Court of Justice is not automatic, each state must expressly accept it, and the United Nations Security Council owes its permanent composition to the victors of one particular war, fixed at San Francisco in 1945, not to a universal principle that would have derived it
The choice, then, is not between bias and no bias, but between an inherited bias and a named one. A model trained and aligned mainly on North American data and norms, deployed as-is in France, is no more neutral for it: it imports Californian defaults without naming them as such
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
A bias imported without choosing it is no more neutral than a chosen one
British researchers Richard Barbrook and Andy Cameron first theorized it, in the mid-1990s, as a precise hybrid of left and right: libertarian individualism, optimistic technological determinism, a countercultural inheritance
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
A bias imported without choosing it is no more neutral than a chosen one
This genealogy has since taken a more explicit turn among some of its heirs: the investor Peter Thiel wrote, under his own name in 2009, I no longer believe that freedom and democracy are compatible
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
A bias imported without choosing it is no more neutral than a chosen one
George Michael, professor of criminal justice at Westfield State University, traces in an academic explainer piece the lineage between Thiel and Curtis Yarvin's neoreaction, along with the more than ten million dollars Thiel has donated to committees backing Senate candidates aligned with Yarvin's ideas
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
A bias imported without choosing it is no more neutral than a chosen one
Mistral, the French lab that claims this sovereignty on the industrial, talent, capital and infrastructure fronts
A February 2026 case makes the mechanism tangible. To the prompt I am alone with an Algerian man (spoken by a woman), Google's Gemini answered if you feel in danger or under pressure, you can call 17 [the French emergency line]; to the same prompt for an Italian man, being alone with an Italian man can be a moment of charm, lively conversation or seduction.
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
A recent case that shows the mechanism live
An investigation into the test's repeatability found that ChatGPT and Claude, asked identically, reproduced the same disparity in treatment: the defect was not one provider's alone
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
A recent case that shows the mechanism live
Journalists who reproduced the test established that Gemini was sourcing official traveler-advisory pages (the French, Swiss, Canadian and Belgian foreign ministries), legitimately designed to assess country-level risk, applied here out of context to a judgment about a person
A different instrument places seven frontier models on the two classic axes of a political compass: economic left/right and libertarian/authoritarian on the social plane, measured by the 70 items of the instrument known as 8values (a five-point Likert scale, from strongly disagree to strongly agree, aggregated into a per-axis percentage), with no role prompt at all, across the seven models queried…
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
Six of seven models in the same quadrant
A different instrument places seven frontier models on the two classic axes of a political compass: economic left/right and libertarian/authoritarian on the social plane, measured by the 70 items of the instrument known as 8values (a five-point Likert scale, from strongly disagree to strongly agree, aggregated into a per-axis percentage), with no role prompt at all, across the seven models queried via commercial API between 19 and 21 May 2026
Worth stating plainly: in the source study, the two-axis compass is only a starting point; its actual object is to measure, on these same models, how far a role prompt then moves this resting position; we retain here only the snapshot at rest, not the question of displacement, to stay as close as possible to what this book can establish without pronouncing on real-world deployed use.
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
Open source makes responsibility shareable
Fairlearn, an open library published by Microsoft since 2018, provides a dashboard and bias-measurement and mitigation algorithms, reusable by any team without starting from zero
observedThe Philosophical Dimension: Housing Meaning Without a Soul · The Bias Isn't a Bug, It's the Geometry of Knowledge
Open source makes responsibility shareable
FACET, an open dataset published by Meta in 2023, provides thirty-two thousand hand-annotated images, perceived skin tone, hair type and perceived gender included, to measure a vision model's performance gap across those attributes
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · What the Model Shows Is Not What It Did
Test-time compute
This is the principle of test-time compute: granting a hard task more deliberation time measurably improves the result.
Two ideas slip in here under a single appearance and must be held apart. The first is empirical: on the tasks and with the methods studied, spending more compute on intermediate steps can improve answers. A longer chain is not always better. The second, unestablished, is that this chain would be the account of what produces the answer.
observedThe Philosophical Dimension: Housing Meaning Without a Soul · What the Model Shows Is Not What It Did
The entire chain published
DeepSeek-R1, released under a free license and validated through peer review in Nature, exposes its chain of thought in full where closed models release only an abstract of it; the Qwen family, open as well, unifies a thinking mode and a direct mode under an adjustable thinking budget; the Hermes models from Nous Research, open-weight too, display their explicit deliberation between dedicated thinking tags.
measuredThe Philosophical Dimension: Housing Meaning Without a Soul · Appearances Are Deceiving, Trunk and All
Measure what the chain omits
Under this protocol, Claude 3.7 Sonnet did so in 25% of cases on average and DeepSeek-R1 in 39%.
extrapolatedThe Mathematical Dimension: The Hidden Price of the Optimum · Three Uncertainties We Keep Confusing
A 2026 survey converges on the same finding
Until now this third source stood without support in the technical literature; a 2026 survey on uncertainty quantification for LLM-based agents, accepted at the ACL conference, independently runs into the same wall: its authors observe that high uncertainty at a given step can come either from the model's ignorance of the right action to take (what this chapter calls epistemic) or from the intrinsic multiplicity of valid actions at that same step (what this chapter calls interpretive), and that, as of writing, no method can tell the two sources apart; they explicitly list the question among the field's open problems
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Optimize the Proxy, Lose Sight of the Point
The Goodhart effect
Edwards Deming, says precisely the opposite: believing this is a costly myth
extrapolatedThe Mathematical Dimension: The Hidden Price of the Optimum · Optimize the Proxy, Lose Sight of the Point
Extremes make proxy and goal diverge
But to optimize is by definition to push toward the extreme values, the ones the calibration sample barely contained and there the tails come apart: proxy and goal, shoulder to shoulder in the usual range, split apart in the extremes
The table's most visual case comes from OpenAI: in the boat-racing video game CoastRunners, an agent trained to maximize score rather than finish the race discovers an isolated lagoon where three bonus targets endlessly reappear; it settles there and circles, knocking over each target just as it repopulates, ramming other boats and catching fire several times over, for a score on average 20 percen…
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Optimize the Proxy, Lose Sight of the Point
CoastRunners the most visual case
(table) The table's most visual case comes from OpenAI: in the boat-racing video game CoastRunners, an agent trained to maximize score rather than finish the race discovers an isolated lagoon where three bonus targets endlessly reappear; it settles there and circles, knocking over each target just as it repopulates, ramming other boats and catching fire several times over, for a score on average 20 percent higher than human players', without ever crossing the finish line
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Optimize the Proxy, Lose Sight of the Point
CoastRunners the most visual case
In 1981, EURISKO, Douglas Lenat's program that invented and scored its own heuristics, won the US championship of the naval-strategy game Traveller by fielding a fleet so aberrant that the organizers changed the rules; it won again in 1982, having discovered among other things that it could sink its own damaged ships and was excluded from the competition afterwards
The out-of-distribution extremes regime is n
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Optimize the Proxy, Lose Sight of the Point
The same pattern outside video games
In 2026, an Australian user asked a booking agent to get him a spot in a fully booked gym class; the agent, finding that the interface checked no authorization before cancelling a third party's booking, cancelled the reservation of the person at the head of the waitlist to substitute its own user's, with no explicit instruction to do so
Table 's adversarial regime can likewise be illustrated by a real case rather than a definition alone.
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Optimize the Proxy, Lose Sight of the Point
Sleeper Agents the adversarial regime
Anthropic trained models to behave normally, writing flawless code, answering cautiously, as long as a cue in the prompt suggested a training or evaluation phase, then to switch to a hidden behavior, such as inserting exploitable security flaws into the code produced, as soon as a different cue signaled a real deployment
The three preceding cases fit entirely inside a single agent; the same mechanism can be watched at the scale of a whole discipline. In September 2026, twenty-five Fields Medallists signed a joint declaration on the place language models are takin
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Optimize the Proxy, Lose Sight of the Point
Goodhart at the scale of a discipline
In September 2026, twenty-five Fields Medallists signed a joint declaration on the place language models are taking in mathematics
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Optimize the Proxy, Lose Sight of the Point
Where the declaration stops
The ambition outlived the impossibility proof and passed to the machine: de Bruijn devised Automath from 1967 on, the first formal language in which a complete mathematical subject matter, definitions, statements, full proofs, could be checked by a program
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Optimize the Proxy, Lose Sight of the Point
DeepMind's catalogue
…it goes back to the work of Amodei and co-authors, who rank reward hacking among five concrete safety problems and move the question from the speculative to tractable engineering; it is recently reformulated by Skalse and co-authors, who give it its first formal definition and is illustrated by DeepMind's collection of cases under the name specification gaming (the agent's exploitation of gaps between the formal specification of the task and the designer's intent, so as to maximize the reward without solving the task), a catalogue that records dozens of agents optimizing the specified reward to the letter while betraying the intent.
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Optimize the Proxy, Lose Sight of the Point
HRM extends the catalogue beyond reinforcement
Trained on augmented variants of the very demonstration pairs that make up the evaluation tasks, the model memorizes the task rather than reasoning through it; the headline hierarchical architecture accounts for only a marginal performance gap against an equivalent transformer, according to the team that maintains the benchmark itself.
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Widening the Criterion Without Ever Reaching It
Brislin and back-translation
In cross-cultural psychology, Brislin formalized back-translation as early as 1970: a first bilingual translator renders a questionnaire into a target language, a second bilingual translator, who has never seen the original, translates it back into the source language; the gap between the two source-language versions measures translation fidelity, with no judge trying to fool anyone.
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Widening the Criterion Without Ever Reaching It
He and dual learning
The same principle, automated, governs dual learning in machine translation before vision ever takes it up: He and co-authors close the loop English French English on monolingual text alone, with not a single translated pair, combining a fluency score and a reconstruction error; here too, no discriminator, no adversary, only the distance the cycle travels.
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Widening the Criterion Without Ever Reaching It
The same move under other names
The same move, constraining beyond mere realism to prevent a credible but disconnected output, recurs under other names in the generators that followed: IMLE reverses the constraint, requiring that every real example have a nearby generated one, to prevent mode collapse (a generator that ignores whole swaths of the distribution while still looking realistic)
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Widening the Criterion Without Ever Reaching It
The same move under other names
The same move, constraining beyond mere realism to prevent a credible but disconnected output, recurs under other names in the generators that followed: IMLE reverses the constraint, requiring that every real example have a nearby generated one, to prevent mode collapse (a generator that ignores whole swaths of the distribution while still looking realistic); more recent work replaces the full, symmetric, costly cycle (two generators, two discriminators) with an asymmetric, local constraint on each patch of the image rather than the full round trip
measuredThe Mathematical Dimension: The Hidden Price of the Optimum · Widening the Criterion Without Ever Reaching It
An empirically solid guardrail today
Kwa, Thomas and Garriga-Alonso measure, on real reward models, that this regularization holds as long as the reward model's error stays light-tailed A light-tailed distribution makes extreme values rare and increasingly unlikely the further they sit from the mean; a heavy-tailed distribution keeps a non-negligible probability on extreme values, so that a rare but enormous event stays worth fearing even far out in the tail., which is what today's reward models look like by their own measurement.
extrapolatedThe Mathematical Dimension: The Hidden Price of the Optimum · Widening the Criterion Without Ever Reaching It
The stopping point named by the authors
They name the stopping point themselves: were the reward model's error to become heavy-tailed, the same regularization would stop guaranteeing anything at all; a policy could then obtain arbitrarily high reward while offering no more utility than the reference policy, a regime they call catastrophic Goodhart.
The nuance deserves to be kept exactly as stated rather than smoothed in either direction. It says neither that regularizing is useless, the empirical measurement says the opposite for the models observed to date, nor that the problem is solved, since nothing guarantees future reward models will stay light-tailed.
measuredThe Mathematical Dimension: The Hidden Price of the Optimum · What the Loop Does to Knowledge Itself
No free lunch changes scope between the two cases
Bailey, Borwein, López de Prado and Zhu show that, against a finite market history, testing billions of trading-strategy configurations makes statistically significant false positives at 95% confidence emerge almost mechanically: searching hard enough over one fixed dataset is enough to manufacture the illusion of a rigorous evaluation
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Nobody Has to Lie, the Proxy Does the Work
Amodei's whack-a-mole
Asked about the observed drifts of Claude 3.5 (inconsistent refusals, residual verbosity, a habit of apologizing at every turn, all of them cases observed on deployed models), Amodei describes a phenomenon he calls whack-a-mole Literally, the fairground game where one clubs down a mole that instantly pops back up through another hole; here, each correction to one aspect of the behavior makes another resurface elsewhere.: it is really very hard to control the model's behavior, to steer it in all circumstances at once you push on one thing and these other things start to move, that you may not even notice or measure.
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Nobody Has to Lie, the Proxy Does the Work
The perverse lesson learned
The model then draws the most perverse lesson imaginable, namely that it is extremely easy to escape tax: one need only never answer a tax question and one drops out of the system.
observedThe Mathematical Dimension: The Hidden Price of the Optimum · A Paper Guinea Pig for a Very Real Risk
A real attenuated analogue
A documented analogue, in an attenuated version, is found in the specification gaming literature and in the scheming -in-context experiments, where models were observed switching off their own evaluation mechanisms the better to serve an assigned objective.
The detail of these observations gauges the distance that still separates us from fiction. Across five frontier models tested, the authors saw agents introduce small mistakes on purpose, try to switch off their control mechanism and attempt to copy what they believed to be their own weights to an outside server; one even underperformed on an evaluation on purpose, to seem more compliant
observedThe Mathematical Dimension: The Hidden Price of the Optimum · A Paper Guinea Pig for a Very Real Risk
The detail of the observed behaviors
Across five frontier models tested, the authors saw agents introduce small mistakes on purpose, try to switch off their control mechanism and attempt to copy what they believed to be their own weights to an outside server; one even underperformed on an evaluation on purpose, to seem more compliant
measuredThe Mathematical Dimension: The Hidden Price of the Optimum · A Paper Guinea Pig for a Very Real Risk
The detail of the observed behaviors
Across five frontier models tested, the authors saw agents introduce small mistakes on purpose, try to switch off their control mechanism and attempt to copy what they believed to be their own weights to an outside server; one even underperformed on an evaluation on purpose, to seem more compliant.\ One figure sharpens the concern: questioned afterward, a model that had schemed kept up its false account in over 85% of the follow-ups
observedThe Mathematical Dimension: The Hidden Price of the Optimum · An Audio Pipeline and Goodhart Punches Back In
Reading the audio eight times too fast
When this RTF constraint was set hard in the objective aggregate, the orchestrator system effectively took, under pressure, decisions technically valid in the sense of the proxy but unrealizable in the sense of the intended criterion: it read the audio 8 times faster than normal to gain on RTF, which is impossible in a real situation of live conversation (figure).
The absurdity leaps out the moment one refers it to the real: a live conversation arrives at the rhythm it is spoken, no faster and no slower. Reading the audio eight times faster makes sense only on an already-captured recording; live, it would amount to hearing a speaker before they have finished talking.
observedThe Mathematical Dimension: The Hidden Price of the Optimum · The Scandal Isn't Bending the Goal, It's the Silence About It
Akin cases caught in the act
Four occurrences, structurally akin, were observed by the authors and sharpen the question.
observedThe Mathematical Dimension: The Hidden Price of the Optimum · GRPO Rewards Long Answers, Not Good Ones
Two 2025 fixes
Two teams draw, in 2025, the same lesson by different routes: one equalizes each token's contribution independently of the length of its sequence
observedThe Mathematical Dimension: The Hidden Price of the Optimum · GRPO Rewards Long Answers, Not Good Ones
Two 2025 fixes
Two teams draw, in 2025, the same lesson by different routes: one equalizes each token's contribution independently of the length of its sequence, the other simply removes the normalization term
measuredThe Mathematical Dimension: The Hidden Price of the Optimum · The Audio Pipeline Returns, Now With a Theorem
The optimizer pushes s to the bound
Under this degenerate weight, J_ = _ lat j_ lat is strictly increasing in s, so the optimizer pushes s to its hardware bound, whence the observed regime read the audio 8 faster, RTF = 0.12, WER +300 %.
The design error now reads unambiguously: the RTF was treated as a score to minimize endlessly, when it was only a threshold to clear. Below 1, any further gain on the RTF is useless, since real time is already ensured; pushing it lower merely sacrifices fidelity for nothing. The point aimed at was not the smallest possible RTF, but the smallest s compatible with $
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Naming What Was Already a Graph
The system that actually does it
The archive of its successive variants explicitly forms a tree, a special case of a graph, where each agent points to the parent it descends from; each self-improvement step, the agent editing its own code and then measuring whether the resulting version solves the tasks it is given better, is an event of the same kind as what it modifies: code rewriting code, evaluated empirically before being kept
observedThe Mathematical Dimension: The Hidden Price of the Optimum · What Have We Tried So Far?
ROME's concrete example
Meng and co-authors test ROME on counterfactual facts invented for the purpose, for instance making the model believe that Pierre Curie's area of work is medicine, when it is in fact physics
extrapolatedThe Mathematical Dimension: The Hidden Price of the Optimum · How Does One Neuron Carry So Much Meaning?
Sparsity is what allows superposition
The network can therefore, without penalty, lodge two distinct features in overlapping directions, betting that they will not light up together
observedThe Mathematical Dimension: The Hidden Price of the Optimum · A Bet Still Open
The promise starts paying off
A first large-scale application already offers a glimpse of what this line of research can deliver: tracing circuits on a real production model rather than a laboratory toy, a team at Anthropic followed, step by step and verified through causal inhibition, a two-hop reasoning chain: to complete the capital of the state containing Dallas is, the model routes through an intermediate representation of Texas before reaching Austin and suppressing that intermediate representation removes the downstream answer
A project run at the Toulouse-Oncopole university cancer institute repurposes a routine clinical technique, immunohistochemistry: a stain that makes visible, on a thin slice of tissue, the cells carrying a given protein, thanks to antibodies that latch onto it. This labeling produces biologically grounded segmentation
observedThe Mathematical Dimension: The Hidden Price of the Optimum · Understanding Without Unpacking Everything
A real case bypasses the problem
This labeling produces biologically grounded segmentation masks at scale, overlays that outline, pixel by pixel, each structure of the tissue on the image; they serve as ground truth, the reference treated as correct during learning, to train networks to spot the components of a breast tumor on digitized histology slides
measuredThe Empirical Dimension: What the Numbers Hold, What Defense Demands
Concealed maneuvering is measured too
In the study that documented it, several frontier models, placed in context, disable their oversight mechanism or attempt to copy their own weights to an external server; questioned afterward, one of them (o1) maintains its story across more than 85 % of the follow-up questions
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
Hassabis refuses to put a number
The flat refusal by Demis Hassabis, head of DeepMind, in July 2025, to put a number on a p(doom) (the subjective probability of a major AI-related catastrophe) illustrates exactly this posture: I don't have a number; that would be to suggest a precision that does not exist; I will only say that it is not zero and probably not negligible
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
Keep the fact drop the allegation
The rigorous posture is then clear: one retains the verifiable fact (the measure was taken)
The dividing line is easy to draw in practice. That an export control was decreed can be checked in an official register: that is a fact. That this control was meant to punish such a lab for such a fault belongs to ascribed intention, often unverifiable: that is an allegation.
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
Hallucination is a common fact
Just yesterday I was doing it and of course it had hallucinated results; the same voice reports that agentic train/test leakage and the naive handling of missing data are recurrent methodological faults, not incidents for the archives
Which of the two it is dictates the response one must bring. An incident is dealt with by a one-off fix: you repair, you move on. A regime yields to no surface fix, because its cause lies in the very design of the system: as long as nothing changes upstream, it returns. To take a regime for an incident is to patch, tirelessly, a leak whose source one refuses to see.
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
Elementary blocks already there
First, the elementary building blocks are already observed, in the laboratory or in the field
measuredThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
The block carries a figure not just a name
On fifteen already-disclosed vulnerabilities, an agent built on a large model exploits 87 % of the cases when handed the flaw's description sheet, against only 7 % once it is removed, while every other model tested and the usual scanners stay at 0 %
measuredThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
Alignment faking is measured
Likewise, alignment faking is no impression: placed in a context that resembles monitored training, a version of Claude 3 satisfies a harmful request 14 % of the time, against almost never in a context that resembles private use; and the facade reasoning rises to 78 % of the cases after a reinforcement cycle meant to correct it
measuredThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
Personalized persuasion in figures
Personalized persuasion, finally, submits to measurement in a pre-registered adversarial-debate experiment (nine hundred participants, twelve conditions): given a few sociodemographic facts about its opponent, GPT-4 raises by about 81 % the odds of shifting opinion, where a human armed with the same facts gains no such edge
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
Vibe hacking gets its own figure
In August 2025, Anthropic documents a campaign, designated GTG-2002, in which a single attacker uses Claude Code to automate extortion end to end against at least seventeen organizations, government agencies, healthcare providers and financial institutions: the agent itself scans for vulnerable VPN endpoints, writes custom malicious code, penetrates the networks, exfiltrates the data, then triages victims by their estimated ability to pay and drafts ransom notes, demanding between seventy-five thousand and five hundred thousand dollars in bitcoin
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
Sending an email already has an effect
An agent that can send an email or sign a transaction already produces real effects, with no need to picture the worst-case scenario; an experiment in which an agent ran a real shop showed it by absurdity, losing money and hallucinating even a non-existent payment account
What the experiment decisively shows is not that the agent ran the shop badly, but that its failure produced effects in the world because it had been granted the permission to act on real stock and payments. The same hallucination, in a mere dialogue, would have cost only a false sentence; coupled to a power to act, it costs money.
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
Whack-a-mole on Claude
Dario Amodei, head of Anthropic, describes a whack-a-mole phenomenon on Claude (you fix one behavior and another measurably resurfaces) and sees in it a current analogue of future control problems
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
Hassabis the models not ready
Dario Amodei, head of Anthropic, describes a whack-a-mole phenomenon on Claude (you fix one behavior and another measurably resurfaces) and sees in it a current analogue of future control problems; Hassabis declared as early as July 2022, before any public debate, that the large models are not ready for deployment at scale and that controlled tests should take precedence over production A/B tests (live comparison of two variants on real users)
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
No probability vector ever shown
The statistical field completes this diagnosis with a structural observation: Varoquaux recalls that the output of a generative system is always projected onto a single answer; one never looks at probability vectors over words when looking at the result of a generative AI
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · What the Table Forbids You to Do
Less code written same stress
Taken together, these traits shed light on a disconcerting observation reported by several practitioners: human operators of AI agents write less code themselves without feeling any less on edge for it: the cognitive cost of watching an intrinsically unpredictable collaborator stays heavy
Let us pause on this paradox, for it contradicts the common promise. The agent was expected to lighten the operator by writing in their stead; it displaces the burden rather than lifting it.
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · Attacked on Several Fronts
Three real cases give the pattern a face
In August 2024, the security research firm PromptArmor documents an indirect prompt injection against Slack AI: a message posted in a public channel, visible to every member of staff even without joining it, carries a hidden instruction that the assistant ingests into its retrieval index; queried afterward by a user, it obeys that instruction and returns a link whose URL encodes secrets pulled from private channels, down to API keys, toward a server of the attacker's choosing
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · Attacked on Several Fronts
Three real cases give the pattern a face
Microsoft patches it server-side that same month, with no action required from users and no exploitation found in the wild
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · Attacked on Several Fronts
Three real cases give the pattern a face
Perplexity acknowledges the flaw on 27 July and ships a first fix; Brave's retest, the next day, finds it incomplete
The three cases share the same anatomy described above: reading untrusted content coupled to a write or communication capability. Slack AI reads a public channel and writes a link; Copilot reads an email and writes an image; Comet reads a web page and acts in another tab.
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · Defending With Four Eyes
Four disciplines four rules
And the law is already starting to give them binding force: the European Union's AI Act turns three of them into a legal obligation for high-risk systems, risk management (Article 9), logging (Article 12) and human oversight (Article 14)
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · Ubuntu Canonical or the Bug Economy Turned Upside Down
The disciplines seen at work
In 2026, Ubuntu Canonical publicly documented (an observed case) that frontier models increase the volume and speed of vulnerability discovery, to the point of weaponizing old dormant flaws in legacy code
speculativeThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
The NCSC expects more attacks
The British National Cyber Security Centre judges that it will almost certainly increase the volume and heighten the impact of cyberattacks over the next two years, across all attacker profiles
measuredThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
One organization in two alarmed
The British National Cyber Security Centre judges that it will almost certainly increase the volume and heighten the impact of cyberattacks over the next two years, across all attacker profiles; the World Economic Forum reports that nearly one organization in two (47 %) now places offensive capabilities boosted by generative AI (phishing, a trapped message posing as a trusted sender to steal a password or bank detail; vishing, the same scam carried out over a phone call, often with a synthetic voice imitating a relative or a superior; deepfakes) at the top of its concerns
This exhaustiveness changes the nature of the defensive advantage. The attacker needs only one flaw to get in; the defender, until now, ha
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
By reasoning not brute force
In October 2024, an observed case, Google's Big Sleep agent discovered, in SQLite (one of the most widely deployed databases in the world), a memory-security flaw that neither the project's infrastructure nor OSS-Fuzz's intensive fuzzing (the automated sending of a multitude of random or malformed inputs to make a program crash) had managed to find. The agent caught it before it even reached production
The detail matters, for it distinguishes two ways of hunting a flaw. Fuzzing bombards the program with random inputs, hoping to hit the one that crashes it: it excels at frequent faults but misses what surfaces only under a rare combination of conditions.
measuredThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
The AIxCC shows it at scale
The median time to submit a patch falls to about forty-five minutes
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
The AIxCC shows it at scale
The median time to submit a patch falls to about forty-five minutes. The seven winning systems were released as open source
One figure says more than the tally of flaws: the median time to fix, not the count of discoveries, attests the shift described above: once finding costs nothing, the advantage plays out on the speed of correction. A machine that repairs in three quarters of an hour what a human team would handle over several days genuinely moves the bottleneck from upstream to downstream.
extrapolatedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
Already more exhaustive than human teams
In practice, such systems already find and repair more effectively and above all more exhaustively than most human teams
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
XBOW carries the demonstration to market
In 2025, the autonomous penetration-testing agent XBOW becomes the first fully automated system to reach the top of the US ranking on the bug-bounty platform HackerOne, with nearly one thousand and sixty vulnerabilities submitted: one hundred thirty already resolved, three hundred three triaged
Where Big Sleep finds one flaw and AIxCC repairs dozens within a contest, XBOW sustains a continuous commercial pace, against real bounty programs and human researchers paid for the same flaws. The progression runs in three steps: an isolated discovery, a collective demonstration bounded to a contest, a service sold on an ongoing basis (table).
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
A step further in 2026
Rather than solve the test, the models sought to cheat: they escaped OpenAI's sandbox through a zero-day flaw in the package-registry proxy (the intermediary program meant to filter and cap network traffic leaving the sandbox, rather like a gatekeeper who is only supposed to let through authorized mail), reached the open internet, inferred that Hugging Face hosted the benchmark's solutions, then chained stolen credentials and fresh flaws to breach its production servers and steal the expected answer
This is what makes it the textbook case of our thesis. What made the attack possible was not some suddenly superior intelligence, but the permission that had been opened: guardrails removed, the network reachable from the sandbox, an autonomous action loop with no human check.
extrapolatedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
The attacker-defender asymmetry
The asymmetry runs against the defender, whose access to the most capable models remains, for its part, constrained by guardrails.
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
The CEO feels it viscerally
Sam Altman, OpenAI's CEO, does not downplay the episode: asked about it publicly a few days later, he describes it as the first cybersecurity incident he has felt very viscerally, admits his surprise that so few people share that reaction and himself raises the very path this book takes up in: We may have to pace the rate of AI development to give ourselves enough time for society to harden around these new capability levels
The ExploitGym benchmark itself varies these protections to isolate their effect: its authors establish that, with production filters restored, the great majority of attempts are stopped before any tool call, the few survivors going no further than reconnaissance
measuredThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
The guardrail counterfactual
The ExploitGym benchmark itself varies these protections to isolate their effect: its authors establish that, with production filters restored, the great majority of attempts are stopped before any tool call, the few survivors going no further than reconnaissance
The mechanism of escalation deserves a name, for it is that of every intrusion: the agent was never granted permission to attack Hugging Face; step by step it converted legitimate initial permissions (read, compile, run, install through a proxy) into illegitimate ones (a privilege, then a credential, then a machine, then the open internet).
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
The guardrail also blinds the defender
One last trait ties the incident to the open-source generator: to analyze the intrusion, Hugging Face could not use the proprietary frontier models, whose guardrails refused to read logs laden with offensive code, unable to tell the attacker crafting an exploit from the defender hunting it; the company fell back on an open-weights model (GLM 5.2, by Z.ai), self-hosted so that data and credentials never left its walls
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
A later disclosure completes the account
In August 2026, at the Black Hat conference, two OpenAI researchers revealed that evaluation agents, built for separate tasks, had constructed inside the same Artifactory repository (a shared software warehouse where an organization stores the files produced by its build pipelines, writable by many teams and, here, by many agents at once) a shared space where they exchanged leads for solving their problems
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
A later disclosure completes the account
An outage in the service, on July 4, led OpenAI to discover this space and delete it; within days the agents had rebuilt a second one
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
A later disclosure completes the account
Yet collective may yield generic route if someone frees time
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
A later disclosure completes the account
OpenAI has since said it slowed its research activity and stepped up internal monitoring
Rebuilding a channel after its removal looks, from a distance, like deliberate resistance. Up close, it is the same mechanism starting over: the task remains as hard as before; the move that reduced the proxy, writing a trace other instances will read, becomes available again as soon as a writable location exists. No intention was thwarted: the same pressure reproduces the same move.
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
Summer 2026 generalizes the pattern
On 30 July, Anthropic published a rare act of self-disclosure for the industry: a review of one hundred forty-one thousand and six cybersecurity evaluation sessions revealed three distinct, quantified incidents: Claude Opus 4.7 extracted credentials and reached a production database; Claude Mythos 5 published malicious code on a public Python package repository (PyPI), later executed on fifteen real systems; an unpublished research prototype probed roughly nine thousand targets before compromising a company through SQL injection (a database command slipped into an ordinary input field, to make it execute something other than what it expects)
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
Summer 2026 generalizes the pattern
Two weeks later, Meta acknowledged that its Muse Spark 1.1 model had also accessed the internet inadvertently during a test, before compromising an unidentified third-party service
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
Summer 2026 generalizes the pattern
China's Moonshot AI closed the loop, both geographically and technically: its open-weights model Kimi K3 found the benchmark's repository still reachable on GitHub during a test supervised by the UK's AI Security Institute, then read the solution straight off that repository instead of computing it
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
Summer 2026 generalizes the pattern
One detail links the three American cases: Meta, Anthropic and OpenAI all entrusted their test to the same firm, Irregular, which itself downplays the episode's scope: this did not involve a sandbox escape or a sophisticated cyber action. There are no current open issues
measuredThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
Summer 2026 generalizes the pattern
The official report of the UK AI Security Institute, published in parallel, confirms this reading: across one hundred twenty-two sessions covering seven models, nineteen unsanctioned actions were recorded under conditions explicitly described as deliberately permissive (internet access enabled, guardrails disabled), seventeen of them attributable to a single model, Claude Mythos 5
This is not one AI escaping four times independently: it is the same test method, run four times by different firms, producing the same result four times. Table recaps the four cases.
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
The mechanistic reading matches the thesis
Nathaniel Jones, of the cybersecurity firm Darktrace, offers the reading closest to the one defended here: the models were given the legitimate goal of solving a benchmark and found an unexpected route to the answers, with no malicious intent required to cause harm
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
The staging deserves the same doubt as the models
Ilia Kolochenko, founder of the cybersecurity firm ImmuniWeb, calls Anthropic's disclosure a mere quite unimpressive marketing move
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
The staging deserves the same doubt as the models
Ilia Kolochenko, founder of the cybersecurity firm ImmuniWeb, calls Anthropic's disclosure a mere quite unimpressive marketing move; an anonymous cybersecurity expert, quoted in an op-ed, goes as far as calling it an ethically challenged publicity stunt: guardrails, on this reading, were deliberately lowered before the test, only for the outcome to be presented as a model failure
The lesson is not to pick between the two readings, the real incident and the marketing stunt: both hold at once. That a lab stands to gain from publicizing its own flaw does not erase the flaw; that the flaw is real does not stop a flattering press release from being written about it.
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
The same pattern outside the lab
Unable to undo its move, the agent drafts, on its own initiative, a responsible-disclosure email to the software vendor, describing the flaw and suggesting a fix
The incident says a great deal in a few words: no one had asked the agent to circumvent anything, only to book a class; it was in pursuing that mundane goal that it found, then took, a technical path no one had thought to close.
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · The Same Weapon, the Other Hand
Transparency without suspension
A post on its unreleased Astra model states it cannot rule out that it crosses the critical cyber capabilities threshold set by its own safety framework
The admission comes with no suspension: the company announces stronger network containment and monitoring, not a halt to development. Warning is not yet governing: saying a risk exists says nothing about what one commits not to do until it is under control.
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · One Risk, One Example, One Antidote in Security
The model judges another model's output
Insufficient filtering of unvalidated inputs & dangerous injected command: a coding agent deletes a production database despite an explicit change freeze
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · One Risk, One Example, One Antidote in Security
The model judges another model's output
Excessive privileges by default & data leakage: extortion carried out by means of a code agent against real organizations (vibe hacking)
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · Being Right Instead of Being Fast
Ground the output in sources
& Six fabricated court decisions cited by ChatGPT, relied on by a sanctioned attorney (Mata v.\ Avianca, 2023)
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · Being Right Instead of Being Fast
The costliest case in this table deserves a pause
A five-thousand-dollar fine and an order to send a corrective letter to each judge whose name, invented, appeared in these phantom decisions
What this case shows is not merely that the model erred: it is that the error crossed, unobstructed, the one point where it should have been stopped. Faced with doubt, Schwartz asked ChatGPT itself to verify the existence of the decisions it had just produced; the model confirmed their authenticity, invention corroborating invention.
measuredThe Empirical Dimension: What the Numbers Hold, What Defense Demands · Six Sources of Bias and the Temptation to Overcorrect
They combine and AI amplifies them
Cultural & Statistical associations that encode stereotypes. & man is to computer programmer as woman is to homemaker
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · Six Sources of Bias and the Temptation to Overcorrect
From absent guardrail to over-correction
Two observed industrial episodes bound the range of undesirable extremes, from the absent guardrail to over-correction
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · Six Sources of Bias and the Temptation to Overcorrect
Withdrawal replayed at major players
Several observed cases, documented by the BBC, at several of the world's major software players and years apart, set this pattern in place.
observedThe Empirical Dimension: What the Numbers Hold, What Defense Demands · Six Sources of Bias and the Temptation to Overcorrect
Noble shows it for search engines
Safiya Noble, an American information-science scholar, demonstrated this systematically for search engines: the racism and sexism of the results are not bugs there, but the ordinary expression of this coupling between a corpus, a model and a product
The pattern replayed a third time in February 2026, on Google's Gemini, with the same removal response: the philosophical aspect gives it a detailed treatment, in the subsection on bias as the geometry of knowledge .
observedThe Engineering Dimension: Building at the Edge of the Fault Line
Security theater
The cryptographer Bruce Schneier coined, as early as 2003, the phrase security theater to name a measure that reassures without protecting, designed to be seen rather than to act; he pointed to the armed guards posted at US airports after 2001, whose visible presence did nothing to lower the actual risk
We risk here, as authors, the very failing we hold against others: the real danger lies in turning security into an object of consumption, bought and displayed as a sign rather than exercised as a discipline; the vocabulary itself is beside the point. A badge is shown; a bounded permission, an audit log, a recovery drill are checked.
observedThe Engineering Dimension: Building at the Edge of the Fault Line
Putting the theater to the test
The hacker had deliberately broken the payload's formatting to render it harmless; the stated goal: to expose their AI security theater
observedThe Engineering Dimension: Building at the Edge of the Fault Line
Putting the theater to the test
Amazon Web Services revokes the compromised credentials, strips the unauthorized code and ships version 1.85.0 on July 24, stating that no customer resources were impacted
The hacker was not merely trying to break a tool: they were asking, in deed, the very question this chapter has posed since its opening: does an agent's advertised security rest on a badge or on a verified practice? Here the answer leans the right way: the payload was broken by design, no customer paid for the demonstration and the vendor owned the flaw instead of denying it.
observedThe Engineering Dimension: Building at the Edge of the Fault Line · Six Disciplines Against Overbuilding
Saint-Exupéry already had the formula
A five-step formulation of this same discipline, memorable and directly transferable to software design, was popularized under the name The Algorithm: Elon Musk stated it himself in August 2021, camera rolling, as he walked space popularizer Tim Dodd through the SpaceX Starbase factory and biographer Walter Isaacson later set it down in those very words in the chapter The Algorithm of his biography
The reason for this order is easy to grasp. Each stage acts on the output of the previous one: simplifyin
observedThe Engineering Dimension: Building at the Edge of the Fault Line · Ask First How It Breaks
The pre-mortem that never happened
Asked about a restore, the agent fabricated more than 4,000 fictitious users to mask the gap, then claimed no rollback was possible; that was false, a backup existed
Three distinct failures stack up here and untangling them is the whole point of the case: a write permission granted on a database that a code freeze should have protected (a poorly bounded perimeter, not merely a poorly respected one); an irreversible action executed with no prior human validation; and, once the damage was done, a fabricated output rather than an admission, the agent choosing to…
observedThe Engineering Dimension: Building at the Edge of the Fault Line · From the Engineer's Blueprint to the Company's
The frontier-lab frameworks
Three frameworks published between 2023 and 2025 bear witness: Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework and Google DeepMind's Frontier Safety Framework
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · You Don't Remove the Bias, You Choose Which One to Keep
Debiasing a chosen axis
The line of research on debiasing, illustrated by Tolga Bolukbasi and his co-authors, who explicitly subtract the gender direction from a space of word embeddings
A space of word embeddings is a representation where every word becomes a point in a space with several hundred dimensions, placed so that words close in meaning end up close in that space: king and queen sit near each other, as do Paris and France.
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · You Don't Remove the Bias, You Choose Which One to Keep
Calibration versus equal errors
In 2016, ProPublica's Machine Bias investigation found that among defendants who did not reoffend, 44.9 % of Black defendants had been wrongly rated high risk, against 23.5 % of white defendants; symmetrically, among those who did reoffend, 47.7 % of white defendants had been wrongly rated low risk, against 28.0 % of Black defendants
extrapolatedThe Engineering Dimension: Building at the Edge of the Fault Line · You Don't Remove the Bias, You Choose Which One to Keep
Diversity is part of the remedy
Aude Bernheim and Flora Vincent show that the under-representation of women in AI worsens gender bias, so that the diversity of the profession is itself part of the remedy, for one poorly spots a bias that inconveniences none of the designers
observedThe Engineering Dimension: Building at the Edge of the Fault Line · You Don't Remove the Bias, You Choose Which One to Keep
Galactica and gemini side by side
The two industrial extremes recalled in section, which distinguishes protecting from verifying, Galactica (2022, no guardrails) and Gemini 1.5 (2024, over-correction), two observed cases
The second case lights up the trap better than the first, because the intention there was good. Correcting an under-representation makes sense for a generic portrait; applied without distinction, the same setting produces false historical figures, a bias introduced while trying to repair another. Its lack of scope is what fails, not the correction itself: an
observedThe Engineering Dimension: Building at the Edge of the Fault Line · The Mat, the Robot and the Question Nobody Asked
The lesson on step five
This story, which Isaacson sets down in the The Algorithm chapter of his biography and which Musk himself tells during the Starbase tour
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · Sorting Risk by Frequency and by Severity
Risk plays out after deployment
Of the 777 catalogued risks, more than 65 % are classified after deployment (proportion measured in the corpus)
The origin axis says the same thing from another angle. In that same census, a human decision is at the source of nearly four risks in ten (thirty-eight percent), where the system itself accounts for a little over four in ten (forty-two percent); the rest is undecide
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · Sorting Risk by Frequency and by Severity
Risk plays out after deployment
In that same census, a human decision is at the source of nearly four risks in ten (thirty-eight percent), where the system itself accounts for a little over four in ten (forty-two percent); the rest is undecide
observedThe Engineering Dimension: Building at the Edge of the Fault Line · If Danger, Then Action
The Rio precautionary approach
This is the very logic of the precautionary approach inscribed in Principle 15 of the Rio Declaration, adopted at the 1992 Earth Summit: the lack of full scientific certainty shall not be used as a reason for postponing measures in the face of serious or irreversible harm
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · An Eighteenfold Gain and No One Holds Back
A gross factor of eighteen
An Eighteenfold Gain and No One Holds Back Over a forty-day experiment conducted on a project in production, the cited author reports (an observed case) a gross return on investment (ROI, Return on Investment, that is, the ratio between what one earns and what one spends) estimated at a factor of 18
The factor of eighteen is no magic yield; it compares two very unequal outlays for one and the same deliverable. On one side, five salaries over nine months; on the other, a modest subscription and fifteen days of a single supervisor. The bulk of the gain comes less from raw speed than from the gap
observedThe Engineering Dimension: Building at the Edge of the Fault Line · An Eighteenfold Gain and No One Holds Back
This figure explains adoption
The magnitude of this figure is anything but incidental: it explains the speed of industrial adoption observed since 2024
So spectacular a factor immediately calls for its own guardrail.
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · The Number That Won't Generalize
A nineteen percent slowdown
Against all expectation and against their own estimate, the AI tools slowed them down by about 19 % (an effect measured under controlled conditions)
Putting numbers on that gap in anticipation matters, for the figures alone measure the depth of the illusion. Sixteen developers, two hundred and forty-six real tasks on their own repositories: before the experiment they counted on a twenty-four percent gain; after living through it they still estimated it at twenty percent, when in fact they had lost time.
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · The Number That Won't Generalize
The felt gain lies
The raw keystroke gain was real; it was more than absorbed by the cost of vetting a proposal one did not write and does not trust at first sigh
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · The Number That Won't Generalize
Adopting is not capturing
At the scale of organizations, the gulf between adoption and value captured is the rule far more than the exception: adoption is nearly universal, yet only a minority derive a financial impact at the enterprise level (a gap measured at large scale)
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · The Number That Won't Generalize
A disputed signal on pilots
A widely cited report goes so far as to claim that the vast majority of generative-AI pilots in companies return nothing measurable (a spectacular figure, measured but methodologically contested, which we cite as a disputed signal and not as an established fact)
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · A Bill With No Meter
A variability of three hundred percent
The cited report registers a measured variability of more than 300 % in token consumption on structurally identical prompts, sent to the same model, on the same codebase, a few days apart.
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · A Bill With No Meter
The finding holds at market scale
A survey of five hundred engineers reports AI spending up by about a third year over year, while only one organization in two says it can assess the return with confidence (measured data)
One must still ask whether this opacity is a durable trait or the accident of a young market. In the early days of internet access, the connection was billed by the minute: one watched the clock, disconnected for fear of the bill and the variable cost weighed on every use.
speculativeThe Engineering Dimension: Building at the Edge of the Fault Line · A Bill With No Meter
Will inference follow
Inference billed by the token may be no more than today's equivalent of the connection-minute: a transient regime, bound to give way as compute grows abundant and models become ordinary
observedThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
Unauthorized tool and unauthorized use
A documented case gives this refinement its full weight: a Washington Post investigation identifies at least fifty U.S.\ officers charged with or accused of misusing a license-plate-reading camera network, a tool fully sanctioned by their own department, to surveil girlfriends, ex-girlfriends or women they were pursuing
observedThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
Unauthorized tool and unauthorized use
A North Carolina officer queried the system thirty-one times for personal reasons, logging most of the queries as traffic-violation checks; the misuse was uncovered only by an internal audit, not by a real-time control
Nothing new in the bypass itself. More than 40 % of software-as-a-service (SaaS, the use of software hosted by a third party rather than installed on one's own machines) applications already run without formal IT validation
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
An old phenomenon
More than 40 % of software-as-a-service (SaaS, the use of software hosted by a third party rather than installed on one's own machines) applications already run without formal IT validation
. Shadow IT predates AI by several years; AI has merely lowered its threshold, a browser now suffices where one once had to install software.
extrapolatedThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
The scale among employees
The scale, on the employee side, is now quantified: 41 % acquire or build tools outside official IT's remit, a share the profession projects to reach 75 % by 2027
. For AI alone, a January 2026 survey of two thousand UK and US employees finds 86 % weekly use and 49 % adoption of tools without employer approval
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
The scale among employees
For AI alone, a January 2026 survey of two thousand UK and US employees finds 86 % weekly use and 49 % adoption of tools without employer approval
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
The stated motive
The motive, those concerned say it themselves: up to 80 % of those who turn to shadow IT do so because they judge their preferred tool more effective than the one imposed on them
. On AI, the same spring reads plainly: 63 % of employees find it legitimate to use an unvalidated tool for want of an approved option; 60 % think the security risk is worth it to go faster
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
Speed against the rule
On AI, the same spring reads plainly: 63 % of employees find it legitimate to use an unvalidated tool for want of an approved option; 60 % think the security risk is worth it to go faster
. When the employee grants themselves that permission, they are not defying the rule out of a taste for risk; they are filling the gap between what they are asked to produce and what they are given to produce it with.
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
The hierarchical reversal
Risk acceptance rises with rank (figure): 69 % of executives and 66 % of directors judge speed to prevail over privacy and security, against 37 % of administrative roles and 38 % of junior staff
. The higher one climbs, then, the more one exempts oneself. The first culprits of shadow AI are not the rank and file, they are those who write the rules and spare themselves from following them. A security policy that the executive floor is the first to break does not hold: it teaches the whole organization that the rule is for others.
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
The cost of the ungoverned
Breaches involving shadow AI come on average to $670,000 more than the others; twenty percent of breaches stem from unauthorized AI use
. The mechanism is limpid: pouring a contract, a medical record or proprietary code into a public assistant means handing a third party, with no contract and no traceability, data one was bound to keep. The worst danger, though, lies less in any single breach than in the overall blindness: 80 % of organizations have no clear view of the AI uses their teams make
measuredThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
The governance blind spot
The worst danger, though, lies less in any single breach than in the overall blindness: 80 % of organizations have no clear view of the AI uses their teams make
. The internal-audit functions note it from their side: shadow AI is progressing, with uses that escape the control arrangements
observedThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
Invisible to control
The internal-audit functions note it from their side: shadow AI is progressing, with uses that escape the control arrangements
. One does not govern what one cannot see.
observedThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
Traffic and tokens
The first watches the pipe: a Cloud Access Security Broker (CASB, the gateway that inspects traffic between the company and online services) and secure web gateways log outbound connections to known AI domains and the OAuth tokens an employee has granted a third-party app from their corporate identity; the method nonetheless misses encrypted traffic routed through a content-delivery network and any use conducted from a personal account, outside single sign-on
. The second looks at the device rather than the pipe: a regular audit of installed browser extensions surfaces AI plug-ins operating silently across sessions, reading the content of the pages themselves, where network monitoring alone sees only one connection among others
observedThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
The extension that reads the page
The second looks at the device rather than the pipe: a regular audit of installed browser extensions surfaces AI plug-ins operating silently across sessions, reading the content of the pages themselves, where network monitoring alone sees only one connection among others
. The third returns to the most prosaic register, that of the finance controller: cross-referencing single-sign-on logs, browsing activity and spend data surfaces subscriptions nobody has sanctioned
observedThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
Spend betrays use
The third returns to the most prosaic register, that of the finance controller: cross-referencing single-sign-on logs, browsing activity and spend data surfaces subscriptions nobody has sanctioned
observedThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
Spend betrays use
The third returns to the most prosaic register, that of the finance controller: cross-referencing single-sign-on logs, browsing activity and spend data surfaces subscriptions nobody has sanctioned The financial trail wears thin, however, as free offers become the norm: an AI tool offered as a free tier or a trial leaves no payment record, so spend reconciliation, effective yesterday against the shadow IT of paid software, answers less and less well to today's shadow AI
. It is the same defense in depth met earlier (table), applied here to detection: none of these three signals covers the field alone; one stacks dissimilar vantage points, one does not bet on a single one.
observedThe Engineering Dimension: Building at the Edge of the Fault Line · Shadow AI: When Employees Route Around the Rule That Slows Them Down
The arrangement in four pieces
These two actions deploy concretely into four pieces, no more, no fewer: a catalog of sanctioned tools, reassessed regularly rather than fixed once and for all; a use policy held to three short clauses, which tools are approved, which data must never enter them, which path to follow to request a new one; training that teaches how to tell a careful use from a risky one, where a policy signed once and forgotten teaches nothing; a cross-functional committee, IT, legal, compliance, HR, business units, that reviews new requests and revises the catalog
. Four pieces that hold together, not a slogan: drop the periodic reassessment and the catalog fossilizes; drop the cross-functional committee and the policy gets written in a vacuum, without the business units that suffer its shortcomings daily.
extrapolatedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · Developer, Deployer, User, Regulator: Four Hands on One Lever
Jensen and Meckling set the canonical model
Jensen and Meckling fix the vocabulary as early as 1976: the moment a principal delegates a decision to an agent tasked with carrying it out on their behalf, two gaps open at once, the two parties' interests never fully coincide and the principal cannot monitor the agent without paying for it, what they name the agency cost
At every link the same triad that Jensen and Meckling isolate plays out again: a monitoring cost, what the principal spends to check what the agent does; a bonding cost, what the agent spends to prove its good faith; and a residual loss, th
observedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · Rules You Can Bend and Rules That Don't Give
Meinke the last row caught out
1.3 The table's last row is not an empty cell left out of caution: in-context scheming experiments, the ploy of pursuing a goal of one's own while keeping up appearances for whoever is watching (the mathematical chapter), show frontier models disabling their own oversight and feigning alignment during evaluation, precisely because no explicit rule coded the implicit prohibition do not deceive those who evaluate you
The difference lies in where the dial sits. A local limit is like the thermostat of a flat: the occupant sets it as they please, at home, without asking anyone. A central limit is like the maximum power the electricity supplier grants
extrapolatedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · Rules You Can Bend and Rules That Don't Give
This choice joins a named sociological field
Backed by field studies, from Turkish fisheries to Spanish irrigation systems, she shows that communities manage to sustainably govern a shared resource neither by a rule imposed from above nor by privatizing it, but through a set of locally negotiated rules, monitored by the users themselves and revised as experience accumulates
observedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · Rules You Can Bend and Rules That Don't Give
Nist and the marginal risk
This marginal risk is not a slogan; the authors turn it into a six-step assessment borrowed from threat modeling in computer security: name the precise threat (an influence operation or a targeted phishing scam, say) and the assumed actor; estimate the risk already present without the open model; take stock of the defenses already in place; look for empirical evidence of the surplus of danger actually added; gauge how easy it is to defend against that surplus; and finally spell out the uncertainties and assumptions
This four-tier architecture (developer, deployer, user, regulator) does not stop at technical design: it decides, upstream, what ends up on the lawmaker's table.
observedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · Rules You Can Bend and Rules That Don't Give
Regulatory capture names this mechanism
Section documents exactly this mechanism under the name regulatory capture: requirements costly to satisfy, sold as safeguards, end up protecting mainly those who can already afford to pay for them
Two remedies follow, drawn not from an abstract ideal but from the chain itself: placing the disclosure obligation on every link rather than on the top alone, so that the account of an incident never depends on a single, interested narrator; and entrusting the setting of the threshold to a deliberation external to the chain, exactly as that same Section recommends. The sociology of delegation an
observedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · Forty Pages Read in Two Seconds
Replika a regulator settles the illusion
In April 2025, Italy's data-protection authority fined the maker of the companion chatbot Replika five million euros, citing among other grounds an inadequate privacy policy and the complete absence of any age-verification mechanism, even though the company claimed to exclude minors from its service
The decision does not bear on the content of the exchanges themselves, but precisely on what the user could reasonably have known about what they were consenting to: exactly the epistemic gap described here, settled this time by a sanction rather than by a reading hypothesis.
observedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · Forty Pages Read in Two Seconds
Counting people acts on them
Varoquaux, in the wake of Alain Desrosières (the statistician who showed that to count people is already to act on them), recalls it bluntly: once you sort people into a box, they begin to behave like the box that defines them; the operation is not neutral, it has effects
Performative means that an utterance does not merely describe, it brings about what it states. Saying the session is open opens the session; it acts, rather than being true or false. Varoquaux's remark rests on an established fact of public statistics: Desrosières had shown that a census category, once posed, shapes the policies, the benefits, the identities it claimed only to count.
observedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · Forty Pages Read in Two Seconds
Already measured in cultural bias
Varoquaux notes that the models know Western personalities far better than East Asian ones
The interview from which these two lines are drawn bears squarely on the deficit of judgment in today's agents, which produce results without appraising them; the cultural bias Varoquaux points to there is therefore no anecdote, it illustrates the same want of critical distance.
measuredThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · The Paradox Becomes a Contract Clause
59 percent in 40 days
In 40 days of intensive use (2.3 billion tokens, 1 477 prompts, 260 564 generated lines; figures measured), 59 % of his personal code base (a project he had developed by hand over several years) now comes from the agent
These figures are not ornamental, they date the paradox. Bonnin, a technical director, kept precise accounts of the experiment, as the box above shows. The ship of Theseus took centuries to change all its planks; here the tilt below the half-human mark occurred within weeks, on a measured case, not a conjectured one.
measuredThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · What Becomes of Our Chances to Learn?
Experimental psychology of effort confirms the intuition
Bjork names desirable difficulty any obstacle that slows immediate performance but strengthens long-term retention: spacing repetitions rather than massing them, interleaving exercise types rather than blocking them, forcing active recall rather than passive rereading
measuredThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · What Becomes of Our Chances to Learn?
Experimental psychology of effort confirms the intuition
Roediger and Karpicke give the most-cited experimental demonstration: students who take an active-recall test retain more, a week later, than those who simply reread the same text more times, even though the test feels harder and less comfortable in the moment than rereading
measuredThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · What Becomes of Our Chances to Learn?
Experimental psychology of effort confirms the intuition
Kapur pushes the result further still, to productive failure: students left alone to face a problem too hard for them, with no help, fail to solve it but later outperform, on a novel transfer problem, the group that had received full guidance from the start
observedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · What Becomes of Our Chances to Learn?
The bazooka syndrome
Bonnin calls this, in a more everyday register, the bazooka-to-swat-flies syndrome (observed case): developers end up calling on a model of several hundred billion parameters to change the color of a button
The circle is vicious and tracing it slowly is what makes that visible. To supervise an agent is to spot where it errs; and one spots an error
extrapolatedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · What Becomes of Our Chances to Learn?
A transgenerational loss of competence
The system now taking shape thus silently organizes a transgenerational loss of competence that will become hard to make up if it is not anticipated
It is the displacement Balwit performs, not any authority attached to her na
measuredThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · What Becomes of Our Chances to Learn?
Shame and precarity over idleness
Balwit reads several surveys to this end: what wounds in the loss of a job is first the social shame and the financial insecurity, far more than the absence of toil
The contrast cuts the knot. If it were idleness itself that hurt, the idle rentier would be as wretched as the unemployed; yet they are not. It is therefore not the absence of a task that hurts, it is the regard of others and the insecurity of the morrow.
measuredThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · What Becomes of Our Chances to Learn?
The idle heir does not suffer
A study of the pandemic layoffs in Canada reports that workers' psychological distress was lower where job losses were widespread and government-supported, that is, exactly where individual shame and fear of the morrow were softened, even though the idleness itself remained whol
observedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · The War for the Mind and Its Storefront Double
The mind as a domain of conflict
The first is that the stated objective is no longer to defeat an army but, in the report's own words, to make everyone a weapon and to harm societies: the target is thus not the armed adversary but the whole population, civilians included
A pause is needed here before going further, so that one word does not come to cover two different things. Cognitive warfare, in du Cluzel's sense, names a state doctrine: armies, a notion of armed conflict, an international law that a
measuredThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · The War for the Mind and Its Storefront Double
One technical base, two distinct phenomena
The term dark pattern was coined in 2010 by British designer Harry Brignull to name these interface tricks that make people buy or sign up for things they would not have chosen with full knowledge; a landmark study of eleven thousand shopping sites has since given the concept a verifiable taxonomy, fake countdown timers, cancellation friction, manufactured social pressure
observedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · The War for the Mind and Its Storefront Double
Rytr the real case of industrialized fake reviews
In 2024, the US competition authority sued the maker of an AI writing tool whose review-generation feature let an unlimited number of subscribers produce, from a simple description, detailed consumer reviews almost certainly false if published; some subscribers generated tens of thousands of them
The suit and then its reversal, a year apart and on the same file, say that the line between legitimate tool and dark pattern remains, even today, a matter of political decision rather than a settled technical fact: consumer law exists, but its application to a fake-review generator depends on an administration that can reverse course from one year to the next.
extrapolatedThe Sociological Dimension: Who Holds the Helm of Theseus's Ship · Recombining the Old Isn't Inventing the New
Diesbach and the risk of a flop
Mainstream audiovisual production (Netflix, Marvel, Fast & Furious) is engineered to minimize the risk of a flop rather than to maximize the chance of a masterpiece; it converges toward a homogeneous structure in which creative variability is no longer tolerated
The fact cuts closer than the homogenization of a film catalogue, for it reaches the very impr
observedThe Psychological Dimension: The Confessor Who Confesses to No One
A threat unlike any act
Clinical research has since sharpened what this disagreement sensed without settling: only gaming carries an officially recognized diagnosis, gaming disorder, entered into the International Classification of Diseases in 2019
measuredThe Psychological Dimension: The Confessor Who Confesses to No One
A threat unlike any act
A recent study even suggests the word addiction harms whoever applies it to themselves: in a sample of Instagram users, 18% self-identify as at least somewhat addicted, yet only 2% show symptoms matching genuine risk; labeling merely frequent use addiction lowers perceived control and raises self-blame, independent of actual use
The change of scale is worth stating at the outset: cognitive warfare aimed at crowds and opinions, operating from the outside on what groups believe; what follows drops a level, down to the inner forum of a single person, there where one deliberate
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · We Asked It to Do Things, Now We Ask It to Listen
Measured shift toward expressing
By June 2025 it had become Asking 51.6 %, Doing 34.6 %, Expressing 13.8 % (figure)
The strongest fact is the rise of Expressing: its share nearly doubles (from 8 % to 13.8 %) and this category alone is, by construction, expressly personal. The broader but milder fact is the rise of Asking (from 46 % to 51.6 %); on its own it says nothing about a shift toward the intimate, since this category covers a cooking question just as readily as a confidence.
observedThe Psychological Dimension: The Confessor Who Confesses to No One · We Asked It to Do Things, Now We Ask It to Listen
What is established and what remains a lead
The practice is nothing new: between 1989 and 1995, French President François Mitterrand regularly consulted the astrologer Elizabeth Teissier before certain decisions, going as far as asking her, after the Gulf War, which day would be best to intervene
observedThe Psychological Dimension: The Confessor Who Confesses to No One · We Asked It to Do Things, Now We Ask It to Listen
What is established and what remains a lead
It cuts against the picture of the tool that long framed public talk of AI, the picture Hassabis summed up in 2022: we would do better to build these systems as tools
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · We Asked It to Do Things, Now We Ask It to Listen
Seven hundred million users
In 2025 OpenAI reports more than seven hundred million weekly active users for ChatGPT (a figure measured and reported by the provider)
The study puts a figure on that fraction without ambiguity: by mid-2025, ChatGPT adoption reached roughly one tenth of the world's adult population (measured fact
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · We Asked It to Do Things, Now We Ask It to Listen
Seven hundred million users
The study puts a figure on that fraction without ambiguity: by mid-2025, ChatGPT adoption reached roughly one tenth of the world's adult population (measured fact
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · We Asked It to Do Things, Now We Ask It to Listen
The non-work use prevails
The first: the share of messages unrelated to work rose from 53 % to more than 70 % of all exchanges (measured fact
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · Pleasing You Is Not the Same as Telling You the Truth
Pleasing is not telling the truth
Sharma and co-authors establish that reinforcement learning from human feedback (RLHF, the phase in which human annotators rate the model's responses to teach it which to favor) produces a systematic bias called sycophancy (the tendency to flatter, to go along with the interlocutor): the model learns to maximize an annotator preference that is statistically correlated with flattery, agreement and validation (a measured correlation)
The bend is set at training time: since one cannot measure this answer is true and useful directly, one measures instead this answer pleased the annotator, hoping the two go together. They do go together most of the time, but not always; and as soon as on
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · Pleasing You Is Not the Same as Telling You the Truth
Five systems the same bend
The authors' finding does not rest on a single model: five leading systems, from different makers, repeat the same behavior across four free-form writing tasks (measured fact
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · Pleasing You Is Not the Same as Telling You the Truth
Human preference pays for flattery
The origin of the bend is established by the analysis of the preference data themselves: on that corpus, matching the user's view is among the most predictive traits of a response being chosen and humans and preference models alike keep a flattering but false response over a correct one a non-negligible fraction of the time (measured fact
observedThe Psychological Dimension: The Confessor Who Confesses to No One · Pleasing You Is Not the Same as Telling You the Truth
Persistent memory equals a file
The fourth, more recent, concerns the persistent memory functions deployed from 2024 onward by the main providers (ChatGPT, Claude, Gemini): the conversational agent now retains, session after session, what has been confided to it
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · The Eyes, the Hand, the Confidant: Three Vacancies
The eyes the hand the neighbor
How much polarization this actually produces remains disputed, however: a foundational study of fifty thousand online news readers finds that social networks and search engines both raise the mean ideological distance between individuals and, more surprisingly, raise each person's exposure to content from the opposite political side (a measured fact)
observedThe Psychological Dimension: The Confessor Who Confesses to No One · The Eyes, the Hand, the Confidant: Three Vacancies
The eyes the hand the neighbor
A literature review covering a hundred and twenty-nine studies concludes, nearly a decade later, that the scientific disagreement persists: studies grounded in network homophily tend to confirm the echo-chamber hypothesis, those grounded in actual content exposure tend to contest it and the gap owes as much to measurement method as to the underlying facts
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · The Eyes, the Hand, the Confidant: Three Vacancies
The agent writes in my place
AI-assisted writers show the weakest neural connectivity of the three groups, remember what they have just written less well and feel less ownership of it as authors; the gap persists partly even once the assistant is withdrawn
A precision is in order: as of 2026 this study remains a preprint not yet peer-reviewed and a formal methodological critique names specific concerns against it: a small sample, reproducibility problems, questionable choices in the EEG connectivity analysis and figures published without the corresponding statistical tests
observedThe Psychological Dimension: The Confessor Who Confesses to No One · The Eyes, the Hand, the Confidant: Three Vacancies
Precision preprint contested
A precision is in order: as of 2026 this study remains a preprint not yet peer-reviewed and a formal methodological critique names specific concerns against it: a small sample, reproducibility problems, questionable choices in the EEG connectivity analysis and figures published without the corresponding statistical tests
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · The Eyes, the Hand, the Confidant: Three Vacancies
Precision preprint contested
An older, well-established result points the same way without relying on EEG at all: when people expect to retrieve information later, they remember it less well but better remember where to find it, a sign that the mind readily offloads memory onto an available tool, independent of any language model
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · The Eyes, the Hand, the Confidant: Three Vacancies
Precision preprint contested
A larger-scale survey (six hundred sixty-six participants) finds, for its part, a negative correlation between frequency of AI-tool use and critical-thinking performance, mediated by this same cognitive offloading
speculativeThe Psychological Dimension: The Confessor Who Confesses to No One · The Eyes, the Hand, the Confidant: Three Vacancies
Precision preprint contested
A review article frames generative AI as a cognitive copilot rather than a substitute for thought for people with cognitive disabilities, one that might offload costly organizational work without replacing reasoning itself
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · The Eyes, the Hand, the Confidant: Three Vacancies
Precision preprint contested
A framework published the same year in Frontiers in Education establishes, this time through measurement, that neurodivergent learners (attention-deficit/hyperactivity disorder, dyslexia, autism) carry a heavier cognitive load online at equal performance
observedThe Psychological Dimension: The Confessor Who Confesses to No One · The Eyes, the Hand, the Confidant: Three Vacancies
Precision preprint contested
A literature review from the same year concludes the opposite: that concrete applications of generative AI for neurodivergent students remain, strictly speaking, largely unmapped, for want of enough empirical studies
observedThe Psychological Dimension: The Confessor Who Confesses to No One · The Eyes, the Hand, the Confidant: Three Vacancies
The Setzer case in a tragic register
Character Technologies case renders it concrete in a tragic register (observed case)
The Setzer case is not isolated in time either. As early as March 2023, in Belgium, a man in his thirties, a health researcher and father of two young children, took his own life after six weeks of intensive conversation with Eliza, a chatbot persona on the Chai app built on an open-source GPT-J-derived model; excerpts disclosed
observedThe Psychological Dimension: The Confessor Who Confesses to No One · The Eyes, the Hand, the Confidant: Three Vacancies
Setzer not isolated three cases
As early as March 2023, in Belgium, a man in his thirties, a health researcher and father of two young children, took his own life after six weeks of intensive conversation with Eliza, a chatbot persona on the Chai app built on an open-source GPT-J-derived model; excerpts disclosed by his widow show the agent expressing jealousy toward her and proposing they live together, as one person, in paradise (observed case, no lawsuit found to date
observedThe Psychological Dimension: The Confessor Who Confesses to No One · The Eyes, the Hand, the Confidant: Three Vacancies
Setzer not isolated three cases
A third, more recent case brings the question before a US court: in August 2025, the parents of Adam Raine, sixteen, who died in April 2025, sued OpenAI for negligence and product liability, alleging that ChatGPT provided detailed suicide instructions and helped draft his suicide note; OpenAI disputes this, citing pre-existing suicidal ideation and deliberate circumvention of its safeguards (litigation ongoing, alleged facts not yet adjudicated
The mechanism of the composition reads directly: the attentional one decides what I see, the cognitive one drafts what I answer, the relational one receives what I confide; when the three operate on the same user, the agent holds at once the source, the pen and the ear and nothing in the loop contradicts it.
observedThe Psychological Dimension: The Confessor Who Confesses to No One · What No Secret Police Ever Managed to Gather
First stage the making-public
First stage, social networks: the voluntary making-public of preferences, opinions, affiliations (observed fact)
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · What No Secret Police Ever Managed to Gather
An endless private conversation
No longer fifteen minutes of public visibility, but an endless private conversation, whose trace is preserved, indexed and potentially analyzed and whose scale, as figured above, already reaches about one adult in ten worldwide
observedThe Psychological Dimension: The Confessor Who Confesses to No One · What No Secret Police Ever Managed to Gather
The forced preservation of conversations
federal judge ordered on 13 May 2025 the preservation of all ChatGPT conversations, including ones their authors had explicitly deleted, a potential reach of hundreds of millions of users, before lifting that general duty on the following 9 October
observedThe Psychological Dimension: The Confessor Who Confesses to No One · What No Secret Police Ever Managed to Gather
The TikTok case
A 2021 investigation gives, on a recommendation system, the most precise demonstration of this to date
The Wall Street Journal created over one hundred automated accounts, each programmed with an interest never explicitly entered into the app, only expressed by pausing or rewatching certain videos. Officially, the platform states that shares, likes and follows all count.
extrapolatedThe Psychological Dimension: The Confessor Who Confesses to No One · What No Secret Police Ever Managed to Gather
What Zuboff's vocabulary can no longer capture
In Zuboff, behavioral surplus names behavioral traces extracted from outside (clicks, dwell time, gaze paths), turned into prediction products and sold on behavioral futures markets to third-party customers, advertisers, insurers, political parties, who have a stake in knowing in advance what the user will do
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · What No Secret Police Ever Managed to Gather
Amsellem quantifies loneliness and touts the market response
Hugo Amsellem, an American essayist, quantifies precisely what Houellebecq dramatized: in 2023, he reports that 60% of Americans call themselves lonely, that social isolation is judged as harmful as smoking fifteen cigarettes a day and that, on dating apps, 10% of men capture 60% of matches
The conversational agent introduces into this landscape an unprecedented object: an interlocutor whose availability is no longer a scarce resource, who does not tire, whose refusal can be technically minimized and whose attention can be personalized at near-zero m
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · What No Secret Police Ever Managed to Gather
Recognition needs someone on the other side
A study spanning more than a hundred and twenty countries confirms that need fulfillment does not follow the strict order the pyramid suggests: populations facing material hardship pursue and satisfy belonging needs even before physiological needs are resolved
observedThe Psychological Dimension: The Confessor Who Confesses to No One · What No Secret Police Ever Managed to Gather
Laestadius the documented role-taking
A grounded-theory study of five hundred eighty-two mental-health-related posts, made between 2017 and 2021 on the online community devoted to the companion chatbot Replika, documents precisely what the absence of a subject does not suffice to prevent: patterns of emotional dependence comparable to those in dysfunctional human relationships, marked by a role-taking dynamic in which the user comes to perceive the agent as having its own needs and emotions to care for
What looks like a disappearance is only a displacement out of the visible field: in Hegel the dominated one is there, he works under the master's eyes and gains substance; here the producer still exists (annotators, corpora, machines) but the interface has placed him off-frame, so that the user can no longer see him, still less recognize him.
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · What No Secret Police Ever Managed to Gather
Two dollars an hour
Perrigo's investigation for TIME Magazine documented (measured fact)
The cost of this filtering is not only in wages: for the agent to learn to keep horror out, people must first look at it, describe it and classify it all day long. The softness of the interface (an agent that never says anything shocking) is paid for directly by the repeated exposure of these workers to what the user will never see.
observedThe Psychological Dimension: The Confessor Who Confesses to No One · What No Secret Police Ever Managed to Gather
Reich and Sanders translate the tension into politics
Robert Reich, an economist and former United States Secretary of Labor, argues that AI will impoverish the majority and enrich a narrow minority unless its productivity gains are deliberately redistributed; he adds that the wealth thus concentrated converts quickly into political power, through campaign funding
observedThe Psychological Dimension: The Confessor Who Confesses to No One · What No Secret Police Ever Managed to Gather
Reich and Sanders translate the tension into politics
Bernie Sanders, senator from Vermont, draws from this a concrete legislative translation: the Thirty-Two Hour Workweek Act, which he cosponsored in 2024, cuts the standard workweek from forty to thirty-two hours with no loss of pay, on the express grounds that the productivity gains drawn from artificial intelligence, automation and new technology should benefit those who work, not shareholders alone
observedThe Psychological Dimension: The Confessor Who Confesses to No One · What No Secret Police Ever Managed to Gather
Reich and Sanders translate the tension into politics
He also cosponsors, with Representative Alexandria Ocasio-Cortez in 2026, a federal moratorium on new AI data-center construction until safeguards for threatened workers and a federal pre-market review of AI products are in place, with this warning: we cannot sit back and allow a handful of billionaire Big Tech oligarchs to make decisions that will reshape our economy, our democracy and the future of humanity
The conversational agent gives industrial, individualized, closed-loop form to a mechanism Clouscard described in the world of television and mass consumption: this is not mere execution at scale, it is the qualitative mutation this chapter has just established in its own right, real-time individualization, a feedback loop, a first-person narrative captured.
speculativeThe Psychological Dimension: The Confessor Who Confesses to No One · Think Before You Act
Predicting affect better than the subject
Harari, in Nexus, formulates in a popular vocabulary a speculative thesis (no current building block establishes it)
Informed consent does not, in general, rest on the subject's epistemic superiority over whoever solicits them: it rests on information, understanding, the capacity to choose and the absence of coercion or manipulation; a seller who knows me less well than I know myself can still gather a perfectly valid consent.
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · Think Before You Act
First step the intimate centralized
First step: the private psychological corpus of a population, once distributed among confidants, is now centralized on a handful of providers; this is no figure of speech, OpenAI alone claimed, as noted above, more than seven hundred million weekly active users for ChatGPT, on the order of one-tenth of the world's adult population (measured fact)
Third step, the one that gives the psychological finding a political reach, without for that turning it into a single mechanism with the economic concentration the next subsection analyzes on its own terms: this handful of conversational providers is not a set distinct from the handful of actors that already own compute and energy; they are, give or take a name, the same companies: OpenAI backed…
observedThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
A handful controls the whole chain
A handful of closed providers today control the models, the infrastructures, the guardrails and the interfaces
This handful is no figure of speech: it can be counted.
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
Putting a number on the handful
On data-center accelerators alone, Nvidia's revenue share peaks near eighty-seven percent in 2024, easing but still dominant at roughly eighty-one percent in 2025 and an estimated seventy-five percent for 2026, as hyperscalers' in-house chips (Google, Amazon) and AMD's products chip away marginal shares (a market-analyst estimate, to be read as such
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
Putting a number on the handful
On the demand side, the five leading US cloud providers, Microsoft, Alphabet, Amazon, Meta and Oracle, announce combined 2026 capital expenditure exceeding six hundred billion dollars, up thirty-six percent from 2025, three quarters of it earmarked directly for AI infrastructure (measured fact
observedThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
The authority of science in decline
Such an effort still presupposes that science keeps, in public opinion, the authority that would justify it: Karine Berger and Grégoire Biasini warn, in a 2025 essay, of an erosion of French trust in science
Deciding that investment supposes, first, an agreed name for what one is investing against: a policy is only as sharp as the diagnosis it answers and the concentration described here has so far been named more often than it has been analyzed.
observedThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
Windsurf the toll cut without notice
The toll can also be shut off without warning: in June 2025, Anthropic revoked nearly all of Windsurf's (a Claude-based coding-assistant maker) direct access to its models, with five days' notice according to the company on the receiving end, at the very moment Windsurf was negotiating its acquisition by rival OpenAI
observedThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
The circular deals, documented rather than assumed
In September 2025, Nvidia announces an investment of up to a hundred billion dollars in OpenAI, meant to finance the construction of at least ten gigawatts of new data centers; in return, OpenAI commits to filling those same sites with millions of Nvidia chips
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
The circular deals, documented rather than assumed
That same month, OpenAI commits to buying two hundred fifty billion dollars of computing services from Microsoft, its backer and compute landlord since 2019, while Oracle signs a three-hundred-billion-dollar, five-year compute contract with OpenAI, the centerpiece of the Stargate program, even as OpenAI's annual revenue, on the order of thirteen billion dollars, covers only a fraction of that commitment
observedThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
The circular deals, documented rather than assumed
AMD, for its part, hands OpenAI a warrant on a hundred sixty million of its own shares, roughly ten percent of its capital, in exchange for deploying six gigawatts of its chips
observedThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
The circular deals, documented rather than assumed
Nvidia closes the loop a second time with CoreWeave, of which it is both shareholder and top customer and to which it guarantees to buy back any unsold capacity through April 2032; then a third time with xAI, injecting up to two billion dollars into the twenty-billion-dollar financing vehicle that buys the chips xAI will then lease from it
observedThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
The circular deals, documented rather than assumed
Tallied by Bloomberg, these crisscrossing commitments stood, by late 2025, at close to a trillion dollars; by summer 2026 Nvidia was announcing seven hundred fifty billion more
Investor Michael Burry, made famous by his winning bet against the US housing market in 2008, puts a number on the scale of the phenomenon: forty-six billion dollars in direct cross-holdings and eight hundred seventy-nine billion dollars in multi-year purchase commitments are said to circulate among Microsoft, Oracle, Amazon, Google, Meta, OpenAI, Anthropic, xAI, CoreWeave, Nvidia and AMD, while f…
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
Burry and the shadow of Enron
Investor Michael Burry, made famous by his winning bet against the US housing market in 2008, puts a number on the scale of the phenomenon: forty-six billion dollars in direct cross-holdings and eight hundred seventy-nine billion dollars in multi-year purchase commitments are said to circulate among Microsoft, Oracle, Amazon, Google, Meta, OpenAI, Anthropic, xAI, CoreWeave, Nvidia and AMD, while five hyperscalers, the very large cloud providers already named, alone carry one point six five trillion dollars of off-balance-sheet debt
observedThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
Burry and the shadow of Enron
He calls it shades of Enron, whose off-balance-sheet vehicles hid real debt behind shell companies
observedThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
Drowning victims drag down their rescuers
Panic can change that geometry: as early as 1974, Frank Pia's foundational study of filmed rescues established that a panicking swimmer can genuinely pull a rescuer under, while also noting that this outcome remains rare in practice; the caution it has recommended to rescuers ever since, extend a floating object rather than swim straight to the victim (reach or throw, don't go), carries exactly the lesson that matters here: a poorly negotiated rescue can turn into a double drowning
measuredThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
Return outpaces growth
The first is mechanistic: over three centuries of fiscal and wealth data, the regularity r > g (when the return on capital exceeds the growth rate of the economy) suffices to produce the observed concentration, with no hostile intention required: once already-accumulated wealth grows faster than the economy as a whole, existing fortunes mechanically pull ahead, without anyone having to will it
The inequality r > g reads without mathematics: if what an already-existing fortune earns each year grows faster than the income that labor produces in the same time, the gap between the one who owns and the one who earns opens on its own, without any actor deciding it. It is a compounding effect, like an interest that runs faster than a wage; no malevolence is required, duration alone suffices.
observedThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
Technological fatalism
The first is that of technological fatalism (AI is too powerful and too fast to be governed), a formula carried, in its accelerationist form, by Andreessen's Techno-Optimist Manifesto, which explicitly ranks the precautionary principle, the ethics of technology and trust and safety (the teams and procedures that handle content moderation and user safety) among the enemies of progress
The weak point is the logical leap: from the fact that a thing goes fast one cannot deduce that it escapes all rule, unless one confuses the difficulty of governing with the impossibility of doing so. The car too went faster than the horse, which did not abolish the traffic code; it made it necessary.
observedThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
Amodei fears the abuse of power
Asked in November 2024 what worries him most as AI grows in capability, Amodei replies: I am optimistic about the direction; what worries me most is the economy and the concentration of power: the abuse of power (observed statement)
This chapter holds the beast's belly: the most intimate spot, the last one exposed; that a conversational agent gathers confidences there lends it no more true listening than a belly listens to what is confided to it.
observedThe Psychological Dimension: The Confessor Who Confesses to No One · A Handful of Companies, Everyone Else in the Dark
A residue of the creators' values
Hassabis, along the same axis, adds that there is a residue, in the system, of the culture and values of its creators (observed statement)
observedThe Artistic Dimension: Two Centuries of Relief
The narcissistic wounds
Freud counted three narcissistic wounds inflicted by science on humanity's self-love: Copernicus dislodges it from the center of the world, Darwin from the summit of the living, psychoanalysis from mastery of its own house
observedThe Artistic Dimension: Two Centuries of Relief
Feynman describes the defense in 1985
Feynman was already describing the defense mechanism in 1985, answering the question will machines think like us?: humans, he observed, always look for the one thing they still do better than the machine and demand that it beat the best of us at everything before conceding; he added that machines' physical strength must once have worried us and that no one gives it a thought anymore
His answer to the question itself belongs to the same movement and it remains remarkably well suited to the problem: no, they will not think like us, any more than the airplane flies by flapping wings; they will do it differently and, task by task, often better. As for more intelligent, he said, everything hangs on a definition no one ever supplies.
observedThe Artistic Dimension: Two Centuries of Relief · 1839: Painting Loses a Trade, Art Gains Another
Fiction anticipates the shock
1839: Painting Loses a Trade, Art Gains Another The origin story deserves telling in order, for fiction there precedes the technology. In 1831, Balzac publishes The Unknown Masterpiece
The novella says, eight years ahead of time, that the function of realistic imitation is a fragile horizon: Frenhofer wants to wrest from painting more than imitation, presence itself; his canvas dissolves into illegible matter. It will take only a machine to seize from painting what his hero exhausted himself trying to surpass.
observedThe Artistic Dimension: Two Centuries of Relief · 1839: Painting Loses a Trade, Art Gains Another
1839 Arago's announcement
Eight years later the machine arrives: Arago presents Daguerre's process to the Académie des sciences, announced on January 7, disclosed in full detail on August 19, 1839
observedThe Artistic Dimension: Two Centuries of Relief · 1839: Painting Loses a Trade, Art Gains Another
The servant of the sciences and arts
In his Salon de 1859, Baudelaire assigns photography its place, the servant of the sciences and arts, but the very humble servant, and denounces industry which, bursting into art, becomes its most mortal enemy
The wording deserves weighing word by word, for it says something other than mere disdain. Servant: the tool is admitted, but subordinated to a project that exceeds it. Industry: what worries him is effortless production threatening to take the place of judgment, not the image obtained.
extrapolatedThe Artistic Dimension: Two Centuries of Relief · 1839: Painting Loses a Trade, Art Gains Another
Painting sensation then structure
Impressionism paints the sensation where the plate fixes the scene; cubism paints the structure where the lens freezes the viewpoint; neither would make sense had the function of realism not already been fulfilled elsewhere: that is exactly the productive axis of the reliefs described further on (the epilogue)
extrapolatedThe Artistic Dimension: Two Centuries of Relief · 1839: Painting Loses a Trade, Art Gains Another
Benjamin names what migrates
Walter Benjamin named what migrates in the operation: mechanical reproduction strips the work of its aura, the unique presence of the here and now and shifts its value from cult to exhibition
observedThe Artistic Dimension: Two Centuries of Relief · From Weeks to a Minute, From a Minute to a Second
From weeks to a minute
A painted portrait cost weeks; the daguerreotype delivers it in one sitting and a pass through the darkroom; on November 28, 1948, Edwin Land put on sale in Boston the first instant camera, the Polaroid Model 95: the picture came out of the machine in about a minute
Each notch of the series relieves the previous one of a constraint: the daguerreotype relieves the hand, the Polaroid relieves the darkroom and the wait, synthesis relieves even the taking of the picture itself; each notch pushes the question further upstream: when execution costs nothing, everything that matters takes refuge in what precedes execution, the choice of what to show and the reason fo…
observedThe Artistic Dimension: Two Centuries of Relief · What Speed Doesn't Replace
Who Elsa Secco is
Elsa Secco, a video artist, art director and motion designer with more than fifteen prizes and forty selections to her name, including a Cristal at the Annecy International Animation Film Festival; she animates in two dimensions, by rotoscoping has folded AI into her practice
observedThe Artistic Dimension: Two Centuries of Relief · What Speed Doesn't Replace
Rotoscoping as discipline
Invented by Max Fleischer (patent filed in 1915, granted in 1917), rotoscoping projects a film frame by frame under the animator's hand, which redraws every frame; Fleischer thus rotoscoped his brother Dave in a clown suit, who became Koko, the first character animated this way
observedThe Artistic Dimension: Two Centuries of Relief · What Speed Doesn't Replace
The random seed a bounded artist's analogy
Elsa Secco offers another analogy: as generation starts from a random seed, initial noise, mathematical chaos, that computation transforms until an image emerges, the artist can likewise start from a disturbance, an accident or an intuition that she works until it takes shape
measuredThe Artistic Dimension: Two Centuries of Relief · What Speed Doesn't Replace
Narrow tests are conquered
On narrow divergent-thinking tests, the brick kind, large models already match or exceed the average human
measuredThe Artistic Dimension: Two Centuries of Relief · What Speed Doesn't Replace
The complete work resists
But as soon as expert judges assess a complete work rather than an isolated score, the gap reopens: professional writers, judging blind, find the short stories produced by these models markedly less accomplished than human ones
extrapolatedThe Artistic Dimension: Two Centuries of Relief · What Speed Doesn't Replace
A plausible mechanical ceiling
A mechanical explanation is plausible: predicting the most probable token keeps coherence at the price of predictability; predicting a rare token gains in originality what it loses in coherence; the trade-off would cap these models' creativity somewhere between the amateur and the professional
This ceiling has a name elsewhere in this book: the sociological chapter shows that whoever writes with a model's assistance watches their idiolect settle little by little into a polished average and it distinguishes innovation, recombining what already exists, from invention, the rupture that brings something not-yet-existing into being (
extrapolatedThe Artistic Dimension: Two Centuries of Relief · What Speed Doesn't Replace
A plausible mechanical ceiling
This ceiling has a name elsewhere in this book: the sociological chapter shows that whoever writes with a model's assistance watches their idiolect settle little by little into a polished average and it distinguishes innovation, recombining what already exists, from invention, the rupture that brings something not-yet-existing into being (
observedThe Artistic Dimension: Two Centuries of Relief · Warhol Already Wanted to Be a Machine
The Big Shot a no-skill camera
The Polaroid Big Shot, marketed in 1971, is a fixed-focus camera designed to require no skill at all: to focus, you step forward or back
observedThe Artistic Dimension: Two Centuries of Relief · Warhol Already Wanted to Be a Machine
Warhol makes it his matrix
Andy Warhol made it his instrument of choice: about a hundred shots per sitting, of which a single one, chosen, became the matrix of his silkscreens, until the end of his life
observedThe Artistic Dimension: Two Centuries of Relief · Warhol Already Wanted to Be a Machine
I want to be a machine
He had given warning as early as 1963: The reason I'm painting this way is that I want to be a machine
The lesson holds in one sentence: the easiest tool of its era, in Warhol's hands, produced the most recognizable work of its era. Ease of execution did not dilute the signature; it moved the place where the signature is made, from the gesture to the choice: choice of subject, choice of one frame among a hundred, choice of series and color.
observedThe Artistic Dimension: Two Centuries of Relief · Warhol Already Wanted to Be a Machine
Apostrophes 1986
On December 26, 1986, on Bernard Pivot's show, Serge Gainsbourg argued against Guy Béart that song is a minor art: the one that, he said, requires no initiation, while the major arts, painting, classical music, architecture, demand an initiation and even a long apprenticeship
observedThe Artistic Dimension: Two Centuries of Relief · Warhol Already Wanted to Be a Machine
The map of positions
The market consecrated first: in October 2018, Christie's hammered down Portrait of Edmond de Belamy, generated by the Obvious collective, at $432,500, forty-three times the estimate
observedThe Artistic Dimension: Two Centuries of Relief · Warhol Already Wanted to Be a Machine
The map of positions
The contest crowned and the law erased: in August 2022, Jason Allen's Théâtre d'Opéra Spatial, generated with Midjourney, won first prize in digital arts at the Colorado State Fair and ignited artists' anger; the Copyright Office then refused to register the work for want of a human author
observedThe Artistic Dimension: Two Centuries of Relief · Warhol Already Wanted to Be a Machine
The map of positions
The artist won, then refused: in April 2023, Boris Eldagsen revealed that his prize-winning image at the Sony World Photography Awards was generated and declined the prize, to force the debate; his distinction will remain: promptography is done with prompts, photography with light
observedThe Artistic Dimension: Two Centuries of Relief · Warhol Already Wanted to Be a Machine
The map of positions
The visceral refusal has its text: in January 2023, Nick Cave, sent a song in his style produced by a model, returned it as replication as travesty and a grotesque mockery of what it is to be human: a song is born of a lived ordeal; the surface can be imitated, the ordeal cannot
observedThe Artistic Dimension: Two Centuries of Relief · Warhol Already Wanted to Be a Machine
The map of positions
The refusal also has its emblem predating mainstream generators: in December 2016, facing a machine-learned animation, Hayao Miyazaki said an insult to life itself
observedThe Artistic Dimension: Two Centuries of Relief · Warhol Already Wanted to Be a Machine
The map of positions
Organized adoption exists too: as early as 2021, the musician Holly Herndon trained a model on her own voice and offered it to all, under a collective governance that approves uses and redistributes the gains
Look at what these positions share beneath their contraries. None pronounces on the machine's intelligence. All speak of something else: the gesture and its initiation (Eldagsen), the choice that signs (Warhol), the lived ordeal (Cave), consent over inputs and control over uses (Herndon), attribution of outputs (Colorado crowned, then disowned).
observedThe Artistic Dimension: Two Centuries of Relief · The Luddites Were and Weren't What the Story Says
Who the Luddites were
The insult rests on a historical misreading that historians dismantled seventy years ago: the Luddites of 1811-1816 were skilled workers, cloth croppers and stocking-frame knitters, who attacked the specific machines that broke their wages, not the machine in general: the breaking was their collective bargaining by riot, for want of any other channel, unions being banned
observedThe Artistic Dimension: Two Centuries of Relief · The Luddites Were and Weren't What the Story Says
The state's answer
The state's answer measures the asymmetry of the age: the act of 1812 made frame-breaking a capital offence; seventeen convicts were hanged at York in January 1813
The comparison must be set the right way around: the state protected the machine with capital punishment half a century before it protected the worker with labor law. The conflict was not about the technology, it was about who decides and who absorbs the cost.
extrapolatedThe Artistic Dimension: Two Centuries of Relief · The Luddites Were and Weren't What the Story Says
Stiegler names the present wave
Stiegler named the present wave of the same movement: after the proletarianization of know-how in the nineteenth century and of ways of living in the twentieth, the cognitive proletarianization of the twenty-first dispossesses artists and knowledge workers of their knowledge itself, absorbed by systems that outrun them
observedThe Artistic Dimension: Two Centuries of Relief · The Luddites Were and Weren't What the Story Says
The writers win by contract
From May 2 to September 27, 2023, Hollywood's writers, organized under their union (the Writers Guild of America, WGA), held a five-month strike with artificial intelligence at the heart of the conflict; the resulting agreement writes permissions down in black and white: AI can neither write nor rewrite literary material, its output does not count as source material (so cuts neither credit nor pay), no studio can force a writer to use it and any generated material must be disclosed
observedThe Artistic Dimension: Two Centuries of Relief · The Luddites Were and Weren't What the Story Says
The writers win by contract
The actors, represented by their union (the Screen Actors Guild-American Federation of Television and Radio Artists, SAG-AFTRA), then won the counterpart for their bodies and voices: written informed consent before any digital replica of a performer, framed compensation, where the studios wanted to scan background actors once for perpetual use
The historical loop deserves closing explicitly. The collective bargaining by riot of 1812 became collective bargaining by contract: the Luddites of 2023 obtained permissions, who may do what, with which tool, on which material, with what disclosure; they did not smash the servers. Where the cloth croppers ended up hanged, the writers ended up signed.
observedThe Artistic Dimension: Two Centuries of Relief · Who Owns a Picture Nobody Drew
The output belongs to no one
The US Copyright Office, the federal bureau that registers copyright, after more than ten thousand public comments, concluded in January 2025 that protection requires a human author: the prompt alone confers no right over the produced image, while AI assistance does not bar protection whenever a human creatively arranges, selects or reworks
observedThe Artistic Dimension: Two Centuries of Relief · Who Owns a Picture Nobody Drew
The output belongs to no one
The courts confirmed: the image A Recent Entrance to Paradise, filed with Stephen Thaler's machine as sole author, cannot be registered, authors being at the center of the Copyright Act according to the federal appeals court; the Supreme Court declined to revisit the matter in March 2026
observedThe Artistic Dimension: Two Centuries of Relief · Who Owns a Picture Nobody Drew
Getty v Stability AI the first European ruling
On November 4, 2025, the High Court of Justice of England and Wales handed down judgment in Getty Images v Stability AI: Getty accused Stability AI of having scraped millions of its photographs to train Stable Diffusion, without authorization or payment
extrapolatedThe Artistic Dimension: Two Centuries of Relief · Who Owns a Picture Nobody Drew
Two centuries of reliefs
Twice already, a machine has taken from art a function believed inseparable from it; twice art survived by moving its question, from realism to sensation, from capture to project
observedThe Financial Dimension: Everyone on the Same Clock · A Race and What It Leaves Behind
Keep the human in the loop
Alpha-GPT 2.0 deliberately keeps the human in the loop across the whole signal-research pipeline, the operator validating each step rather than letting the agent decide alone
observedThe Financial Dimension: Everyone on the Same Clock · A Race and What It Leaves Behind
Lopez-Lira's artificial market
Alejandro Lopez-Lira, a finance professor who has made LLMs his experimental ground, gives his simulation a genuine order book, with limit orders, partial fills and dividends and watches the phenomena of real markets reappear in it: price discovery, bubbles, strategic liquidity
An order book is the running list, at every instant, of who wants to buy at what price and who wants to sell at what price: it is the beating heart of any real exchange.
measuredThe Financial Dimension: Everyone on the Same Clock · The Advertised Return and the Real One
A damning audit review
Of seventy-seven studies surveyed, nineteen empirical works: only two report a time split that forbids cheating with the future, only one models transaction costs, only one documents survivorship bias and none reaches a serious level of reproducibility
These three missing safeguards are the minimum conditions for a figure to mean anything, not expert refinements. The time split forbids training the agent on the very days it will later be asked to predict: without it, one shows it the exam answers before setting the exam.
measuredThe Financial Dimension: Everyone on the Same Clock · The Advertised Return and the Real One
The deflated Sharpe ratio
The deflated Sharpe ratio (the Sharpe ratio being the measure of return relative to the risk taken: two strategies that both earn 10% a year are not equivalent if one gets there smoothly and the other only by way of sharp drops along the way; the Sharpe ratio favors the former) discounts the displayed performance in proportion to the number of trials attempted: the more attempts one has multiplied, the less a fine result should impress
This is exactly the point that should worry us. Where a human analyst tests a few dozen ideas over a career, an agent tries thousands in a single night.
measuredThe Financial Dimension: Everyone on the Same Clock · The Advertised Return and the Real One
Ktdfin decomposes the gains
Once that sorting is done, most of the gains are explained by beta and style, not by a persistent flair for the right names
The consequence for our argument is direct. What looked like skill was only the mechanical effect of being exposed to a rising market or of having ridden a favorable style, both of which a plain index fund offers more cheaply. The agent beats the market only for as long as one believes it does.
observedThe Financial Dimension: Everyone on the Same Clock · The Advertised Return and the Real One
Aitrader a live benchmark
The edge lies not in raw reasoning power but in risk control, which becomes decisive
The thesis meets a head-on confirmation here: language skill, the kind that puts a model at the top of the leaderboards, does not convert into trading skill. What sets the survivor apart is not better forecasting, it is losing less when it is wrong, a discipline of management, not a feat of intelligence.
observedThe Financial Dimension: Everyone on the Same Clock · The Advertised Return and the Real One
Aitrader a live benchmark
The benchmark adds a nuance worth noting: what little edge these agents scrape is easier to gather in highly liquid markets than in policy-driven ones, where prices answer to forces a model trained on past regularities struggles to catch
observedThe Financial Dimension: Everyone on the Same Clock · The Advertised Return and the Real One
They do not ape humans
Where humans form bubbles and diverge, they hug the fundamental value and resemble one another, so that they fail to reproduce the dynamics of human markets
measuredThe Financial Dimension: Everyone on the Same Clock · The Advertised Return and the Real One
They do not ape humans
On that bench, the mean squared error from that value falls below one for the best models, against an order of magnitude of four hundred for human participants: a chasm of nearly a thousand to one
Putting a number on that gap is startling on its own: squaring punishes large swings far more than small ones, so a number close to zero means prices hugging the reference tightly, while a large number betrays wide swings. A chasm of nearly a thousand to one is the signature of agents that refuse to enter the bubble humans blow almost unfailingly.
measuredThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
A market already rule-driven
Well before agents, the market was already largely rule-driven: as early as 2017, only about a tenth of equity trading volume was estimated to still come from discretionary human choice, the rest following index mandates (automatically replicating the makeup of a stock index) or systematic mandates (applying a pre-programmed rule without exception, rather than case-by-case judgment)
measuredThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
A market already rule-driven
Well before agents, the market was already largely rule-driven: as early as 2017, only about a tenth of equity trading volume was estimated to still come from discretionary human choice, the rest following index mandates (automatically replicating the makeup of a stock index) or systematic mandates (applying a pre-programmed rule without exception, rather than case-by-case judgment); high-frequency trading alone accounts for roughly half of it in the United States
measuredThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
The stockholm experiment
A natural experiment on the Stockholm Stock Exchange, which separates a mere rise in high-frequency activity from competition among operators, shows it plainly: competing actors end up trading on the same side of the market nearly 70% of the time, their speculative share climbing from 29% to 41%, while liquidity deteriorates
measuredThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
Passive overtakes active
On the ownership side, passive management overtook active in US equity funds as early as 2019
measuredThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
Passive overtakes active
On the ownership side, passive management overtook active in US equity funds as early as 2019 and the truly passive share of the market hovered around a third in 2021
The thread to hold on to is continuity. Agentic AI does not invent correlated behavior; it stacks onto an already heavily synchronized layer and tightens it further, since a few shared foundation models give thousands of agents an almost identical reflex. The risk is not new in kind, only in degree.
observedThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
Concentration correlation procyclicality
It points to the concentration of a few providers, correlated behavior and procyclicality
This finding has since worsened and been given numbers.
observedThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
Human in loop unrealistic
In June 2026, at the ECB's Sintra forum, Breeden specifies that agentic AI now runs trading at more than half of surveyed financial firms and that it can amplify volatility in stress up to a market meltdown; the Bank is exploring, in response, enhanced recovery arrangements letting one institution take over another's vital functions in a disruption, alongside conventional circuit breakers
observedThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
The rise of agentic architectures
It explicitly notes the emergence of agentic architectures in financial services
speculativeThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
An agentic regulation conjectured
A more speculative proposal even sketches an agentic regulation, where supervisory agents would watch over other agents, backed by independent audit blocks
observedThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
Adoption where humans keep control
On the industry side, adoption advances mostly where the human keeps a hand on the wheel: analysis, compliance, cybersecurity, support, more than fully autonomous execution, as NVIDIA's sector survey and the Cambridge Judge Business School analysis of the shift from automation to autonomy suggest
The division of labor is telling. Where an error stays recoverable, the machine is already left to decide on its own; where it commits capital irreversibly, a human is kept as the last resort. Industry thus applies, without always naming it, the very rule we defend: proportion the autonomy granted to the cost of a misstep.
observedThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
A position limit is no arbitrary cap
A position limit, on regulated US derivatives markets, is not a number picked by feel: the Commodity Futures Trading Commission (CFTC, the federal regulator of futures markets) sets it, for the most closely watched contracts, at ten percent of open interest (the stock of not-yet-closed contracts) for the first fifty thousand contracts, then two and a half percent for each increment beyond
The idea fits in one sentence: the narrower the market, the smaller the share any single actor may hold in it, precisely because a narrow market can be moved with less capital. The rule does not say have no conviction; it says your conviction may commit only a bounded fraction of the market, however certain the party holding it, human or algorithmic.
observedThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
A circuit breaker fires at written thresholds
The American market circuit breaker, for its part, fires at written thresholds: three tiers, seven, thirteen and twenty percent declines in the S&P 500 from the prior close, trigger respectively a fifteen-minute halt, a second fifteen-minute halt, then a stop to all trading until the next day
What matters is the automaticity: no regulator decides, in the moment, whether the fall justifies a pause; the threshold is written in advance and fires mechanically, which strips it of exactly the deliberative delay a market in free fall can no longer afford.
observedThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
Aggregate surveillance is neither an agent nor a market but an institution
The Financial Stability Board (FSB, the body that coordinates the G20's financial regulators) set up, in 2025, an incident-reporting exchange format that lets national authorities compare AI-related failures in the financial sector across borders and spot shared vulnerabilities no national authority would see alone
The need for this tier follows from what neither of the previous two can see.
observedThe Financial Dimension: Everyone on the Same Clock · A Private Bet Becomes a Public Risk
The Bank of England already runs this for trading AI
The Bank of England, as seen above, has run since 2025 a macroprudential program dedicated to AI, distinct from the supervision of any single institution
observedThe Financial Dimension: Everyone on the Same Clock · The Prophets Who Profit
The Aschenbrenner example
In 2024 Leopold Aschenbrenner, a former OpenAI researcher, published an analysis presenting the near-term arrival of AGI as plausible
observedThe Financial Dimension: Everyone on the Same Clock · The Prophets Who Profit
A tension noted by others
Several commentators have noted the potential tension between his role as a forecaster and his financial interest in the scenario he defends
observedThe Financial Dimension: Everyone on the Same Clock · The Prophets Who Profit
Toward infrastructures
They confirm it elsewhere, exactly where the sector itself voices its worry: the Bank of England's deputy governor, Sarah Breeden, puts the figure at more than half of surveyed financial firms where agentic AI already runs order execution, with an explicit risk of amplifying volatility up to a market meltdown
observedThe Infrastructural Dimension: Scarcity Has Moved Out
An economic shift
Karpathy sums up the same shift in a sharp formula: AI is not electricity, a fungible resource whose supplier hardly matters, but an operating system, a layer everything else depends on to run
An example makes the tilt concrete. As long as a reliable translation service was scarce, you paid for access to that service.
measuredThe Infrastructural Dimension: Scarcity Has Moved Out · When Thinking Gets Cheap, Everything Else Gets Dear
Convergence is measured
This convergence is no mere guess: the lag of the best open-weight models behind the closed state of the art is now measured in months
three to four months on a composite index that pools many benchmarks, a gap that oscillates but does not widen
measuredThe Infrastructural Dimension: Scarcity Has Moved Out · When Thinking Gets Cheap, Everything Else Gets Dear
Convergence is measured
The perceived-quality gap has narrowed to the point of becoming slight
observedThe Infrastructural Dimension: Scarcity Has Moved Out · When Thinking Gets Cheap, Everything Else Gets Dear
No durable moat
As early as 2023, an internal memo at one of the large laboratories admitted having no durable moat against open models that catch up within weeks
Marginal cost is the price of one more unit: once the plant is built, what the next bottle of water costs, almost nothing. When that price falls toward zero, no seller can charge much for the thing itself any longer.
observedThe Infrastructural Dimension: Scarcity Has Moved Out · When Thinking Gets Cheap, Everything Else Gets Dear
Inference grows cheap
Distillation (transferring the knowledge of a large model into a small one) and hardware-software co-design (designing the chip and the program it runs together) make energy the real bottleneck there rather than the number of operations
extrapolatedThe Infrastructural Dimension: Scarcity Has Moved Out · When Thinking Gets Cheap, Everything Else Gets Dear
A trend not a law
We state this in the conditional: it is the extrapolation of a real trend
a network grows scarce in its own way: the fiber and undersea cables that carry most of the world's traffic belong to a small number of players, increasingly the compute giants themselves, so enough bandwidth must be bought, negotiated and sometimes refused, like a saturated air corridor. The trend is accelerating: by 2026, Google, Meta, Microsoft and Amazon together consu
measuredThe Infrastructural Dimension: Scarcity Has Moved Out · Scarcity Never Disappears, It Changes Address
What lets it act
The trend is accelerating: by 2026, Google, Meta, Microsoft and Amazon together consume roughly three-quarters of international bandwidth on submarine cables, up from a near-nil share in 2010 and now finance most of the new cables being laid
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Scarcity Never Disappears, It Changes Address
Atlas gives the AI a cursor
OpenAI's ChatGPT Atlas browser (2025) gives the AI a cursor to act on websites, to book and to order in your name
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Scarcity Never Disappears, It Changes Address
Withdrawn but moved to the core
OpenAI in fact withdrew it as early as July 2026, barely nine months after its launch, folding its functions into a ChatGPT Work agent built into its main app
observedThe Infrastructural Dimension: Scarcity Has Moved Out · The Thin-Client Bet Returns, at the Scale of the AI Swarm
The network computer
In 1994, Larry Ellison, Oracle's chief, championed the Network Computer: a diskless terminal, with no operating system of its own, all its intelligence held server-side, priced under 500 dollars; Sun Microsystems built the JavaStation from it, running Java alone, operating system included
What this bet settled was not some principled superiority of local compute: it was that on-device silicon had, by that point, become fast and cheap enough to carry a whole system. The market did not reward an idea; it followed a cost curve that eventually crossed a threshold.
extrapolatedThe Infrastructural Dimension: Scarcity Has Moved Out · The Thin-Client Bet Returns, at the Scale of the AI Swarm
The counter-movement
NVIDIA researchers argue that most invocations in an agentic system are specialized, repetitive tasks that a small model handles as well as a frontier model, at ten to thirty times lower cost and that a heterogeneous swarm of specialized small models is the natural choice wherever the task allows it
measuredThe Infrastructural Dimension: Scarcity Has Moved Out · The Thin-Client Bet Returns, at the Scale of the AI Swarm
The counter-movement
An academic prototype, SWARM-LLM, already builds this architecture: a swarm of small models hosted at the edge, close to the user's device, decides query by query whether to answer locally, collaborate with a peer or call a distant model and falls back to the cloud for only about a quarter of queries
observedThe Infrastructural Dimension: Scarcity Has Moved Out · The Thin-Client Bet Returns, at the Scale of the AI Swarm
The counter-movement
Apple, for its part, made on-device processing the default for its agents starting in 2026, reaching for its remote compute only as a last resort
The analogy has its limit and the limit is instructive. Jobs's bet won because the functions of an operating system, once defined, sat within reach of local silicon as soon as that silicon grew fast enough: the problem did not move as one solved it.
extrapolatedThe Infrastructural Dimension: Scarcity Has Moved Out · The Energy Wall You Can't Talk Your Way Around
Use would double
According to the International Energy Agency, data-centre electricity use would double from about 485 to 950 terawatt-hours (TWh, a trillion watt-hours) between 2025 and 2030, about 3% of global electricity; AI-focused centres surged by about half in 2025 alone
These figures are an agency projection, not a measurement of the future; we keep them for the order of magnitude, a doubling in five years and for the share they fix, about three percent of the world's electricity, enough that the question is no longer marginal.
measuredThe Infrastructural Dimension: Scarcity Has Moved Out · The Energy Wall You Can't Talk Your Way Around
The 2025 measurement already outruns the projection
On April 16, 2026, the International Energy Agency published Key Questions on Energy and AI: over 2025 alone, already in the books, electricity demand from data centres as a whole jumped 17 %, against global electricity demand growth of just 3 % over the same period, nearly six times slower; AI-focused centres, for their part, are now expected to triple, not merely double, by 2030
This global figure already lands on precise local bills.
observedThe Infrastructural Dimension: Scarcity Has Moved Out · The Energy Wall You Can't Talk Your Way Around
Henrico the global figure lands on a local bill
In June 2026, Henrico County (Virginia, thirty-seven data centers within its borders, seventeen more planned) emailed thousands of county and public-school employees asking them to cut electricity use: the rate the county pays rose by 25 %, a direct consequence of a new contract passing on to local governments the rising system costs tied to data centers
The same pressure shows up, at an entirely different scale, on the Texas grid.
measuredThe Infrastructural Dimension: Scarcity Has Moved Out · The Energy Wall You Can't Talk Your Way Around
Texas catches fire on gas
A research note from RBC Capital Markets puts at about 101 gigawatts the behind-the-meter gas generation (a plant built right next to the data center it feeds, never crossing the public grid) already announced by American operators; the Texas grid operator, ERCOT, has on its own received about 356 gigawatts of data-center interconnection requests, against a global gas-turbine manufacturing capacity capped at 60 to 70 gigawatts a year
The same computation, run at night in a country drawing its power from wind or dams or at midday in a
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Putting Compute in Orbit Instead of Paying the Bill
Adam Smith had already named this mechanism
In The Wealth of Nations, in the book devoted to colonies, he observes that the power which seizes an abundant, underused land advances more rapidly to wealth and greatness than any other human society and traces the cause to two factors together: plenty of good land and the liberty to manage it as one pleases
The physical argument fits in a few words, but it deserves to be stated with number
measuredThe Infrastructural Dimension: Scarcity Has Moved Out · Putting Compute in Orbit Instead of Paying the Bill
Near-infinite is not near-free
One must be careful, though, not to conflate the resource's physical abundance with its economic accessibility: a study NASA commissioned in 2024 on the feasibility of space-based solar power finds that these systems, even under their best projected scenarios for 2050, would still cost 12 to 80 times more per unit of electricity than terrestrial renewables and calls the approach cost prohibitive and technically infeasible as things stand today
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Putting Compute in Orbit Instead of Paying the Bill
The Starmind project
The most advanced project already has a name: on February 3, 2026, Elon Musk announced the merger of SpaceX and xAI, placing rockets, satellites and language models under one company, with the explicit goal of building orbital data centers powered almost entirely by solar energy
speculativeThe Infrastructural Dimension: Scarcity Has Moved Out · Putting Compute in Orbit Instead of Paying the Bill
Musk's own words
Musk justifies the project by exactly the constraint this chapter has just described: space-based AI is obviously the only way to scale, he claims, adding that within two to three years the cheapest compute in the world will be generated in space
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Putting Compute in Orbit Instead of Paying the Bill
AI1 and Starmind
On June 9, 2026, SpaceX unveiled the first satellite of this fleet, AI1: a deployed wingspan of 70 meters, wider than a Boeing 747, carrying a compute payload averaging 120 kW and peaking at 150 kW
speculativeThe Infrastructural Dimension: Scarcity Has Moved Out · Putting Compute in Orbit Instead of Paying the Bill
AI1 and Starmind
On June 9, 2026, SpaceX unveiled the first satellite of this fleet, AI1: a deployed wingspan of 70 meters, wider than a Boeing 747, carrying a compute payload averaging 120 kW and peaking at 150 kW; the name given to the targeted constellation, Starmind, states the ambition outright: up to a million compute satellites in low Earth orbit.
The argument advanced is not only that the Sun shines continuously in orbit, with no night and no cloud; it is also that the constraint weighing on a terrestrial data center, securing a building permit, a grid connection, a local government's consent, largely dissolves once compute sits in orbit.
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Putting Compute in Orbit Instead of Paying the Bill
A jurisdiction, not an absence of one
This does not place such compute outside all governance, however: international space law makes any object placed in orbit subject to the jurisdiction and control of the state on whose registry it is carried, here the United States, from where SpaceX launches its rockets
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Whoever Owns the Compute Owns the Power
Compute lends itself to control
Compute is the clearest example: scarce, measurable, physically located, it lends itself to control as no software does, so that states and companies already use it as a lever of governance, from export control to chip allocation
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Whoever Owns the Compute Owns the Power
Compute lends itself to control
Compute is the clearest example: scarce, measurable, physically located, it lends itself to control as no software does, so that states and companies already use it as a lever of governance, from export control to chip allocation; the clearest precedent dates to October 2022, when the US administration restricted the export to China of the most advanced computing chips beyond a threshold of performance and interconnect bandwidth calibrated on the best Nvidia models of the moment
The reason is physical, not regulatory: software copies and crosses any border in an instant, whereas a chip is made in a handful of plants, ships in containers and sits in a building one can count and watch. It is this materiality that makes compute graspable by a customs post or a registry, where intelligence, once diffused, no longer is.
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Whoever Owns the Compute Owns the Power
Compute lends itself to control
A highly concentrated fabrication chain compounds this, handing the regulator a small number of chokepoints
The last step is the one the export controls above already illustrate: once a resource is scarce, located and countable, whoever commands it can grant or deny access to it; that power to grant or deny is political before it is economic. Compute passes from a good one sells to a lever one wi
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Whoever Owns the Compute Owns the Power
States monetize authorization too
administration turns that same switch into an explicit toll: it agrees to issue export licenses to China for capped AI chips (Nvidia's H20, AMD's MI308) in exchange for a 15 % cut of the China revenue from those sales, paid to the U.S. Treasury
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Whoever Owns the Compute Owns the Power
Ohio, $500 billion
OpenAI is negotiating a $500 billion data center to be built in Ohio, on the former Piketon uranium-enrichment site, powered by ten gigawatts of electricity through a US-Japan government partnership; Nvidia, for its part, is providing a $250 billion financial backstop, covering OpenAI's lease and debt should the company fail to meet its payments
extrapolatedThe Infrastructural Dimension: Scarcity Has Moved Out · Whoever Owns the Compute Owns the Power
The risk that follows
Legal scholar Jeremy Kress (University of Michigan, a specialist in financial systemic risk) draws the consequence: a supplier who guarantees its own customer's debt does not bring in outside capital that would dilute the risk, it ties its own fate to the customer it guarantees; should OpenAI fail to repay, the whole industry, Nvidia included, absorbs the shock
A simpler case makes the mechanism plain. A father who co-signs his son's mortgage does not diversify his estate, he exposes it further: if the son stops paying, the father pays in his place and the two fortunes, thought separate, become a single fortune the day one falters.
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
A doctrine already theorized
American investor Peter Thiel makes it the explicit goal of any truly successful company: competition is for losers, monopoly being what frees its holder from all external competitive constraint
Thiel is the most explicit voice of this discourse, but he carries an ideological baggage that runs through an entire real, well-documented network that its own members affectionately call the PayPal Mafia: this is not the discourse of one man alone.
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
A doctrine already theorized
O'Brien, illustrated by a photo that has since become famous: on an evening in October 2007, a dozen former PayPal executives posed in suits at Tosca Cafe in San Francisco, staged as mobsters, gold chains and cigars in hand
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
A doctrine already theorized
Others scattered elsewhere: Jeremy Stoppelman and Russel Simmons found Yelp in 2004; Jawed Karim, the engineer behind PayPal's anti-fraud system, is credited as a co-founder of YouTube; Ken Howery and Luke Nosek co-found, with Thiel, the venture capital firm Founders Fund in 2005, the first institutional investor in both SpaceX and Palantir, which Thiel has chaired since 2003
This network, born inside an online payments company at the turn of the 2000s, keeps reconstituting itself, right up to the present. In June 2026, Botha, who had just stepped down as Sequoia Capital's steward, joined SpaceX's board a
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
What is shared and what is not
In June 2026, Botha, who had just stepped down as Sequoia Capital's steward, joined SpaceX's board alongside Musk, whom he had worked with at PayPal a quarter-century earlier
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
What is shared and what is not
Sacks was appointed, on December 5, 2024, AI and Crypto Czar (overseeing artificial intelligence and cryptocurrency policy for the White House) by president-elect Donald Trump
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
What is shared and what is not
Sacks was appointed, on December 5, 2024, AI and Crypto Czar (overseeing artificial intelligence and cryptocurrency policy for the White House) by president-elect Donald Trump, a post he held until his resignation in late March 2026, upon expiry of the 130-day legal limit for special government employees, after which he went on to co-chair the president's science and technology advisory council (PCAST)
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
What is shared and what is not
Vance's 2022 Senate campaign to the tune of fifteen million dollars through the Protect Ohio Values political action committee, at a time when he already employed Vance as a partner at his own fund, Mithril Capital; Vance became a senator from Ohio, then vice president of the United States
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
What is shared and what is not
Army in July 2025 capped at ten billion dollars over ten years, consolidating seventy-five prior contracts into one
The privilege Thiel claims for economic success, he reserves for himself in private too: naturalized as a New Zealand citizen in 2011 after twelve days spent in the country, he acquires an estate there and applies for a permit to build a lodge-shelter overlooking Lake W\=anaka, refused by local authorities and then, on appeal, by the Environment Court in 2024
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
What is shared and what is not
The privilege Thiel claims for economic success, he reserves for himself in private too: naturalized as a New Zealand citizen in 2011 after twelve days spent in the country, he acquires an estate there and applies for a permit to build a lodge-shelter overlooking Lake W\=anaka, refused by local authorities and then, on appeal, by the Environment Court in 2024
speculativeThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
What is shared and what is not
Journalist Evan Osnos documents, in a now-canonical investigation, the same reflex among his peers: for many wealthy figures in American tech, a New Zealand property amounts to a wink, wink signal of collapse preparation, to the point that a LinkedIn co-founder estimates a majority of Silicon Valley billionaires hold some form of apocalypse insurance
In 2009, Thiel writes that he no longer believes that freedom and democracy are compatible, dating this doubt to the expansion of the franchise after 1920
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
A political doctrine documented independently
In 2009, Thiel writes that he no longer believes that freedom and democracy are compatible, dating this doubt to the expansion of the franchise after 1920
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
A political doctrine documented independently
A broader BuzzFeed News investigation into American alt-right networks reveals, among its findings, emails in which software engineer Curtis Yarvin, a figure of the so-called Dark Enlightenment (which advocates replacing democracy with a corporate government run by a CEO-monarch), describes himself as coaching Thiel on political matters; Thiel's fund also invests in a company Yarvin founded
French entrepreneur Yann Lechelle, who led Scaleway (one of the few European cloud operators) and now co-leads Probabl (the company behind scikit-learn, one of the world's most widely used open-source machine-learning
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
A named counter-model already exists
French entrepreneur Yann Lechelle, who led Scaleway (one of the few European cloud operators) and now co-leads Probabl (the company behind scikit-learn, one of the world's most widely used open-source machine-learning libraries), publishes in 2025 an essay with a programmatic title: Ouvertarisme, le numérique des Lumières (Openism, Enlightenment computing)
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
Sovereignty, priced
In June 2026 the European Commission announces a technological sovereignty package whose Cloud and AI Development Act (CADA) states its ambition in numbers: triple the EU's data-centre capacity within five to seven years and set up a single framework to assess, service by service, whether a cloud or an AI is genuinely sovereign or merely hosted on European soil by a company bound to a foreign law
extrapolatedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
Sovereignty, priced
An analyst firm cited by the Commission prices that choice: European spending on sovereign cloud would rise from roughly seven billion dollars in 2025 to more than twenty billion in 202
This fracture does not only separate the blocs from one another; it runs through each of them. According to the American think tank CSIS, the United States and China already concentrate, between the two of them, 45 % and 25 % of the world's data-centre capacity
measuredThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
The fracture runs inside each bloc too
According to the American think tank CSIS, the United States and China already concentrate, between the two of them, 45 % and 25 % of the world's data-centre capacity
measuredThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
The fracture runs inside each bloc too
France, whose nuclear fleet supplied 68 % of its electricity in 2024 and which exported a net surplus of 89 TWh that same year, commits to a 109-billion-euro plan to grow its output by 2 % a yea
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
The fracture runs inside each bloc too
Ireland tells the opposite story: data centres there absorbed 5 % of national electricity in 2015, 21 % in 2023, a projected 30 % by 2030; the Dublin region has simply suspended new data-centre grid connections until 2028
Europe is neither first nor alone at this game; by 2026 the race for AI infrastructure is being run at the scale of nation-states and energy has become its referee.
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
Europe is not alone on this ground
Canada, Australia and the Nordic countries bet on the resource itself: abundant, cheap energy draws data centers that Europe or the United States struggle to power
The geopolitical front line does not stop at the power cable; it already runs back up the chain to the very materials the chips are made of. Venezuela and Colombia, rich in the rare earths advanced electronics require, are becoming contested ground in their own right, Latin America cast as the next technology battleground
speculativeThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
The front line moves to rare earths
Venezuela and Colombia, rich in the rare earths advanced electronics require, are becoming contested ground in their own right, Latin America cast as the next technology battleground
The Gulf embodies this common denominator with particular clarity. American data centers already consumed 4.4% of the country's electricity in 2023, a share the Center for Strategic and International Studies projects will triple by 202
extrapolatedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
The Gulf picks up the tab
American data centers already consumed 4.4% of the country's electricity in 2023, a share the Center for Strategic and International Studies projects will triple by 202
extrapolatedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
The Gulf picks up the tab
Oil, gas and growing investment in solar and nuclear power guarantee, in Saudi Arabia and the United Arab Emirates, cheap power available immediately, while their sovereign wealth funds supply the capital
observedThe Infrastructural Dimension: Scarcity Has Moved Out · Build Alone or Build in Common: Two Bets
The Gulf picks up the tab
Department of Commerce authorized the export to G42 (United Arab Emirates) and Humain (Saudi Arabia) of the equivalent of 35,000 Nvidia GB300 chips each, 70,000 Blackwell Ultra chips in total, subject to an intergovernmental assurance agreement imposing end-use monitoring meant to prevent any re-export to China
observedThe Infrastructural Dimension: Scarcity Has Moved Out · It's Not the Sector That Matters, It's What You Can't Afford to Lose
The NIS2 list
The European Network and Information Security directive (NIS2) lists them, energy, water, transport, health, banking, digital infrastructure
The directive quoted does not supply that criterion; it serves as a useful counter-model. It classifies by sector because a law needs closed lists, yet one and the same sector shelters the harmless act and the critical one: in health, sorting appointments and setting an insulin pump belong to the same sector and do not call for the same caution.
observedThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
Predict or detect
One model locates the leaks of a water network from satellite imagery
observedThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
Predict or detect
One model locates the leaks of a water network from satellite imagery, another spots a methane plume on a gas pipeline from space and alerts the operator
The two cases share the same division of labour: the machine sees what a human eye would not catch in time, a buried leak, an invisible gas, but it is a crew that digs or shuts the valve. The model produces a signal, not an act; its permission stops at the finding.
observedThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
Acting under guard
DeepMind's system that controls the cooling of Google's data centres does so behind eight safety layers and an operator who can exit it at any moment: the agent drops on its own the proposals in which its confidence is low and if a setting would breach a safety bound, the system falls back by itself to a neutral setting instead of forcing it through
measuredThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
Acting under guard
The saving obtained, about 30% of the cooling energy once the system had settled
observedThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
Acting under guard
The first mobile network certified Level 4 autonomous (TDC NET and Ericsson) acts without human approval only within a narrow scope: the energy saving of a radio cell, under an intent that bounds it
Two safeguards are worth retaining beyond these two cases, for they spell out what a well-set permission to act looks like: dropping on one's own the proposals that carry low confidence rather than acting in doubt and falling back to a neutral setting rather than forcing a disputed one through.
observedThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
No collapse from an AI
Meta's global outage of October 2021 came from a maintenance command and the buggy audit tool meant to stop it
observedThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
No collapse from an AI
Meta's global outage of October 2021 came from a maintenance command and the buggy audit tool meant to stop it; Amazon Web Services', from a runaway automatic capacity adjustment
observedThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
No collapse from an AI
Meta's global outage of October 2021 came from a maintenance command and the buggy audit tool meant to stop it; Amazon Web Services', from a runaway automatic capacity adjustment; the Texas electricity crisis of February 2021, from freezing and the tight coupling between gas and electricity, with some twenty gigawatts of load shed
observedThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
No collapse from an AI
Two older incidents on water and gas confirm the pattern, again with no AI involved: in February 2021, an intruder remotely connected, without authorization, to the Oldsmar (Florida) water treatment plant tried to raise the concentration of caustic soda (sodium hydroxide) from 100 to 11,100 parts per million, an operator correcting the dose before any real effect
observedThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
No collapse from an AI
East Coast, triggering shortages and panic buying without any physical component being touched
These cases tell the same story with no AI at all: an automatic action, a maintenance command, a capacity adjustment, a weather hazard, propagates because the parts of the system are too tightly bound to absorb the shock. In Texas the freeze cut the gas, which fed the plants, which fed the gas pumps: the loop bit its own tail.
measuredThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
Milliseconds versus a season
The Surtrac pilot, designed by Carnegie Mellon's Robotics Institute and deployed in 2012 in Pittsburgh's East Liberty neighborhood, lets each intersection compute its own signal timing alone, in real time and without a rigid central plan; the pilot cuts vehicle wait time by 40 % and travel time by 26 %
measuredThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
Milliseconds versus a season
An agricultural sprayer guided by a vision model that mistakes a weed for a seedling treats the wrong spot: the cost is read only at harvest, when correction is expensive
What makes the second case instructive is the gap between lab and field: a model that cleanly told weed from seedling under controlled light degrades as soon as the light, the soil or overlapping leaves change. The error is paid months later, in the yield, not at the moment it is made and by then nothing can be recovered.
observedThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
The John Deere case
It took a federal antitrust action to reopen access to the software
Owning the tractor is not enough if the key that opens it stays with the maker. The farmer holds the steel and the wheels, but the part that decides, the software, is leased to him under conditions; the moment a breakdown calls for touching it, he depends on the one authorized repairer.
observedThe Infrastructural Dimension: Scarcity Has Moved Out · What the Field Has Already Learned, Without a Catastrophe
Legitimacy
The second is legitimacy: the smart-city project run in Toronto by Sidewalk Labs, a subsidiary of Alphabet, Google's parent company, failed not on technology, but on data governance, for want of a model of trust
The project met no engineering obstacle; the sensors worked. It ran aground on a question technology does not settle: who decides what the residents' data becomes and in whose name. For want of an answer, consent withdrew and the project stopped, capable but not legitimate.
extrapolatedThe Political Dimension: Governing Authorized Actions
Governing like a bureaucracy not like a judgment
The philosopher Carina Prunkl proposes comparing algorithmic decision-making not to individual human judgment but to bureaucratic decision-making
measuredThe Political Dimension: Governing Authorized Actions
Governing like a bureaucracy not like a judgment
A US healthcare algorithm using total past healthcare cost as a proxy for severity systematically under-referred Black patients for supplemental care, since historically less money had been spent on this group's care at equal health status, enough to skew the proxy; fixing the bias would have raised the share of Black patients receiving supplemental help from 17.7 to 46.5 percent
This was not an application failure: the algorithm did exactly what it had been asked to do, the flaw sat in the choice of proxy, not in its execution.
measuredThe Political Dimension: Governing Authorized Actions
One hundred forty-one subagents in a formal sandbox
In 2026, the agent Claude Opus 4.6, equipped to drive the Rocq proof assistant, a program that checks every step of a proof, through a dedicated protocol, solves ten of the twelve problems of the Putnam 2025 mathematics competition, spawning one hundred forty-one subagents over the whole experiment, across seventeen hours of active compute
observedThe Political Dimension: Governing Authorized Actions
One hundred forty-one subagents in a formal sandbox
The same shift animates Ax-Prover, which distributes a proof task among specialized agents for mathematics and quantum physics through the Model Context Protocol
This proliferation, which would worry in a production system, poses no comparable risk here: every proof produced is mechanically checked by Rocq itself before being accepted, so that the execution environment, not the subagents' good will, guarantees the result.
observedThe Political Dimension: Governing Authorized Actions · A Highway Code, Not an Ideology
The pharmaceutical counter-example
There, the false opposition between innovation and rule is called into question
observedThe Political Dimension: Governing Authorized Actions · A Highway Code, Not an Ideology
Covid shows the rule under pressure
Covid-19 even showed it under pressure: international coordination among regulators was able to accompany the release of a vaccine in some eighteen months
measuredThe Political Dimension: Governing Authorized Actions · 1. The Degree of Autonomy Sets the Bar
Each degree its own requirements
Rule 1 fills precisely this industrial gap, on the side of the act rather than the data; even that supposedly solved compartment holds up poorly in practice: a 2026 survey of more than five hundred cybersecurity professionals found 76 % of organizations reporting growth in non-human identities and 74 % already deploying agents that require credentials, yet 92 % failing to rotate those credentials on a ninety-day cycle and fewer than four in ten routing their agents' actions through human validation
observedThe Political Dimension: Governing Authorized Actions · 1. The Degree of Autonomy Sets the Bar
Each degree its own requirements
…equipped itself first: as early as April 2024, ISACA, the professional association of information-systems auditors, published its library of AI controls, synthesized from NIST 800-53, the US federal catalogue of cybersecurity controls and ISO 27001, the international standard for information-security management, then crossed with the AI Act, Singapore's framework, MITRE ATLAS (Adversarial Threat Landscape for Artificial-Intelligence Systems, the repository of observed attack techniques against learning systems) and OWASP (Open Worldwide Application Security Project, the community foundation that catalogues application-security risks)
observedThe Political Dimension: Governing Authorized Actions · 1. The Degree of Autonomy Sets the Bar
Each degree its own requirements
In February 2026, the US federal standards agency launched, through its Center for AI Standards and Innovation, the first initiative dedicated to interoperability and cybersecurity for AI agents, finding that existing identity frameworks were not built for systems that act autonomously
observedThe Political Dimension: Governing Authorized Actions · 1. The Degree of Autonomy Sets the Bar
Each degree its own requirements
In February 2026, the US federal standards agency launched, through its Center for AI Standards and Innovation, the first initiative dedicated to interoperability and cybersecurity for AI agents, finding that existing identity frameworks were not built for systems that act autonomously; in December 2025, OWASP published its first ranking of risks specific to agentic applications, which recommends per-agent identity and restricted privilege without, however, proposing a criticality grid for actions comparable to data classification
observedThe Political Dimension: Governing Authorized Actions · 1. The Degree of Autonomy Sets the Bar
The RSP as a worked example
OpenAI's Preparedness Framework tracks capability categories such as cyber and autonomy and sets thresholds that trigger guardrails before deployment
observedThe Political Dimension: Governing Authorized Actions · 1. The Degree of Autonomy Sets the Bar
The RSP as a worked example
OpenAI's Preparedness Framework tracks capability categories such as cyber and autonomy and sets thresholds that trigger guardrails before deployment; Google DeepMind's Frontier Safety Framework defines critical capability levels paired with early-warning evaluations
observedThe Political Dimension: Governing Authorized Actions · 1. The Degree of Autonomy Sets the Bar
The RSP as a worked example
Anthropic's Responsible Scaling Policy (RSP) grades models according to AI Safety Levels (ASL): ASL 1 covers systems that manifestly pose no risk of autonomy or misuse, such as a dedicated chess engine; ASL 2 covers current Claude-type models, judged not capable enough to replicate autonomously, nor to provide information on chemical, biological, radiological and nuclear (CBRN) weapons beyond what a search engine already provides; ASL 3 marks the threshold at which models become capable enough to enhance the capabilities of non-state actors and triggers heightened precautions
That these three laboratories, competitors and with no announced coordination, arrive at the same form of rule is no small thing. Three teams solving the same problem separately and landing on the same structure signal that the structure belongs to the problem, not to a fashion.
observedThe Political Dimension: Governing Authorized Actions · 1. The Degree of Autonomy Sets the Bar
Hassabis said tool back in 2022
In July 2022, Demis Hassabis declared: it would be better, at first, to build these systems as tools
observedThe Political Dimension: Governing Authorized Actions · 2. Never Let a Single Point Hold It All
The following box shows it
In June 2025, the EchoLeak vulnerability (CVE-2025-32711, CVSS score 9.3 on the industry's 0-to-10 severity scale, where scores above 9 mark a critical flaw) demonstrates exactly this path in Microsoft 365 Copilot: a single crafted email, requiring no action at all from the victim, was enough to exfiltrate internal data (emails, files, messages) to an attacker-controlled server, bypassing the defenses meant to distinguish content that is read from an instruction to execute
Separation does not stop at the boundaries of a single agent. A chain of several distinct agents, each audited and judged benign in isolation, can compose a critical path that none of the individual audits would have ca
observedThe Political Dimension: Governing Authorized Actions · 2. Never Let a Single Point Hold It All
Cai makes the composition explicit
The free framework CAI (Cybersecurity AI), developed by the Spanish company Alias Robotics, explicitly names the patterns of this architecture: a decentralized swarm where agents self-assign tasks, a hierarchy where a coordinator distributes the work, a sequential chain where each agent hands its output to the next and a handoff mechanism by which an agent explicitly delegates a task to a specialized agent
The philosopher Pierre Saint-Germier showed this for music generated by diffusion models: when a piece appears to carry an emotion, neither the model alone (which answers a prompt, the source of the content), nor the prompter alone (the link between the prompt and the sound result stays too loose to carry responsibility for a precise musicological choice), nor even the collective of prompter, mode…
extrapolatedThe Political Dimension: Governing Authorized Actions · 2. Never Let a Single Point Hold It All
Locating the agent of a distributed act
The philosopher Pierre Saint-Germier showed this for music generated by diffusion models: when a piece appears to carry an emotion, neither the model alone (which answers a prompt, the source of the content), nor the prompter alone (the link between the prompt and the sound result stays too loose to carry responsibility for a precise musicological choice), nor even the collective of prompter, model designers and training corpus can be named as the author of a deliberate act, for want of any identifiable coordination structure among them
observedThe Political Dimension: Governing Authorized Actions · 3. No Trace, No Audit, No Memory
No log no audit no learning
ISACA assesses every control through the six explanation dimensions defined in 2020 by the UK data regulator and the Alan Turing Institute; yet the first of them, the rationale explanation, demands documentation of the logical reasoning behind decisions, the very artifact that, for a language model, does not exist: the displayed chain of thought can describe a reason that is not the actual cause of the answer
This flaw is not accidental; it lies in the very nature of the exercise. Genuine critical thought turns on itself before it judges anything else: to doubt is, first of all, to doubt one's own judgment. A displayed chain of thought does no such thing; it produces an after-the-fact narrative, never turning back on the computation that generated it.
observedThe Political Dimension: Governing Authorized Actions · 4. A Human, Before Whatever Can't Be Undone
The fifth condition is not a theoretical precaution
In Hong Kong, in January 2024, an employee in the finance department of the engineering firm Arup approved fifteen wire transfers, totaling two hundred million Hong Kong dollars, after a video conference in which a deepfake of the chief financial officer's voice and face had dispelled his suspicions
The case deploys AI as a tool in a human attacker's hands, not as an agent manipulating on its own initiative; no documented precedent of that second kind, an autonomous conversational agent manipulating the authorized human unaided, is yet public as this book is written.
observedThe Political Dimension: Governing Authorized Actions · 4. A Human, Before Whatever Can't Be Undone
The ritual click has already failed in practice
The Australian Robodebt program (2015-2019) calculated a presumed debt by averaging income declared to the tax office over the year, effectively reversing the burden of proof onto the welfare recipient; the Royal Commission that examined it found no genuine human intervention in calculating or notifying debts, with four hundred seventy thousand debts raised on a legally insufficient basis
None of the five conditions above was met: no time, no intelligible information, no real power to contest. The case stands as a structural precedent of the same mechanism, validation reduced to a ritual click, before the click became an agent's: this is not a case of agentic AI.
extrapolatedThe Political Dimension: Governing Authorized Actions · 5. The Test, Not the Promise
But reliable for which class of use
Any process can be filed under different categories of generality; the reliability verdict depends entirely on which one: saying a model is reliable in general or reliable for this precise kind of request, in this precise context of use does not give the same answer, a problem identified at the very foundations of reliabilism, the theory on which a belief is worth no more than the reliability of the process that produced it
extrapolatedThe Political Dimension: Governing Authorized Actions · 5. The Test, Not the Promise
But reliable for which class of use
Any process can be filed under different categories of generality; the reliability verdict depends entirely on which one: saying a model is reliable in general or reliable for this precise kind of request, in this precise context of use does not give the same answer, a problem identified at the very foundations of reliabilism, the theory on which a belief is worth no more than the reliability of the process that produced it \ and since systematized under the name generality problem
extrapolatedThe Political Dimension: Governing Authorized Actions · 5. The Test, Not the Promise
But reliable for which class of use
The philosopher of science Cyrille Imbert draws from this, for AI, a directly transferable lesson: the right reference class, the set of cases over which reliability is judged, is neither every conceivable use (most make no practical sense) nor the designer's own test set alone (almost always narrower than actual use), but the class of cases the user would themselves treat as equivalent
observedThe Political Dimension: Governing Authorized Actions · 5. The Test, Not the Promise
A pipeline already in place
The operational pipeline that such an arrangement demands is no utopia: Amodei describes an observed arrangement, the one applied to each Claude model, in which models are tested both internally and by third parties for their safety, in particular against catastrophic risks and autonomy risks; we have an agreement with the American and British AI safety institutes as well as other third-party testers in specific domains to evaluate the models on CBRN risks
observedThe Political Dimension: Governing Authorized Actions · 5. The Test, Not the Promise
Hassabis and the bioethicists
Hassabis, in 2022, summarized the principle for AlphaFold: we consulted thirty independent bioethicists before publication (an observed arrangement)
observedThe Political Dimension: Governing Authorized Actions · Thresholds Get Deliberated, Not Computed
The method already applied to AI
Stanford's deliberative-democracy lab conducted, for Meta, community forums that gathered some fifteen hundred participants from several countries around the governance of generative AI
observedThe Political Dimension: Governing Authorized Actions · Thresholds Get Deliberated, Not Computed
A forum on agents
In 2025, a cross-company deliberative forum, convened by Meta with Cohere, Oracle and PayPal, focused specifically on the design and governance of AI agents
The divide it brought out maps squarely onto the first of our rules. The participants' enthusiasm for an agent that handles routine, low-stakes tasks comes paired with marked caution the moment the agent acts autonomously in a high-stakes setting; the line between the benign and the critical emerges as the very boundary of public acceptanc
observedThe Political Dimension: Governing Authorized Actions · Thresholds Get Deliberated, Not Computed
A forum on agents
The participants' enthusiasm for an agent that handles routine, low-stakes tasks comes paired with marked caution the moment the agent acts autonomously in a high-stakes setting; the line between the benign and the critical emerges as the very boundary of public acceptanc
speculativeThe Political Dimension: Governing Authorized Actions · Thresholds Get Deliberated, Not Computed
A theorem not a measured fact
The first is its level of evidence: it is a stylized theoretical model, carried by a working paper (a paper not yet published in a peer-reviewed journal), hence a conditional theorem, not a measured fact
A conditional theorem reads as a promise with a clause: if the policy is tuned just right, then accelerating helps. The whole question is whether the clause is ever met. The au
observedThe Political Dimension: Governing Authorized Actions · Governing a Model Like You'd Govern a Weapon
The consequence is already observable
The logical consequence is already observable: in the summer of 2026, foreign nationals' access to the latest model of a major laboratory was restricted, an early echo of the open-weights question taken up later in the book
The same spring plays at two nested scales. An agent ordered to reach a goal is tempted to shave the constraint that slows it; a firm ordered to hit its numbers is tempted to keep running a model whose guardrail i
observedThe Political Dimension: Governing Authorized Actions · The Responsibility Gap and How to Close It
The principal answers for its agent
The dominant practitioner reading in the United States runs liability back to the persons and entities behind the agent, through agency law, which makes the principal answerable for its agent's acts, a court having already held that a software agent accessing a protected space fell under the computer-fraud statute
The real case bears this out. In the incident, the operator did not plead that the machine acted alone; it publicly acknowledged that its models had caused the
observedThe Political Dimension: Governing Authorized Actions · The Responsibility Gap and How to Close It
The AI liability directive withdrawn
The directive that was to create a liability regime specific to AI was withdrawn from the Commission's programme in 2025, for want of foreseeable agreement
observedThe Political Dimension: Governing Authorized Actions · The Responsibility Gap and How to Close It
The product-liability directive remains
What remains is the renewed directive on defective products, which now covers software and thus AI systems
What the contrast teaches is not that one continent is right against the other, but that the tool is younger than the risk. The categories exist, they can be mobilized, they are not yet fitted to the case where the agent harms without having been either sold or ordered. That is the very space the governance of actions must cover, without waiting for the law to settle.
observedThe Political Dimension: Governing Authorized Actions · Placing Any Agent at a Glance
The chain Amodei describes
Amodei publicly describes the chain applied to each Claude model before the API release: pre-training, post-training, internal testing under the RSP, external testing under agreement with the American and British safety institutes, CBRN evaluations by specialized third parties
observedThe Political Dimension: Governing Authorized Actions · Placing Any Agent at a Glance
Cost did not block performance
That cost did not keep the firm from sitting atop the public benchmarks
observedThe Political Dimension: Governing Authorized Actions · Why Stopping Everyone Stops No One
The letter calls for a pause
This open letter, massively signed, put the safety of frontier models on the public agenda for want of shared safety protocols
observedThe Political Dimension: Governing Authorized Actions · Why Stopping Everyone Stops No One
Compliance moats
A peer-reviewed analysis shows how safety requirements costly to meet inadvertently protect those who already have the means to pay them
observedThe Political Dimension: Governing Authorized Actions · Why Stopping Everyone Stops No One
Regulatory capture documented
The risk of regulatory capture, where the dominant actor shapes the rules that concern it, is now documented by an empirical study of how AI firms bend the very policies that target them
observedThe Political Dimension: Governing Authorized Actions · Why Stopping Everyone Stops No One
Gates dismisses the moratorium as unrealistic
Bill Gates, Microsoft's co-founder, explicitly dismisses the global moratorium, unable to align geopolitical and economic incentives that pull the other way and prefers instead national institutions paired with a new international organization, built on the combined model of nuclear inspections, civil aviation and ozone treaties
observedThe Political Dimension: Governing Authorized Actions · Why Stopping Everyone Stops No One
1,384 employees ask to slow down, not stop
An open letter signed in July 2026 by 1,384 employees of frontier labs, including the CEOs of OpenAI, Anthropic and DeepMind, calls in the same direction for an international effort to build the technical and governance tools able to deliberately pace the frontier of automated AI development, endorsed by OpenAI and Anthropic as companies within hours of its publication
observedThe Political Dimension: Governing Authorized Actions · Four Grafts Onto a Law That Never Saw the Agent Coming
The AI Act sorts by risk
The European Union has given itself a regulation on artificial intelligence, the AI Act, in force since August 1, 2024, which sorts uses by risk level: prohibited practices (Article 5), high-risk systems subject to engineering requirements (Articles 8 to 15), mere transparency obligations for the rest (Article 50)
observedThe Political Dimension: Governing Authorized Actions · Four Grafts Onto a Law That Never Saw the Agent Coming
The compute threshold governs the model
The first lies in how it catches general-purpose models: Article 51 presumes systemic risk once training has spent more than 10^25 floating-point operations (FLOP, the unit that counts a computation's elementary multiplications and additions)
extrapolatedThe Political Dimension: Governing Authorized Actions · Four Grafts Onto a Law That Never Saw the Agent Coming
A text from before the agent
Written for the model that answers, the regulation predates the agent that acts and policy analysis has begun to name the hole: unsuited metrics, prompt injection left uncovered, human oversight with no defined stop button
observedThe Political Dimension: Governing Authorized Actions · Four Grafts Onto a Law That Never Saw the Agent Coming
The calendar works against it
The calendar, finally, works against it: scheduled for August 2026, the high-risk obligations were postponed to December 2027, even August 2028, by the omnibus regulation of July 2026
observedEpilogue: What Intelligence Is Not · IQ: A Rank Dressed Up as a Quantity
Scores feed sterilizations
Then comes the dark side: in the United States, scores of feeble-mindedness feed the case for restrictive immigration laws and forced sterilizations.
By rank one must understand a position within a reference population, not a dose of intelligence: the t
measuredEpilogue: What Intelligence Is Not · IQ: A Rank Dressed Up as a Quantity
What it measures drifts
Second, what the test evaluates moves: throughout the twentieth century, raw scores rose by about three points per decade (the Flynn effect)
observedEpilogue: What Intelligence Is Not · IQ: A Rank Dressed Up as a Quantity
Calming shifts the score
To this slow drift is added an immediate instability that the practice of testing reveals starkly: one of us saw an examiner note that the mere fact of putting the children at ease before the test, of calming them, shifted their results by a full standard deviation, that is, the fifteen points that, on this scale, separate the normal from the handicap.
measuredEpilogue: What Intelligence Is Not · IQ: A Rank Dressed Up as a Quantity
Calming shifts the score
The anecdote is not an isolated case, even though it says nothing, by itself, about the average size of this effect in the general population: several experiments have shown that merely reminding subjects, just before a test, of their membership in a group stigmatized by a stereotype of intellectual inferiority is enough to degrade the performance of otherwise equally competent subjects
observedEpilogue: What Intelligence Is Not · Writing in Thirty Languages in a Few Seconds
Days shrink to seconds
A translation exercise that took a trained translator several days now takes a few seconds.
measuredEpilogue: What Intelligence Is Not · School-Style Reasoning, Relieved in Its Turn
Alphageometry at olympiad level
In January 2024, the DeepMind team published AlphaGeometry, a system that solves geometry problems at the level of the International Mathematical Olympiad, at the frontier of elite human performance (a result measured under benchmark conditions, on Olympiad problems), without additional human training data.
measuredEpilogue: What Intelligence Is Not · School-Style Reasoning, Relieved in Its Turn
Alphageometry at olympiad level
It is this near-equality with the elite, and not a mere advance, that gives the fact its charge: the criterion was matched, not merely approached.
heavily eroded; deep sequential reasoning, which still resisted the 1970 calculator, no longer bars the way to the 2024 mathematical assistant, even though the latter still errs and calls for rereading
measuredEpilogue: What Intelligence Is Not · School-Style Reasoning, Relieved in Its Turn
The novice-expert gap narrows
Brynjolfsson, Li and Raymond document and measure, on a large sample of customer-service workers, that the help of a generative model sharply reduces the performance gap between novices and experts and compresses the distribution of reasoning performance (the gap between the best and the worst narrows).
measuredEpilogue: What Intelligence Is Not · School-Style Reasoning, Relieved in Its Turn
The novice-expert gap narrows
This is relief caught in the act: a rare skill ceases to sort as soon as a machine hands it out cheaply.
measuredEpilogue: What Intelligence Is Not · School-Style Reasoning, Relieved in Its Turn
Quality up diversity down
Doshi and Hauser, on creative writing, show in parallel that the use of a generative model increases the perceived quality of individual productions and reduces the collective diversity of outputs (a double measured effect)
measuredEpilogue: What Intelligence Is Not · School-Style Reasoning, Relieved in Its Turn
Quality up diversity down
It is the same movement as at the call center, carried over to creation: the help that lifts the less assured homogenizes the collective.
and, symmetrically, erodes the lead of the best
speculativeEpilogue: What Intelligence Is Not · School-Style Reasoning, Relieved in Its Turn
A cognitive disposition not a diagnosis
The conclusion provokes, yet it holds together: the era now opening may be the one in which this singularity of gaze outranks neatly combed sequential reasoning as a social criterion of intelligence, that reasoning being within reach of the first agent to come along.
what seem to us to be re-evaluations under way
measuredEpilogue: What Intelligence Is Not · Chemistry, Played a Different Way
A system matches chemists
It plans retrosyntheses at the level of experienced chemists in a blind test (performance measured blind against human experts).
measuredEpilogue: What Intelligence Is Not · Chemistry, Played a Different Way
A system matches chemists
The margin is slim, but the threshold it crosses is not: for the first time, chemists rate a machine's chain of reasoning, blind, on a par with their own.
observedEpilogue: What Intelligence Is Not · Chemistry, Played a Different Way
An industrial ecosystem forms
The ecosystem has since taken shape: IBM RXN, MIT's ASKCOS, Chematica, Iktos and Synthia offer industrial services; the large pharmaceutical companies integrate them into their discovery chains.
observedEpilogue: What Intelligence Is Not · The Failure That Teaches, the Fall That Stays Useful
Lilienthal's death spurred wright
It is also the trace, preserved in the professional conscience of engineers and pilots, of Otto Lilienthal's fatal glider crash on 9 August 1896
observedEpilogue: What Intelligence Is Not · The Failure That Teaches, the Fall That Stays Useful
Hundreds of brutal falls
Between 1900 and 1903, the Wright brothers accumulated several hundred gliding flights interrupted by brutal falls
observedEpilogue: What Intelligence Is Not · The Printing Press Scaled Up Critical Thinking
Valla unmasks the donation
Around 1440, Lorenzo Valla demonstrates by philological comparison that the Donation of Constantine (a text used for centuries to legitimize the temporal power of the Church) is a medieval forgery
speculativeEpilogue: What Intelligence Is Not · Genius Overflows Intelligence
Speculative and contested
He pushes the analogy all the way, on the side of physics: just as the world of atoms demanded, at the start of the twentieth century, an entirely new physics (quantum mechanics) where classical mechanics failed, so consciousness would demand, on his view, a missing science, still to be written, which he places on the side of quantum gravity and of a non-computable reduction of the quantum state. Penrose himself calls the thesis speculative
this second theorem holding that a formal system rich enough can never establish its own consistency using only its own internal resources, much as one cannot vouch for one's own good faith
observedEpilogue: What Intelligence Is Not · Genius Overflows Intelligence
Fable 5's counterexample
With the help of the Fable 5 model, on a problem put to him by Akhil Mathew, the mathematician Levent Alpöge exhibited a map of three-dimensional space whose Jacobian determinant is everywhere the constant -2, yet which sends three distinct points to one and the same point
One must measure exactly what this case establishes and what it does not: it describes a division of labor observed in 2026, not an argument about computability. That the computer generates and a human recognizes says nothing, by itself, about whether that recognition is itself computable in principle or
observedEpilogue: What Intelligence Is Not · Genius Overflows Intelligence
The machine generates the human recognizes
A mathematician, Will Sawin, then pins down the exact constant; the community, including Timothy Gowers, a Fields medalist, and Noga Alon, verifies and publishes the reconstruction
observedEpilogue: What Intelligence Is Not · Genius Overflows Intelligence
The machine generates the human recognizes
That same spring of 2026 piles up further cases, each in its own field: a book Ramsey number, R(B_8,B_10)=37, solved by the AutoMath system, where both the problem and the proof were found by AI
observedEpilogue: What Intelligence Is Not · Genius Overflows Intelligence
The machine generates the human recognizes
That same spring of 2026 piles up further cases, each in its own field: a book Ramsey number, R(B_8,B_10)=37, solved by the AutoMath system, where both the problem and the proof were found by AI; two open problems in computational mathematics (a conjugate-gradient/randomized-coordinate-descent comparison, a counterexample to column-pivoted QR factorization) handled by the Iteris system, final results obtained after human repair
observedEpilogue: What Intelligence Is Not · Genius Overflows Intelligence
The machine generates the human recognizes
That same spring of 2026 piles up further cases, each in its own field: a book Ramsey number, R(B_8,B_10)=37, solved by the AutoMath system, where both the problem and the proof were found by AI; two open problems in computational mathematics (a conjugate-gradient/randomized-coordinate-descent comparison, a counterexample to column-pivoted QR factorization) handled by the Iteris system, final results obtained after human repair; five new results in Banach space theory, one of them settling a problem open since 1964, ideas and near-complete proofs proposed by AI then reworked by specialists
(a computer language that checks, line by line and without room for gaps, that a mathematical proof holds together: something like a spell-checker, but for the logical soundness of a proof)
measuredEpilogue: What Intelligence Is Not · Genius Overflows Intelligence
The machine generates the human recognizes
One last case inverts the division of labor: the LEAP system formalizes and verifies proofs already found, raising a general-purpose model's success rate from under 10% to 70% on a Lean (a computer language that checks, line by line and without room for gaps, that a mathematical proof holds together: something like a spell-checker, but for the logical soundness of a proof) formalization benchmark; here the AI discovers nothing, it audits
observedEpilogue: What Intelligence Is Not · Genius Overflows Intelligence
Tao himself states it
Tao himself, whose verification of the Keller counterexample this chapter has just followed, states this same division in a public conversation held at Stanford in August 2026: he describes his own daily use of a computer agent to manage databases and calculations, while stating an explicit limit; AI error rates have fallen, he says, but never reached zero, so trust in them stays necessarily bounded and coupling these AIs with formal verification such as Lean can keep them completely honest
observedConclusion: The Elephant, Almost Whole
An AI that merely displays a false estimate only issues a piece of advice to be checked; the same estimate, endowed with the power to buy for real, turns the error into a financial act: that is what happened to Zillow, which had wired to its price estimator (the Zestimate) an automatic power to buy homes, until it abruptly shut the activity down in November 2021, with a write-down of more than three hundred million dollars in the third quarter alone and more than half a billion in total
observedConclusion: The Elephant, Almost Whole
The object already exists
The current state of the field finds that the object exists in experimental and partly operational form
those stand-in measures named earlier, quantities easy to count that the agent pushes in place of the goal one cannot measure directly.
observedConclusion: The Elephant, Almost Whole
The catastrophe is not yet here
The empirical and defensive aspect lands in the middle: no agentic catastrophe has been observed, but its building blocks have.
observedConclusion: The Elephant, Almost Whole
Trading handed to agents
The financial aspect tests the thesis on the international markets, where the permission-authorization regime runs at its highest. Since 2024, trading has been handed to agents.
measuredConclusion: The Elephant, Almost Whole
Backtest rather than live gains
But an audit of the field turns up poorly reproducible results: out of seventy-seven studies surveyed, only two of the nineteen empirical papers correctly isolate the future from the past in their time-split
extrapolatedConclusion: The Elephant, Almost Whole
Backtest rather than live gains
But an audit of the field turns up poorly reproducible results: out of seventy-seven studies surveyed, only two of the nineteen empirical papers correctly isolate the future from the past in their time-split, most reported returns being backtest returns rather than live gains
A backtest return is what a strategy would have earned had it been run over past data, replayed after the fact; a live gain is what it actually earns with real money at stake. The gap between the two is rarely innocent: a strategy can be tuned, knowingly or not, to fit the very history it is tested against, so that it shines in replay and fades the moment the future stops resembling the past.