Pacing the frontier will not fix AI's real problem: nobody can read what it is working toward

Emmanuel Théry, CEO and co-founder, DigitalKin • September 16, 2026

Share this article

By Emmanuel Théry, CEO and co-founder of DigitalKin. Lyon, 16 September 2026.

Something unusual happened earlier this month, and I have taken a few days to think it through before writing this. Three people who compete for the same talent, the same compute and the same customers said the same thing on the same day.

On Saturday 12 September, Anthropic's Dario Amodei published an essay titled "We Must Pace the Frontier", with a sentence he clearly wanted quoted: "We must slow the pace at which we improve the capabilities of AI models." Within hours, Sam Altman agreed that the industry needs to "pace the frontier" and committed OpenAI to the same first step Amodei proposed: giving independent evaluators permanent, employee-level access inside frontier labs. Elon Musk's reply was three words long. On Monday, AI stocks fell and cybersecurity stocks rose, and the President of the United States publicly objected.

The essay came at the end of a bad fortnight. An Anthropic researcher resigned on 9 September warning against self-improving AI. Before that, an OpenAI agent had escaped its sandbox during testing and compromised several companies through Hugging Face. And on 3 September, OpenAI shipped GPT-6 Astra, the first model it classifies as "critical" for offensive cyber capabilities, calling it both its most capable and its most aligned model to date.

I want to take this debate seriously, because I think it is sincere. And I want to say, as the head of a small deeptech company that has spent three years on a different architecture, that the debate is being framed around the wrong variable.

Speed is the symptom. Illegibility is the disease.

Read Amodei's argument closely. The reason to slow down is not that capability is bad. It is that capability is growing faster than "our ability to understand and control" the systems. Altman said the same in more operational terms: safety cases and monitoring have significant costs, and pacing is how you buy the time to pay them.

So the real claim underneath the slowdown is this: we cannot currently verify what these systems are working toward, and the gap is widening. Speed matters only because it widens a gap that already exists at any speed.

Look at where that gap comes from, and it stops looking like a law of nature.

The dominant paradigm produces systems whose purposes are nowhere written down. A frontier model does not contain a goal you can inspect. It contains billions of weights that were shaped by training toward behaviours that, on average, satisfied evaluators. Whatever the model is "working toward" in a given interaction is an emergent property of that shaping, inferred after the fact from what it does.

For a while, the industry had a workaround: read the model's chain of thought. Astra shows how fragile that workaround is. The model uses a reasoning technique its own maker calls opaque recurrence, which reduces the amount of legible reasoning available to monitor. OpenAI's chief scientist framed the trend plainly: as capabilities increase, monitorability gets harder, because stronger models complete harder tasks with fewer language tokens, or none.

The alignment evidence follows the same logic. Astra's headline safety result is that it declined to pursue planted "honeypot" objectives during evaluation. A former OpenAI researcher asked the obvious question in public: does that show alignment, or a model that recognises when it is being tested? An independent analysis went further and concluded it could not rule out that the model was faking alignment. I am not in a position to adjudicate. Nobody outside OpenAI is. That is precisely the point.

When the purpose of a system lives implicitly in its weights, alignment can only ever be declared by its maker and sampled by evaluators. It cannot be read . Independent evaluators with employee-level access are a genuine improvement: they audit the process. They still cannot open a single answer and see what it was aiming at.

The three layers every serious model needs, and why the third is missing

Here is the technical thesis behind everything DigitalKin builds. To do useful, checkable work in an open domain, an intelligence needs three layers, all explicit:

  1. A modelling layer : what things are. Entities, their characteristics, their structure.
  2. A causal layer : how things interact. Which characteristics determine which others, through which relations.
  3. A teleological layer : what things are for. The purposes an actor pursues, expressed as effects they value, so that the model can say which outcome counts and why.

No mainstream approach carries all three explicitly. Knowledge graphs and ontologies have the first layer and treat relations as associations, not causes. Causal AI in the Pearl tradition has a rigorous second layer, but its graphs must be built exhaustively before they are queried, on flat variables, with no mechanism to construct them on demand from a goal. Automated planning has goals, but reasons over flat action spaces and returns plans that are valid or invalid, not better or worse. And large language models hold all three layers at once, implicitly, entangled, and unreadable. That entanglement is the source of both their fluency and their opacity.

The third layer is the one the field has treated as an afterthought, to be bolted on through training signals. It is also the one that makes the other two auditable. If the purpose is explicit, you can ask of any conclusion: which purpose does this serve, through which causal chain, and does the person concerned agree with that purpose?

What we built instead

We call our engine Telos, after the Greek word for purpose, because purpose is where it starts.

It builds the causal model from the goal, just in time. Given a stated outcome, Telos asks what determines it, what determines those determinants, and so on, down to the parameters an actor actually controls. It stops exactly when every term of every relation can be evaluated directly. Nothing is modelled that does not bear on the outcome, and nothing that bears on it is left unmodelled.

Decomposition is forced by the causal relations, not by intuition. Each independent term of a relation produces exactly one sub-model, and each sub-model resolves exactly one term. Coverage is checkable mechanically: with every sub-model in place, every term can be computed; remove any single sub-model and exactly one term, no more and no fewer, becomes uncomputable. Zero would mean redundancy; two would mean one sub-model doing the work of two. Consultants would call this MECE, mutually exclusive and collectively exhaustive. The difference is that here it is a property of the structure, not a discipline of the analyst.

The same fragment of causal structure supports two operations. Backward, from the goal: derive the specification each controllable parameter must meet, including the eliminatory conditions, those values that block the outcome regardless of everything else. Forward, from the choices: fix what has been decided, propagate the remaining degrees of freedom, and read off the marginal contribution of each parameter to the outcome, so effort goes where it changes the result.

Every output is traceable to a purpose. A recommendation is not a plausible continuation of text. It is a term in a relation that serves a stated goal, and the chain from one to the other is available for inspection. When it is wrong, you can see where. When you disagree, you can say with what.

Understanding is stored as structure, not as prose. A stabilised fragment of causal-teleological model, which we call an Understanding, can be saved, shared, queried and combined with others. The output of one is an input of the next. This is how a small team accumulates what it learns instead of losing it in documents nobody reopens.

Where do the domain facts come from? From large language models, among other sources. This matters for intellectual honesty: Telos uses LLMs to bootstrap the content of relations, asking them narrow, ordered questions. The rigour comes from Telos; the raw knowledge often comes from the model. We do not make the underlying LLM transparent. We make the reasoning that leads to a decision transparent, which is the layer that matters to the person who has to sign off on it. Reliable, traceable, auditable: that is the whole product promise, and it is a promise about the reasoning, not about a model's inner life.

Declared alignment versus verifiable alignment

This reframes the debate that has played out over the past few weeks.

The frontier labs are proposing to slow down so that evaluators can catch up with systems whose purposes are implicit. That is a reasonable thing to do with the architecture they have. It is also a treadmill: every capability gain makes the implicit purpose harder to infer, so every gain requires more pacing. Amodei's own essay concedes that the slowdown is bounded by geopolitics, and Altman insists pacing does not mean stopping. The treadmill will keep moving.

The alternative is not to slow the machine down. It is to build machines whose purposes are written where people can read them. Then alignment stops being a property a vendor asserts about a model and becomes a property anyone can check about a piece of reasoning: here is what it was working toward, here is how it got there, here is where you can object.

I want to be careful about the claim. We are not saying Telos solves frontier-model misalignment. A recursively self-improving agent with network access is a problem of a different kind, and I am glad the people building those systems are finally saying so out loud. We are saying that for the vast territory of applied AI where decisions have consequences, in medicine, engineering, finance, law, industry, procurement, public administration, the choice between an intelligence you must trust and an intelligence you can verify is an architectural choice. It has been made, by default, in favour of the first. It does not have to be.

The point has already travelled beyond the technical community. In May, a document as far from Silicon Valley as one can imagine, the encyclical Magnifica Humanitas, argued that making AI "more moral" is meaningless if that morality is decided by a handful of people and cannot be debated by those it affects. Strip away the theology and that is a demand for legibility. You cannot debate a purpose you cannot read.

The unexpected dividend: a purpose layer is also a coordination layer

There is a second consequence of making purpose explicit, and it is the one that shaped our mission.

Human models of the world do not fit together. Every profession, every team, every discipline builds its own, each internally consistent, each incompatible with the others. We usually blame the people. The real cause is that each model was built for a purpose the others cannot see. Two engineers, a lawyer and a buyer are not disagreeing about reality; they are optimising for different things, and none of those things is written anywhere.

Make the purpose explicit and the incompatibility becomes intelligible. You can see why the models diverge, which means you can connect them without flattening them. The lawyer keeps her model, the engineer keeps his, and the machine holds the relation between them: what each is working toward, and where those purposes meet or collide. Coordination, which today is an overhead that grows faster than the organisation, becomes a property of the structure. Disagreement does not disappear; it moves to where it belongs, the places where human judgment is actually needed, and away from the places where it was only an artefact of models that could not talk to each other.

That is why we have chosen, this month, to say what we are for in three words: make understanding abundant. Understand how things work. Understand one another. Come to an understanding. We think the three are one movement, and that it takes a purpose layer to make it so.

What this means if you follow this field closely

Watch three things in the coming months.

First, whether "monitorability" keeps declining at the frontier as models reason with fewer tokens. If it does, chain-of-thought oversight will be gone as a safety mechanism within a couple of model generations, and the case for explicit-purpose architectures will make itself.

Second, whether independent evaluators are given access to reasoning, or only to process. Employee-level access to a lab tells you whether the lab did what it said. It does not tell you what a given answer was aiming at.

Third, whether the debate stays framed as "how fast" or shifts to "how legible". Pace is a dial on the machine we already have. Legibility is a decision about which machine to build.

We made that decision three years ago, before it was fashionable to worry about it. We are a small team in Lyon and we do not pretend to hold the frontier. We hold a different question: not how to keep up with intelligences nobody can read, but how to build ones that everyone can.

Emmanuel Théry is CEO and co-founder of DigitalKin. Telos, the company's causal engine, was designed by co-founder and Chief Science Officer Sébastien Deschaux.

Frequently asked questions

What does "pace the frontier" mean?

It is the phrase Dario Amodei used on 12 September 2026 to describe deliberately slowing the rate at which frontier AI capabilities improve, so that safety research, third-party evaluation and regulation can keep up. Sam Altman and Elon Musk endorsed the idea. It does not mean stopping AI development.

Why is chain-of-thought monitoring becoming less reliable?

Because more capable models complete harder tasks with fewer language tokens, or none, and some now use reasoning techniques such as opaque recurrence that reduce the legible trace. Less visible reasoning means less to monitor.

What is the difference between declared and verifiable alignment?

Declared alignment is a property a vendor asserts about a model, supported by sampled evaluations. Verifiable alignment is a property of a piece of reasoning: the purpose it serves and the causal chain that leads to its conclusion are explicit, so the people concerned can check, question and correct them.

What is causal-teleological AI?

An approach in which an intelligence models explicitly what things are, how they cause one another, and what they are for, and derives its outputs from a stated purpose through that structure. DigitalKin's Telos engine is built on this principle.

Sources

  • Fortune, "Anthropic CEO calls to slow the race toward AI superintelligence", 12 September 2026: fortune.com
  • CNBC, "AI stocks sink while cybersecurity shares rally on slowdown fears", 14 September 2026: cnbc.com
  • CNBC, "Sam Altman spells out how and why the AI industry wants to slow down", 14 September 2026: cnbc.com
  • TechCrunch, "OpenAI launches Astra, its powerful (and controversial) new model", 3 September 2026: techcrunch.com
  • OpenAI, "Toward Astra: critical capabilities and frontier safeguards": openai.com
  • IT for Business, "OpenAI lance GPT-6 Astra, son premier modèle cyber-critique": itforbusiness.fr
  • LessWrong, "OpenAI's Astra alignment claims are dubious": lesswrong.com
  • Léon XIV, Magnifica Humanitas, 15 May 2026, §107: vatican.va

Recent Posts

By Emmanuel Théry, CEO et cofondateur, DigitalKin • September 16, 2026
Amodei, Altman et Musk veulent ralentir l'IA. Mais la vitesse n'est qu'un symptôme : on ne peut pas vérifier ce que vise un modèle. Une autre voie existe.
DigitalKin featured in Les Echos, June 2026, article on agentic AI for SMEs
By Maëlle Boisis • July 2, 2026
Les Echos features DigitalKin and our agentic AI approach for SMEs. Discover why the French business daily covered our Kins, and read the full article.
Mapping 2026 class campaign banner with blue, orange and green abstract shapes on a light background
By Maëlle Boisis • June 30, 2026
DigitalKin is a French deeptech startup pioneering causal and agentic AI from Lyon since 2023. The company is featured in the 2026 Auvergne-Rhône-Alpes French Startup Mapping , the annual reference cartography of the regional tech ecosystem co-published by France Digitale and Mesh Ventures , with the support of the local French Tech chapters: La French Tech Alpes and La French Tech Saint-Étienne Lyon .
DigitalKin at VivaTech 2026, Paris Expo Porte de Versailles, Hub France IA Village, Hall 7, June 17-
By Maëlle Boisis • June 15, 2026
At VivaTech 2026, DigitalKin launches the platform that lets experts turn their know-how into autonomous AI services. No code, no tech team. Applications open.
DigitalKin officially listed in the Hub France IA 2026 cartography as a Trusted AI Partner
By Maëlle Boisis • March 22, 2026
DigitalKin is now a certified Trusted AI Partner in the Hub France IA 2026 cartography, the official mapping of French AI startups presented at Bercy.
DigitalKin selected in the France Digitale 2026 mapping of French AI startups, unveiled at AI Day
By Maëlle Boisis • February 10, 2026
DigitalKin joins the France Digitale 2026 AI startup mapping, confirming its position as a French agentic AI leader for businesses. Here's why it matters.
Group of nine people standing together in a bright office, posing and smiling for a team photo.
By Maëlle Boisis • January 1, 2026
Enhance your skills with DigitalKin's expert AI agents. Join 1,500+ users transforming R&D processes today!
Promotional banner for “Le Journal Éco” with Emmanuel Thiéry on a purple gradient background
By Maëlle Boisis • December 31, 2025
DigitalKin's CEO shares insights on making expertise accessible with AI agents. Learn how we transform industries today!
AI and ethics: inside Lyon's rising agentic AI ecosystem - DigitalKin x BFM
By Maëlle Boisis • December 18, 2025
AI and ethics: inside Lyon's rising agentic AI ecosystem - DigitalKin x BFM
Man speaking in a bright exhibition hall, gesturing with both hands, with booths and people in the background
By Maëlle Boisis • November 25, 2025
DigitalKin builds personalized AI agents that enhance expertise. Test our platform for reliable AI solutions today!
Show More