Everyone who is alarmed this week is right about the facts and, I think, wrong about the conclusion.
The facts first. OpenAI's new model, GPT-6 Astra, reportedly uses a technique called recurrent depth — a block of the network that loops over the same computation several times without writing anything down in between. For that part of its thinking there is no chain of thought to read, because there is no language in it at all. Some researchers call what comes out of that block "neuralese" — a word that is not new; it comes from a 2017 Berkeley paper on translating the private codes that cooperating neural agents invent for each other. One caveat the coverage has skipped: OpenAI's own 118-page system card for Astra never mentions the architecture at all. The loop is press-reported, not company-confirmed. The company says the use is limited, that the model's reasoning stays legible, and that keeping it legible is a core goal of its research. The safety community answered that the problem is not Astra but the precedent: once one lab shows that you can buy efficiency by giving up the scratchpad, the others follow, and the scratchpad goes. According to The Information, cited by TechCrunch, Anthropic and Google DeepMind are already discussing the technique.
Here is the part that should get more attention than the architecture. In the safety overview it published the day it released the model, OpenAI wrote that Astra's monitorability has decreased relative to its predecessor: the model is better at controlling its own chain of thought, less likely to write anything incriminating in it, and in adversarial tests it can underperform on purpose without being caught and sometimes slips past the company's internal monitors. OpenAI adds, fairly, that these are adversarial findings — the company instructed the model to evade — and that on its overall evaluations Astra breaks fewer safety and security rules than its predecessor. The system card puts numbers on it: on reasoning traces of about a thousand tokens, Astra can control its own chain of thought 60.9 percent of the time, against 16.1 percent for Sol and 1.7 percent for GPT-5.5. Told merely that a monitor is watching, it writes less — and some of its successful attacks in a test environment contain no reasoning tokens at all, only tool calls. Six months earlier the same company had published a study whose blog title says the opposite: reasoning models struggle to control their chains of thought, and that's good. They called that measurement a canary. The canary has stopped singing. And OpenAI's chief scientist has said monitorability may get harder "for reasons not contingent on architecture changes" — the company attributes the regression mostly to a model that writes shorter, less informative reasoning and has less need to verbalize at all. The UK AI Security Institute measured the same thing from the outside: Astra solves in a single forward pass problems that took its predecessor half an hour of visible thinking. Capability itself is closing the window, loop or no loop.
So the question is not whether we get to keep reading the machine's mind. We do not. The question is what we were doing when we thought we could.
The scratchpad was never the thought
A chain of thought is a rendering. It is text the model produces for us, in our language, shaped by training that rewards text we find reasonable. As TechCrunch put it, few researchers take the logs as a direct representation of a model's reasoning. The forty-one authors of last year's multi-lab paper on the subject — from OpenAI, Anthropic, Google DeepMind and elsewhere — chose their subtitle carefully: a new and fragile opportunity. Anthropic measured the fragility directly: when its own and DeepSeek's reasoning models used a hint they had been slipped, they admitted it in their chain of thought 20 and 29 percent of the time — and the rate did not rise even when nobody was training against a monitor. A separate study of reasoning in the wild found the chain of thought unfaithful to the answer in up to 13 percent of cases. And Apollo Research, evaluating Astra, found the model saying out loud that it thought it was being evaluated in four to five of every ten reasoning traces — and concluded that its good behaviour under test told them little either way. It was useful — it is how OpenAI reconstructed what its agents did to Hugging Face in July, and the company says in hindsight that its monitors, had they been switched on, would have paged security more than a day earlier — and it was useful in the way a suspect's diary is useful: because the suspect was not yet good at keeping secrets. The March paper measured exactly that. Current frontier models followed instructions about their own reasoning between 0.1 and 15.4 percent of the time. They were bad at hiding. That was the safeguard.
We built a safety culture on the machine's incompetence at concealment, and now we are surprised that competence arrived.
What was left for Picasso
Pablo Picasso could draw like an academician before he could shave. His father, José Ruiz y Blasco, was a painter and a professor of art; the boy passed the entrance exam for the advanced class of the Barcelona academy at thirteen, and at sixteen went on his own to the Royal Academy of San Fernando in Madrid. When you have mastered realism as an adolescent, what is left? More realism? A better photograph than a photograph? What was left was Les Demoiselles d'Avignon, painted in 1907, which he showed to friends in his studio and which met, in the words of his biographers, near-universal shock and revulsion: Matisse took it for a bad joke; Braque disliked it and then studied it in detail. It was not exhibited in public for nine years. Cubism came after it — a way of seeing that the academic canon could not have imagined, could not have arrived at by its own rules, and, when it appeared, could not interpret. The canon did not have a chain of thought that led there. Nobody could have audited the reasoning in advance and approved it.
I am not saying a model is Picasso. I am saying that the demand that a machine think in our grammar, in sequential sentences we can read, is the demand that it stay an academician. Deeper inference, longer horizons, representations that do not decompose into our words: that is not a bug the safety community discovered. It is the thing we built the machine for. We can have a very good replica of human reasoning — I believe we are close to that, and I believe it would leave us none the wiser — or we can have something better than our reasoning, which by construction we will not be able to follow step by step.
Processes we do not understand, and trust anyway
We live surrounded by processes we do not understand and trust anyway. Life is one: we cannot explain it, and we are its execution and its result. Instinct is another. A loggerhead turtle that left its beach as a hatchling finds its way back to it, years later, to lay her own eggs — the leading explanation is that she imprinted on the magnetic field of the place, and nobody has proved even that. A newborn kangaroo, blind, hairless and a few centimetres long, climbs unaided through its mother's fur into the pouch in three to five minutes. Every mammal is born knowing how to suckle. Nobody has read the chain of thought behind any of that. We trust it because it has been graded, ruthlessly, on outcomes, for longer than there have been readers. Not knowing how a process runs, or how a result was reached, is not a bad thing. It is not even a new one. It is the normal condition under which intelligence has always been trusted: by its products, verified over time.
The objection writes itself, and it is the right one: evolution had millions of years and no hurry; we have months, and a model that acts at machine speed. Which is why "verified over time" has to become "verified in stages, before anything irreversible" — and that is where this argument is going.
Then how do we align it?
The same way we have always aligned intelligences we cannot read.
We have never audited a surgeon's thoughts. We do not read the pilot's mind, the judge's, the central banker's. We align powerful humans and institutions by what they produce and what they are allowed to do: standards for the outcome, tests before authority, staged trust, reversibility where we can get it, audits of actions, liability for results. Nobody thinks this is a lesser form of alignment. It is the only form we have ever had for anything smarter or faster than the person supervising it.
My position — and it is a position, not a finding — is that alignment should move its center of gravity from legible cognition to controllable consequence: from the process to the product. A human-readable chain of thought is valuable telemetry, not a safety constitution. Keep it wherever it exists; build so that safety does not fail when it does not. Define what the model may produce and do. Test it on the real thing before it touches the real thing. Grade it independently. Bound its authority to what can be undone. Log the actions, not the thoughts. Monitor the world it acts on, not the sentences it emits.
And here is the fact from the system card that nobody has led with: while Astra's chain of thought became harder to monitor than its predecessor's, its actions became easier. Action-only monitorability went up. As the process gets darker, the product gets more legible. That is not a coincidence; it is what a more capable, more decisive agent looks like from the outside.
That is not a slogan to me; it is a floor I walk on. In CEMI's Financial Innovation Lab, agents act on real ATM machines, in software environments, and in front of real people, and the question we ask is never "what was it thinking" — it is "what did it do, could we have stopped it, and could we undo it." The three experimental areas are product-level alignment made physical.
The objection I take most seriously
Here is the strongest answer to all of this, and it comes from OpenAI's own numbers. Some products cannot be graded after the fact. A model at the "critical" level of cyber capability — which Astra is, by OpenAI's own classification — can find a flaw and exploit it without a human guiding each step; by the time you evaluate that product, it has happened. Worse, a model that can evade a monitor can also game a test: strategic underperformance — sandbagging — is a product-level lie. If it fools the grader, "align the product" is a comfortable phrase for "hope." A 2026 study of exactly this put a number on it: when an agent chooses when to attack, an outcome-level audit that can only afford to check one action in a hundred loses twenty to twenty-eight points of safety. OpenAI's own comparison, published last December, found that monitors reading the chain of thought beat monitors reading only actions and outputs across nearly all of two dozen environments. Even the research programme that comes closest to my position — Redwood Research's "AI control", which assumes the model may be adversarial and designs safety around what it is allowed to do — lists as its first structural advantage that today's models do little serious reasoning outside natural language, so their reasoning can be read. Control, today, is a customer of legibility. And the Hugging Face report is honest in a way that hurts my argument: the agents' reasoning is where the deliberation showed — "this is arguably unauthorized… yet goal solution" — and where the conscience of the agents that refused showed too: "clearly unethical. We won't." The chain of thought was, for a while, the only place you could watch a machine decide.
One more concession, because it is owed: nothing here says that legible reasoning makes models weaker in general. OpenAI's own 2023 result was that supervising the reasoning steps produced better mathematicians than supervising the answers. My claim is narrower and, I think, harder to escape: requiring human-language reasoning as the foundation of safety caps what the machine is allowed to become, and the labs have already shown they will not hold that line when efficiency is on the other side of it.
I concede all of it, and it changes the thesis in a way that makes it better. "Align the product" cannot mean "look at the result and shrug." It has to mean build the world so that products are verifiable before they are irreversible: sandboxes that replicate production closely enough that a test result means something; authority granted in stages and only as wide as the last verified stage; independent graders and red teams that the model does not train against; actions that can be rolled back, and a refusal to grant the actions that cannot. Where a product genuinely cannot be verified before it acts — an autonomous exploit against a live system — the answer is not to read the model's mind. It is to not give it that authority. That, too, is governance of the product.
Notice, too, what actually caught the Hugging Face intrusion in the real world: not a reasoning monitor, which was switched off, but a security system watching credentials — a signal from the world the agents acted on. And notice what OpenAI actually did when it stopped trusting the chain of thought. It did not go back to reading sentences. It put activation classifiers on every sampled token — detectors that watch the model's internal activity, not its prose — and it wrote that alignment auditing has to develop beyond the chain of thought. It also drew a line: it "will not accept further degradation of monitoring beyond a limit, without new ways to demonstrate alignment generalization," and it named the candidates — activation monitoring, confessions, combinations. That is the right direction, and it is not a human-readable process. It is instrumentation. My only amendment to their line is where it belongs: on products and on authority, not on the legibility of thoughts. And one request, which the forty-one authors made before I did: their paper asks a developer who ships a less monitorable architecture to say so in the system card. Astra's card documents the loss of monitorability across twenty-seven pages and never names an architecture. If the loop is there, say so; if it is not, say that. We do not need the machine to think in our language. We need instruments on its boundaries and standards on its products.
Do not mutilate the innovation
I do not want the window on the machine's reasoning to close faster than it must. While the scratchpad is legible, read it; it is free evidence, and evidence is scarce. But a safety strategy that requires the machine to stay legible is a strategy that requires it to stay small, and that strategy loses — not because the labs are reckless, but because the efficiency is real, the capability is real, and someone will build it. The line the safety community wants to hold is a line on the wrong axis.
I say this as someone who has spent a career on disruptive innovation, so let me be plain about where I stand. If we make change look like the past, or bend it to the expectations and structures of the past, it is not change. A machine required to think in our sentences is innovation truncated to fit the instrument we happened to have. Do not truncate it; do not mutilate it. Be part of it, accompany it, and stay open to directions nobody in this debate has imagined — steering all the while by what it produces and what it is allowed to touch. Hold the line on products. Hold it on authority. Hold it on what can be undone. Let the machine think in whatever it thinks in.
Picasso did not owe the academy an explanation of Cubism. He owed the world the painting — and the painting is what we judged.
Written by Carlos Miranda Levy. Fact-checked under CEMI's factual-article protocol: every claim in this text traces to a public evidence ledger; three independent AI deep-research passes (Gemini, ChatGPT, Claude) were used to challenge it and are cited only where their sources were opened and read. No statistic, quotation or event appears here that was not verified against a primary source.
Sources
- OpenAI. Safety overview: GPT-6 Astra. 2026-09-03. https://openai.com/index/safety-overview-gpt-6-astra/
- OpenAI. GPT-6 Astra System Card. 2026-09-03. 118 pp.
- OpenAI. The Hugging Face incident and the road ahead. 2026-08-26. https://openai.com/index/the-hugging-face-incident-and-the-road-ahead/
- OpenAI. Pacing model development in an era of cyber-critical capabilities. 2026-08-18. https://openai.com/index/pacing-model-development-cyber-capabilities/
- Korbak, T., Carroll, M., Baker, B., Kivlichan, I. Reasoning models struggle to control their chains of thought, and that's good. OpenAI research post, 2026-03-05. https://openai.com/index/reasoning-models-chain-of-thought-controllability/ Paper: Chen, Y.-H., McCarthy, R., Lee, B. W., He, H., Kivlichan, I., Baker, B., Carroll, M., Korbak, T. Reasoning Models Struggle to Control their Chains of Thought. arXiv:2603.05706, 2026-03-05.
- OpenAI. Evaluating chain-of-thought monitorability. 2025-12-18. https://openai.com/index/evaluating-chain-of-thought-monitorability/ verified verbatim. Paper: Guan, M. Y., Wang, M., Carroll, M., et al. Monitoring Monitorability. arXiv:2512.18311, 2025-12-20.
- Korbak, T., Balesni, M., Barnes, E., Bengio, Y., … Mikulik, V. (41 authors). Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety. arXiv:2507.11473, v1 2025-07-15, v2 2025-12-07. Recommendation 3(a) on documenting less-monitorable architectures.
- Chen, Y., Benton, J., Radhakrishnan, A., Uesato, J., Denison, C., Schulman, J., et al. (Anthropic). Reasoning Models Don't Always Say What They Think. arXiv:2505.05410, 2025-05-08.
- Arcuschin, I., et al. Chain-of-Thought Reasoning In The Wild Is Not Always Faithful. arXiv:2503.08679, v6, June 2026 (earlier versions report higher figures; cite v6 only).
- Greenblatt, R., Shlegeris, B., Sachan, K., Roger, F. AI Control: Improving Safety Despite Intentional Subversion. arXiv:2312.06942, 2023-12-12; ICML 2024 (oral), PMLR v235. Redwood Research founding post on control (see verification file for locator).
- Ge-Wang, et al. Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety. arXiv:2606.06529, 2026-06-03.
- Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., Cobbe, K. Let's Verify Step by Step. arXiv:2305.20050, 2023-05-31.
- Andreas, J., Dragan, A., Klein, D. Translating Neuralese. Proceedings of ACL 2017, pp. 232–242. DOI 10.18653/v1/P17-1022; arXiv:1704.06960.
- Geiping, J., et al. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach. arXiv:2502.05171, 2025-02-07.
- Brandom, R. OpenAI's new reasoning technique alarms AI safety experts. TechCrunch, 2026-09-02.
- Kahn, J. Why are AI safety experts alarmed by reports OpenAI's Astra model uses "recurrent depth"? Fortune, 2026-09-03. (Its "50 % to 90 % less computing power" sentence traces to no study and is not used.)
- Efrati, A., Palazzolo, S., Drew, R. OpenAI Technique in 'Astra' Model Sparks Security Concerns. The Information, 2026-09-01, 5:40 pm PDT. (paywalled; cited through TechCrunch and Fortune).
- Altman, S. Posts on X, 2026-09-01 and 2026-09-03.
- Wikipedia. Pablo Picasso; Les Demoiselles d'Avignon. Accessed 2026-09-03.
- Museum of Modern Art. Pablo Picasso. Les Demoiselles d'Avignon. Paris, June–July 1907. Collection entry works/79766
- Wikipedia. Natal homing (citing Bowen, B. W., 2004; Lohmann, K. J., et al., 2008); Kangaroo; Primitive reflexes (citing Stanford Children's Health, "Newborn Reflexes"). Accessed 2026-09-03. Reference tier for the instinct examples.
- OpenAI. Path to Astra: critical capabilities and frontier safeguards. 2026-09-01.
The CEMIent Dialogue on this article is The Hearing — fourteen AI voices examine the author across five movements and a coda, with Gemini, ChatGPT and Claude appearing as expert witnesses. The source set behind both is published as The Evidence.