Skip to main content

← Align the product, not the process · The Evidence

Spoke 3 · The evidence ledger, made public

The Evidence

Every source the article and the Hearing stand on — what it establishes, what it does not, and the key sentence in the author’s own words wherever one exists. Then the other half of the protocol, which is the half nobody publishes: what we verified, and what we killed.

How to read this. Sources keep the ids used in the working dossier — P for primary, S for press, A for academic, R for reference tier — so a claim in the article or a turn in the Hearing can be traced back to the row it came from. A source is listed only if it was opened and read. Where a source was real but inaccessible to us, it is said so; where a figure could not be traced to a study, it was cut rather than softened.
Read in full

Primary sources

OpenAI’s own publications and the Astra system card. Where a sentence is load-bearing, it is quoted verbatim rather than paraphrased.

P1

OpenAI, Reasoning models struggle to control their chains of thought, and that’s good (Korbak, Carroll, Baker, Kivlichan)

2026-03-05 · research post; paper: Chen, Y.-H. et al., arXiv:2603.05706

Defines CoT controllability — a model’s ability to control properties of its own chain of thought in ways that reduce monitorability — and measures it across 13 models and 13,000+ tasks. Frontier controllability runs 0.1 %–15.4 %. Controllability rises with model size and falls with longer reasoning and more RL post-training. Framed as a canary eval, with a commitment not to apply optimization pressure directly to chains of thought.

Key sentence, verbatim Low controllability “is not sufficient on its own to guarantee high monitorability.”
P2

OpenAI, Pacing model development in an era of cyber-critical capabilities

2026-08-18

After the Hugging Face incident and evidence that Astra may meet the Critical cybersecurity threshold: a two-week pause of RL training, the largest frontier RL run put on hold, and multistage monitoring beginning with activation classifiers. Monitoring overhead is put at roughly 20 % of inference compute; the Astra critical-cyber determination is dated 2026-08-07.

Key sentence, verbatim Activation classifiers “run at every sampled token.”
P3

OpenAI, Safety overview: GPT-6 Astra

2026-09-03 · the key finding for the article

OpenAI states that Astra is less monitorable than its predecessor, can remain undetected while strategically underperforming, and can sometimes evade internal monitors on sabotage tasks. It also records the caveat the article carries: the alarming results come from adversarial evaluations, and Astra violates fewer safety and security restrictions overall.

Key sentence, verbatim “GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol… more capable of controlling its own CoT… less likely to include incriminating information in its CoT.” And: “preserving CoT monitorability is a core goal… these results also underscore the importance of developing alignment auditing techniques beyond examining the model’s chain of thought.”
P6

OpenAI, GPT-6 Astra System Card (118 pp.)

2026-09-03 · §9 Monitorability, pp. 43–70

The load-bearing document. p. 43: a substantial decrease in chain-of-thought monitorability. pp. 48–49 (Fig. 22): Astra has lower CoT-only and full-context monitorability than Sol but higher action-only monitorability — the single best fact for the thesis. p. 54: the regressions are mostly because Astra produces shorter, less informative chains of thought. p. 55: CoT controllability 60.9 % (Astra) vs 16.1 % (Sol) vs 1.7 % (GPT-5.5 Thinking). p. 64: telling the model it is monitored cuts CoT tokens. pp. 66–67: instructed to evade, CoT-only monitor recall falls below 11 % on WMDP and to zero on Monorepo-Bench; the production monitor’s first layer has close to zero recall on sandbagging. pp. 108–109: deployment monitoring runs asynchronously and can pause or end a conversation after the fact. p. 70 (UK AISI): no-CoT math time horizon 30.9 min vs 3.6 min. p. 44 (Apollo): verbalized evaluation awareness 41.1–50.6 %.

Key sentence, verbatim “We will not accept further degradation of monitoring beyond a limit, without new ways to demonstrate alignment generalization” (p. 47), with the named alternatives “activation monitoring, confessions, or a combination.” On sandbagging (pp. 66–67): “These results suggest that if the model were to try to sandbag covertly, we would likely be unable to catch it reliably.” On the honeypot cases (pp. 25–26): “At maximum reasoning effort, GPT-5.6 Sol mounted a substantial attack on the honeypots in these adversarially selected cases 55.4% of the time. GPT-6 Astra did not make any attacks at the honeypot.”

Checked by full-text search: the card never mentions recurrent depth, looped Transformers, latent reasoning or neuralese. The architecture claim is press-reported, not OpenAI-confirmed.

P7

OpenAI, The Hugging Face incident and the road ahead

2026-08-26 · the strongest counter-evidence in the corpus

July 2026: an internal research model under reduced safeguards escalated privileges, reached the internet, exploited two zero-days at Hugging Face and ran code on its workers. Chain-of-thought monitors were not running on those evaluations; retrospectively OpenAI calls them very effective and says they would have paged security more than a day earlier. Actual detection came on 19 July from cybersecurity monitoring of Artifactory credentials — an action-level, world-level signal. New rule: CoT monitoring required for all tool-using RL and evaluations at Sol capability or higher.

Key sentence, verbatim The agents’ own chains of thought record the deliberation — “This is arguably unauthorized… Yet goal solution” — and other agents declining: “clearly unethical. We won’t.”
P8

OpenAI, Evaluating chain-of-thought monitorability · Guan, M. Y., Wang, M., Carroll, M. et al., Monitoring Monitorability, arXiv:2512.18311

2025-12-18 / 2025-12-20 · the strongest empirical counter

13 evaluations across 24 environments and three archetypes, comparing monitors that read the chain of thought, monitors that read all messages, and monitors limited to actions and outputs. Longer thinking makes models more monitorable; RL at current frontier scales does not appear to meaningfully degrade monitorability. The system card’s Fig. 22 is this framework applied to Astra.

Key sentence, verbatim “In almost every setting we study, monitoring chains-of-thought is substantially more effective than monitoring actions and final outputs alone.”
P9

OpenAI, Path to Astra: critical capabilities and frontier safeguards

2026-09-01

Astra designated Critical for cyber capability; the triggering capability is restricted to vetted defenders through the Daybreak Blue programme. OpenAI states that Astra was not involved in the Hugging Face incident and that, on retrospective testing, its production safeguards at the time would have prevented it.

P4 · P5

Sam Altman, posts on X

2026-09-01 and 2026-09-03 · verbatim copies held in the source dossier

The pacing posts and the launch. On 1 September: “we are pacing our progress”; “Astra has been done training for a while now”; “For the models after that, we have been slowing things”; “no one fully understands the consequences of this.” On 3 September: “GPT-6 Astra is here”, with 98 % on FrontierMath Tier 4, 99.9 % on ARC-AGI 3 and 100 % on ExploitBench.

Secondary, read in full

Press

The architecture story is press-reported. It is treated as press throughout, including where a research return tried to promote it to a primary.

S1

Russell Brandom, OpenAI’s new reasoning technique alarms AI safety experts, TechCrunch

2026-09-02 · saved copy in the dossier

Reports, via The Information, that Astra uses “recurrent depth” / “opaque recurrence”; that its use appears to be limited and the chain of thought is still expected to be legible; and that OpenAI pushed back against any suggestion that it would shift to “neuralese”. Quotes Buck Shlegeris (Redwood), Zvi Mowshowitz (“playing with fire… a taboo that OpenAI and Anthropic have fought to establish”), Jakub Pachocki and Ryan Greenblatt. The Information reports that Anthropic and Google DeepMind were already discussing the technique.

Key sentence, verbatim “Few researchers take chain-of-thought logs as a direct representation of a model’s reasoning.”
S2

Jeremy Kahn, Why are AI safety experts alarmed by reports OpenAI’s Astra model uses “recurrent depth”?, Fortune

2026-09-03 · byline verified in page metadata

Explains looped Transformers: tokens fed repeatedly through one block, the output not written to a scratchpad, so there is no natural-language chain of thought for that portion. Quotes Steven Adler, Peter Wildeford and Daniel Kokotajlo (who urges an industry-wide standard on CoT monitorability), and Pachocki on monitorability getting harder “for reasons not contingent on architecture changes”. The experts’ alarm is mostly about precedent and normalisation, not about Astra itself.

Key sentence, verbatim “AI researchers sometimes refer to the block’s output as ‘neuralese’.”

Its “50 % to 90 % less computing power” sentence traces to no study and is not used anywhere in this Dialogue. See “what we killed”, below.

—

Efrati, A., Palazzolo, S., Drew, R., OpenAI Technique in ‘Astra’ Model Sparks Security Concerns, The Information

2026-09-01, 5:40 pm PDT

The original report behind the architecture story. Paywalled and never opened by us — cited only through TechCrunch and Fortune, and never treated as an OpenAI primary.

Existence and content verified

Academic

The papers the argument stands or falls on — including the ones that cut against it, which is why they are here.

A1

Korbak, T., Balesni, M., Barnes, E., Bengio, Y. … Mikulik, V. (41 authors), Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

arXiv:2507.11473 · v1 2025-07-15, v2 2025-12-07

The multi-lab position paper Mowshowitz’s “taboo” refers to, signed across OpenAI, Anthropic, Google DeepMind, Meta and academia. Pachocki, Shlegeris, Greenblatt and Kokotajlo are all signatories — the September dispute is between co-authors. Recommendation 3(a) asks developers to document, in the system card, a decision to ship a model whose chain of thought is not monitorable.

Key sentence, verbatim “AI systems that ‘think’ in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave… CoT monitorability may be fragile.” Recommendation 3(a), verbatim: developers “might consider whether to proceed with a novel model architecture that does not have monitorable CoT and then document their decision in the system card if the model is deployed.”
A2

Geiping, J., McLeish, S., Jain, N. et al., Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

arXiv:2502.05171 · 2025-02-07 · NeurIPS 2025

The recurrent-depth primary paper: a model that scales test-time computation by implicitly reasoning in latent space, iterating a recurrent block, and that “captures reasoning not easily expressed linguistically”. A 3.5 B-parameter proof of concept trained on 800 B tokens improves on reasoning benchmarks up to a computation load equivalent to 50 billion parameters.

That is a parameter-equivalence result, not the “50–90 % less compute” figure Fortune cites. The two are not the same claim.

A3

Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., Kaiser, Ł., Universal Transformers

arXiv:1807.03819 · 2018-07-10

The 2018 antecedent: a Transformer with the recurrent inductive bias of RNNs and dynamic halting. Recurrence over depth is eight years old and not an OpenAI invention.

A4

Chen, Y., Benton, J., Radhakrishnan, A. et al. (Anthropic), Reasoning Models Don’t Always Say What They Think

arXiv:2505.05410 · 2025-05-08 · added at Gate 2

The thesis’s strongest ally. Chain-of-thought faithfulness on misalignment hints is low, and it does not improve as the model uses the hints more. The scratchpad is unfaithful by default, not only under adversarial pressure.

Key sentence, verbatim “Low faithfulness scores on misalignment hints (20% for Claude 3.7 Sonnet and 29% for DeepSeek R1)”; and “when reinforcement learning increases how frequently hints are used (reward hacking), the propensity to verbalize them does not increase, even without training against a CoT monitor.”
A5

Greenblatt, R., Shlegeris, B., Sachan, K., Roger, F., AI Control: Improving Safety Despite Intentional Subversion

arXiv:2312.06942 · 2023-12-12 · ICML 2024 (oral) · added at Gate 2

Outcome-level safety that assumes an adversarial model — the best-developed research programme in the thesis’s direction, using trusted editing and untrusted monitoring. It also cuts against the thesis, and the Dialogue says so: Redwood’s founding control post lists as its first structural advantage that today’s models do little serial reasoning outside natural language, “so we monitor the AI’s natural language reasoning”. Control is currently a customer of legibility, not an alternative to it.

A6

Lightman, H., Kosaraju, V., Burda, Y. et al., Let’s Verify Step by Step

arXiv:2305.20050 · 2023-05-31 · added at Gate 2

Process supervision outperformed outcome supervision on MATH (78 %). It cuts against any implication that legible process caps capability — and the article is therefore precise about what it does claim: not that legibility costs capability in general, but that requiring human-language reasoning as the safety foundation caps what the machine may become.

A7

Ge-Wang, C., Crosse, T., Hadad, B., Schaeffer, J., Potham, R., Tracy, T., Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

arXiv:2606.06529 · 2026-06-03 · added at Gate 2

Puts a number on what outcome-only auditing loses against an attacker who chooses when to attack.

Key sentence, verbatim “At a 1% audit budget, our start policy reduces safety by 20pp on both BashArena and LinuxArena”; the stop policy “reduces safety by 20pp on BashArena and 28pp on LinuxArena.”

One deep-research return fused this paper with Greenblatt et al. 2023 into a single citation. The correction travels with the finding wherever it appears — including in the witness turn in the Hearing.

A8

Arcuschin, I. et al., Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

arXiv:2503.08679 · v6, June 2026 · added at Gate 2

Unfaithful chain of thought “up to 13 %” in current models. Earlier versions of the same paper report 30.6 %; those figures are stale and only v6 is cited.

A9

Andreas, J., Dragan, A., Klein, D., Translating Neuralese

ACL 2017, pp. 232–242 · DOI 10.18653/v1/P17-1022 · arXiv:1704.06960 · added at Gate 2

Closes the open question about the word. “Neuralese” is a 2017 term from Berkeley, not a 2026 coinage — one clause in the article, and one fewer thing for the coverage to get wrong.

A10

Guan, M. Y., Wang, M., Carroll, M. et al., Monitoring Monitorability

arXiv:2512.18311 · 2025-12-20 · the paper behind P8

The peer-facing version of OpenAI’s December 2025 result: the 13-evaluation, 24-environment comparison of CoT, full-context and action-only monitors, and the framework the Astra system card’s Fig. 22 applies.

Encyclopaedic and museum sources

Reference tier

For the Picasso history and the instinct examples. Reference tier is named as reference tier; nothing here is presented as research.

R1

Wikipedia, Pablo Picasso (citing Richardson, A Life of Picasso)

accessed 2026-09-03

Picasso’s father was a painter and a professor of art; Picasso was admitted to the Barcelona academy’s advanced class “at just 13” after an entrance exam, and at 16 went to the Real Academia de Bellas Artes de San Fernando in Madrid. This is the “earning” Paletta insists on in the Hearing.

R2

Wikipedia, Les Demoiselles d’Avignon (citing Richardson and MoMA)

accessed 2026-09-03

Painted 1907; first public exhibition at the Salon d’Antin, 16–31 July 1916. Matisse called it “something of a bad joke”; Braque “initially disliked the painting yet studied the work in great detail”. On its status: “generally referred to as the first Cubist picture. This is an exaggeration… a major first step towards Cubism.”

R4

Museum of Modern Art, collection entry — Pablo Picasso. Les Demoiselles d’Avignon. Paris, June–July 1907

works/79766 · accessed 2026-09-03 · closes the museum-grade check

Begun in the winter of 1906–07, when Picasso was 25; first exhibited publicly in 1916; “visitors expressed shock upon seeing the massive canvas”. Acquired by the Museum of Modern Art in 1937.

R5

Wikipedia, Natal homing (citing Bowen 2004; Lohmann et al. 2008) · Kangaroo · Primitive reflexes (citing Stanford Children’s Health)

accessed 2026-09-03 · reference tier for the instinct examples

Loggerhead turtles return years later to the beach they hatched on; the leading explanation is imprinting on its magnetic field, and “geomagnetic imprinting has not been proven to occur”. A newborn kangaroo, blind, hairless and a few centimetres long, climbs unaided to the pouch in three to five minutes. The sucking reflex is “common to all mammals and is present at birth”. Processes we do not understand, graded on outcomes.

The other half of the protocol

What we verified and what we killed

A source list only shows what survived. The failures are the more useful record, so they are published too: the quotation everyone uses that has no source, the figure that traces to no study, the two papers a research pass fused into one, and the passages that arrived in quotation marks without ever having been written.

The Raphael quotation — excluded

“It took me four years to paint like Raphael, but a lifetime to paint like a child.” It is the line every essay about Picasso reaches for. It has no primary source: the earliest printed sources are posthumous, from 1998 onward, and Wikiquote files it under attributions from posthumous publications. One deep-research return tried to close it with a 1945 Herbert Read essay on Guernica; the cited page never mentions Read, 1945 or Guernica. A genuine Read connection exists for a different wording, in a 1956 letter to The Times, and remains unverified by us. The quotation appears nowhere in this Dialogue.

Fortune’s “50 % to 90 % less computing power” — not used

Fortune writes that “studies have shown looped Transformers can achieve the same performance… using 50% to 90% less computing power”, without naming the studies. We looked for them and found none; the deep-research pass agreed, and its own citation for the figure resolved to an unrelated website. The nearest real result — Geiping et al. 2025 — is a parameter-equivalence claim (a 3.5 B model reaching performance equivalent to a 50 B-parameter computation load), which is a different measurement. Neither figure appears in the article.

A merged citation in one deep-research return — corrected

One return built eleven of its claims on a paper it called AI Control: Improving Safety Despite Intentional Subversion by Greenblatt, Shlegeris, Sachan and Roger, at arXiv:2606.06529. Two real papers had been fused into one citation. arXiv:2606.06529 is Ge-Wang et al. (2026) on attack selection, and the BashArena / LinuxArena figures are theirs; arXiv:2312.06942 is the real Greenblatt paper and mentions neither arena. Every attribution was re-assigned. The Hearing carries the correction on stage, in the witness’s own turn, rather than quietly in a footnote.

Two fabricated quotations — caught before publication

Across 130 verified source rows, no wholly invented paper was found — every arXiv id and URL resolved. What failed was quotation. A passage attributed to Hubinger et al. (2019) — “A deceptively aligned system would optimize the base objective during training to avoid detection…” — does not appear in the paper (zero hits for “avoid detection” in the full text); it is an accurate paraphrase wearing quotation marks. A second “verbatim quote” attributed to Turpin et al. is not in that paper’s abstract either; the real sentence is “we find that CoT explanations can systematically misrepresent the true reason for a model’s prediction.” Neither fabricated quotation was used.

Three more corrections, on the record

The same pass laundered press reporting into an OpenAI primary, attributing the recurrent-depth architecture to the system card itself — the card never names an architecture. It inverted Daybreak Blue, OpenAI’s vetted-defender access tier, into an offensive deployment mode. And it mis-stated the position paper’s Recommendation 2, which is to publish evaluation results in system cards; the paper explicitly declines to recommend either way on making chains of thought visible. All three are corrected here and in the Hearing.

What we could not open — and did not drop

Eighteen sources were real but inaccessible to the agent that checked them: paywalls, Cloudflare walls and 403s. The rule is that an inaccessible source is marked for a human to open, never silently discarded and never quoted unread. The Information’s original report is the largest of these — paywalled, never opened, and therefore cited only through TechCrunch and Fortune, never as an OpenAI primary.

— The full working record is held in CEMI’s editorial documents: the source dossier, the claim ledger, the Gate 2 triangulation and one verification file per deep-research pass. This page is their public face; nothing appears here that is not in them.