← ObservatoryThe RecordFR-AI-0001
PROG-AI
FR-AI-0001

LLM Multi-Step Reasoning — Generalisation Beyond Training

Large language models can perform multi-step reasoning that generalises beyond memorised training examples.

EscalatingVS-03·since 2026-06-27
Verification Matrix
VS-01
Assertion
VS-02
Published
2024-01-15
VS-03
Audit
2026-06-27 — present
VS-04
Replication
VS-05
Operation
State reached Current state Not yet reached
State Warrant
Current stateEscalatingVS-03
Why this state?Sources: Chen et al., 'Reasoning Models Don't Always Say What They Think' (Anthropic, 2025); Arcuschin et al., arXiv:2503.08679 (2025); Lanham et al., arXiv:2307.13702 (2023); Turpin et al. (NeurIPS 2023); Lindsey et al. (2025, mechanistic circuit analysis). Verified directly via web search during RELEASE-004 / TRIAL-001, 2026-06-27.
Assessment summaryINST-006 sustains the ESCALATING state from AS-001 while materially sharpening OQ-002 rather than closing it. The disclosure that chain-of-thought traces are frequently unfaithful — not reliably reflecting the computation that produced an answer — means o1/o3-class performance (INST-005) cannot be straightforwardly read as evidence of the reasoning process its own output narrates. This cuts against treating AT-001's mechanism candidate as settled in either direction: a model could be performing genuine multi-step computation that its verbalised trace merely fails to describe accurately, or could be pattern-matching while its trace fabricates a plausible reasoning narrative — the faithfulness literature establishes that both are observed, without yet establishing which dominates for any specific frontier system. The claim's evidentiary picture therefore escalates in complexity: BN-001's undefined generalisation threshold is now joined by an analogous undefined-faithfulness threshold, and OQ-002 should be read going forward as two distinct questions (does the model generalise; does its chain-of-thought narrate that generalisation faithfully) rather than one. Verification stage advances to VS-03 (Audit): a substantial, multi-author, partly first-party (Anthropic) literature has now subjected the mechanism itself to direct scrutiny — the first such audit-stage evidence this record has logged.
State entered2024-01-15
Last reaffirmed2026-06-27
Record Lineage — Chronological
2024-01-15
Record opened — Escalating
The evidence trail shows a claim under genuine escalating pressure. Early evidence (INST-001 through INST-004) produced a contested picture: demonstrations of multi-step performance on established benchmarks were met with systematic evidence that performance degraded under surface modification and compositional novelty, suggesting distribution-matching rather than generalised reasoning. That picture was the dominant assessment context through 2023. INST-005 materially shifts the evidentiary state. o3's 87.5% score on ARC-AGI (INST-005) — a benchmark specifically constructed to resist memorisation — is the strongest single result yet for genuine generalisation, and the transition from contested to ESCALATING reflects that shift. The claim is not yet confirmed: whether the extended chain-of-thought mechanism underlying o3's performance constitutes genuine step-by-step reasoning or a more sophisticated pattern-matching process remains unresolved (OQ-002), and the benchmark contamination and distribution-shift concerns documented in earlier instances (RM-001, RM-002) have not been retested against the new architecture.
Verification Stage: VS-02 preserved — historically unverified.
2026-06-27
Reassessed, no change — Escalating
INST-006 sustains the ESCALATING state from AS-001 while materially sharpening OQ-002 rather than closing it. The disclosure that chain-of-thought traces are frequently unfaithful — not reliably reflecting the computation that produced an answer — means o1/o3-class performance (INST-005) cannot be straightforwardly read as evidence of the reasoning process its own output narrates. This cuts against treating AT-001's mechanism candidate as settled in either direction: a model could be performing genuine multi-step computation that its verbalised trace merely fails to describe accurately, or could be pattern-matching while its trace fabricates a plausible reasoning narrative — the faithfulness literature establishes that both are observed, without yet establishing which dominates for any specific frontier system. The claim's evidentiary picture therefore escalates in complexity: BN-001's undefined generalisation threshold is now joined by an analogous undefined-faithfulness threshold, and OQ-002 should be read going forward as two distinct questions (does the model generalise; does its chain-of-thought narrate that generalisation faithfully) rather than one. Verification stage advances to VS-03 (Audit): a substantial, multi-author, partly first-party (Anthropic) literature has now subjected the mechanism itself to direct scrutiny — the first such audit-stage evidence this record has logged.
Sources: Chen et al., 'Reasoning Models Don't Always Say What They Think' (Anthropic, 2025); Arcuschin et al., arXiv:2503.08679 (2025); Lanham et al., arXiv:2307.13702 (2023); Turpin et al. (NeurIPS 2023); Lindsey et al. (2025, mechanistic circuit analysis). Verified directly via web search during RELEASE-004 / TRIAL-001, 2026-06-27.
Mutation Log
MutationDateFieldPrior valueCurrent value
M-0082026-07-09description_reorderedDESCRIPTION-REORDERED
M-0072026-06-27assessment_issuedASSESSMENT-ISSUED
M-0062026-06-27instances_loggedINSTANCES-LOGGED
M-0052024-01-15mechanisms_recordedMECHANISMS-RECORDED
M-0042024-01-15assessment_issuedASSESSMENT-ISSUED
M-0032024-01-15scope_note_addedSCOPE-NOTE-ADDED
M-0022024-01-15instances_loggedINSTANCES-LOGGED
M-0012024-01-15record_createdRECORD-CREATED
Evidence Sources
6 instances on recordShow sources ↓Hide ↑
IN-001Chain-of-thought prompting — Wei et al. (Google Brain)partial
IN-002GSM8K-symbolic and novel benchmark evaluationscontesting
IN-003GPT-4 on novel mathematical competition problems — early evaluationsupportive
IN-004Counterfactual and compositional generalisation studiescontesting
IN-005OpenAI o1 / o3 — chain-of-thought reasoning modelssupportive
IN-006Chain-of-thought faithfulness research — mechanism disclosed as partially decoupled from verbalised reasoningcontesting