← ObservatoryThe RecordFR-AI-0004
PROG-AI
FR-AI-0004

Scaling Laws — Emergent Performance on Unseen Tasks

Scaling language model training increases performance on previously unseen tasks without task-specific optimisation.

FragmentingVS-03·since 2026-06-29
Verification Matrix
VS-01
Assertion
VS-02
Published
VS-03
Audit
2024-01-15 — present
VS-04
Replication
VS-05
Operation
State reached Current state Not yet reached
State Warrant
Current stateFragmentingVS-03
Why this state?Sourced from: Medium, "The State of Large Language Models: Latest Updates & Trends (2025–2026)" (Feb 2026) for the inference/tooling consensus-shift framing; general field knowledge of o1 (Sept 2024), o3 (Dec 2024), and DeepSeek-R1 (Jan 2025) release timing and capability framing. The Medium source is a secondary roundup, not a primary research paper — adequate for establishing that a shift occurred, not for citing specific benchmark figures. Primary literature (e.g. the test-time-compute scaling papers referenced in FR-AI-0005's evidence trail) should be consulted before this assessment is extended with specific numbers.
Assessment summaryThe claim's core assertion remains supported, and the pressure state remains FRAGMENTING — IN-007 adds to the fragmentation rather than resolving it. Test-time compute and reasoning-model architectures (o1/o3, DeepSeek-R1) demonstrate that scaling inference-time computation, not only training-time parameters and data, improves performance on previously unseen reasoning tasks. This is a genuinely new mechanism for the claim's core phenomenon, not merely a third data point alongside Kaplan et al. and Chinchilla: the original claim statement ("scaling language model training") describes training-time scaling specifically, and IN-007's mechanism operates at inference time. By early 2026, field commentary describes a broader shift in where capability gains are expected to come from — inference and tooling rather than raw training-scale increases — which bears directly on BN-001 (no agreed definition of "previously unseen") and on the record's account of what "scaling" means well past the boundary AS-001 anticipated. This assessment does not propose a reclassification; it records that the claim's mechanism account is now materially incomplete without IN-007, two years into the record's life, in a field moving fast enough that the gap itself is notable.
State entered2024-01-15
Last reaffirmed2026-06-29
Stage provenanceRatified VS-03; stored historical code VS-03 preserved.
Record Lineage — Chronological
2024-01-15
Record opened — Fragmenting
The claim is supported in its core assertion: scaling language model training does increase performance on previously unseen tasks without task-specific optimisation. This is documented across multiple model families, task types, and evaluation methodologies. The few-shot performance documented in INST-002, the smooth scaling curves in INST-001 and INST-006, and the emergent task capabilities in INST-003 all constitute positive evidence for the claim as stated. The evidence trail is nonetheless complicated by two interior disputes that do not threaten the claim's truth but substantially complicate its mechanism: whether apparent emergent abilities (INST-003) are genuine discontinuities or artefacts of metric choice (INST-004), and whether benchmark performance gains reflect genuine generalisation or training-data contamination (INST-005). The pressure state is FRAGMENTING: the claim's core assertion holds, but the evidence quality disputes over how and why it holds have not converged, and no agreed definition of 'previously unseen' yet exists to resolve them (BN-001).
Verification Stage: VS-03 after ratified review (stored code VS-03 preserved).
2026-06-29
Reassessed, no change — Fragmenting
The claim's core assertion remains supported, and the pressure state remains FRAGMENTING — IN-007 adds to the fragmentation rather than resolving it. Test-time compute and reasoning-model architectures (o1/o3, DeepSeek-R1) demonstrate that scaling inference-time computation, not only training-time parameters and data, improves performance on previously unseen reasoning tasks. This is a genuinely new mechanism for the claim's core phenomenon, not merely a third data point alongside Kaplan et al. and Chinchilla: the original claim statement ("scaling language model training") describes training-time scaling specifically, and IN-007's mechanism operates at inference time. By early 2026, field commentary describes a broader shift in where capability gains are expected to come from — inference and tooling rather than raw training-scale increases — which bears directly on BN-001 (no agreed definition of "previously unseen") and on the record's account of what "scaling" means well past the boundary AS-001 anticipated. This assessment does not propose a reclassification; it records that the claim's mechanism account is now materially incomplete without IN-007, two years into the record's life, in a field moving fast enough that the gap itself is notable.
Sourced from: Medium, "The State of Large Language Models: Latest Updates & Trends (2025–2026)" (Feb 2026) for the inference/tooling consensus-shift framing; general field knowledge of o1 (Sept 2024), o3 (Dec 2024), and DeepSeek-R1 (Jan 2025) release timing and capability framing. The Medium source is a secondary roundup, not a primary research paper — adequate for establishing that a shift occurred, not for citing specific benchmark figures. Primary literature (e.g. the test-time-compute scaling papers referenced in FR-AI-0005's evidence trail) should be consulted before this assessment is extended with specific numbers.
Verification Stage: VS-03 after ratified review (stored code VS-03 preserved).
Mutation Log
MutationDateFieldPrior valueCurrent value
M-0102026-07-09description_reorderedDESCRIPTION-REORDERED
M-0092026-06-29open_question_raisedOQ-RAISED
M-0082026-06-29assessment_issuedAS-001AS-002
M-0072026-06-29instances_loggedINSTANCES-LOGGED
M-0062024-01-15programme_panel_addedPROGRAMME-PANEL-ADDED
M-0052024-01-15null_boundary_condition_metNULL-BOUNDARY-CONDITION-MET
M-0042024-01-15mechanisms_recordedMECHANISMS-RECORDED
M-0032024-01-15assessment_issuedASSESSMENT-ISSUED
M-0022024-01-15instances_loggedINSTANCES-LOGGED
M-0012024-01-15record_createdRECORD-CREATED
Evidence Sources
7 instances on recordShow sources ↓Hide ↑
IN-001Kaplan et al. — Neural scaling laws for language modelssupportive
IN-002GPT-3 few-shot performance and emergent task generalisationsupportive
IN-003Wei et al. — Emergent abilities of large language modelspartial
IN-004Schaeffer et al. — Are emergent abilities a mirage?contesting
IN-005Scaling law limitations — benchmark saturation and contaminationcontesting
IN-006Chinchilla scaling laws and compute-optimal trainingpartial
IN-007Test-time compute and reasoning models — a second scaling axis emergespartial