Recent Activity
2026-08-01FR-AI-0002IN-007 appended — ORCA-bench (arXiv:2607.28545v1), authorised through Post-Scout flag 2026-08-01-02. Bounded CONTESTING evidence strengthens the existing low-supervision agentic boundary with operational root-cause-analysis measurements. Instance only: AS-001 remains current; Pressure State ESCALATING, Verification Stage VS-02, mechanisms, and open questions unchanged. No new assessment issued.Escalating
2026-08-01FR-AI-0007IN-008 appended — open-ended AI research shadow evaluation (arXiv:2607.27191v1), authorised through Post-Scout flag 2026-08-01-01. Bounded CONTESTING evidence: agents completed substantial research engineering but failed the tested open-ended scientific-judgment task. Instance logged before AS-002; no prior corpus content modified.Fragmenting
2026-07-17FR-AI-0007IN-007 appended — Distributed Denial of Science (arXiv:2607.10712v1). Bounded CONTESTING instance on the correctness component: under the tested adversarial retrieval setting, correctness in autonomous scientific workflows is conditional on evidential provenance. The existential claim and previously demonstrated correct discoveries are preserved; the autonomous problem-identification boundary (BN-001/OQ-001/AT-001) is unchanged. No assessment issued; pressureState FRAGMENTING, verificationStage VS-03, mechanisms, and openQuestions unchanged. The separately authorised NEW RECORD PATH component is not implemented by this mutation.Fragmenting
Evidence Trajectories - PROG-AI
View Artificial Intelligence trajectoriesAI records are brought forward within the full archive.Frontier Records — PROG-AI
FR-AI-0001LLM Multi-Step Reasoning — Generalisation Beyond TrainingEscalating2 assessments2024-01-15FR-AI-0002LLM Knowledge-Work Utility — Economically Valuable Task PerformanceEscalating1 assessment2024-01-15FR-AI-0003RLHF Preference Generalisation — Behaviour Beyond Training DistributionFragmenting2 assessments2024-01-15FR-AI-0004Scaling Laws — Emergent Performance on Unseen TasksFragmenting2 assessments2024-01-15FR-AI-0005AGI Through Scaling — LLM Architecture as the Path to General IntelligenceFragmenting2 assessments2024-01-15FR-AI-0006Scaling Mechanism Coherence — Continuity Across Model SizesFragmenting1 assessment2024-01-15FR-AI-0007Autonomous AI Scientific Discovery — Novel, Correct, IndependentFragmenting2 assessments2024-01-15FR-AI-0008AI Medical Imaging Diagnosis — Specialist-Level Accuracy on Defined TasksFragmenting1 assessment2024-01-15
Programme Notes — PROG-AI
Landscape Essays — PROG-AI
No current programme-level diagnosis.