← ObservatoryThe RecordFR-AI-0007
PROG-AI
FR-AI-0007

Autonomous AI Scientific Discovery — Novel, Correct, Independent

AI systems can autonomously conduct scientific research that produces novel, correct discoveries.

FragmentingVS-03·since 2026-08-01
Verification Matrix
VS-01
Assertion
VS-02
Published
VS-03
Audit
2024-01-15 — present
VS-04
Replication
VS-05
Operation
State reached Current state Not yet reached
State Warrant
Current stateFragmentingVS-03
Why this state?AS-002 was issued following the operator-approved Post-Scout review of flag 2026-08-01-01. It is triggered by IN-008 (arXiv:2607.27191v1) and updates the current autonomy-boundary rationale without modifying AS-001, BN-001, OQ-001, AT-001, or any prior instance. The source is a primary preprint with two principal case studies and five total runs; its limits are preserved in IN-008.
Assessment summaryFRAGMENTING remains the correct pressure state, but the reason for fragmentation is now more precisely described. AS-001 treated autonomous problem identification as the decisive missing component because the strongest supportive examples — GNoME and FunSearch — solve human-framed problems. IN-008 shows that supplying the problem does not isolate the remaining difficulty: frontier agents can execute substantial literature review, coding, debugging, and experimentation while still failing at scientific judgment, including evidential prioritisation, abandonment of weak approaches, project-level backtracking, and recognition of publishable progress. The autonomy boundary therefore has at least two separable dimensions: who identifies the research problem, and whether the system can exercise adequate scientific judgment after the problem is specified. This new contesting evidence does not reverse the existential support supplied by IN-002 and IN-004, and its two-case preprint design is too bounded to justify a stronger negative state. It nevertheless changes the assessment's structure because autonomous research engineering can no longer be treated as evidence that only autonomous problem identification remains unresolved. IN-007 separately shows that correctness can fail through compromised evidential provenance. Together the 2026 evidence deepens fragmentation across autonomy and correctness while preserving the record's verified bounded discoveries. Pressure State and Verification Stage remain FRAGMENTING and VS-03.
State entered2024-01-15
Last reaffirmed2026-08-01
Mechanisms
BottleneckBN-001

"Autonomously conduct" lacks an agreed boundary. The claim requires autonomous research conduct, but the boundary between autonomous AI research and AI-assisted human research is contested. All current leading examples involve human-framed problems solved autonomously. Whether the claim requires only autonomous problem-solving (satisfied by GNoME and FunSearch) or also autonomous problem-identification (not yet demonstrated) is the critical definitional gap. This is the fifth lexical bottleneck in the corpus — and notably the second in PROG-AI within two records, following the same pattern identified at FR-AI-0006.

BottleneckBN-002

Novelty assessment is itself a research task. The claim requires that discoveries be novel, but establishing novelty requires surveying the accessible scientific literature — which is itself an incomplete and poorly indexed object. For fast-moving fields, a result that appears novel may have been anticipated in preprints, conference talks, or unpublished work. For large, old literatures, a result that appears novel may rediscover forgotten work. Novelty is not directly measurable from the discovery alone; it requires a comparison to the state of knowledge, which is itself uncertain. This is a measurement validity bottleneck of the same type as FR-BT-0002 BN-001: the measurement tool (literature survey) may not reliably track the thing it purports to measure (genuine novelty).

AttractorAT-001

Autonomous problem identification with verified novel correct results. The resolution path is a demonstration where an AI system identifies a previously unrecognised scientific problem, generates hypotheses about it, designs or conducts experiments, and produces results that are independently verified as correct and novel — without a human specifying the problem space. FunSearch and GNoME satisfy part of this; the problem-identification component is the remaining gap. Several AI research systems in development are explicitly targeting this boundary. The attractor is clearly defined and closer than analogous attractors in other records — the current evidence is within one component of satisfaction.

Assessment History
2024-01-15
Record opened — Fragmenting
The evidence is fragmenting across the three component claims. The correctness and novelty components are most strongly evidenced: GNoME (INST-002) and FunSearch (INST-004) both demonstrate AI systems producing results that are verified correct and independently novel in their domains. The autonomy component is more contested: in both cases, the research question was human-framed; the AI system discovered answers within a human-specified problem space rather than identifying the problem itself.
Verification Stage: VS-03 after ratified review (stored code VS-03 preserved).
2026-08-01
Reassessed, no change — Fragmenting
FRAGMENTING remains the correct pressure state, but the reason for fragmentation is now more precisely described. AS-001 treated autonomous problem identification as the decisive missing component because the strongest supportive examples — GNoME and FunSearch — solve human-framed problems. IN-008 shows that supplying the problem does not isolate the remaining difficulty: frontier agents can execute substantial literature review, coding, debugging, and experimentation while still failing at scientific judgment, including evidential prioritisation, abandonment of weak approaches, project-level backtracking, and recognition of publishable progress. The autonomy boundary therefore has at least two separable dimensions: who identifies the research problem, and whether the system can exercise adequate scientific judgment after the problem is specified. This new contesting evidence does not reverse the existential support supplied by IN-002 and IN-004, and its two-case preprint design is too bounded to justify a stronger negative state. It nevertheless changes the assessment's structure because autonomous research engineering can no longer be treated as evidence that only autonomous problem identification remains unresolved. IN-007 separately shows that correctness can fail through compromised evidential provenance. Together the 2026 evidence deepens fragmentation across autonomy and correctness while preserving the record's verified bounded discoveries. Pressure State and Verification Stage remain FRAGMENTING and VS-03.
AS-002 was issued following the operator-approved Post-Scout review of flag 2026-08-01-01. It is triggered by IN-008 (arXiv:2607.27191v1) and updates the current autonomy-boundary rationale without modifying AS-001, BN-001, OQ-001, AT-001, or any prior instance. The source is a primary preprint with two principal case studies and five total runs; its limits are preserved in IN-008.
Claim Lineage
1955–90
Early AI discovery systems. DENDRAL (1965) and AM (1976) demonstrate early AI systems generating hypotheses in chemistry and mathematics. The claim's aspirational form is established; the capability is far from practical demonstration.
2020–22
AlphaFold and domain-specific breakthroughs. AlphaFold demonstrates AI-enabled discovery at unprecedented scale in structural biology. The claim transitions from aspiration to active frontier. Autonomy remains limited to execution within human-framed problems.
2023
GNoME, FunSearch, and generative discovery. Systems demonstrating autonomous generation of novel, experimentally verified results in materials science and mathematics. The correctness and novelty components are strongly evidenced in constrained domains. Autonomy at the problem-generation level remains partial.
2024
End-to-end autonomous systems and institutional restructuring. AI Scientist demonstrates the full research loop; major institutions begin restructuring around AI-assisted discovery. The claim enters FRAGMENTING as component claims diverge in evidential strength.
Open Questions
OQ-001

Does "autonomously conduct scientific research" require autonomous problem identification, or is autonomous problem-solving within human-framed domains sufficient? BN-001 cannot close until this is resolved. The claim's satisfaction hangs on this distinction.

Raised 2024-01-15
OQ-002

INST-005 is the sixth occurrence of anticipatory institutional evidence and the first within PROG-AI. Does it fit the existing taxonomy of act types (commercial commitment, regulatory preparation, community standards tightening), or does institutional reorganisation constitute a fourth act type? The Broad Institute and EMBL restructuring is neither a commercial contract nor a regulatory act — it is a scientific workflow redesign. This may be relevant to a fourth act-type option within that developing taxonomy.

Raised 2024-01-15
OQ-003

BN-002 (novelty assessment as a measurement validity bottleneck) is structurally similar to FR-BT-0002 BN-001 (biological age measurement validity). Both are cases where the measurement tool may not reliably track the thing it purports to measure. Two occurrences of this specific bottleneck structure across two programmes. Has measurement validity as a distinct resistance/bottleneck type now reached watchlist elevation?

Raised 2024-01-15
Mutation Log
MutationDateFieldPrior valueCurrent value
M-0132026-08-01assessment_issuedAS-001AS-002
M-0122026-08-01instance_appendedIN-007IN-008
M-0112026-07-17instance_appendedIN-007
M-0102026-07-14vector_correctedneutral--constrained-autonomy-boundary-untouchedNEUTRAL
M-0092026-07-14instance_appendedIN-006
M-0082026-07-09reference_correctedREFERENCE-CORRECTED
M-0072026-07-09description_restoredDESCRIPTION-RESTORED
M-0062026-07-09description_reorderedDESCRIPTION-REORDERED
M-0052024-01-15programme_panel_addedPROGRAMME-PANEL-ADDED
M-0042024-01-15mechanisms_recordedMECHANISMS-RECORDED
M-0032024-01-15assessment_issuedASSESSMENT-ISSUED
M-0022024-01-15instances_loggedINSTANCES-LOGGED
M-0012024-01-15record_createdRECORD-CREATED
Evidence Sources
8 instances on recordShow sources ↓Hide ↑
IN-001AlphaFold2 — protein structure prediction at scalepartial
IN-002GNoME — graph neural network materials discoverysupportive
IN-003AI Scientist (Sakana AI) and early autonomous research systemspartial
IN-004FunSearch and mathematical discovery — verified novel results in combinatoricssupportive
IN-005Major lab restructuring around AI researchers — anticipatory institutional evidencepartial
IN-006Structured Concept Evolution — LLM-driven discovery of qLDPC code familiesNEUTRAL
IN-007Distributed Denial of Science — indirect data poisoning of autonomous research agentsCONTESTING
IN-008Open-ended AI research case studies — engineering competence without successful scientific judgmentCONTESTING