← ObservatoryThe RecordFR-AI-0002
PROG-AI
FR-AI-0002

LLM Knowledge-Work Utility — Economically Valuable Task Performance

Large language models can perform economically valuable knowledge-work tasks with limited human supervision.

EscalatingVS-02·since 2024-01-15
Verification Matrix
VS-01
Assertion
VS-02
Published
2024-01-15 — present
VS-03
Audit
VS-04
Replication
VS-05
Operation
State reached Current state Not yet reached
State Warrant
Current stateEscalatingVS-02
Why this state?The claim is supported by the current evidence in a qualified but meaningful sense. Three independent lines of evidence — controlled experiments (INST-001, INST-002), commercial deployment at scale (INST-003), and natural experiment in deployed settings (INST-005) — all find that LLMs produce measurable economic value in knowledge-work contexts under conditions approximating limited supervision. The effect sizes are not marginal: 14–55% productivity improvements in relevant task domains, with quality improvements accompanying rather than trading off against speed in at least two of the three studies (INST-002, INST-005). Contesting evidence is concentrated at the boundary of the claim rather than at its core: documented hallucination failures in high-stakes domains (INST-004) and agentic multi-step task failures (INST-006) show that the 'limited human supervision' condition holds reliably in single-turn, reviewed contexts but not yet in extended autonomous workflows. The pressure state is ESCALATING: the core claim is well supported within a supervision boundary that has not yet been precisely defined (OQ-001).
In this state since2024-01-15
Stage provenanceStored VS-02; historically unverified after legacy review.
Record Lineage — Chronological
2024-01-15
Record opened — Escalating
The claim is supported by the current evidence in a qualified but meaningful sense. Three independent lines of evidence — controlled experiments (INST-001, INST-002), commercial deployment at scale (INST-003), and natural experiment in deployed settings (INST-005) — all find that LLMs produce measurable economic value in knowledge-work contexts under conditions approximating limited supervision. The effect sizes are not marginal: 14–55% productivity improvements in relevant task domains, with quality improvements accompanying rather than trading off against speed in at least two of the three studies (INST-002, INST-005). Contesting evidence is concentrated at the boundary of the claim rather than at its core: documented hallucination failures in high-stakes domains (INST-004) and agentic multi-step task failures (INST-006) show that the 'limited human supervision' condition holds reliably in single-turn, reviewed contexts but not yet in extended autonomous workflows. The pressure state is ESCALATING: the core claim is well supported within a supervision boundary that has not yet been precisely defined (OQ-001).
Verification Stage: VS-02 preserved — historically unverified.
Mutation Log
MutationDateFieldPrior valueCurrent value
M-0072026-08-01instance_appendedIN-006IN-007
M-0062026-07-09description_reorderedDESCRIPTION-REORDERED
M-0052024-01-15programme_panel_addedPROGRAMME-PANEL-ADDED
M-0042024-01-15mechanisms_recordedMECHANISMS-RECORDED
M-0032024-01-15assessment_issuedASSESSMENT-ISSUED
M-0022024-01-15instances_loggedINSTANCES-LOGGED
M-0012024-01-15record_createdRECORD-CREATED
Evidence Sources
7 instances on recordShow sources ↓Hide ↑
IN-001GitHub Copilot productivity study — Peng et al. (Microsoft Research)supportive
IN-002Noy and Zhang — Experimental evidence on productivity effects of generative AI in professional writingsupportive
IN-003AI in legal practice — contract review and due diligence deploymentsupportive
IN-004Hallucination and reliability failure documentation — legal and medical contextscontesting
IN-005Brynjolfsson, Li, and Raymond — Generative AI at work (customer service study)supportive
IN-006Agentic deployment failures — early autonomous task completion attemptspartial
IN-007ORCA-bench — live-system root-cause analysis under limited supervisionCONTESTING