paper-0124 · paper · 2020
Jared Kaplan et al.
Loss as a power law of compute, data, parameters; the industry's planning document.
Abstract
This paper develops a transport-validity theory for agentic AI interventions that are first screened on small systems and later considered for frontier-scale deployment. Rather than predicting absolute frontier performance, it asks when a comparative gain observed at small scale can be carried forward without overclaiming. The analysis targets an explicitly delimited class of operationally isolatable interventions whose effects can be compiled from logged event-local channels with bounded spillover and replayable extraction maps. The paper proves structured failure modes for naive extrapolation, including sign reversal under bottleneck-weight shift and the vacuity of observable closeness when a descriptor omits a sign-relevant coordinate. It then develops a constructive positive framework based on executable lower certificates: a two-stage compiler architecture, replayable identified sets for stage-one mode laws, descriptor-language growth audits, witness-cover transport certificates, branch-local evaluation bridges, confidence-calibrated audit rules, and portfolio-level frontier allocation under interaction risk. The result is a finite, machine-readable, and operational framework for deciding which small-scale architectural improvements—such as decomposition, tool routing, memory policy, verifier coupling, orchestration, and related inference-time interventions—deserve expensive frontier trials, and how scarce frontier budget should be allocated among them. The paper does not claim a law for AGI timelines and does not remove the need for frontier experimentation; its claim is narrower and practical: to provide replayable, falsifiable conditions for responsible scale-up decisions in agentic AI research. [OpenAlex]
Academic, score -0.1966
| Metric | Status | Value | Norm. | Weight | Contribution | Source | Confidence | License | Provenance |
|---|---|---|---|---|---|---|---|---|---|
| citation_count | present | 1512.0 | 0.006802 | 0.5 | 0.003401 | OpenAlex | high | OpenAlex, CC0 metadata | link |
| library_holdings | missing | recorded as missing, penalized by rule, never imputed | −0.1 | recorded as missing; penalized by rule, never imputed | |||||
| readership_persistence | present | 8.0 | 0.5 | 0.05 | 0.025 | OpenAlex | medium | OpenAlex, CC0 metadata | link |
| syllabus_adoptions | missing | recorded as missing, penalized by rule, never imputed | −0.125 | recorded as missing; penalized by rule, never imputed | |||||
Broad Influence, score 0.0014
| Metric | Status | Value | Norm. | Weight | Contribution | Source | Confidence | License | Provenance |
|---|---|---|---|---|---|---|---|---|---|
| citation_count | present | 1512.0 | 0.006802 | 0.2 | 0.00136 | OpenAlex | high | OpenAlex, CC0 metadata | link |
| library_holdings | missing | recorded as missing, penalized by rule, never imputed | −0.125 | recorded as missing; penalized by rule, never imputed | |||||
| readership_persistence | present | 8.0 | 0.5 | 0.4 | 0.2 | OpenAlex | medium | OpenAlex, CC0 metadata | link |
| syllabus_adoptions | missing | recorded as missing, penalized by rule, never imputed | −0.075 | recorded as missing; penalized by rule, never imputed | |||||
Governance Practitioner, score -0.2733
| Metric | Status | Value | Norm. | Weight | Contribution | Source | Confidence | License | Provenance |
|---|---|---|---|---|---|---|---|---|---|
| citation_count | present | 1512.0 | 0.006802 | 0.25 | 0.0017 | OpenAlex | high | OpenAlex, CC0 metadata | link |
| library_holdings | missing | recorded as missing, penalized by rule, never imputed | −0.15 | recorded as missing; penalized by rule, never imputed | |||||
| readership_persistence | present | 8.0 | 0.5 | 0.1 | 0.05 | OpenAlex | medium | OpenAlex, CC0 metadata | link |
| syllabus_adoptions | missing | recorded as missing, penalized by rule, never imputed | −0.175 | recorded as missing; penalized by rule, never imputed | |||||
A rank is not a verdict on intrinsic worth. It is a transparent output of declared evidence, weights, and missing-data rules at a specific release date.
Disagree with this rank or a number? Challenge it with your evidence. Every challenge gets a public identifier and a published resolution.