paper-0044 · paper · 2001
Leo Breiman
Named the split between data modeling and algorithmic prediction; prophetic for ML's rise.
Abstract
There are two cultures in the use of statistical modeling to reach conclusions from data. One assumes that the data are generated by a given stochastic data model. The other uses algorithmic models and treats the data mechanism as unknown. The statistical community has been committed to the almost exclusive use of data models. This commitment has led to irrelevant theory, questionable conclusions, and has kept statisticians from working on a large range of interesting current problems. Algorithmic modeling, both in theory and practice, has developed rapidly in fields outside statistics. It can be used both on large complex data sets and as a more accurate and informative alternative to data modeling on smaller data sets. If our goal as a field is to use data to solve problems, then we need to move away from exclusive dependence on data models and adopt a more diverse set of tools. [OpenAlex]
Academic, score -0.1720
| Metric | Status | Value | Norm. | Weight | Contribution | Source | Confidence | License | Provenance |
|---|---|---|---|---|---|---|---|---|---|
| citation_count | present | 1345.0 | 0.00605 | 0.5 | 0.003025 | OpenAlex | high | OpenAlex, CC0 metadata | link |
| library_holdings | missing | recorded as missing, penalized by rule, never imputed | −0.1 | recorded as missing; penalized by rule, never imputed | |||||
| readership_persistence | present | 15.0 | 1.0 | 0.05 | 0.05 | OpenAlex | medium | OpenAlex, CC0 metadata | link |
| syllabus_adoptions | missing | recorded as missing, penalized by rule, never imputed | −0.125 | recorded as missing; penalized by rule, never imputed | |||||
Broad Influence, score 0.2012
| Metric | Status | Value | Norm. | Weight | Contribution | Source | Confidence | License | Provenance |
|---|---|---|---|---|---|---|---|---|---|
| citation_count | present | 1345.0 | 0.00605 | 0.2 | 0.00121 | OpenAlex | high | OpenAlex, CC0 metadata | link |
| library_holdings | missing | recorded as missing, penalized by rule, never imputed | −0.125 | recorded as missing; penalized by rule, never imputed | |||||
| readership_persistence | present | 15.0 | 1.0 | 0.4 | 0.4 | OpenAlex | medium | OpenAlex, CC0 metadata | link |
| syllabus_adoptions | missing | recorded as missing, penalized by rule, never imputed | −0.075 | recorded as missing; penalized by rule, never imputed | |||||
Governance Practitioner, score -0.2235
| Metric | Status | Value | Norm. | Weight | Contribution | Source | Confidence | License | Provenance |
|---|---|---|---|---|---|---|---|---|---|
| citation_count | present | 1345.0 | 0.00605 | 0.25 | 0.001512 | OpenAlex | high | OpenAlex, CC0 metadata | link |
| library_holdings | missing | recorded as missing, penalized by rule, never imputed | −0.15 | recorded as missing; penalized by rule, never imputed | |||||
| readership_persistence | present | 15.0 | 1.0 | 0.1 | 0.1 | OpenAlex | medium | OpenAlex, CC0 metadata | link |
| syllabus_adoptions | missing | recorded as missing, penalized by rule, never imputed | −0.175 | recorded as missing; penalized by rule, never imputed | |||||
A rank is not a verdict on intrinsic worth. It is a transparent output of declared evidence, weights, and missing-data rules at a specific release date.
Disagree with this rank or a number? Challenge it with your evidence. Every challenge gets a public identifier and a published resolution.