paper-0172 · paper · 2023
Suriya Gunasekar et al. (Microsoft Research)
Primary technical report for a notable AI model (verified primary source).
Abstract
We introduce phi-1, a new large language model for code, with significantly smaller size than competing models: phi-1 is a Transformer-based model with 1.3B parameters, trained for 4 days on 8 A100s, using a selection of ``textbook quality" data from the web (6B tokens) and synthetically generated textbooks and exercises with GPT-3.5 (1B tokens). Despite this small scale, phi-1 attains pass@1 accuracy 50.6% on HumanEval and 55.5% on MBPP. It also displays surprising emergent properties compared to phi-1-base, our model before our finetuning stage on a dataset of coding exercises, and phi-1-small, a smaller model with 350M parameters trained with the same pipeline as phi-1 that still achieves 45% on HumanEval. [OpenAlex]
Academic, score -0.2105
| Metric | Status | Value | Norm. | Weight | Contribution | Source | Confidence | License | Provenance |
|---|---|---|---|---|---|---|---|---|---|
| citation_count | present | 99.0 | 0.000441 | 0.5 | 0.000221 | OpenAlex | high | OpenAlex, CC0 metadata | link |
| library_holdings | missing | recorded as missing, penalized by rule, never imputed | −0.1 | recorded as missing; penalized by rule, never imputed | |||||
| readership_persistence | present | 5.0 | 0.285714 | 0.05 | 0.014286 | OpenAlex | medium | OpenAlex, CC0 metadata | link |
| syllabus_adoptions | missing | recorded as missing, penalized by rule, never imputed | −0.125 | recorded as missing; penalized by rule, never imputed | |||||
Broad Influence, score -0.0856
| Metric | Status | Value | Norm. | Weight | Contribution | Source | Confidence | License | Provenance |
|---|---|---|---|---|---|---|---|---|---|
| citation_count | present | 99.0 | 0.000441 | 0.2 | 8.8e-05 | OpenAlex | high | OpenAlex, CC0 metadata | link |
| library_holdings | missing | recorded as missing, penalized by rule, never imputed | −0.125 | recorded as missing; penalized by rule, never imputed | |||||
| readership_persistence | present | 5.0 | 0.285714 | 0.4 | 0.114286 | OpenAlex | medium | OpenAlex, CC0 metadata | link |
| syllabus_adoptions | missing | recorded as missing, penalized by rule, never imputed | −0.075 | recorded as missing; penalized by rule, never imputed | |||||
Governance Practitioner, score -0.2963
| Metric | Status | Value | Norm. | Weight | Contribution | Source | Confidence | License | Provenance |
|---|---|---|---|---|---|---|---|---|---|
| citation_count | present | 99.0 | 0.000441 | 0.25 | 0.00011 | OpenAlex | high | OpenAlex, CC0 metadata | link |
| library_holdings | missing | recorded as missing, penalized by rule, never imputed | −0.15 | recorded as missing; penalized by rule, never imputed | |||||
| readership_persistence | present | 5.0 | 0.285714 | 0.1 | 0.028571 | OpenAlex | medium | OpenAlex, CC0 metadata | link |
| syllabus_adoptions | missing | recorded as missing, penalized by rule, never imputed | −0.175 | recorded as missing; penalized by rule, never imputed | |||||
A rank is not a verdict on intrinsic worth. It is a transparent output of declared evidence, weights, and missing-data rules at a specific release date.
Disagree with this rank or a number? Challenge it with your evidence. Every challenge gets a public identifier and a published resolution.