Rules the ranking cannot break
- Deterministic scoring. Identical inputs and weights produce identical ranks; reproducible from the audit package with one command.
- Provenance on every number: source, retrieved_at, confidence, licence note. A number without provenance does not exist.
- No silent imputation. Missing evidence is recorded as missing and penalized by a published rule, never estimated.
- Domains never cross-rank. Books, papers, reports, and standards are scored within their own domain.
- Each language ecosystem scores within itself first. Coverage gaps are declared, not hidden.
- People are context, not contestants. Persons, organizations, and platforms carry no score, ever.
- Manual decisions are records. Every override carries a written rationale and is published; Apparens-authored works are flagged.
- Humility on rank. A rank is a transparent output of declared evidence, weights, and missing-data rules at a release date, not a verdict on intrinsic worth.
Ontology v0.2 (frozen)
Canonical entities (book, paper, report, standard) are scored within their domain. Context entities (person, organization, platform) are described, never ranked: structurally, they carry no score field. Governance records (releases, challenges, overrides) are append-only.
Weighting scenarios
| Scenario | citation_count | library_holdings | readership_persistence | syllabus_adoptions |
|---|---|---|---|---|
| academic | 0.5 | 0.2 | 0.05 | 0.25 |
| broad_influence | 0.2 | 0.25 | 0.4 | 0.15 |
| governance_practitioner | 0.25 | 0.3 | 0.1 | 0.35 |
Missing-data penalty factor: 0.5. Normalization: per_domain_min_max. method_version 0.1-pilot. These are pilot placeholder weights; every change ships with a changelog entry.
What each signal means
- citation_count: all-time citations from OpenAlex (CC0). The scale of scholarly impact.
- readership_persistence: the number of distinct years a work keeps being cited (from OpenAlex counts_by_year). A longevity proxy: a work cited across many years scores higher than a one-year spike. It rewards enduring use, not recent volume.
- library_holdings, syllabus_adoptions: declared but not yet harvested for the pilot (WorldCat / Open Syllabus drops pending). Works are penalized for them by rule, never imputed.
Declared deferred capabilities
The method names these now and does not pretend they are done. Each is deferred openly, not silently stubbed:
- Per-ecosystem normalization (rule 5): scoring runs per domain today. Per-language normalization activates only once works from more than one ecosystem enter a scored domain. Until then the site does not claim worldwide or present-tense multilingual coverage. The Chinese-language spine is now a curated 63 books, but it carries no harvested metrics yet (the Chinese citation ecosystem is a separate, deferred harvester), so it is browsable, not scored.
- A fuller longevity proxy: library holdings over time, edition count, and continued availability, to complement readership_persistence.
- Book scoring: books are curated and browsable now but not yet scored; the pilot ranks papers only.
What is not here, and why
A reference is defined as much by what it excludes as by what it lists. These gaps are deliberate and declared, not oversights.
- The closed frontier ships no papers. Many of the most capable 2025 models, including the latest GPT, Claude, Gemini, Grok, and Llama releases, are documented only by a system card or a blog post, not a paper. A canon of the literature cannot rank what was never written down. We note this not as a complaint but as a finding: the most-discussed models are increasingly the least-documented, and open-weight and Chinese labs now carry most of the published record.
- Models are not entities; their papers are. The Canon ranks texts, not products. A model enters only through a primary paper or technical report. Where a model has none, it is absent by design, however important it is.
- Stable sources are preferred. We cite arXiv or a DOI wherever possible, because those are permanent and versioned. A few significant reports exist only as a PDF on a company's own site, such as Baidu's ERNIE 4.5. We include those sparingly and flag them, since vendor links can change or disappear.
- New entries are candidates, not verdicts. A freshly added paper is in the corpus but not yet scored. Scoring waits on harvested evidence, so a 2025 model report sits unranked until that evidence accrues. Candidacy asserts nothing.
- The corpus is still partial. Coverage is a pilot. The Chinese-language section in particular is openly under construction, and the paper set leans English. We would rather say so than pretend completeness.
How this was made
The Canon is curated and computed, with AI used as a drafting aid, never as the authority. To be exact:
- Ranks and scores are computed deterministically from declared evidence and weights. They are not generated by a language model, and they rebuild bit-identically from the audit package.
- Sources are real and human-checked. Every voice links to a source you can open, and the bibliographic record is reconciled against OpenLibrary and Crossref.
- Book and entry descriptions: many are AI-drafted from public sources, written to be neutral and factual, with no claims beyond what the work is about; each carries a confidence flag in the data. If one is wrong, challenge it and we will fix it.
- Voice biographies: AI-drafted from each voice's own cited source and the verified affiliation, written to be neutral and factual with no claims beyond the sources. If one is wrong, challenge it and we will correct it.
- Paper significance notes: the one-line notes on the paper shelf are AI-drafted, anchored to each paper's own record, neutral and factual. The verbatim abstracts are NOT AI text: they are the authors' own words, quoted with a source.
- The frontier map: the research-frontier review that surfaced the recent candidate papers was AI-run and human-checked, with every admitted paper verified against its arXiv record. It nominates candidates only; it scores nothing.
- The cover image (the person holding a phone) is AI-modified and labelled as such on the social card, EU AI Act style.
Where AI helped draft text, a human checked it against the evidence. Where evidence is missing, we say so rather than let a model fill the gap.
Cite this method
The method is documented in a citable note (Corpus Cognitivum), archived with a DOI: doi.org/10.5281/zenodo.21042034 (concept DOI, always the latest version). It is licensed CC BY 4.0.
Janssen, J. (2026). The AI Canon: a method for auditable knowledge curation (Corpus Cognitivum). Apparens. https://doi.org/10.5281/zenodo.21042034