Append-only

Changelog

Append-only. Scoring-logic or weight changes must land with an entry here (rule 10).

pilot-v0.3: full-provenance hash, the override channel, fail-loud penalty (2026-07-02) <a id="pilot-v03"></a>

A new frozen release. Weights and both signals are UNCHANGED; ranks are identical to

pilot-v0.2. What changes is what the hash protects and what the machinery enforces:

provenance_url, retrieved_at, confidence, license_note), so tampering any provenance field

fails verification exactly like tampering the value would. New corpus_hash cd8459a7…,

verified bit-identical on rebuild. `release --verify` additionally re-compares every

per-work breakdown and coverage.json, and checks the raw evidence cache against its

harvest-time manifest first.

basis of its source, e.g. "OpenAlex, CC0 metadata"), and the adversarial review's

provenance check requires confidence and license_note downstream, not just at ingestion.

(similarity 0.85 to 0.95) is BLOCKED until a human decision exists in data/overrides/,

written through the ontology's Override model, whose mandatory-rationale validator now

actually runs; records are append-only. Tested end to end.

instead of silently disabling the penalty with a 0.0 default.

gate check: no function over 60 lines, no module over 600 in src/canon), the corpus loaders

are public API instead of underscore-imports, and the domain min-max no longer rebuilds

per-work dicts quadratically. None of this changes any score.

Hardening and decomposition: render boundary, toolchain, god-module split (2026-07-02)

No scoring change; ranks, weights, and both released corpus_hash values are unchanged, and

the generated site was verified byte-identical across the refactor (A/B tree hash).

also HTML-escaped (safe_url passes quotes through by design; seven sites fixed); model page

ids, frontier statistics, canonical/og:url, and the mailto challenge link (now URL-encoded)

are escaped; work-page titles no longer double-escape; the footer's inline style attributes,

which the site's own CSP silently blocked, became classes. Hostile-fixture tests inject

attack strings through every renderer and assert no breakout.

dash-rewritten, so an author's em-dash survives to the page; the no-em-dash rule [S3] is

scoped to generated copy and both directions are tested (no em-dash in copy, em-dash kept

in quotes). The cover image is now consistently described as AI-modified, matching the

card's own label; the social card was re-rendered at pilot-v0.2.

manifest sha256 and refuses altered bytes; `release --verify` checks the whole raw cache

first; raw record names reject path separators; harvester responses are size-capped.

`make install`, CI, and the audit bundle's reproduce harness; the gate fails on venv drift.

A one-off scan of the full git history found no secret signatures; the secrets guard now

also knows modern key shapes and scans the local pilot-frontier workspace when present.

package (theme, shell, page renderers, JS assets), verified byte-identical A/B; the

adversarial-review harness is one function per check; corpus counts in page copy are

computed from the seed data instead of hand-typed (they used to require a manual sweep on

every corpus change); the release version is bumped in one place; the governance record now

validates through the frozen ontology's Release model before it is written.

no check fails, and so does an undocumented check); the pytest bucket is labelled with what

it actually asserts; [S11] also rejects cloud-sync duplicate files in the deployable.

Evidence currency: registers, abstract provenance, disclosure completeness (2026-07-02)

No scoring change. This pass makes the trust surface current and every quoted abstract

checkable, and adds gate [S17] so it stays that way.

(327 OK; 12 year flags, all edition-vs-first-publication artifacts plus two bad OpenLibrary

records, kept for adjudication) and 269/269 papers against Crossref and arXiv (173 OK,

21 OK-without-DOI; 67 DOIs and 106 arXiv ids attached). The parallel run degraded under API

throttling, so verified rows from the prior run were kept (a verified hit does not rot) and

a sequential, polite arXiv pass resolved 72 more. The remaining REVIEW rows are database

coverage gaps (mostly the Chinese spine and pre-digital classics), recorded, not hidden.

a license note. 145 OpenAlex texts were matched exactly against the write-once raw cache

(the abstract must reconstruct from the cached record, or no URL is recorded); all 64 arXiv

texts were verified against the live arXiv abstract (similarity 0.93 or better, most exact);

the 7 Chinese journal abstracts already carried their URLs. Zero unverified URLs, zero gaps.

verbatim abstracts contradicted. The abstracts are now explicitly carved OUT of the CC BY

grant (they remain the authors'/publishers', quoted for identification and scholarship, with

a rights-holder removal route), and the exception is stated on the data page, the papers

page, and the footer.

significance notes and the AI-run, human-checked frontier review (the papers page says so at

point of display); abstracts are explicitly labelled as the authors' words, not AI text, and

now link to their source. The social card is re-rendered at PILOT v0.2 and the og:image alt

text notes the AI-use label.

against their live sources; Qiang Yang's source upgraded to his personal HKUST page, which

also covers federated learning and the WeBank role; Andrew Yao's bot-gated ACM source

replaced with the live Tsinghua IIIS page). Zero QUEUED remain.

control character plus "3", breaking the expand markers on abstracts and library

descriptions; and 3 abstracts carried OpenAlex indexing artifacts (JATS sup tags on Shannon,

an IEEE end-of-text marker on Rabiner, an HTML entity in Toolformer), now stripped

surgically while leaving author-written angle-bracket text intact.

from an unchecked source, every abstract carries provenance or a declared gap, and the

license carve-out cannot silently disappear.

Governance correction: the change process now enforces itself (2026-07-02)

Scoring-logic change recorded retroactively (rule 10). On 2026-06-29 (commit 8845a59) the

second signal was redefined: `sustained_readership` (citations in 2023-2025, a recency proxy

wearing a longevity name) became `readership_persistence` (the number of distinct years a work

keeps being cited, a genuine longevity proxy). That commit renamed the metric in scenarios.yaml,

the fixtures, and the harvester, and should have landed with an entry here. It did not. This

entry is the retroactive record; the definition itself does not change today.

audit bundle on 2026-06-30), violating the append-only promise (rule 11). The files are kept

exactly as they stand: `data/releases/pilot-v0.1/ADDENDUM.md` records the incident, the

original bytes remain in git history, and the bundle still self-verifies offline

(corpus_hash MATCH, re-checked 2026-07-02).

pilot-v0.2 to tree hashes; `release --check-frozen` verifies every pin, `release.build()` and

`bundle()` refuse a frozen version, and the registry itself is append-only (no re-freeze).

scenarios.yaml or src/canon/score.py must carry a CHANGELOG.md entry in the same commit.

run, which is how v0.1 was silently rewritten. Release tests now build a scratch version;

the red-team test reviews the committed release as an outside reader would.

release.py and the reviewer note in redteam.py) and regenerated reports/red_team_findings.md.

Corrected an overclaim on the press page and in ARCHITECTURE.md: the conflict-flagged book is

subject to the same rules as everything else, but books carry no metrics yet, so it is not

"scored" and the copy no longer says it is.

No scoring change today: weights, method_version, and both released corpus_hash values are

unchanged.

No scoring change: the new papers are candidates in the Chinese citation ecosystem, which

has no harvester yet, so they are browsable, not scored.

Chinese spine deepened (2026-06-30)

No scoring change: these are curated candidates, browsable not scored. The Chinese citation

ecosystem needs its own harvester, which stays deferred.

pilot-v0.2: paper evidence harvest (2026-06-30)

A new frozen release. The method, weights, and ontology are unchanged (method_version

0.1-pilot); pilot-v0.1 stays on disk as the prior frozen record. What changed is the

evidence base: more papers now carry harvested metrics, so the ranking is computed over a

larger, stronger corpus. corpus_hash c379e0a8…, verified bit-identical on rebuild, GATE A pass.

Post-launch enrichment (2026-06-29)

No scoring-logic or weight changes. The pilot ranking is unchanged; the work below adds

provenance, descriptions, and corpus coverage, all as candidates pending evidence harvest.

Provenance: every context entity is now checkable

Voices: stale roles corrected, biographies added

Papers: model coverage, Chinese-first (162 to 214)

Models index

Library: every book now described (250 to 573)

Bibliographic reconciliation

Transparency and security

seed v0.3: Stage A skeleton + harvest layer (2026-06-29)

Sprint 1 + CAN-07: method package & seed import

weights in `scenarios.yaml` (3 placeholder weightings).

250 descriptions; 573 categorized_as + 172 authored_by edges; Trap conflict-flagged.

Sprint 2: harvest layer (CAN-09 / CAN-10 / CAN-12)

cached snapshot (offline-deterministic); no match / offline => declared gap, never imputed.

into `data/resolved/metrics.json` with a `coverage.json` report.

without evidence are an honestly-declared coverage gap, not a fabricated zero.

Sprint 3: release builder, audit package, adversarial review (CAN-15/16/17)

(domain, scenario), full per-work breakdowns, divergence summary, a `Release`

governance record with a deterministic `corpus_hash` (date is metadata, not hashed),

coverage.json, and REPRODUCE.md. `--verify` rebuilds and asserts bit-identical (rule 3).

domain isolation, no-imputation, conflict-flag surfacing, declared coverage, ranking

sanity, divergence honesty → `reports/red_team_findings.md` + a GATE-A verdict.

`sustained_readership` = citations in 2023-2025 (recent momentum), distinct from all-time

`citation_count`. Assembled 174 metrics (88 citation_count + 86 sustained_readership).

(`scenario_divergence: observed`). The method's central claim is demonstrated, not merely asserted.

`ranking_sanity` check (false positive); fixed to assert composite-score monotonicity; iteration 2 clean.

drops; ~74 papers await the next OpenAlex daily-budget window.

Stage C: public site (CAN-21..25)

`site/` from the release JSON + seeds: Canon-50 (3 scenario views), per-work breakdown

pages (the trust surface: every metric + provenance + missing-data penalty), papers shelf,

method, challenges, changelog, and a downloadable audit package under `site/audit/`.

is GENERATED (top-3 injected by the builder, idempotent) so the manifesto never carries

hand-typed ranking data that can drift.

`apparens.nl/ai-canon/` (Cloudflare Pages, static).

Design pass: align to apparens.nl + house style

the white Apparens logo + serif wordmark, white body, orange `#B8430A`, DM Serif + DM Sans.

visually consistent and rebuilds from one place. Its Canon-50 teaser is the live top-3.

Acceptance audit response (decisions 1 to 6)

provenance, descriptions where written and "Description pending" otherwise, conflict-of-interest flag

shown inline, labelled candidacy not canonical. Books are curated and browsable but not yet scored.

grouped by category, alphabetical within category, labelled "described, never ranked" (no score).

page; significance lines added to the papers shelf.

(rule 5) activates only when more than one ecosystem enters a scored domain, so the site makes no

worldwide / present-tense multilingual claim and the Chinese spine (28 works) is a declared gap; a fuller

longevity proxy (holdings over time, editions, availability); and book scoring. The pilot ranks papers

only, behind honest framing, and that scored view passed GATE A.

the build on any em-dash in generated HTML.

byte-deterministic archive carrying the pipeline code, weights, pinned data snapshot, release outputs,

and a one-command `reproduce.sh`. Verified: extracted into a clean directory with no repo, it rebuilds

the release and reports corpus_hash MATCH. This is what makes the package archival and time-invariant.

Security hardening to the app's bar (v1.2, the [S##] guardrails)

Derived from the AI Control Index app's posture and adapted for a static site.

plus X-Content-Type-Options, Referrer-Policy, X-Frame-Options DENY, COOP, CORP,

Permissions-Policy, HSTS. [S5]

style remains in any page. [S6]

no Google Fonts. [S7]

(javascript:/data: collapse to `#`); adversarial XSS fixtures prove hostile titles,

descriptions, and URLs cannot become markup or script. [S8]

build if ARCHITECTURE.md and the checks drift. ARCHITECTURE.md added with the [S##] system.

Fixed body links to be underlined (distinguishable without color) and footer text contrast

(a global `p` rule was rendering footer text dark-on-dark); promoted a heading to fix order.

A static a11y lint (lang, single h1, img alt, heading order) keeps it from regressing in CI.

Published

**10.5281/zenodo.21042034**. The method note, README badge, CITATION.cff, and the site Method

page all cite it.

Not yet

Book metric harvesting (title collisions: deferred), CN verification toward 60-90,

more harvested metrics (next OpenAlex daily window + WorldCat/Open Syllabus drops),

deploy the site to Cloudflare Pages.