What changed, and when
Every change to the protocol, the data, published findings and the site itself —
newest first, each entry linking to the evidence it rests on. Entries older than this page live in the
notes and the dataset release notes on Zenodo; nothing before 2026-09-03 was retroactively
reconstructed.
2026-09-26protocolsite
Protocol errata: measurement algorithm pinned, arm order a MUST, references pinned by checksum
Four clarifications to the specification, folded into the
site edition (the Zenodo snapshot predates them; they merge into v1.2):
- §4.1 measurement algorithm is now fully specified — trapezoidal integration over the sampler
thread's clock, energy spanning first-to-last sample with no edge extrapolation, dropped reads counted and
reported (fewer than two usable samples ⇒ discard), and the pooled-ratio iteration aggregation
E/token = ΣE/Σtokens. All five items describe exactly what the reference container does.
- §4.6.2 arm order: SHOULD → MUST. Order and thermal state demonstrably move results, so
randomised or counterbalanced order is now mandatory; a fixed order is a disclosed deviation.
- Level C renamed "Standard-candidate" (was "v1.0-grade" — confusing next to protocol v1.1), and
the document status is now "stable core / candidate specification" rather than plain "stable".
- §5.1 pinned artifacts — release tag
schema-1.3-r1, commit, SHA-256 checksums of the
schema and validator, and the image digest, so a conformance claim is a claim against immutable pins, not a
moving main. The container quickstart gains ECOCOMPUTE_REF and a
checksummed release variant replaces curl | bash on main
for reproducible runs.
Site polish in the same pass: the homepage mobile nav is now a proper menu button (nothing truncated),
button time-notes sit below their titles instead of side-by-side, and two absolute phrasings are softened
("Nobody measures…" → "Few open tools measure…"; "Every number is reproducible" → "every published
measurement is accompanied by its reproducibility artifacts").
2026-09-26protocolsite
The spec–schema loop is closed: schema-valid is no longer mistaken for protocol-conformant
The public energy.schema.json checked structure only — a report could pass it while
violating Protocol v1.1 (no thermal block, no same-session FP16 baseline, arbitrary version strings,
sample_rate_hz below 10 Hz). Two layers now close that gap in the
container repo:
the schema itself is tightened (version-enum and const fields, minimum 10 Hz sampling, and at
schema 1.3 the workload/measurement/software MUST-fields become required — no existing valid report
breaks), and a new semantic validator
(ecocompute validate --profile v1.1-core) checks every machine-checkable §4 MUST with
clause-numbered violations and CI exit codes. The compliance levels are renamed to say what they actually
check: Schema-valid / Protocol-conformant / Dataset-eligible (A/B/C unchanged in substance; see
/schema/ §3 and the
spec §5). One honest consequence, stated rather than
hidden: the 2026-09-25 two-card re-test reports themselves grade as schema-valid but not
protocol-conformant — the container does not yet emit the thermal block (tracked issue), and the
validator's test suite pins exactly that verdict.
2026-09-25protocol
The spec layer becomes a standalone, DOI-tracked specification document
Protocol v1.1 is extracted from the method page into an independent normative
document — EcoCompute Protocol v1.1 — written in the
W3C-Note / IETF-RFC style: explicit MUST / SHOULD / MAY requirements for the power source
(NVML GPU-package, ≥ 10 Hz), the same-session FP16 baseline, the workload (batch 1, 256 tokens),
the generation window and the report fields; SHOULD-level replication (n ≥ 2, n travels with the
data); MAY-level extras (additional modes/batches/formats, the draft environment block), each labelled and
never pooled. The document binds the protocol to the machine contract — a conforming submission
MUST pass ecocompute-energy/1.2+ validation — and carries the protocol ↔ schema ↔
dataset version table. The document is archived on Zenodo at
DOI 10.5281/zenodo.22958675 (concept:
10.5281/zenodo.22958674) so papers can cite
"EcoCompute Protocol v1.1, DOI: 10.5281/zenodo.22958675" instead of a moving web page; dataset DOIs and the
protocol DOI stay separate. The
method page remains the human-readable guide and now states that the normative text
is the document. Source: spec/ in the site repository.
2026-09-25siteprotocol
The site is restructured into four layers, and a number bug is fixed
The single mixed homepage is split into Spec (the protocol, normative),
Evidence (everything measured), Changelog (this page) and
Findings (exploratory results, explicitly not counted in the coverage matrix). No existing
URL changed; /changelog/ is the only new one. A number glossary now
defines every count on the site with its denominator, and the homepage hero shows a single number (26 of 56
coverage-matrix cells). Also fixed: the meta description claimed "26 of 56 cells are still empty" — the empty
count is 30; the measured count is 26. Search snippets had been propagating the wrong number.
2026-09-25data
RTX 4090 NF4 re-tested across two physical cards: July's +0.8% anchor holds
Qwen2.5-3B NF4 vs FP16, three independent sessions across two physical RTX 4090s (different UUIDs),
current stack (torch 2.14.0, CUDA 13.0, bitsandbytes 0.50.2 — the same stack as the 2026-09-20 RTX 5090
re-test): +3.5% / +0.7% / +0.5%, mean +1.6%. Conclusions: the July n = 1 +0.8% anchor
is confirmed (Ada 3B is break-even to a small penalty); the −15.1% from the 2026-08-19 paired-quality
session was not reproduced in three current-stack sessions — its stack and conditions differ from the
present run and no specific error was identified, so it is retained as a divergent historical observation and
excluded from current-stack conclusions; and at the 3B anchor the RTX 5090's stack flip
(crossover ≈5B → ≈1.8B) did not appear on the RTX 4090 — only that one anchor was re-measured,
and whether the full Ada crossover curve moves under the current stack awaits the 1.5B and 7B points.
Observed gaps: the two cards differ by ≈ 2.8 points, the two sessions on card 2 by ≈ 0.2 points
(card 1 was measured once, so this is a description, not a variance decomposition; for card-population questions
ncard = 2).
Kept as a separate versioned archive
(data/rtx4090_bnb_2026-09-25.csv,
with power-trace sidecars and the generator script) — not pooled into the seed counts or the fitted curves,
per the same versioning policy as the RTX 5090 re-test. Stack differs from the published pins (bitsandbytes
0.50.2 ≠ 0.43.3), so vs_fp16 is not directly comparable with image runs; the reports say so.
2026-09-24dataprotocol
All energy numbers unified on the generation measurement window
The coverage matrix and every llama.cpp number now state one window: generation only — model load,
on-load quantization and warm-up excluded. The bitsandbytes rows always measured this way (the container starts
its sampler after warm-up; the grid once mislabelled them whole-process), and the llama.cpp GGUF row was
re-cut from its archived 100 Hz power traces: at 576 output tokens it reads
−62.9% on the counter basis (the previously displayed −63.0 was a rounding artefact of the full-precision
−62.9481). Window choice is load-bearing: one session re-integrated over whole-process vs generation vs
decode-only windows moves its 64-token figure by up to 12.5 percentage points (576-token figure: within 2).
Every energy claim now states its window and output length.
Evidence: data/rtx4090_llamacpp_window_comparison_2026-09-24.csv
(generated from the 45 archived traces by build/make_window_comparison.py, assertion-guarded).
2026-09-20datafinding
The RTX 5090 crossover moved from ≈5B to ≈1.8B on a current stack
Re-measuring the same RTX 5090 months later, on bitsandbytes 0.50.2 / torch 2.14 / CUDA 13.0, moved the NF4
break-even size from ≈5B to ≈1.8B — nothing about the hardware changed. A crossover is a property of the
software stack, not only of the architecture. 22 runs, n = 2 across a full instance restart; the two
sessions agree to 2.1 points on average, 5.2 at worst. INT8 stayed positive at every size (+55% at 7B, +256% at
0.5B). A torchao FP8 companion sweep found both FP8 code paths (weight-only, dynamic activation+weight) more
expensive than FP16 at every size, failing in opposite directions.
Maintained as a separate archive — not pooled into the seed counts or the fitted
curves, which still describe the 0.4x-era stack: pooling the two would produce a crossover neither session
measured. Data: rtx5090_bnb_2026-09-20.csv,
rtx5090_fp8_torchao_2026-09-20.csv;
raw reports archived at 10.5281/zenodo.22855133.
2026-09-03finding
The kernel decides the sign, not the bit width
Inside one runtime on one RTX 4090: llama.cpp GGUF Q4_0 decoding Llama-3.1-8B-Instruct costs
61.9% less energy per token than GGUF F16 (for +5.60% perplexity) — the quantized arms draw
more power (296–319 W vs 273 W) and win purely on throughput. On the same card, bitsandbytes
LLM.int8() costs 106% more. Same idea, opposite sign: format-level claims ("4-bit saves
energy") are meaningless without the kernel. Corrected the same day over 45 runs (n = 5, randomized
order, cooldown before every run, hardware energy counter): the provisional −63.6% was an overestimate.
Note: the note and raw data
(10.5281/zenodo.22295184)
predate the 2026-09-24 window unification; where the two disagree on a number, the window-comparison CSV is
authoritative.
Older history — the v1.1.0 dataset release, the July RTX 4090 deep dive, the August INT8 repeats —
is documented in the notes and in each Zenodo record's release notes, not reconstructed here.