EcoCompute.

EcoCompute Protocol v1.1 — an open micro-benchmark for quantization energy · 量化能耗微基准协议

Every quantization tool tells you how to quantize. Few open tools measure whether it saves energy on your card — so this is a protocol for measuring it, not a table to trust.

26 / 56 cells of the coverage matrix have a measurement. The other 30 are not an oversight — they are the reason this site exists. What these counts mean →

Scope of every number on this site: GPU-package power (NVML), not whole-system draw — no PSU losses, CPU, DRAM, cooling, PUE or CO₂e · single-stream decode, batch 1, 256 tokens · generation window (model load, quantization and warm-up excluded) · replication stated per point: main dataset n = 2 (CV < 2%), deep dives n = 1, re-tests and their stacks versioned separately · thermal mode: the seed coverage grid is unknown (pre-thermal-block sessions, spec §4.6.5 — unknown is not cold), the September re-tests are cold and modes are never crossed. In one line: bitsandbytes LLM.int8() never saved energy in any cell we measured, and NF4 changes sign with model size and GPU — the evidence layer holds every number, its n and its DOI. Limits and how to disprove us →

Public replications →

Four layers · 这个网站怎么读

The site is deliberately split so that what we require, what we measured, what changed and what we are still exploring never blur into one page.

1 · Spec — 规范

What makes two submissions comparable. The normative protocol, the container and schema that enforce it.

Protocol v1.1 · Number glossary · Container guide · Run it yourself · Schema

2 · Evidence — 证据

Everything measured, at its replication level — and every count with its denominator.

Evidence hub · Coverage matrix · Latest: two-card NF4 re-test · Public replications · Cite

3 · Changelog — 版本更新

What changed in the data, protocol and site — dated, typed, linked to its evidence.

All entries · Window unification (09-24) · NF4 two-card re-test (09-25)

4 · Findings — 探索性结果

Deeper write-ups and tools whose output is a model, not a measurement — never counted in the matrix.

Notes & findings · Variable matrix · Estimator · Optimizer

Inference energy per 1M tokens

Paste the verdict + citation into your tech spec, or share this exact config.

Crossover curve · in what we measured, the bigger the model, the more quantization saves

NF4 INT8 Your model Filled = n ≥ 2 repeated trials Hollow = n = 1 single trial Above zero = penalty (more energy) · below = savings

Solid lines connect measured points only (direct NVML). Marker fill encodes replication, not accuracy: filled dots are points repeated at least twice in the main dataset (v1.1.0, n = 2, CV < 2%); hollow dots on a faded line were measured once (n = 1) — the July 2026 RTX 4090 deep dive and one early RTX 4090D point — so they carry no measured spread and should be read as a single observation, not as a distribution. INT8 was measured on A800 and RTX 4090 only, and sizes outside each GPU's measured range aren't drawn here. For fitted extrapolation beyond the measured range, use the Your model tab.

Recent updates · 最近更新

2026-09-25 · data RTX 4090 NF4 re-tested across two physical cards — July's +0.8% anchor holds (mean +1.6%, n=3); the −15.1% historical reading was not reproduced and stays a divergent observation; no direction flip at the 3B anchor under the current stack. →
2026-09-24 · data All energy numbers unified on the generation measurement window; the llama.cpp row re-cut from its 100 Hz traces (−62.9% at 576 tokens, counter basis). →
2026-09-25 · site The site split into four layers — spec, evidence, changelog, findings; every count now has a defined denominator. →
Full history, each entry linked to its evidence: the changelog →