Inference energy per 1M tokens
Crossover curve · in what we measured, the bigger the model, the more quantization saves
Solid lines connect measured points only (direct NVML). Marker fill encodes replication, not accuracy: filled dots are points repeated at least twice in the main dataset (v1.1.0, n = 2, CV < 2%); hollow dots on a faded line were measured once (n = 1) — the July 2026 RTX 4090 deep dive and one early RTX 4090D point — so they carry no measured spread and should be read as a single observation, not as a distribution. INT8 was measured on A800 and RTX 4090 only, and sizes outside each GPU's measured range aren't drawn here. For fitted extrapolation beyond the measured range, use the Your model tab.
Advanced options · GPU, precision, batch & context
Estimated energy change vs FP16
Fitted crossover curve · your model marked with its uncertainty
Hollow anchors were measured once (n = 1, the RTX 4090 deep dive) and filled anchors come from the replicated main dataset (n = 2, CV < 2%); both feed the fit, but a hollow point carries no measured spread. The curve is a fitted extrapolation calibrated to the measured anchors (dots) — not new measurements. Beyond the largest measured anchor the slope is not measurement-anchored, and INT8 on non-A800 architectures is fitted. Treat large-model (≈≥7B) points as directional — see About → How we measure. Each fitted curve also inherits the software stack its anchors were measured on — the crossover is sensitive to the stack as well as the architecture (on one unchanged RTX 5090, a newer stack moved it from ≈5B to ≈1.8B; the re-test), so treat the fitted crossing as card-plus-stack, not hardware destiny.
What each fitted curve rests on: Ada NF4 9 anchors (RTX 4090 ×5, RTX 4090D ×4) · Ada INT8 6 (RTX 4090 only) · Blackwell NF4 4 · Turing NF4 4 · Ampere NF4 3 and INT8 3 (A800). Ada is the only class that pools two different cards, and they disagree by roughly 20 points at the same model size, so the Ada band is the widest of the four — see the methodology page for the leave-one-out error per class.
Report schema output (JSON)
From static query to constrained optimization: give your GPU, model size, an objective and optional budgets — the tool exhaustively searches precision × batch × context, filters by your constraints and returns the objective-optimal config, alternatives and the energy↔latency Pareto frontier. Energy is measured-anchored; latency/throughput are a roofline model; VRAM is computed — each field is labelled below.
Advanced constraints — latency / VRAM / throughput budgets (optional)
Recommended configuration
Alternatives
Energy ↔ latency Pareto frontier
Non-dominated trade-offs: no other config is both lower-energy and lower-latency. Down-left is better; ★ is the recommended config.
Every published measurement on this site is accompanied by its reproducibility artifacts. The EcoCompute energy MLCube container runs the same
measurement on your GPU — direct NVML power sampling, no TDP guesses — and writes an
energy.json with the exact same fields the charts here use. Run it, then drop the result
below to overlay your measurement on the crossover curve. No GPU of your own? The
Colab notebook
does the same measurement on a free T4.
Full guide, output schema, scope and how to contribute a run: the container page.
1 · No GPU? Run the notebook in Colab
A free Colab T4 is enough — no GPU of your own, no Docker, nothing installed locally. The notebook runs the same NVML measurement as the container and writes the same energy.json, then prints an overlay link and a block of text you can paste into an issue.
It measures your chosen precision and its own FP16 baseline, in randomized order, waiting for the card to come back to idle before each arm — the two things we got wrong the first time and wrote up. Both arms have to fit: on a 16 GB T4 that means models up to about 3B, and the notebook refuses to start rather than failing with an out-of-memory error twenty minutes in. Budget 20–35 minutes; most of it is cooldown and the model download.
One pass is n=1 and energy is integrated from sampled power — on our RTX 4090 the GPU's hardware energy counter read on average 14.9% higher than integrating the same traces, so read these figures at that resolution. The last cell repeats both arms 5× if you want a repeatability figure. Same scope as everything else here: GPU-package power, not wall AC; not a certified benchmark result.
2 · Have a GPU? Run the container
One command, nothing to configure. Needs an NVIDIA GPU. With Docker it runs the prebuilt image; on a rented GPU instance that cannot nest containers (AutoDL, vast.ai …) the same command installs into a venv instead and runs the identical code. No GPU? It still runs and returns dataset-derived reference values, clearly flagged (never faked).
Measures NF4 and its own FP16 baseline on TinyLlama-1.1B, writes a schema-validated energy.json to ./ecocompute-out/ and prints an overlay link. The measurement itself takes a few minutes; the first run is dominated by downloads — the one-time image pull (or dependency install) plus a ~2 GB model download — and later runs reuse the caches. Our one timed end-to-end run took 57 minutes on a China-hosted rented RTX 4090 (mostly fetching torch/CUDA wheels; since fixed to use a domestic index, but not re-timed). No flags: the architecture comes from the NVML device name, so your card cannot be mislabelled. Other models/precisions and the plain docker run form: the container page.
Repo & full guide: github.com/hongping-zh/ecocompute-mlcube. This is a supplemental energy-methodology container (MLCube-compatible), not a certified benchmark.
The script ends with a link like quantenergy.tech/?tab=run&overlay=… — the point is encoded in the URL, so opening it restores your marker below with nothing uploaded.
3 · What you get — the same structure as this site
Example energy.json. The two fields the charts read are highlighted: vs_fp16_energy_pct (ΔE% vs FP16) and basis (measured / interpolated / extrapolated).
Shown here is the shipped example (examples/energy.measured.illustrative.json): its power/throughput numbers are synthetic only to illustrate the output shape. When you run on a real GPU, the same fields carry your actual NVML measurement (basis: "measured").
4 · Overlay your result on the crossover curve
Your point is placed at (model size, ΔE%) from the uploaded report and coloured by its basis. The fitted curve is calibrated to measured anchors for the report's GPU architecture — filled dots are replicated (n = 2), hollow dots were measured once (n = 1) — and beyond the largest anchor it is extrapolated. Nothing is uploaded to a server; parsing is 100% in your browser.
The copy is Markdown — paste it into a new issue at ecocompute-mlcube/issues/new, or use the full submission form (handle, n, consent checkboxes) on the replications page.
Recent updates · 最近更新
2026-09-25 · data RTX 4090 NF4 re-tested across two physical cards — July's +0.8% anchor holds (mean +1.6%, n=3); the −15.1% historical reading was not reproduced and stays a divergent observation; no direction flip at the 3B anchor under the current stack. →
2026-09-24 · data All energy numbers unified on the generation measurement window; the llama.cpp row re-cut from its 100 Hz traces (−62.9% at 576 tokens, counter basis). →
2026-09-25 · site The site split into four layers — spec, evidence, changelog, findings; every count now has a defined denominator. →
Full history, each entry linked to its evidence: the changelog →