The same measured data this site plots, as a conversation. Watt lives in the WorkBuddy expert center; this page is just how you get there.
This site shows you what was measured. Watt turns that into what you should do: hand it a GPU, a model size and an accuracy target, and it says whether weight-only NF4 or INT8 will raise or lower your inference energy — labelled with where the number came from and how far to trust it.
It is not a chatbot with opinions. It answers from the published EcoCompute dataset, states its scope, and says “no data” instead of guessing.
FreeNo GPU needed to askWorks in English & 中文
How to open it · 如何打开
Open WorkBuddy and go to the expert center — the 专家 entry in the left sidebar.
Search Watt, or 大模型能耗, or LLM energy.
Give it three inputs: GPU model · model size · accuracy target. That is the whole form.
Nothing to install, no account beyond WorkBuddy. If the search box shows several experts, Watt is the one whose answers carry a measured / estimated label on every number.
What it does that this site does not
A provenance label on every number — measured, interpolated, estimated or extrapolated. Never blended together.
A confidence rating — n=1 anchor points are down-weighted; repeated trials with CV < 2% are not.
Refusal over guessing — outside the data it answers “no data”, then hands you a reproduction plan.
Boundary conditions attached — GPU, driver, CUDA, framework, model revision, batch size, and the fact that this is GPU-package power, not wall power.
Three things worth asking it
“Will NF4 actually save energy on an RTX 4090 for a 3B model at batch size 1?”
“Is load_in_8bit=True a trap on Ada for a 0.5B model — and what should I set instead?”
“The crossover sits around 3.7B for me. What would it take to re-measure it on an A800?”
Submit a measurement Watt does not have and it walks you through the required fields — your point can join the published dataset under CC BY 4.0.
Scope of every number
GPU-package power (NVML), not whole-system draw — no PSU losses, no CPU, no DRAM, no PUE or CO₂e. Single-stream decode, batch 1. Some anchors are n = 1; the main dataset is n = 2 (CV < 2%). This is a reference methodology, not a certified benchmark — no LoadGen, no accuracy target, not an MLPerf result.
Stating these limits is the point. A number without its scope is not evidence. How to disprove us →