EcoCompute · Watt

Ask Watt — the decision, not just the curve

The same measured data this site plots, as a conversation. Watt lives in the WorkBuddy expert center; this page is just how you get there.

This site shows you what was measured. Watt turns that into what you should do: hand it a GPU, a model size and an accuracy target, and it says whether weight-only NF4 or INT8 will raise or lower your inference energy — labelled with where the number came from and how far to trust it.

It is not a chatbot with opinions. It answers from the published EcoCompute dataset, states its scope, and says “no data” instead of guessing.

Free No GPU needed to ask Works in English & 中文

How to open it · 如何打开

  1. Open WorkBuddy and go to the expert center — the 专家 entry in the left sidebar.
  2. Search Watt, or 大模型能耗, or LLM energy.
  3. Give it three inputs: GPU model · model size · accuracy target. That is the whole form.

Nothing to install, no account beyond WorkBuddy. If the search box shows several experts, Watt is the one whose answers carry a measured / estimated label on every number.

What it does that this site does not

Three things worth asking it

“Will NF4 actually save energy on an RTX 4090 for a 3B model at batch size 1?”

“Is load_in_8bit=True a trap on Ada for a 0.5B model — and what should I set instead?”

“The crossover sits around 3.7B for me. What would it take to re-measure it on an A800?”

Submit a measurement Watt does not have and it walks you through the required fields — your point can join the published dataset under CC BY 4.0.

Scope of every number

GPU-package power (NVML), not whole-system draw — no PSU losses, no CPU, no DRAM, no PUE or CO₂e. Single-stream decode, batch 1. Some anchors are n = 1; the main dataset is n = 2 (CV < 2%). This is a reference methodology, not a certified benchmark — no LoadGen, no accuracy target, not an MLPerf result.

Stating these limits is the point. A number without its scope is not evidence. How to disprove us →