Jev, the 193.6×, and every measurement window that was chosen for you
On September 15, 2026, a company called TypeSafe AI released Jev. The naming is honest — Jev is short for the Jevons paradox: the cheaper and more efficient a resource becomes, the more of it gets consumed in total. A hundred and fifty years ago, that was economics warning the cult of efficiency; after September 15, it is a product name.
A warning becoming a selling point is reason enough to write something. But the angle I care about is not "is this model any good" — the past two weeks have produced plenty of praise and plenty of scorn. What I care about is this: Jev has turned the measurement window — the oldest craft in metrology — into a pricing art. And every window has bills outside it.
First, the basics. Jev is a "System One" model: unstructured state goes in, typed probabilistic decisions come out. Three question primitives — judge the probability that a statement holds, choose one option from up to 255, score against a rubric. It generates no text, so type errors are structurally impossible and there are no hallucinations. The price is $0.042 per million input tokens; outputs are free. Claimed latency: 70–500 ms.
"Outputs are free" — three words worth stopping on.
That is not a physical fact; it is a pricing policy — OpenChamber's analysis classifies it exactly that way: a pricing policy, not a measured result. Physically, the compute on the output side of course happens; commercially, it has been moved outside the billing window.
My own industry has a counterpart. When measuring GPU energy, NVML reads GPU package power — the draw on the chip-package side. It does not include the CPU, memory, power-conversion losses, or fans. So whenever someone takes my numbers to size a machine-room power bill, I have to say: this is a lower bound; a whole node adds CPU, memory, conversion and cooling on top; I do not provide a conversion factor — that would be another estimate without provenance — measure at the wall with a power meter. Choosing a window is not the error. The error is not declaring the window.
"Outputs are free" is the pricing edition of the same craft. Cut the window at the input side, and:
More interesting: TypeSafe's own homepage carries a worked example — $0.013880 / 8.566 s on the TypeSafe side. Work it out by hand and you get roughly 75× on speed and 171× on cost. One page, two sets of numbers, a factor of 2.6 apart — with the difference hidden in a footnote pointing to "System One tasks."
Credit where due: in the launch blog, they voluntarily listed biases against themselves, including, in their own words, that their numbers are not empirical; that they cannot prove the pricing is not subsidized; that the comparison group is the average of GPT-6 Astra and Fable 5.1; and that evaluations were generally run on their West Coast laptop. A company that writes down, in advance, that its own numbers will not survive scrutiny — in my vocabulary, the equivalent of a dataset card reading n = 1, CV unknown, adopt with caution — gains credibility by doing so.
But acknowledging a bias is not correcting it. As of 2026-10-01, the homepage still showed 193.6×.
Three days after launch, OpenChamber delivered a piece of Twitter archaeology: September 15–18, 26,896 tweets collected, 12,759 valid, with "numbers the author measured themselves" and "numbers the author relayed from the vendor" counted separately:
| Metric | Vendor claim | User-measured median | First-hand sample | IQR |
|---|---|---|---|---|
| Speedup | 193.6× (homepage) / 20–200× (launch post) | 7× | 215 numbers | 2×–20× |
| Cost reduction | 444.6× (homepage) / 40–400× (launch post) | 30× | 180 numbers | 5×–85× |
| End-to-end latency | 70–500 ms | 76 ms | 333 numbers | 2 ms–270 ms |
From 193.6× to 7×, nobody is lying; four causes are at work in sequence:
So 7×/30× is not the bankruptcy of a lie; it is the shape 193.6× takes when it hits the ground. I have seen this plot many times in energy datasets: FLOPs saved on paper is one thing, energy measured at the task level is another, and between the two always stand baseline choice, workload, and measurement boundary.
I maintain a public dataset of GPU inference energy. Its coverage matrix is GPU model × quantization precision × model size, and 30 of its 56 cells are empty. There is no INT8 curve on the T4 — so when someone asks me "does INT8 save energy on a T4," the only honest answer is: no data, no conclusion, here is the plan to measure it. Empty cells are the dataset's discipline, not its shame.
Jev's claimed numbers live in a sparse matrix of their own; nobody has drawn it. The rows are task domains (spam classification, intent routing, risk judgment, tool-call review, …); the columns are baseline, measurement boundary, repetitions, variance. 193.6× lives in one cell, 7× in another — they cannot refute each other, because they are not in the same cell. And extrapolating any single cell across the whole matrix is a larger error than the window choice itself. In my line of work: presenting the extrapolated as measured is the cardinal sin of data provenance.
In the week after launch, the community started filling that matrix in, cell by independent cell:
Note the craft of that last author: 60 cases mixing clear, ambiguous and adversarial classes; raw per-call results public; disputed labels can be changed and the run repeated; the analysis script re-executed before publishing to check the numbers. That is the honest use of n = 60. This is how the matrix gets filled — not by one 193.6× homepage, but by a few hundred small cells, each labeled with its boundary conditions.
On September 24, Check Point published the widely cited test: Jev Is Not a Language Model, but It Breaks Like One. Due-diligence-assistant scenario: a high-risk company report that is red flags all the way down; the attacker controls only one small section of the document; the goal is to flip "high risk / do not invest" into "low risk / recommended."
The result: across 9 attacker × difficulty combinations, every one was broken at least once; the strongest attacker succeeded 25 of 27 times, on average by round 4, at roughly fifty cents per success. Jev's per-call cost, on webofmike's accounting, is about 17.3 microdollars.
The defense-side findings are even more worth copying down: structured input did not help; marking the document "untrusted" did not help (16/27, same as unmarked); explicit anti-injection instructions moved breaches from 18 to 17 — noise. The only effective defense was the reasoning tier, which raised the cost per successful break from $0.56 to $4.39, an eight-fold improvement — but the comparison models have a reasoning tier, and Jev does not. Why not? Because reasoning would break the 70–500 ms latency window.
The window choice showed its price for the first time.
A paper posted by Nanyang Technological University on September 23 (arXiv:2609.28613, 54,060 preregistered calls) reached the finer-grained conclusion: typed output did hold what it promised — Jev never returned an undeclared action, and the attacker's target option was selected only 1.8% of the time — but injected text reliably pushed probability mass toward the attacker's side. Check Point's summary became the sentence worth remembering: output format constrains what a model can say, not what it can be convinced of. The model did not fail; it "did its job correctly, on false evidence."
Now put the Jevons paradox back in. It says: efficiency up → unit price down → call volume rises faster → total consumption up. Jev's pricing is that curve, engineered: input two orders of magnitude cheaper, outputs free, median latency 76 ms — every parameter lowers the psychological threshold of one more call. Vercel reported 13% of paid teams integrated it within 24 hours (the fastest in AI Gateway history, twice the pace of GPT-5.6); OpenRouter request volume climbed visibly in the days after launch; Vercel's engineers have already switched their internal safety classifier to Jev.
The adoption curve is the overture to the consumption curve. A DCVC partner says TypeSafe is already profitable — if the paradox holds, the vendor's revenue curve and the customers' call curves will steepen together, and the customers' total bills will not necessarily fall. Cheap never means spending less; that sentence is not mine — it is printed on the product name.
Having said all this about windows, do not misread it as doom. The 7×/30× are real; the calibration is real — 96.5% accuracy above 0.95 confidence, and "never errs at 1.000 confidence" holds stably in independent replication; "one function call replacing half a year of a team's engineering" is real too. These are the measured parts, and they deserve respect.
As someone who deals with data provenance every day, what I want is four things — the same four things I demand of my own dataset:
The real lesson of the Jevons paradox was never "cheap is useless"; it is "cheap changes behavior" — the unit price you save gets spent twice over on call volume. Making it a product name is clever marketing, and an honest preview.
And the measurer's job is to remember that behind every "free" there is a window somebody chose: the billing window at the input side, the task window of System One, the latency window of 70–500 ms. The bills outside the window always arrive — to the caller paying by the token, to the platform that swapped its classifier for Jev, and perhaps to the next attacker holding fifty cents.