<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>EcoCompute — Notes</title>
    <link>https://quantenergy.tech/blog/</link>
    <atom:link href="https://quantenergy.tech/blog/rss.xml" rel="self" type="application/rss+xml" />
    <description>Evidence-based notes on the energy cost of LLM inference, from direct NVML GPU power measurements.</description>
    <language>en</language>
    <lastBuildDate>Sat, 25 Jul 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>We ran our own container on a rented RTX 4090: NF4 crosses over at 7B, INT8 never wins</title>
      <link>https://quantenergy.tech/blog/we-ran-our-own-container-rtx4090.html</link>
      <guid isPermaLink="true">https://quantenergy.tech/blog/we-ran-our-own-container-rtx4090.html</guid>
      <pubDate>Sat, 25 Jul 2026 00:00:00 GMT</pubDate>
      <description>15 of 15 configurations measured with direct NVML on a rented RTX 4090 using our own open MLCube container. NF4 flips from a +36% penalty at 0.5B to a -29% saving at 7B; INT8 costs more energy at every size because decode throughput collapses to 9-19 tok/s. Data archived at DOI 10.5281/zenodo.21528103.</description>
    </item>
    <item>
      <title>When not to quantize: the small-model energy penalty</title>
      <link>https://quantenergy.tech/blog/when-not-to-quantize.html</link>
      <guid isPermaLink="true">https://quantenergy.tech/blog/when-not-to-quantize.html</guid>
      <pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate>
      <description>Weight-only NF4/INT8 quantization shrinks memory but can raise LLM inference energy by 25–55% for small models. What 360+ real GPU measurements show about the crossover point, and when quantization is still the right call.</description>
    </item>
  </channel>
</rss>
