ExANS – Lossless KV cache compression at 622 GB/s on H100
Hi HN, We are the developers of OpenLake, an open source storage engine for KV cache offloading to remote disk and memory. Once we offloaded to local disk, we realized the bottleneck is the PCIe or NIC bandwidth. We wondered whether on GPU lossless compression is viable for fast reads and lower TTFT. BF16 is usually very hard to compress, (high entropy of sign/mantissa). What surprised us is that real world KV blocks are very different. The exponent byte has a very low entropy and barely populated. Instead of compressing the whole tensor, we compress only the exponent stream on the GPU.…
In plain words
ExANS is a GPU-based lossless compression tool for KV cache data that achieves 1.51× compression ratios and 622 GB/s decompression speeds on H100 GPUs. Developed by the OpenLake team, it targets machine learning engineers and systems builders working with large language models who need to offload KV cache to remote storage. The software works by selectively compressing only the exponent stream of BF16 tensors rather than entire blocks, enabling faster data retrieval while maintaining lossless quality without network bandwidth limitations.
written from the facts on this page · September 2026
From the sources
In the maker’s words, at launch
Hi HN, We are the developers of OpenLake, an open source storage engine for KV cache offloading to remote disk and memory. Once we offloaded to local disk, we realized the bottleneck is the PCIe or NIC bandwidth. We wondered whether on GPU lossless compression is viable for fast reads and lower TTFT. BF16 is usually very hard to compress, (high entropy of sign/mantissa). What surprised us is that real world KV blocks are very different. The exponent byte has a very low entropy and barely populated. Instead of compressing the whole tensor, we compress only the exponent stream on the GPU. We see the following results: (H100, production KV snapshot): - 1.51× lossless compression - 622 GB/s median GPU decode Decompression is ~10× faster than a 400 Gb/s NIC bandwidth delivering data losslessly without quality change. We've are open sourcing this as: ExANS which will be available through our vLLM and SGLang connectors on OpenLake v0.8 version. No changes are required in the inference engine. I'm curious how others are handling KV transfer today. Are you using KV compression or is bandwidth not a bottleneck yet? Thanks! GitHub: https://github.com/openlake-project/openlake Technical Blog: https://theopenlake.com/blog/exans-lossless-gpu-compression-...
More dev tools this month
the category →



Open-source GTM skills for technical founders
Dev tools · 29d ago · gtmcofounder.com

OpenTrailPaper is open-source bike computer firmware for the LilyGO T5S3 4.7" E-Paper PRO. It supports offline maps, GPX routes, FIT recording and Bluetooth sensors.
Dev tools · 2d ago · opentrailpaper.com

Launched alongside, August 2026
the whole month →- TL
Life & fun · 10d ago · louisabraham.github.io


- SA
Hello HN! I found that picking out plausible but diverse skin tones for my digital art and game development projects was kind of difficult, and I got curious about if there was a way to define a color space that made it easy. I've built a color picker and procedural generation algorithm based on the space as well as a bunch of other fun js features and demos throughout the page that use the equations. If you find it interesting, I have lots of explanations of how I built it and what properties the space has. The methodology might be a bit shaky, but hopefully the result is as helpful for…
Life & fun · Aug 2026 · toneyalexander.github.io


I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com