Paritok
Spend up to 85% less and run 3× longer coding agent sessions
In plain words
Paritok is a compression tool that reduces token usage for coding agents by up to 85% and extends session length threefold. It sits between a coding agent and language model, non-destructively compressing tools, files, and conversation history on the fly. The system filters and stubs unnecessary data while preserving identifiers, paths, and errors, and summarizes older conversation turns when needed. It runs locally with two commands and uses an open-source 4-billion-parameter compression model built specifically for code.
written from the facts on this page · September 2026
From the sources
Paritok compresses the tools, files, and history your coding agent sends. Save up to 85% on your token bill and run 3× longer sessions. Two commands, nothing lost, fully local.
Paritok drops in between your coding agent and the LLM, non-destructively compressing tools, files, and history on the fly — for longer sessions and smaller bills. Powered by our open-source code-native 4B compression model.
Paritok drops in and non-destructively compresses tools, files, and history on the fly for longer sessions and smaller bills. Filters, compresses, summarizes. Tags everything it touches. Agents ship 70+ tools in full JSON on every request. We keep the relevant ones, stub the rest. Our 4B model knows a function signature from a debug line. Identifiers, paths and errors survive. Turns beyond a recent window get summarized once your context budget fills. Recent turns are left untouched. Ours runs on a budget you set, not when the model runs out of room. Lossy on the wire, recoverable when it counts. The agent asks for the exact bytes and gets them — locally, without burning a turn. Races the…from paritok.com
More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 16d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 18d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 26d ago · cactuscompute.com


Launched alongside, August 2026
the whole month →- TL
Life & fun · 9d ago · louisabraham.github.io


- SA
Hello HN! I found that picking out plausible but diverse skin tones for my digital art and game development projects was kind of difficult, and I got curious about if there was a way to define a color space that made it easy. I've built a color picker and procedural generation algorithm based on the space as well as a bunch of other fun js features and demos throughout the page that use the equations. If you find it interesting, I have lots of explanations of how I built it and what properties the space has. The methodology might be a bit shaky, but hopefully the result is as helpful for…
Life & fun · Aug 2026 · toneyalexander.github.io


I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 16d ago · simedw.com