Dev tools · alternatives · 2026
24 alternatives to Jevstiller – Distill Jev into a local model, with a disagreement bound
September 2026. Every number here is from the benchmarks, and bash experiments/bench.sh --no-record reruns them without an API key.
Jevstiller is a developer tool that creates a local machine learning model trained to replicate Jev's responses. It reduces latency from 300ms to approximately 15ms on CPU for text classification tasks by answering… Below are 24 products that do a similar job, ranked by how close each is in meaning and then by launch-day votes.
- 1
Jev▲544Fast, structured AI decisions for software automation
12d ago · console.typesafe.ai · its alternatives →
- 2

Agent evals and guardrails in one request. Built on Jev, Kev and Laya. - openlayer-ai/jevals
12d ago · github.com · its alternatives →
- 3

Interactive JSON filter using jq. Contribute to ynqa/jnv development by creating an account on GitHub.
2024 · github.com · its alternatives →
- 4

I built LocalGPT over 4 nights as a Rust reimagining of the OpenClaw assistant pattern (markdown-based persistent memory, autonomous heartbeat tasks, skills system). It compiles to a single ~27MB binary — no Node.js, Docker, or Python required. Key features: - Persistent memory via markdown files (MEMORY, HEARTBEAT, SOUL markdown files) — compatible with OpenClaw's format - Full-text search (SQLite FTS5) + semantic search (local embeddings, no API key needed) - Autonomous heartbeat runner that checks tasks on a configurable interval - CLI + web interface + desktop GUI - Multi-provider:…
Feb 2026 · github.com · its alternatives →
- 5

Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly. - Andyyyy64/whichllm
May 2026 · github.com · its alternatives →
- 6

I kept slamming into Claude Code limits mid-session and couldn’t find a quick way to see how close I was getting, so I hacked together a tiny local tracker. Streams your prompt + completion usage in real time Predicts whether you’ll hit the cap before the session ends Runs 100 % locally (no auth, no server) Presets for Pro, Max × 5, Max × 20 — tweak a JSON if your plan’s different GitHub: https://github.com/Maciek-roboblog/Claude-Code-Usage-Monitor It’s already spared me a few “why did my run just stop?” moments, but it’s still rough around the edges. Feedback, bug…
2025 · github.com · its alternatives →
- 7

Explore large language models in 512MB of RAM. Contribute to jncraton/languagemodels development by creating an account on GitHub.
2023 · github.com · its alternatives →
- 8
oqoqo▲340Most benchmarks today exist in curated environments and do not translate well to the real world. We built Oqoqo to bridge this gap. Oqoqo makes it super simple to build realistic evals and custom benchmarks for tasks users actually care about. Oqoqo can: - Reliably measure how agent friendly your product surfaces are against Codex, Claude Code, OpenClaw, Hermes, Pi, Opencode, Cursor, GitHub Copilot - Regression test MCP, CLI, skills, SDK, and any agent facing interface (we are continuously using Oqoqo to dogfood and improve our own MCP/CLI) - Create and share custom benchmarks for how…
Aug 2026 · oqoqo.ai · its alternatives →
- 9

A JVM assembler for the modern age. Contribute to roscopeco/jasm development by creating an account on GitHub.
2022 · github.com · its alternatives →
- 10
JevForAgents▲59Explore real Jev agent builds, demos, and patterns
8d ago · jevforagents.com · its alternatives →
- 11

Hey Folks! I've been building an open source benchmark for measuring local LLM performance on your own hardware. The benchmarking tool is a CLI written on top of Llamafile to allow for portability across different hardware setups and operating systems. The website is a database of results from the benchmark, allowing you to explore the performance of different models and hardware configurations. Please give it a try! Any feedback and contribution is much appreciated. I'd love for this to serve as a helpful resource for the local AI community. For more check out: - Website:…
2025 · localscore.ai · its alternatives →
- 12

Jev returns a decision in 227 ms. The chat models take 2.5 to 3.5 seconds. Pong where the ball moves one step per model decision. Slow model, slow ball.
14d ago · jev-pong.ably.dev · its alternatives →
- 13

I'm Vivek, co-founder/CEO of HackerRank (YC S11); You may know us as a hiring tool for developers/companies. Over the years, we have built up deep expertise in generating programming challenges, and we are now using that to make coding models better. Our first launch is Model Kombat -- an arena where you can directly compare anonymized coding models, side by side, on real problems. * Pick an arena (Java, Python, etc.) * Each battle has 3 rounds: see the problem + two model outputs -> vote on which you’d actually prefer * Leaderboards + problem statements are updated weekly. We…
2025 · astra.hackerrank.com · its alternatives →
- 14

A web-based application for quick, scalable, and automated hyperparameter tuning and stacked ensembling in Python. - reiinakano/xcessiv
2017 · github.com · its alternatives →
- 15

Optimizer library for tail recursive calls in Java bytecode - Sipkab/jvm-tail-recursion
2020 · github.com · its alternatives →
- 16

- 17

- 18

Jev-shaped (TypeSafe System One) classification wrapper over OpenAI-like clients: probabilities and confidence instead of prose - zhulinchng/jevper
9d ago · github.com · its alternatives →
- 19
jebi▲140A supercharged terminal for Mac with built-in local AI
Jun 2026 · jebi.sh · its alternatives →
- 20AT
I recently built a small open-source tool to benchmark different LLM API endpoints — including OpenAI, Claude, and self-hosted models (like llama.cpp). It runs a configurable number of test requests and reports two key metrics: • First-token latency (ms): How long it takes for the first token to appear • Output speed (tokens/sec): Overall output fluency Demo: https://llmapitest.com/ Code: https://github.com/qjr87/llm-api-test The goal is to provide a simple, visual, and reproducible way to evaluate performance across different LLM providers, including…
2025 · llmapitest.com · its alternatives →
- 21

I built this out of frustration as I lead the development of AI features at Yola.com. Prompt testing should be simple and straightforward. All I wanted was a simple way to test prompts with variables and jinja2 templates across different models, ideally somthing I could open during a call, run few tests, and share results with my team. But every tool I tried hit me with a clunky UI, required login and API keys, or forced a lengthy setup process. And that's not all. Then came the pricing. The last quote I got for one of the tools on the market was $6,000/year for a team of 16 people in a…
2025 · langfa.st · its alternatives →
- 22

Hi HN community, I have been working on benchmarking publicly available LLMs these past couple of weeks. More precisely, I am interested on the finetuning piece since a lot of businesses are starting to entertain the idea of self-hosting LLMs trained on their proprietary data rather than relying on third party APIs. To this point, I am tracking the following 4 pillars of evaluation that businesses are typically look into: - Performance - Time to train an LLM - Cost to train an LLM - Inference (throughput / latency / cost per token) For each LLM, my aim is to benchmark them for…
2023 · github.com · its alternatives →
- 23

Semantic search over your own Claude Code session history
Aug 2026 · github.com · its alternatives →
- 24

Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. - d
May 2026 · github.com · its alternatives →
Also compare
- Jev alternatives
- Jevals – replacing LLM judges with typed Jev decisions alternatives
- jnv: interactive JSON filter using jq alternatives
- LocalGPT – A local-first AI assistant in Rust with persistent memory alternatives
- Find the best local LLM for your hardware, ranked by benchmarks alternatives
- Claude Code Usage Monitor – real-time tracker to dodge usage cut-offs alternatives
Ranked by how close each launch is in meaning, then by votes. Prices were read from each product’s own site when checked and can change. Refine with your own description →