nowfound

AI · alternatives · 2026

24 alternatives to Jevals – replacing LLM judges with typed Jev decisions

Agent evals and guardrails in one request. Built on Jev, Kev and Laya. - openlayer-ai/jevals

Jevals is a tool for evaluating and protecting AI agents by replacing traditional language model judges with typed decisions. It combines agent evaluation and guardrail functionality in a single request, built on the… Below are 24 products that do a similar job, ranked by how close each is in meaning and then by launch-day votes.

  1. 1
    Jev▲544

    Fast, structured AI decisions for software automation

    12d ago · console.typesafe.ai · its alternatives →

  2. 2

    Hi HN, I'm Antoine Zambelli, AI Director at Texas Instruments. I built Forge, an open-source reliability layer for self-hosted LLM tool-calling. What it does: - Adds domain-and-tool-agnostic guardrails (retry nudges, step enforcement, error recovery, VRAM-aware context management) to local models running on consumer hardware - Takes an 8B model from ~53% to ~99% on multi-step agentic workflows without changing the model - just the system around it - Ships with an eval harness and interactive dashboard so you can reproduce every number I wanted to run a handful of always-on agentic systems…

    May 2026 · github.com · its alternatives →

  3. 3

    Hey HN, we’re building an open specification that lets agents discover and invoke APIs with natural language, built on the OpenAPI standard. agents.json clearly defines the contract between LLMs and API as a standard that's open, observable, and replicable. Here’s a walkthrough of how it works: https://youtu.be/kby2Wdt2Dtk?si=59xGCDy48Zzwr7ND. There’s 2 parts to this: 1. An agents.json file describes how to link API calls together into outcome-based tools for LLMs. This file sits alongside an OpenAPI file. 2. The agents.json SDK loads agents.json files as tools for an LLM that…

    2025 · github.com · its alternatives →

  4. 4
    Venn.ai▲337

    Delegate real work to AI agents with safety guardrails

    Mar 2026 · venn.ai · its alternatives →

  5. 5

    September 2026. Every number here is from the benchmarks, and bash experiments/bench.sh --no-record reruns them without an API key.

    3d ago · jevstiller.pages.dev · its alternatives →

  6. 6

    An open benchmark for AI agents that test APIs

    May 2026 · resources.kusho.ai · its alternatives →

  7. 7

    Hello! We just released freeact (https://github.com/gradion-ai/freeact), a lightweight agent library that empowers language models to act as autonomous agents through executable code actions. By enabling agents to express their actions directly in code rather than through constrained formats like JSON, freeact provides a flexible and powerful approach to solving complex, open-ended problems that require dynamic solution paths. * Supports dynamic installation and utilization of Python packages at runtime * Agents learn from feedback and store successful code actions as…

    2025 · github.com · its alternatives →

  8. 8

    Explore real Jev agent builds, demos, and patterns

    8d ago · jevforagents.com · its alternatives →

  9. 9

    Jev-AI CLI tool that evaluates text against custom rulesets

    12d ago · github.com · its alternatives →

  10. 10
    Cyris▲98

    "Cyris is the orchestration layer your AI agents are missing. Whether you're a two-person startup or a 10,000-employee health system — connect agents from OpenClaw, Claude, GPT-4o, Ollama, EPIC, or any platform and have them coordinate, hand off, and escalate across your entire organization. Self-hosted. Auditable. Human-governed" Just a side project I've been working on, and was curious whether people would actually find something like this useful! I'd recommend checking out the sandbox as well.

    Apr 2026 · cyrisai.dev · its alternatives →

  11. 11

    A test runner for agentskills.io-style AI agent skills - darkrishabh/agent-skills-eval

    May 2026 · github.com · its alternatives →

  12. 12

    Jev-shaped (TypeSafe System One) classification wrapper over OpenAI-like clients: probabilities and confidence instead of prose - zhulinchng/jevper

    9d ago · github.com · its alternatives →

  13. 13

    Built an open JSON Schema for defining AI agent teams. Multi-agent systems are becoming a real deployment pattern — not single assistants, but teams with roles, handoffs, and human checkpoints. But there's no shared way to define one that travels across frameworks. Every implementation is scattered, locked to whichever tool you picked first. Built the schema to fix that. The schema lives at schema.openenvelope.org and is registered in SchemaStore, so if you drop a .envelope.json file in VS Code you get autocomplete and validation without installing anything. It's also on npm as…

    May 2026 · openenvelope.org · its alternatives →

  14. 14

    An information-flow policy engine for LLM agents.

    4d ago · openappa.com · its alternatives →

  15. 15FL

    Hi HN! We just launched Codacy Guardrails, an IDE extension with a CLI for code analysis and MCP server that enforces security & quality rules on AI-generated code in real-time. It hooks into AI coding assistants (like VS Code Agent Mode, Cursor, Windsurf), silently scanning and fixing AI-suggested code that has vulnerabilities or violates your coding standards, while the code it’s being generated. We built this because coding agents can be a double-edged sword. They do boost productivity, but can easily introduce insecure or non-compliant code. One recent research team at NYU found that 40%…

    2025 · its alternatives →

  16. 16IN

    Tl;dr: I trained a classifier to route to the least expensive model and reasoning depth to complete the request. Coupling that with additional automated token efficiency techniques has yielded 3x usage for the same spend. For anyone interested in trying it themselves: https://nerfguard.com Various teammates and I switched over to Codex from Claude Code recently. We still bounce between the tools, but Codex’s speed and steerability coupled with performance gains were hard to ignore. One of the downsides was that the per token pricing kicked in way sooner. This is happening across…

    Jun 2026 · its alternatives →

  17. 17

    We build runtime security for AI agents. The playground started as an internal tool that we used to test our own guardrails. But we kept finding the same types of vulnerabilities because we think about attacks a certain way. At some point you need people who don't think like you. So we open-sourced it. Each challenge is a live agent with real tools and a published system prompt. Whenever a challenge is over, the full winning conversation transcript and guardrail logs get documented publicly. Building the general-purpose agent itself was probably the most fun part. Getting it to reliably use…

    Mar 2026 · github.com · its alternatives →

  18. 18

    Build autonomous Python agents with native Agent-to-Agent (A2A) communication - protolink/examples/ai_courtroom at main · nMaroulis/protolink

    Aug 2026 · github.com · its alternatives →

  19. 19

    Hi HN! Run it: OPENROUTER_API_KEY="sk" npx bff-eval --demo We built a tool to help people take LLM outputs and easily grade them / eval them to know how good an assistant response is. We've built a number of LLM apps, and while we could ship decent tech demos, we were disappointed with how they'd perform over time. We worked with a few companies who had the same problem, and found out scientifically building prompts and evals is far from a solved problem... writing these things feels more like directing a play than coding. Inspired by Anthropic's constitutional ai concepts, and amazing…

    2025 · github.com · its alternatives →

  20. 20

    10-Model AI Consensus Council for elite exam preparation.

    May 2026 · jeeai.netlify.app · its alternatives →

  21. 21

    Deterministic offline release evidence for AI agents

    Jul 2026 · iisacc-justmoong.github.io · its alternatives →

  22. 22

    Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. - d

    May 2026 · github.com · its alternatives →

  23. 23BA

    I'm one of the creators of The Edge Agent (TEA). We built this because we needed a way to deploy agents that was verifiable and robust enough for production/edge cases, moving away from loose scripts. The architecture aims to solve critical gaps in deterministic orchestration identified by *Prof. Claudionor Coelho Jr. (Stanford alum, ML/DL Faculty at Santa Clara Univ., and Senior Fellow for AI at Majestic Labs)* during our work on the Kiroku project. *Key Technical Features:* * *Neurosymbolic Native:* We integrated Prolog to logically validate LLM outputs. This combines neural…

    Jan 2026 · fabceolin.github.io · its alternatives →

  24. 24

    Tool-agnostic TypeScript library for observing coding agent sessions. Built-in Cursor provider, extensible to Claude Code, Codex, and OpenCode. - vtemian/agentprobe

    Mar 2026 · github.com · its alternatives →

Also compare

Ranked by how close each launch is in meaning, then by votes. Prices were read from each product’s own site when checked and can change. Refine with your own description →