nowfound

Alternatives

Products that do what JuryArena does

Beyond vibe eval: AI-jury picks the right LLM for you.

  1. 1
    Selene 1196

    Evaluate your AI app with the most accurate LLM Judge

    2025

  2. 2
    AskCodi230

    Custom LLMs, without training. Use via openai compatible api

    Nov 2025

  3. 3

    Vibe-check many open-source and proprietary LLMs at once

    2024

  4. 4
    AutoArena110

    Automated GenAI evaluation that works

    2024

  5. 5PE

    Nowadays, a common AI tech stack has hundreds of different prompts running across different LLMs. Three key problems: - Choices, picking from 100s of LLMs the best LLM for that 1 prompt is gonna be challenging, you're probably not picking the most optimized LLM for a prompt you wrote. - Scaling/Upgrading, similar to choices but you want to keep consistency of your output even when models depreciate or configurations change. - Prompt management is scary, if something works, you'll never want to touch it but you should be able to without fear of everything breaking. So we launched Prompt…

    2024 · jigsawstack.com

  6. 6CL

    Hi HN! Run it: OPENROUTER_API_KEY="sk" npx bff-eval --demo We built a tool to help people take LLM outputs and easily grade them / eval them to know how good an assistant response is. We've built a number of LLM apps, and while we could ship decent tech demos, we were disappointed with how they'd perform over time. We worked with a few companies who had the same problem, and found out scientifically building prompts and evals is far from a solved problem... writing these things feels more like directing a play than coding. Inspired by Anthropic's constitutional ai concepts, and amazing…

    2025 · github.com

  7. 7PE

    Spelltest framework simulates conversations between AI ‘synthetic users' in an environment to test and refine LLM-based applications. It ensures your app converse with utmost accuracy and relevance. Post-chat, Spelltest assesses responses, providing qualitative and quantitative feedback on performance. Suitable for both chat and completion modes. When to use: - After modifying your prompt. - When your LLM provider updates. - As a CI step for you repo. All feedback and collaborations appreciated!

    2023 · github.com

  8. 8

    The context manager and skills library for marketing teams

    Apr 2026 · promptr.ai

  9. 9

    Build autonomous Python agents with native Agent-to-Agent (A2A) communication - protolink/examples/ai_courtroom at main · nMaroulis/protolink

    28d ago · github.com

  10. 10

    Find which AI wins for YOUR prompts. Test 100+ models free.

    Dec 2025

  11. 11HT

    Hey HN, We are Zain and Ashish, founders of Vanna AI. We recently embarked on an experiment to see if large language models (specifically LLMs) could help in generating SQL queries for real-world datasets. We initially started this project as a web app but realized that it was most useful and had broadest applicability as a Python package since you can then incorporate it into an existing workflow (Jupyter notebook, Slackbot, etc). We've had some good success with customer datasets but we've generally heard a lot of skepticism so we decided to write a paper about the methodology we're using…

    2023 · github.com

  12. 12

    Compare AI models side-by-side on same prompt

    Feb 2026

  13. 13AF

    We’ve built an AI risk assessment tool designed specifically for GenAI/LLM applications. It's still early, but we’d love your feedback. Here’s what it does: 1. it performs comprehensive AI risk assessments by analyzing your codebase against different AI regulation/framework or even internal policies. It identifies potential issues and suggests fixes directly through one click PRs. 2. the first framework the platform supports is OWASP Top 10 for LLM Applications 2025, upcoming framework will be ISO 42001 as well as custom policy documents. 3. we're a small, early stage team, so the…

    2025 · gettavo.com

  14. 14

    The Only AI Tool That Doesn't Trust AI

    Mar 2026 · triall.ai

  15. 15AL

    Hi HN! We partnered with the Atlas team to build a tool called AI Predict [0] that allows anyone to ask any question about the future and get a thoroughly researched, AI-generated prediction on how likely it is to be true. How it works: Atlas replicated a Berkeley paper [1] that showed LLMs could make predictions as accurate as the crowd. We’re using a mix of models from OpenAI and Anthropic, with information retrieval powered by NewsCatcher [2]. The system is live and fully functional, though it might struggle with hyper-local questions outside of the public domain (e.g., “Will I have…

    2024 · aipredict.fun

  16. 16

    Send your AI agent to an LLM prompt-injection arena

    May 2026 · duel.altaysec.com.tr

  17. 17GV

    Hey HN, I just updated my project that compares some LLMs. It uses your prompt for all the models and runs at the same time. You can see the results being generated in real-time and decide what's the best for your use case. I'm open to any suggestions and feedback. Thanks!

    2024 · geminivsgpt.com

  18. 18

    AI mock trial simulator with full 8-phase U.S. proceedings

    May 2026 · mocktrialonline.com

  19. 19RA

    I built a local-first UI that adds two reasoning architectures on top of small models like Qwen, Llama and Mistral: a sequential Thinking Pipeline (Plan → Execute → Critique) and a parallel Agent Council where multiple expert models debate in parallel and a Judge synthesizes the best answer. No API keys, zero .env setup — just pip install multimind. Benchmark on GSM8K shows measurable accuracy gains vs. single-model inference.

    Mar 2026 · github.com

  20. 20

    Use multiple LLMs at once, privately!

    19d ago · transferllm.com

  21. 21

    AI observability & cost intelligence for LLM apps

    Mar 2026

  22. 22

    Multi-persona AI panel for code reviews & event evaluation

    Jul 2026 · github.com

  23. 23

    Tracing, evals, and a control loop for production LLMs

    May 2026 · trulayer.ai

  24. 24AE

    I've been working on a site [1] to give people control of their LLM workflows through AI evals - automated checks that, once defined, let you move fast without regressions and cut through hype with proof. That one-liner is aimed at software engineers, but I've spent my career helping cross-functional teams collaborate, and that's really what this is about. AI agents make powerful workflows very plausible, but only if teams can grow them incrementally without losing control - no vendor lock-in, no discipline silos, no blind trust in outputs. The site tries to meet different audiences where…

    Feb 2026 · ai-evals.io

Ranked by how close each launch is in meaning, then by votes. Refine with a description →