nowfound

Alternatives

Products that do what Selene 1 does

Evaluate your AI app with the most accurate LLM Judge

  1. 1

    An open benchmark for AI agents that test APIs

    May 2026

  2. 2
    Stax179

    Move your LLM evals from vibes to data

    2025

  3. 3
    Taylor AI118

    Fine-tune open source LLMs in minutes

    2023

  4. 4
    AskCodi230

    Custom LLMs, without training. Use via openai compatible api

    Nov 2025

  5. 5
    Sup AI103

    AI ensemble that scored #1 on Humanity's Last Exam

    Apr 2026

  6. 6
    AutoArena110

    Automated GenAI evaluation that works

    2024

  7. 7
    Mercury 2152

    Fastest reasoning LLM built for instant production AI

    Feb 2026

  8. 8

    Your site scores X/100 for AI agents with next steps

    May 2026

  9. 9
    Gradient153

    Developer API for building private LLMs that you own

    2023

  10. 10
    Pioneer113

    Fine-tune any LLM in minutes, with one prompt

    Apr 2026

  11. 11

    Test-driven development for LLMs

    2023

  12. 12

    Version, test, and collaborate on LLM prompts— like code

    2025

  13. 13

    Transform generic AI models into specialized solutions

    2025

  14. 14SE

    Hey HN! I built self-driving sim and eval at Waymo. Now I’m building Scorecard to bring that approach to agent eval: reproducible, automated scoring for AI. Scorecard lets you: - Run LLM-as-judge evals on agent workflows: test tool usage, multi-step reasoning, and task completion in CI/CD or in a playground. - Debug failures with OpenTelemetry traces: see which tool failed, why your agent looped, and where reasoning went wrong. - Collaborate on datasets, simulated agents, and evaluation metrics. Try it out → https://app.scorecard.io (free tier, no payment required!) Docs →…

    Oct 2025 · docs.scorecard.io

  15. 15RA

    I built a local-first UI that adds two reasoning architectures on top of small models like Qwen, Llama and Mistral: a sequential Thinking Pipeline (Plan → Execute → Critique) and a parallel Agent Council where multiple expert models debate in parallel and a Judge synthesizes the best answer. No API keys, zero .env setup — just pip install multimind. Benchmark on GSM8K shows measurable accuracy gains vs. single-model inference.

    Mar 2026 · github.com

  16. 16AE

    I've been working on a site [1] to give people control of their LLM workflows through AI evals - automated checks that, once defined, let you move fast without regressions and cut through hype with proof. That one-liner is aimed at software engineers, but I've spent my career helping cross-functional teams collaborate, and that's really what this is about. AI agents make powerful workflows very plausible, but only if teams can grow them incrementally without losing control - no vendor lock-in, no discipline silos, no blind trust in outputs. The site tries to meet different audiences where…

    Feb 2026 · ai-evals.io

  17. 17EA

    A few months ago I was working on a flight search engine that would include pet transport costs (I know a few by hearth but storing them and make the calculations in the UI would be nice) While I was collecting pet pricing from several airlines I strugled to extract data in a common format without hallucinated values. That's when I thought: What if I use multiple LLMs and take the most common response to improve accuracy? This idea became this new project. You provide your documents, an SQLModel schema, an LLM provider, plus what you'd like to extract and Extrai does the rest. Including…

    Nov 2025 · github.com

  18. 18SA

    We've built SMELL (Subject-Matter Expert Language Liaison), a new framework that combines human expertise with LLMs to create feedback-informed, domain-specific LLM evaluators. One of the biggest issues with current evaluation methods (heuristics, assertions, LLM-as-a-judge etc.) is that it's difficult for them to match up with and capture human preferences. SMELL addresses this by putting human feedback at the core of the evaluation process. It scales up a small set of human-provided feedback into evaluators that reflect the standards and nuances of specific industries or use-cases. Instead…

    2024 · quotientai.co

  19. 19IB

    Hey HN, I've been working on something cool that I wanted to share with you all. It's called Viewpoint, an analytics tool for LLMs like OpenAI, Anthropic models, and Gemini. The idea came from the constant flood of new LLM models and the need to figure out which ones work best for my projects without breaking the bank. With viewpoint, I can track token usage, costs, latency(WIP), and traffic over time, making it easier to compare different models and see which ones perform best and save money. The tool works asynchronously, so it doesn't add any latency to your LLM requests, and you have…

    2024 · viewpointhq.com

  20. 20VA
  21. 21

    Use multiple LLMs at once, privately!

    19d ago · transferllm.com

  22. 22HL

    At testup.io we have been working for a while to bring artificial intelligence to the field of test automation. Just a few years ago, the primary challenge laid in accurately identifying UI elements following minor structural changes, such as updates to IDs or paths. The emergence of Large Language Models (LLMs) raised the bar for what it meant to be smart. Now, we anticipate the robot to do lots of things autonomously, such as retry in cases of unresponsiveness or handle minor error reports. A more challenging, but soon expected feature, would involve the test robot navigating your web shop…

    2024 · github.com

  23. 23AA

    Looking for feedback on how Props can make your life easier as an LLM application developer.

    2024 · wwww.getprops.ai

  24. 24AG

    I’ve been building LLM tooling for a small VC fund and found myself explaining the same mental model over and over to non-technical people around me: how a stateless LLM becomes a chatbot, how tool use works, what an agent is mechanically, and why context windows shape all of it. I never found a guide that covered that full chain at the level I wanted, so I wrote one. It’s nine short chapters, each building on the last. Deliberately simplified: the goal is a useful mental model, not a textbook. Feedback, corrections, and contributions welcome: github.com/ymyke/aiaiai

    Apr 2026 · aiaiai.guide

Ranked by how close each launch is in meaning, then by votes. Refine with a description →