Alternatives
Products that do what RagMetrics does
From Guesswork to Ground
- 1

Compare LLMs on your data, measure, and pick the best.
Apr 2026 · trismik.com
- 2

- 3
- 4FT
May 2026 · github.com
- 5

- 6

- 7

- 8

- 9

- 10

- 11LS
Hi, I was a corporate lawyer for many years working with a lot of financial services and insurance companies. In practicing law, I noticed there was a lot of repetition in the tasks I was working on even as a highly paid attorney that could be automated. I wanted to solve the problem of dealing with a lot information and data in a practical way, using AI. This motivated me to start AI Bloks/LLMWare with my husband, who had a deep background in software and is a very early adopter of AI. We have been on this journey with our open source project LLMWare for the past 4 months, producing a…
2024 · github.com
- 12FT
2024 · github.com
- 13IL
I have been working in AI space for a while now, first at FAANG with ML since 2021, then with LLM in start-ups since early 2023. I think LLM Application development is extremely iterative, more so than any other types of development. This is because to improve an LLM application performance (accuracy, hallucinations, latency, cost), you need to try various combinations of LLM models, prompt templates (e.g., few-shot, chain-of-thought), prompt context with different RAG architecture, different agent architecture, and more. There are thousands of possible combinations and you need a process…
2024 · github.com
- 14

- 15

- 16EL
2024 · github.com
- 17CL
Hi HN! Run it: OPENROUTER_API_KEY="sk" npx bff-eval --demo We built a tool to help people take LLM outputs and easily grade them / eval them to know how good an assistant response is. We've built a number of LLM apps, and while we could ship decent tech demos, we were disappointed with how they'd perform over time. We worked with a few companies who had the same problem, and found out scientifically building prompts and evals is far from a solved problem... writing these things feels more like directing a play than coding. Inspired by Anthropic's constitutional ai concepts, and amazing…
2025 · github.com
- 18

- 19LP
2023 · retool.com
- 20

- 21

Pick the best LLM. Compare costs and performance.
Mar 2026 · loopthink.ai
- 22

- 23

- 24

Ranked by how close each launch is in meaning, then by votes. Refine with a description →