nowfound

Alternatives

Products that do what Prompt Eval does

Stop guessing if your prompt is good. Measure it.

  1. 1PO

    Hey HN! We’re Kevin and Steve. We’re building PromptTools (https://github.com/hegelai/prompttools): open-source, self-hostable tools for experimenting with, testing, and evaluating LLMs, vector databases, and prompts. Evaluating prompts, LLMs, and vector databases is a painful, time-consuming but necessary part of the product engineering process. Our tools allow engineers to do this in a lot less time. By “evaluating” we mean checking the quality of a model's response for a given use case, which is a combination of testing and benchmarking. As examples: - For generated…

    2023 · github.com

  2. 2

    Design, compare & deploy production-grade prompts

    2024

  3. 3

    Instantly test and compare AI prompts results across models

    2025

  4. 4
    Flapico149

    Prompt versioning, testing, and evaluation

    2025

  5. 5

    Everything you need to work on your prompts

    2023

  6. 6
    Selene 1196

    Evaluate your AI app with the most accurate LLM Judge

    2025

  7. 7

    Turn messy prompts into powerful AI instructions.

    Mar 2026

  8. 8
    Iterate84

    Prompt management made easy

    2024

  9. 9

    Your best prompts built for you. Using the best LLM.

    Feb 2026

  10. 10

    test the performance of different models with the prompts

    2024

  11. 11PE

    Nowadays, a common AI tech stack has hundreds of different prompts running across different LLMs. Three key problems: - Choices, picking from 100s of LLMs the best LLM for that 1 prompt is gonna be challenging, you're probably not picking the most optimized LLM for a prompt you wrote. - Scaling/Upgrading, similar to choices but you want to keep consistency of your output even when models depreciate or configurations change. - Prompt management is scary, if something works, you'll never want to touch it but you should be able to without fear of everything breaking. So we launched Prompt…

    2024 · jigsawstack.com

  12. 12RS

    Couldn't find a reliable, free place to share & rate AI prompts so I thought I'd take a stab at it Already has 500+ prompts generated by AI using the latest model prompting guidelines 5 different supported prompt types: full prompt, enhancement, template, system, chain 20+ categories: coding, writing, marketing, business, creative, etc. Every prompt gets evaluated automatically by multiple AI models (Claude 3 + GPT-4 Mini, more to come) Then humans can rate and there is an overall score that takes both AI & humans into account AI eval prompt here:…

    2025 · josh.ing

  13. 13

    AI that builds you a deterministic evaluation in minutes

    2025

  14. 14SP

    I built a system that lets LLMs automatically learn and improve problem-solving strategies over time, inspired by Andrej Karpathy's idea of a "third paradigm" for LLM learning. The basic idea: instead of using static system prompts, the LLM builds up a database of strategies that actually work for different problem types. When you give it a new problem, it selects the most relevant strategies, applies them, then evaluates how well they worked and refines them. For example, after seeing enough word problems, it learned this strategy: 1) Read carefully and identify unknowns, 2) Define…

    2025

  15. 15PI

    I was working on a project that involved indexing GitHub repos that used really long prompts. Iterating over each section and figuring out which parts of the prompt led to which parts of the output was a quite painful. As a frontend dev, I kept thinking it would be nice if I could just 'inspect element' on particular sections of the prompt. So I built this prompt debugger with visual mapping that shows exactly which parts generate which outputs. Now that context windows are getting huge, working with massive prompts should not be as difficult. Planning to open source this soon, but I'd love…

    2025 · inspectmyprompt.com

  16. 16CL

    Hi HN! Run it: OPENROUTER_API_KEY="sk" npx bff-eval --demo We built a tool to help people take LLM outputs and easily grade them / eval them to know how good an assistant response is. We've built a number of LLM apps, and while we could ship decent tech demos, we were disappointed with how they'd perform over time. We worked with a few companies who had the same problem, and found out scientifically building prompts and evals is far from a solved problem... writing these things feels more like directing a play than coding. Inspired by Anthropic's constitutional ai concepts, and amazing…

    2025 · github.com

  17. 17PD
  18. 18PE

    Test your prompt engineering skills by writing prompts and battling them against each other! Your prompt template can access the board state, the move history, and a list of legal moves, and the game engine selects the first legal move in the string response from the LLM you query. My best prompt so far ignores the board state and the move history and just tries to play mates, make captures, and promote pawns. Can you do better?

    2023 · github.com

  19. 19AB

    tl;dr - today i'm launching agents.blue, where you can learn prompting and master working with LLMs, for free, instantly (no sign in required) Last week at Law x LLM Hackathon, I met a lot of amazing engineers, lawyers and more who want to build with LLMs, but they didn’t know how to prompt effectively. The online resources out there are a lot of reading, but not a lot of doing, and in my experience as an engineer and TA, doing is the best way to learn. agents.blue is the free, fast, and interactive tutorial to go from zero to the cutting edge of prompting in under 10 minutes, so you can…

    2023 · agents.blue

  20. 20
    Verbito11

    Generate expert AI prompts — then score them 0–100

    Jul 2026 · verbito.ai

  21. 21

    200+ battle-tested AI prompts for writing & coding

    May 2026 · generalistprogrammer.gumroad.com

  22. 22IB

    Hi HN, I'm pleased to share Promptspot, an open-source (Apache License 2.0) project that helps automate testing of large language model (LLM) prompts against an array of input data. Modern LLMs offer an enormous amount of leverage if you "teach the bot to fish" — i.e. simply prompt it with both a "system prompt" (which typically doesn't change often) and a dynamic input, which is often application state, search results, recent activity, user profile data, etc. Existing playgrounds and prompt management systems often lack the rigor and flexibility required for this dynamic approach — and as…

    2023 · github.com

  23. 23IM

    Also inspired by this HN submission: https://www.chiark.greenend.org.uk/~sgtatham/quasiblog/findl... The model is gpt-4o-mini-2024-07-18.

    2024 · app4.hc11.org

  24. 24

    Score your AI prompts before using them.

    Feb 2026

Ranked by how close each launch is in meaning, then by votes. Refine with a description →