Alternatives
Products that do what Promptfoo – CLI for testing & improving LLM prompt quality does
- 1

- 2AP
2023 · promptperfect.jina.ai
- 3

- 4ML
2025 · simonwillison.net
- 5PO
Hey HN! We’re Kevin and Steve. We’re building PromptTools (https://github.com/hegelai/prompttools): open-source, self-hostable tools for experimenting with, testing, and evaluating LLMs, vector databases, and prompts. Evaluating prompts, LLMs, and vector databases is a painful, time-consuming but necessary part of the product engineering process. Our tools allow engineers to do this in a lot less time. By “evaluating” we mean checking the quality of a model's response for a given use case, which is a combination of testing and benchmarking. As examples: - For generated…
2023 · github.com
- 6

- 7CA
Hi HN! We’re been working hard on this low-code tool for rapid prompt discovery, robustness testing and LLM evaluation. We’ve just released documentation to help new users learn how to use it and what it can already do. Let us know what you think! :)
2023 · chainforge.ai
- 8PR
2017 · github.com
- 9

- 10PL
https://github.com/elijah-potter/ofc
2025 · elijahpotter.dev
- 11

- 12

- 13

- 14

- 15PE
Nowadays, a common AI tech stack has hundreds of different prompts running across different LLMs. Three key problems: - Choices, picking from 100s of LLMs the best LLM for that 1 prompt is gonna be challenging, you're probably not picking the most optimized LLM for a prompt you wrote. - Scaling/Upgrading, similar to choices but you want to keep consistency of your output even when models depreciate or configurations change. - Prompt management is scary, if something works, you'll never want to touch it but you should be able to without fear of everything breaking. So we launched Prompt…
2024 · jigsawstack.com
- 16

- 17PC
a CLI agent for prompt evaluation loopsw
May 2026 · github.com
- 18

- 19PP
We are excited to show Promptly (https://trypromptly.com), a prompt management platform for LLM apps that makes it easy to experiment, share and manage prompts in production. With Promptly, users can: - Try out different prompts and model parameters for various providers - Quickly share prompt snippets together with parameters and generated output. Think of it as CodePen or JSFiddle for prompts - Create high level endpoints on top of provider APIs (Open AI, DreamStudio etc) with templated and versioned prompts - Use built-in caching for endpoints that will help save on Open AI…
2023 · trypromptly.com
- 20PE
Spelltest framework simulates conversations between AI ‘synthetic users' in an environment to test and refine LLM-based applications. It ensures your app converse with utmost accuracy and relevance. Post-chat, Spelltest assesses responses, providing qualitative and quantitative feedback on performance. Suitable for both chat and completion modes. When to use: - After modifying your prompt. - When your LLM provider updates. - As a CI step for you repo. All feedback and collaborations appreciated!
2023 · github.com
- 21

- 22PS
2022 · prompthero.com
- 23PL
Hi HN, I've been working on an experimental tool that helps you use GPT to work on your codebase. I'd love to improve the tool if there's interest. New ideas welcome! I think this could also be useful for experimenting with other types of recursive prompts. It’s a little bit Swiss Army knife and a little bit skynet: https://github.com/ferrislucas/promptr From the README: Promptr is a CLI tool for operating on your codebase using GPT. Promptr dynamically includes one or more files into your GPT prompts, and it can optionally parse and apply the changes that GPT suggests to…
2023 · github.com
- 24PT
Hello HN! Pierre and Paul here. We are building an open source text analytics tool for user inputs and LLM app outputs The repo is https://github.com/phospho-app/phospho and landing is https://phospho.ai Most people building with LLMs today don’t have quantified evaluation and usage metrics on the interactions between users and their product. The only solution is to read every message (or a sample) to get a sense of what is going on. You can't improve your product without understanding who your users are and how they are using it. Nobody would launch a website…
2024 · github.com
Ranked by how close each launch is in meaning, then by votes. Refine with a description →