nowfound

Alternatives

Products that do what Leo: The Prompt Engineering SDK does

Optimize, benchmark & evaluate LLM prompts with 1 command.

  1. 1
    Latitude685

    The open-source prompt engineering platform

    2024 · latitude.so

  2. 2AP

    2023 · promptperfect.jina.ai

  3. 3

    Open Source LLM Engineering Platform

    2024

  4. 4PE

    Nowadays, a common AI tech stack has hundreds of different prompts running across different LLMs. Three key problems: - Choices, picking from 100s of LLMs the best LLM for that 1 prompt is gonna be challenging, you're probably not picking the most optimized LLM for a prompt you wrote. - Scaling/Upgrading, similar to choices but you want to keep consistency of your output even when models depreciate or configurations change. - Prompt management is scary, if something works, you'll never want to touch it but you should be able to without fear of everything breaking. So we launched Prompt…

    2024 · jigsawstack.com

  5. 5

    Helps you write better prompts

    2022

  6. 6RY

    Hey HN, we've just finished building a dynamic router for LLMs, which takes each prompt and sends it to the most appropriate model and provider. We'd love to know what you think! Here is a quick(ish) screen-recroding explaining how it works: https://youtu.be/ZpY6SIkBosE Best results when training a custom router on your own prompt data: https://youtu.be/9JYqNbIEac0 The router balances user preferences for quality, speed and cost. The end result is higher quality and faster LLM responses at lower cost. The quality for each candidate LLM is predicted ahead of time…

    2024 · unify.ai

  7. 7

    AI that builds you a deterministic evaluation in minutes

    2025

  8. 8

    Evaluate & optimize your LLM performance with DSPy

    2024

  9. 9PO

    Hey HN! We’re Kevin and Steve. We’re building PromptTools (https://github.com/hegelai/prompttools): open-source, self-hostable tools for experimenting with, testing, and evaluating LLMs, vector databases, and prompts. Evaluating prompts, LLMs, and vector databases is a painful, time-consuming but necessary part of the product engineering process. Our tools allow engineers to do this in a lot less time. By “evaluating” we mean checking the quality of a model's response for a given use case, which is a combination of testing and benchmarking. As examples: - For generated…

    2023 · github.com

  10. 10
    Prompts123

    LLMOps and prompt engineering

    2023

  11. 11

    The context manager and skills library for marketing teams

    Apr 2026 · promptr.ai

  12. 12PL

    https://github.com/elijah-potter/ofc

    2025 · elijahpotter.dev

  13. 13RR

    I built a single-file Python script that lets you run LLM prompts from the command line with templating, structured outputs, and the ability to chain prompts together. When I discovered Google's Dotprompt format (frontmatter + Handlebars templates), I realized it was perfect for something I'd been wanting: treating prompts as first-class programs you can pipe together Unix-style. Google uses Dotprompt in Firebase Genkit and I wanted something simpler - just run a .prompt file directly on the command line. Here's what it looks like: --- model: anthropic/claude-sonnet-4-20250514 output:…

    Nov 2025 · github.com

  14. 14PC
  15. 15RA

    Hey everyone! Along with my team, I've developed a reinforcement learning system that automatically optimizes LLM prompts, complete with a visualization feature to track both prompt structure and learning progress over time. Take a look here: https://nomadic-ml.github.io/nomadic/cookbooks/Nomadic_Promp... Check out our website too:https://www.nomadicml.com/ In terms of how this visualization works: The RL Prompt Optimizer employs a reinforcement learning framework to iteratively improve prompts used for language model evaluations. At each episode, the…

    2024 · nomadic-ml.github.io

  16. 16

    100% free prompt optimizer

    Sep 2025

  17. 17RA

    Hi HN, we are the founders of Relari (https://www.relari.ai). We launched our LLM evaluation stack on HN a few months ago (https://news.ycombinator.com/item?id=39641105), which is now used in production by AI teams at companies like Vanta and PwC. We have since expanded to directly optimizing parts of an LLM pipeline using a data-driven approach. In particular, we see a lot of potential in the Auto Prompt Optimization—which could be an attractive alternative to fine-tuning in many cases—to use data to align LLMs for domain-specific tasks. Here’s a demo video:…

    2024

  18. 18

    Professional prompt engineering without the learning curve!

    Nov 2025 · getpromptoptimizer.com

  19. 19CL
  20. 20CL

    Hi HN! Run it: OPENROUTER_API_KEY="sk" npx bff-eval --demo We built a tool to help people take LLM outputs and easily grade them / eval them to know how good an assistant response is. We've built a number of LLM apps, and while we could ship decent tech demos, we were disappointed with how they'd perform over time. We worked with a few companies who had the same problem, and found out scientifically building prompts and evals is far from a solved problem... writing these things feels more like directing a play than coding. Inspired by Anthropic's constitutional ai concepts, and amazing…

    2025 · github.com

  21. 21PP

    We are excited to show Promptly (https://trypromptly.com), a prompt management platform for LLM apps that makes it easy to experiment, share and manage prompts in production. With Promptly, users can: - Try out different prompts and model parameters for various providers - Quickly share prompt snippets together with parameters and generated output. Think of it as CodePen or JSFiddle for prompts - Create high level endpoints on top of provider APIs (Open AI, DreamStudio etc) with templated and versioned prompts - Use built-in caching for endpoints that will help save on Open AI…

    2023 · trypromptly.com

  22. 22PE

    Spelltest framework simulates conversations between AI ‘synthetic users' in an environment to test and refine LLM-based applications. It ensures your app converse with utmost accuracy and relevance. Post-chat, Spelltest assesses responses, providing qualitative and quantitative feedback on performance. Suitable for both chat and completion modes. When to use: - After modifying your prompt. - When your LLM provider updates. - As a CI step for you repo. All feedback and collaborations appreciated!

    2023 · github.com

  23. 23

    Customize powerful AI agents in seconds.

    Nov 2025

  24. 24

    Ship prompt changes without touching your codebase

    Jun 2026 · promptvlt.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →