nowfound

Alternatives

Products that do what Pipevals does

Evaluation pipelines for every LLM application

  1. 1PA

    Hey HN! Pipevals is early and rough (this is a learning project), but usable. It currently lets you: - build evaluation pipelines as graphs - run them against datasets - track how output quality changes over time

    Mar 2026 · github.com

  2. 2PD

    We’re Robin, Louis, and Thomas. Pipelex is a DSL and a Python runtime for repeatable AI workflows. Think Dockerfile/SQL for multi-step LLM pipelines: you declare steps and interfaces; any model/provider can fill them. Why this instead of yet another workflow builder? - Declarative, not glue code: you state what to do; the runtime figures out how. - Agent-first: each step carries natural-language context (purpose, inputs/outputs with meaning) so LLMs can follow, audit, and optimize. Our MCP server enables agents to run pipelines but also to build new pipelines on demand. - Open…

    Oct 2025 · github.com

  3. 3

    Open-source evaluations and observability for LLM apps

    2024

  4. 4AO

    I've been obsessed for the past ~year with the possibilities of talking to LLMs. I built a bunch of one-off prototypes, shared code on X, started a Meetup group in SF, and co-hosted a big hackathon. It turns out that there are a few low-level problems that everybody building conversational/real-time AI needs to solve on the way to building/shipping something that works well: low-latency media transport, echo cancellation, voice activity detection, phrase endpointing, pipelining data between models/services, handling voice interruptions, swapping out different…

    2024 · github.com

  5. 5
    buildpipe116

    Compose, run and automate multi step AI developer workflows

    May 2026 · buildpipe.com

  6. 6
    LLM Stats308

    Compare API models by benchmarks, cost & capabilities

    Oct 2025

  7. 7PO

    Hey HN! We’re Kevin and Steve. We’re building PromptTools (https://github.com/hegelai/prompttools): open-source, self-hostable tools for experimenting with, testing, and evaluating LLMs, vector databases, and prompts. Evaluating prompts, LLMs, and vector databases is a painful, time-consuming but necessary part of the product engineering process. Our tools allow engineers to do this in a lot less time. By “evaluating” we mean checking the quality of a model's response for a given use case, which is a combination of testing and benchmarking. As examples: - For generated…

    2023 · github.com

  8. 8
    Taylor AI118

    Fine-tune open source LLMs in minutes

    2023

  9. 9

    Validate, monitor, and safeguard LLM-based apps

    2023

  10. 10

    Test-driven development for LLMs

    2023

  11. 11

    An open benchmark for AI agents that test APIs

    May 2026 · resources.kusho.ai

  12. 12PL
  13. 13OS

    Our goal with this project is to build a completely open source, state of the art turn detection model that can be used in any voice AI application. I've been experimenting with LLM voice conversations since GPT-4 was first released. (There's a previous front page Show HN about Pipecat, the open source voice AI orchestration framework I work on. [1]) It's been almost two years, and for most of that time, I've been expecting that someone would "solve" turn detection. We all built initial, pretty good 80/20 versions of turn detection on top of VAD (voice activity detection) models. And…

    2025 · github.com

  14. 14
    pipe46

    Pipe coldpress datasets straight into your pipeline

    2024

  15. 15

    3D data prep made easy

    2024

  16. 16OS

    Hey HN! We are building *open source infrastructure for deploying customer-facing data pipelines.* Here’s our repo https://github.com/pipebird/pipebird and website https://pipebird.com/. Pipebird (YC W22) is designed to enable companies that generate important data to offer secure data pushes to their customers’ warehouses, directly from their products. Our team was previously building in fintech, where we heard from many of our peers that their customers wanted data pushed directly to their warehouses. Customers wanted to bring data into their source of…

    2022 · github.com

  17. 17
    Seeknal57

    Data & AI/ML CLI for pipelines and NL queries

    Apr 2026 · seeknal.exe.xyz

  18. 18PB
  19. 19AT

    I recently built a small open-source tool to benchmark different LLM API endpoints — including OpenAI, Claude, and self-hosted models (like llama.cpp). It runs a configurable number of test requests and reports two key metrics: • First-token latency (ms): How long it takes for the first token to appear • Output speed (tokens/sec): Overall output fluency Demo: https://llmapitest.com/ Code: https://github.com/qjr87/llm-api-test The goal is to provide a simple, visual, and reproducible way to evaluate performance across different LLM providers, including…

    2025 · llmapitest.com

  20. 20
    IOpipe107

    Application operations platform for serverless

    2017

  21. 21DP
  22. 22RE

    Hey HN, Kyle here, one of the co-founders of OpenPipe. Reinforcement learning is one of the best techniques for making agents more reliable, and has been widely adopted by frontier labs. However, adoption in the outside community has been slow because it's so hard to implement. One of the biggest challenges when adapting RL to a new task is the need for a task-specific "reward function" (way of measuring success). This is often difficult to define, and requires either high-quality labeled data and/or significant domain expertise to generate. RULER is a drop-in reward function that works…

    2025 · openpipe.ai

  23. 23CL

    Hi HN! Run it: OPENROUTER_API_KEY="sk" npx bff-eval --demo We built a tool to help people take LLM outputs and easily grade them / eval them to know how good an assistant response is. We've built a number of LLM apps, and while we could ship decent tech demos, we were disappointed with how they'd perform over time. We worked with a few companies who had the same problem, and found out scientifically building prompts and evals is far from a solved problem... writing these things feels more like directing a play than coding. Inspired by Anthropic's constitutional ai concepts, and amazing…

    2025 · github.com

  24. 24

    AI code reviews & Pipeline debugging

    2024

Ranked by how close each launch is in meaning, then by votes. Refine with a description →