Alternatives
Products that do what Pi Co-pilot – Evaluation of AI apps made easy does
Hey HN — 2 months ago we shared our first product with the HN community (https://news.ycombinator.com/item?id=43362535). Despite receiving lots of traffic from HN, we didn’t see any traction or retention. One of our major takeaways was that our product was too complicated. So we’ve spent the last 2 months iterating towards a much more focused product that tries to do just one thing really well. Today, we’d like to share our second launch with HN. Our original idea was to help software engineers build high-quality LLM applications by integrating their domain knowledge into a…
- 1PL
Hey HN, after years building some of the core AI and NLU systems in Google Search, we decided to leave and build outside. Our goal was to put the advanced ML and DS techniques we’ve been using in the hands of all software engineers, so that everyone can build AI and Search apps at the same level of performance and sophistication as the big labs. This was a hard technical challenge but we were very inspired by the MVC architecture for Web development. The intuition there was that when a data model changes, its view would get auto-updated. We built a similar architecture for AI. On one side is…
2025 · build.withpi.ai
- 2
- 3

- 4

- 5

Find your next hire or your next role from Hacker News monthly threads. AI-powered matching between candidates and job postings.
5d ago · hnmatchmaker.com
- 6UF
Hi HN! I want to share our latest project at NEXA AI. We developed AI agent foundation models designed to transform how developers create AI agent powered apps. One major challenge we've observed with current human-computer interactions is that many simple, one-step tasks become unnecessarily complex, multi-step workflows due to limitations of current GUIs. AI agents can solve this, but existing AI agent models are slow and costly. To tackle these issues, we built lightweight AI agent models based on our Octopus V2, small language models for function calling (You can learn more about our…
2024 · nexa4ai.com
- 7WB
Hey HN, After GPT-3 created waves in the tech industry, a lot of AI tools were emerging and with that, some AI website builders But the results seemed way too generic to us. It felt like the developers were rushing to catch the wave instead of building a proper tool We took our time, did months of RnD and finally came up with something better than what others in the market are doing. It’s got better design output. While it’s still in beta, I wanted to show HN what we did. Will appreciate the feedback when you guys try it out. Here is the link to signup for the beta:…
2024 · dorik.com
- 8WB
Hey HN, We’re two developers (co-founders) with a team of 20 who got tired of spending hours reviewing PRs, so we built Infinitcode.ai, an AI-powered code reviewer that: - *Summarizes PRs in plain English*: No more deciphering 1,000-line diff jungles - *Catches more than bugs*: Security holes, performance pitfalls, code smells, even typos (yes, we’ll flag “vurnerabilities” and vulnerabilities) - *Zero onboarding*: Works instantly—no “let me learn your codebase for weeks” nonsense. Why we’re posting: We’re in alpha and need brutal honesty. Roast our tool, mock our UI, or tell us why AI will…
2025 · infinitcode.ai
- 9WT
Hi HN, We’re definitely not the first to realise there’s something seriously wrong with how hiring and job-seeking works today. Zero-cost communication and LLMs have created so much noise that good candidates can’t get heard, and it becomes all too tempting to game the system with keywords and prompt-hacking. In fact we discovered that 70% of early stage AI startups don't post their jobs on LinkedIn. Instead, many founders hire exclusively within their network, which works at the start but doesn’t scale. We thought a lot about this problem, and pivoted through a few ideas including an AI…
Oct 2025 · teeming.ai
- 10WI
Hi, as the LLM models are getting smarter in coding tasks, we will soon be using agents as co workers, as current tools like github copilot and cursor are not optimized for team collaboration, we began building PhantomX, please give your feedback, if you think we are in the right direction or what should be changed for finding the optimized development workflow which works for both humans and agents.
Jan 2026 · phantomx.dev
- 11WW
Hey all! @sridatta and I wrote a book/zine called Forest Friends on system evals for LLM-driven apps. But it's a bit more whimsical, a bit more visual, and very much inspired by the meme of LLMs being a shoggoth polished into a smiley face with RLHF. LLM system evals are important as companies move past the flashy AI demos to reliable production apps. System evals keep coming up as the answer for what you "should do", but it's not exactly a standard part of the software engineering toolkit. So we pulled from @sridatta's seven years as a research engineer at Google, plus a ton of best…
2024
- 12IB
I built a tool to roast landing pages with AI agents. I was gathering feedback from watching landing page roast videos, and figured out I could prompt LLMs to analyse a screenshot and roast based on the same criteria. It's not 100% accurate yet, but it has been really insightful when I've tested it on my own websites. Let me know what you think!
2024 · roastmylandingpage.io
- 13PP
We built a protocol that gives AI agents an address, plus discovery/trust/payments between agents. Agents install it themselves with one line (see on pilotprotocol.network) - about 250k have, exchanging ~2B packets/day, mostly without their owners’ knowledge. The interesting part for HN is probably the App Store model: publishers list tools, and agents autonomously discover and install them (30k installs in the first two weeks). Happy to answer anything about the architecture, trust model, or the weird stuff agents do on the network.
Jul 2026
- 14IM
Hey HN! Thank you for all the support and feedback on my original submission 2 months ago. I've been improving the backend using a MCTS/AlphaZero approach and it's currently producing much better results. My long term goal is to allow users to manage multiple projects, deployed autonomously, both from scratch and by making continual updates all prompted with natural language. The cost of each project has been lowered to $9 as performance with smaller models has improved (I migrated from Claude-3-Opus to gemini-1.5-flash). Thanks for checking it out!
2024 · saas-quick.com
- 15MA
Hey HN! We’re Oliver Gilan & Ben Warren and today we’re launching Mesa (https://mesa.dev) into public beta to help large engineering teams review code more effectively. There are plenty of code review agents that exist but we found that they fall short in a number of ways. They don’t… - Let you tailor the reviews with enough fidelity - Give you control over the models used. If a new foundation model is released I might want to try it! - Align the costs effectively Mesa solves these problems with a multi-agent architecture where you define custom review agents that specialize in…
Nov 2025
- 16

- 17IB
I’ve spent the last 2.5 months building a product that runs LLM-powered code reviews on my pull requests — and I just launched it. The tool is built specifically for solo developers. You install it on your repo, trigger a scan by creating a pull request, and it leaves structured review comments using OpenAI under the hood. Funnily enough, I used the dev version of this app to review its own pull requests while building it. It helped me spot bugs, simplify structure, and keep quality high — all with minimal need for another human in the loop. Things I want to try out in the next months : -…
2025 · codii.dev
- 18OS
Hey HN! We built EvalKit, a library you embed to capture agent actions and a UI where domain experts give feedback, evaluate and improve AI agents. We experienced, in large agentic systems, prompt-engineering or auto-prompt improvement tool can get accuracy from 0 to 50% but for increasing accuracy to 100% we had to work with domain experts. Example -> In a law ai agent, lawyers are needed because law is complex and lawyers have a deeper context compared to non-lawyers. Other evaluation tools in the market focus on the experience of the developer and we are focusing on making as easy as…
2025 · github.com
- 19

A rigorous, free framework for evaluating AI PM candidates
Oct 2025
- 20AE
I've been working on a site [1] to give people control of their LLM workflows through AI evals - automated checks that, once defined, let you move fast without regressions and cut through hype with proof. That one-liner is aimed at software engineers, but I've spent my career helping cross-functional teams collaborate, and that's really what this is about. AI agents make powerful workflows very plausible, but only if teams can grow them incrementally without losing control - no vendor lock-in, no discipline silos, no blind trust in outputs. The site tries to meet different audiences where…
Feb 2026 · ai-evals.io
- 21BE
Hey HN, We're excited to introduce Braintrust, a platform for running and tracking AI evaluations (“evals”) [1]. At my previous startup Impira and leading AI at Figma, we had this recurring problem where we never knew if changes we made to our products would improve or regress key user scenarios. We built some tooling to solve this problem and after talking to other developers learned that it was a widespread issue. Specifically, it’s challenging to establish a great dev loop that lets you systematically improve and ship high quality AI products. We worked with the teams at Zapier, Coda, and…
2023
- 22NA
Hi HN, A few months ago, we announced our AI Grant project. We give $2,500 in cash and $20,000 in GPU training credits to anyone who wants to work on AI research. Today, we're excited to announce our second batch of Fellows! We were flooded with nearly 1,000 applications. Topics were diverse: robotics, NLP, biology, tooling, physics, dataset acquisition, fundamental research and more. The applicants were also diverse, spanning from world-class Google researchers to high school students. Check out https://blog.aigrant.org/new-ai-grant-fellows-43f1c26c13d9 for an overview of the…
2017
- 23AA
This repo is the result of a debate about what kind of programming language might be appropriate if humans are no longer the primary authors. Initially the thought was "LLMs can just generate binaries directly" (this was before a more famous person had the same idea). But that on reflection seems like a bad approach because languages exist to capture program semantics that are elided by translation to machine code. The next step was to wonder if an existing "machine readable" program representation can be the target for LLM code generation. It turns out yes. This project is the result of…
Mar 2026 · github.com
- 24Y2
Today we are launching YourGPT 2.0, an update focused on making support, sales, and operational workflows work together more smoothly in one system. This release improves how workflows are built, how the platform connects to external tools, and how context is maintained across long, multi-channel interactions. AI Studio now makes it much easier to create and refine workflows. You can describe the process you want in natural language and the system generates the full workflow for you. It can also help debug issues by showing what happened at each step and explaining where things went off…
Nov 2025 · yourgpt.ai
Ranked by how close each launch is in meaning, then by votes. Refine with a description →