nowfound

AI · November 13, 2025

VA

Verse AI – Catch the AI failures your evals miss

Our AI recruitment pipeline was auto-rejecting anyone who'd worked at companies founded after 2023. It didn't recognize names like Harvey or Snorkel AI, or didn't realize how important they'd become because of training data cutoffs. We had traces, evals, Langfuse dashboards - everything looked fine - but we kept finding failures we should have caught earlier. The pattern kept repeating: - ship an improvement - it works for a while - hit an edge case that breaks it - don't notice until we've lost good candidates That's when we realized - the problem wasn't just our recruitment pipeline -…

What it does

In the maker’s words, at launch

Our AI recruitment pipeline was auto-rejecting anyone who'd worked at companies founded after 2023. It didn't recognize names like Harvey or Snorkel AI, or didn't realize how important they'd become because of training data cutoffs. We had traces, evals, Langfuse dashboards - everything looked fine - but we kept finding failures we should have caught earlier. The pattern kept repeating: - ship an improvement - it works for a while - hit an edge case that breaks it - don't notice until we've lost good candidates That's when we realized - the problem wasn't just our recruitment pipeline - almost every AI product has blind spots that evals miss. So we built Verse, a tool that surfaces issues directly from real AI interactions - whether that's candidates talking to your recruitment pipeline, users interacting with your agent, or any AI making decisions. Instead of relying solely on evals, we cluster conversations, identify the key ones to review, and flag the ones that show failure patterns. We use OpenTelemetry for trace ingestion, so it's compatible with Langfuse, Langsmith, Braintrust, and other AI observability tools - you can add it right alongside your existing setup. I'm posting this because I'm curious whether other teams are hitting the same wall. If you want, I'm happy to audit your AI implementation for free and show you where things commonly break - even if you never use Verse. Happy to answer any technical questions.

Does the same job

all alternatives →
  • Progress AI ObservabilityAug 2026 · telerik.com · ▲168

    Trace, evaluate, and improve AI agents in production

  • HireSweet2020 · ▲156

    AI-powered search to help startups hire passive candidates

  • APIEval-20May 2026 · ▲121

    An open benchmark for AI agents that test APIs

  • TraceaMay 2026 · ▲80

    Datadog for AI agents with traces, RCA, and team memory

  • AE
    Agent-evals – Claude skill to build your own evalsMay 2026 · github.com · ▲9

    I’ve spent the past 10 years working on AI in finance, with much of that time focused on building evaluation systems for production environments. As agents become more widely adopted, more software engineering and product people have start building them. But I’ve noticed that many teams are not yet fluent in systematic evaluation, or in the processes needed to keep agent quality high over time. For large organizations, that gap is rarely the bottleneck due to dedicated teams. But after speaking with a number of startups, it became clear that building strong, up-to-date evals is much harder…

  • Tracely26d ago · tracely-ai.com · ▲9

    Production failures become regression tests for AI agents

More ai this month

the category →
  • I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.

    AI · 17d ago · simedw.com

  • Astute585

    Automate your B2B brand going viral, with new media creators

    AI · 18d ago · company-app.joinastute.com

  • Grok Bot547

    AI teammates that you can give real work to

    AI · 25d ago · x.ai

  • Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…

    AI · 27d ago · cactuscompute.com

  • Turn website visitors into qualified pipeline

    AI · 19d ago · clarasdr.ai

  • Kane CLI446

    Natural language browser & mobile app tests from terminal

    AI · 24d ago · testmuai.com

Launched alongside, November 2025

the whole month →
  • Guideflow1,341

    The AI demo automation platform for SaaS

    AI · Nov 2025 · guideflow.com

  • IB

    Life & fun · Nov 2025 · bitsnpieces.dev

  • Welltory1,030

    Stop energy drain

    Work · Nov 2025 · welltory.com

  • Gemini 31,007

    Bring any idea to life with multimodal capabilities

    AI · Nov 2025 · blog.google

  • TrustMRR836

    The database of verified startup revenues

    Growth · Nov 2025 · trustmrr.com

  • B
    Boing782

    Life & fun · Nov 2025 · boing.greg.technology