nowfound

Alternatives

Products that do what Metrik does

Real-Time LLM Latency Tracker and Fastest-Model Router

  1. 1
    Metrik6

    Real-time llm performance monitoring

    Nov 2025 · metrik-dashboard.vercel.app

  2. 2IB

    I built a voice agent from scratch that averages ~400ms end-to-end latency (phone stop → first syllable). That’s with full STT → LLM → TTS in the loop, clean barge-ins, and no precomputed responses. What moved the needle: Voice is a turn-taking problem, not a transcription problem. VAD alone fails; you need semantic end-of-turn detection. The system reduces to one loop: speaking vs listening. The two transitions - cancel instantly on barge-in, respond instantly on end-of-turn - define the experience. STT → LLM → TTS must stream. Sequential pipelines are dead on arrival for natural…

    Mar 2026 · ntik.me

  3. 3AT

    I recently built a small open-source tool to benchmark different LLM API endpoints — including OpenAI, Claude, and self-hosted models (like llama.cpp). It runs a configurable number of test requests and reports two key metrics: • First-token latency (ms): How long it takes for the first token to appear • Output speed (tokens/sec): Overall output fluency Demo: https://llmapitest.com/ Code: https://github.com/qjr87/llm-api-test The goal is to provide a simple, visual, and reproducible way to evaluate performance across different LLM providers, including…

    2025 · llmapitest.com

  4. 4
    Vapi498

    Voice AI infrastructure for the internet

    2024

  5. 5
    traceAI273

    Open-source LLM tracing that speaks GenAI, not HTTP.

    Apr 2026 · github.com

  6. 6AR

    Hey it’s Hassaan & Quinn – co-founders of Tavus, an AI research company and developer platform for video APIs. We’ve been building AI video models for ‘digital twins’ or ‘avatars’ since 2020. We’re sharing some of the challenges we faced building an AI video interface that has realistic conversations with a human, including getting it to under 1 second of latency. To try it, talk to Hassaan’s digital twin: https://www.hassaanraza.com, or to our "demo twin" Carter: https://www.tavus.io We built this because until now, we've had to adapt communication to the limits of…

    2024

  7. 7
    Bifrost575

    The fastest LLM gateway in the market

    2025

  8. 8MO

    I wanted to share our new speech to text model, and the library to use them effectively. We're a small startup (six people, sub-$100k monthly GPU budget) so I'm proud of the work the team has done to create streaming STT models with lower word-error rates than OpenAI's largest Whisper model. Admittedly Large v3 is a couple of years old, but we're near the top the HF OpenASR leaderboard, even up against Nvidia's Parakeet family. Anyway, I'd love to get feedback on the models and software, and hear about what people might build with it.

    Feb 2026 · github.com

  9. 9

    Evaluate & optimize your LLM performance with DSPy

    2024

  10. 10

    Aggregate uptime monitoring across OpenAI, Claude, and more

    Apr 2026 · tools.lamatic.ai

  11. 11OS

    Our goal with this project is to build a completely open source, state of the art turn detection model that can be used in any voice AI application. I've been experimenting with LLM voice conversations since GPT-4 was first released. (There's a previous front page Show HN about Pipecat, the open source voice AI orchestration framework I work on. [1]) It's been almost two years, and for most of that time, I've been expecting that someone would "solve" turn detection. We all built initial, pretty good 80/20 versions of turn detection on top of VAD (voice activity detection) models. And…

    2025 · github.com

  12. 12
    Chert210

    Vapi for FaceTime: AI video agents in a few lines

    22d ago · trychert.com

  13. 13

    Flat rate to the best LLMs for OpenClaw, Hermes Agent, etc.

    Apr 2026 · wafer.ai

  14. 14

    Real-time text-to-speech model you can self-host

    May 2026 · kugelaudio.com

  15. 15

    Trace LLM requests + costs with OpenTelemetry monitoring

    Oct 2025

  16. 16UD

    Hey HN! I’m the founder of Unify, and we’ve just released our Model Hub, which provides a collection of LLM endpoints with live runtime benchmarks all plotted across time: https://unify.ai/hub A key finding is that static tabular runtime benchmarks for LLMs simply do not work. It’s necessary to take a time-series perspective, and plot the variations through time. We currently have 21 models provided by: Anyscale, Perplexity AI, Replicate, Together AI, OctoAI, Mistral AI and OpenAI, with more on the roadmap. We test across different regions (Asia, US, Europe), with varied…

    2024

  17. 17

    Measure real-world AI response speed from your country in 3m

    Feb 2026 · github.com

  18. 18

    Optimize Performance, Cost, Speed & Carbon for each prompt

    Nov 2025 · modelpilot.co

  19. 19LI

    LLM Inference Calculator — Estimate throughput, latency, TTFT, TPOT, and GPU memory usage for large language model inference. LLM 推理计算器 — 估算大模型推理吞吐量 (throughput)、延迟 (latency)、TTFT、TPOT 与 GPU 显存占用。

    10d ago · llm-inference-calculator-delta.vercel.app

  20. 20
    Perssua61

    Real-time guidance from any LLM (including local ones)

    Nov 2025

  21. 21

    Track and improve your visibility on AI Search

    Dec 2025 · llmpulse.ai

  22. 22OS

    Hey HN, it’s Russ - cofounder of LiveKit. An open source stack for building realtime AI applications. We’re sharing our first homegrown AI model for turn detection. Here’s a live demo: https://cerebras.vercel.app/ Voice AI has come a long way in the last year. We now have end-to-end systems that can generate a response to user input in 300-500ms — human level speeds! As latency reduces, a common problem that surfaces is the LLM responds too quickly. Any time there’s a short pause in a user’s speech, it ends up interrupting them. This is largely due to how voice AI applications…

    2024

  23. 23

    Catch LLM quality drift before your users do

    Jun 2026 · regtrace-docs.vercel.app

  24. 24

    Measure Your LLM's Speed, Right on Your Device!

    2025

Ranked by how close each launch is in meaning, then by votes. Refine with a description →