nowfound

Alternatives

Products that do what Deepmark AI does

LLM benchmarking tool for task-specific metrics on your data

  1. 1DA
  2. 2
    LLM Stats308

    Compare API models by benchmarks, cost & capabilities

    Oct 2025 · llm-stats.com

  3. 3

    Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly. - Andyyyy64/whichllm

    May 2026 · github.com

  4. 4LL

    Hey Folks! I've been building an open source benchmark for measuring local LLM performance on your own hardware. The benchmarking tool is a CLI written on top of Llamafile to allow for portability across different hardware setups and operating systems. The website is a database of results from the benchmark, allowing you to explore the performance of different models and hardware configurations. Please give it a try! Any feedback and contribution is much appreciated. I'd love for this to serve as a helpful resource for the local AI community. For more check out: - Website:…

    2025 · localscore.ai

  5. 5

    Test-driven development for LLMs

    2023

  6. 6
    LLMrefs315

    AI SEO Keyword Rank Tracker for LLM Search Engines

    2025

  7. 7AT

    I recently built a small open-source tool to benchmark different LLM API endpoints — including OpenAI, Claude, and self-hosted models (like llama.cpp). It runs a configurable number of test requests and reports two key metrics: • First-token latency (ms): How long it takes for the first token to appear • Output speed (tokens/sec): Overall output fluency Demo: https://llmapitest.com/ Code: https://github.com/qjr87/llm-api-test The goal is to provide a simple, visual, and reproducible way to evaluate performance across different LLM providers, including…

    2025 · llmapitest.com

  8. 8
    LLMWare358

    Dev tool to make AI apps to deploy privately or locally

    2024

  9. 9

    Open-source stack for industrial-grade LLM applications

    2025 · github.com

  10. 10

    Find your best LLM for a local inference

    2023

  11. 11
    ReachLLM214

    Dominate the AI Search Era

    2025 · reachllm.com

  12. 12

    LLMs price comparison tool developed and updated by LLM

    2024

  13. 13

    Validate, monitor, and safeguard LLM-based apps

    2023

  14. 14

    all-in-one LLM evaluation platform

    2024

  15. 15

    Evaluate & optimize your LLM performance with DSPy

    2024

  16. 16

    Reproducible benchmarks for evaluating AI models

    14d ago · github.com

  17. 17

    Improve your LLM apps with open-source observability tool

    2024

  18. 18FT

    2024 · github.com

  19. 19FL

    Hi HN community, I have been working on benchmarking publicly available LLMs these past couple of weeks. More precisely, I am interested on the finetuning piece since a lot of businesses are starting to entertain the idea of self-hosting LLMs trained on their proprietary data rather than relying on third party APIs. To this point, I am tracking the following 4 pillars of evaluation that businesses are typically look into: - Performance - Time to train an LLM - Cost to train an LLM - Inference (throughput / latency / cost per token) For each LLM, my aim is to benchmark them for…

    2023 · github.com

  20. 20

    An open source Next app template to monitor your AI apps

    2025

  21. 21UD

    Hey HN! I’m the founder of Unify, and we’ve just released our Model Hub, which provides a collection of LLM endpoints with live runtime benchmarks all plotted across time: https://unify.ai/hub A key finding is that static tabular runtime benchmarks for LLMs simply do not work. It’s necessary to take a time-series perspective, and plot the variations through time. We currently have 21 models provided by: Anyscale, Perplexity AI, Replicate, Together AI, OctoAI, Mistral AI and OpenAI, with more on the roadmap. We test across different regions (Asia, US, Europe), with varied…

    2024

  22. 22IL

    I have been working in AI space for a while now, first at FAANG with ML since 2021, then with LLM in start-ups since early 2023. I think LLM Application development is extremely iterative, more so than any other types of development. This is because to improve an LLM application performance (accuracy, hallucinations, latency, cost), you need to try various combinations of LLM models, prompt templates (e.g., few-shot, chain-of-thought), prompt context with different RAG architecture, different agent architecture, and more. There are thousands of possible combinations and you need a process…

    2024 · github.com

  23. 23

    Translate text with native accuracy

    2024

  24. 24KB

Ranked by how close each launch is in meaning, then by votes. Refine with your own description →