nowfound

Alternatives

Products that do what Capicú Edge ML Inference does

Turn instrument measurements into actionable results

  1. 1

    Ultra-fast 309B MoE model for coding & agents

    Dec 2025

  2. 2

    Deploy fast, unmetered embedding inference in your own VPC

    2024

  3. 3
    Openlit152

    One click observability & evals for LLMs & GPUs

    2024

  4. 4
    Lumi194

    Read smarter, not harder

    Nov 2025

  5. 5CT

    We are excited to announce Cedille, the largest language model for French (6b parameters). Demo: https://cedille.ai Language models are general purpose AI systems that are able to solve a range of tasks by simply being prompted for it. It can be used for example to summarize text, do translations, or for idea generation & overcoming writer's block. You may know GPT-3, the humongous model from OpenAI. Cedille is a similar model targeting the French demographic - but smaller, as we don’t yet have $1b in the bank like they do. Although GPT-3 supports multiple languages including…

    2021

  6. 6
    Banana235

    Serverless GPUs for Machine Learning inference

    2022

  7. 7
    Groq®237

    Hyperfast LLM running on custom built GPUs

    2024

  8. 8

    1.6T MoE trained entirely on AI ASICs

    Jul 2026 · longcat.chat

  9. 9

    Trace LLM requests + costs with OpenTelemetry monitoring

    Oct 2025

  10. 10
    LFM2.5134

    The next generation of on-device AI

    Jan 2026

  11. 11
    InternVL3135

    Open MLLMs excelling in vision, reasoning & long context

    2025

  12. 12

    The open sparse MoE model for agentic coding

    Apr 2026

  13. 13OP
  14. 14IG

    2020 · gradiohub.com

  15. 15
    OAK 416

    A complete robotic vision system in a single device

    Dec 2025

  16. 16SS

    SMILE Serve is a production-ready inference server built on [Quarkus](https://quarkus.io/) that brings together three complementary inference capabilities on the JVM: - **Classic ML**: `/api/v1/models` for serialized SMILE models (`.sml`) - **ONNX Runtime**: `/api/v1/onnx` for any model in the ONNX open format (`.onnx`) - **LLM Chat**: `/api/v1/chat` for Llama 3 chat completions A React-based web UI is bundled and served from the same process.

    May 2026 · github.com

  17. 17IE

    Hey HN, when building ML systems for industrial AI, we have learned that data inspection is critical during the ML development process. We are also big fans of the Hugging Face ecosystem. That is why we built an integration to our data exploration tool Spotlight that allows you to interactively explore Hugging Face datasets with one line of code. Spotlight lets you leverage model results such as predictions and embeddings to gain a deeper understanding in data segments and model failure modes. Currently, many many NLP, CV, Audio and multimodal datasets are supported both locally and on the…

    2023 · huggingface.co

  18. 183I

    Hello HackerNews, I am Paul, and I would like to get some feedback on the tool we are releasing as beta today. 3LC is an ML tool that gives detailed insights, real-time data-centric iterative workflows for training/finetuning, and data quality improvements for your Machine Learning datasets and models. 3LC serves as a visualizer, editor, and debugger, focusing on how models learn from the training data. Key Features of 3LC: • Detailed Data Analysis: 3LC enables users to dive into model performance beyond typical labeling errors. It offers the capability to analyze intricate false…

    2024 · pypi.org

  19. 19LO

    Hey HN! I built Lumina – an open-source observability platform for AI/LLM applications. Self-host it in 5 minutes with Docker Compose, all features included. The Problem: I've been building LLM apps for the past year, and I kept running into the same issues: - LLM responses would randomly change after prompt tweaks, breaking things - Costs would spike unexpectedly (turns out a bug was hitting GPT-4 instead of 3.5) - No easy way to compare "before vs after" when testing prompt changes - Existing tools were either too expensive or missing features in free tiers What I Built: Lumina is…

    Jan 2026 · github.com

  20. 20IO
  21. 21WM

    Hi HN, I’ve spent the last decade building hardware products like humanoid robots, 3D printers, and self-driving tractors. I needed a tool to navigate technical documents faster, so I created one with friends. This tool helps with component search, cross-referencing, comparison, and debugging. We’d love your feedback, whether you find it useful or not. Thank you! Try it here: www.convergelab.ai

    2024 · convergelab.ai

  22. 22LH

    I work on inference scheduling — KV cache-aware routing, load balancing across GPU workers, that kind of thing. I wanted something like k9s but for my inference stack. Nothing existed, so I built it. llmtop is a real-time terminal dashboard for LLM inference workers. It scrapes the Prometheus /metrics endpoints that vLLM, SGLang, and LMCache already expose and shows everything in one view: KV cache usage, queue depth, TTFT/ITL latencies (P50/P99 from histogram buckets), token throughput, prefix cache hit rates. Color-coded — red means go fix it. ``` brew install…

    Mar 2026 · github.com

  23. 23IB

    Hi HN, I'm the author of Kumi, a project I've been working on and would love to get your feedback on. The Original Problem: The original idea for Kumi came from a complex IAM problem I faced at a previous job. Provisioning a single employee meant applying dozens of interdependent rules (based on role, location, etc.) for every target system. The problem was deeper: even the data abstractions were rule-based. For instance, 'roles' for one system might just be a specific interpretation of Active Directory groups. This logic was also highly volatile; writing the rules down became a discovery…

    Oct 2025 · kumi-play-web.fly.dev

  24. 24EV

Ranked by how close each launch is in meaning, then by votes. Refine with a description →