nowfound

Alternatives

Products that do what Pip install inference, open source computer vision deployment does

Deploying vision models is time consuming and tedious. Setting up dependencies. Fixing conflicts. Configuring TRT acceleration. Flashing (and re-flashing) NVIDIA Jetsons. A streamlined, developer-friendly solution for inference is needed. We, the Roboflow team, have been hard at work open sourcing Inference, an open source vision deployment solution. Our solution is designed with developers in mind, offering a HTTP-based interface. Run models on your hardware without having to write architecture-specific inference code. Here's a demo showing how to go from a model to GPU inference on a video…

  1. 1
    Inferless749

    Deploy any machine learning models in minutes

    2025 · inferless.com

  2. 2

    AI models that run on an inference cloud optimized for speed

    May 2026 · generalcompute.com

  3. 3
    Banana235

    Serverless GPUs for Machine Learning inference

    2022

  4. 4
    ZeroGPU309

    The compute efficient layer for AI inference

    Jun 2026 · zerogpu.ai

  5. 5

    Fast multimodal-native inference at scale

    Dec 2025

  6. 6OS

    Hi HN, I built a specialized inference engine for running 4-bit Gemma 4 26B-A4B-IT on any M-series Mac using about 2 GB of RAM. It is called TurboFieldfare and is written in Swift and Metal. I have always adored on-device AI. It feels like magic that you can run a powerful NN on your Mac or iPhone. So I wanted to push the limits a bit and run a model whose weights don’t fit in memory. The model’s 4-bit quantized weights occupy roughly 14 GB, which makes running it with conventional inference tools almost impossible on an 8 GB or even 16 GB Mac once the OS, applications, and KV cache are…

    Jul 2026 · github.com

  7. 7WM

    We wrote our inference engine on Rust, it is faster than llama cpp in all of the use cases. Your feedback is very welcomed. Written from scratch with idea that you can add support of any kernel and platform.

    2025 · github.com

  8. 8

    MoE vision-language, now easier to access

    2025

  9. 9
    Zro428

    Private inference for coding agents

    Jul 2026 · zro.moonmath.ai

  10. 10
    Zoo234

    A free, open-source playground for AI image models

    2023

  11. 11OS

    Tom from Tensil here - happy to answer questions! We developed Tensil to bring custom ML accelerators to people who don't have the resources of companies like Google, Facebook and Tesla. Currently, we're focused on supporting convolutional neural network inference on edge FPGA (field programmable gate array) platforms, but we aim to support all model architectures on a wide variety of fabrics for both training and inference. Tensil is different from other ML accelerators in that it is open source and really easy to use. For example, you can generate a custom accelerator with one command: $…

    2022 · tensil.ai

  12. 12

    Deploy fast, unmetered embedding inference in your own VPC

    2024

  13. 13NT

    Hello HackerNews! I’m excited to share what we’ve been working on at nCompass Technologies: an AI inference* platform that gives you a scalable and reliable API to access any open-source AI model — with no rate limits. We don't have rate limits as optimizations we made to our AI model serving software enable us to support a high number of concurrent requests without degrading quality of service for you as a user. If you’re thinking, well aren’t there a bunch of these already? So were we when we started nCompass. When using other APIs, we found that they weren’t reliable enough to be able to…

    2024 · ncompass.tech

  14. 14TR
  15. 15
    Qwen3.5307

    The 397B native multimodal agent with 17B active params

    Feb 2026

  16. 16GA

    2021 · inferrd.com

  17. 17VI

    Most inference UIs that I've come across pretty much just give us a chat-like interface to toy around with models in a single visual conversation thread. Given the fact that we are limited to seeing only one output at a time, it's kind of hard to compare outputs from different models, adjustments made to the prompting, and sampler settings. But even when keeping the generation parameters the same (e.g., to test for reliability in the output) and just going for multiple passes, there is no easy way to have a side-by-side comparison to keep track of the outputs from the multiple "rounds". I…

    2024 · github.com

  18. 18IR

    Private inference app that lets you see the token entropy, explore and change the token probabilities. Just released on macOS, iOS version next then other platforms. Here's a demo of it in action running DeepSeek Terminus: https://youtu.be/kts098EL2PQ Would love to hear any feedback or feature requests from the community.

    Sep 2025 · inferencer.com

  19. 19IB

    We wanted to do something very challenging to prove to ourselves that we can do anything we put our mind to. The reasoning for why we chose to build a toy TPU specifically is fairly simple: - Building a chip for ML workloads seemed cool - There was no well-documented open source repo for an ML accelerator that performed both inference and training None of us have real professional experience in hardware design, which, in a way, made the TPU even more appealing since we weren't able to estimate exactly how difficult it would be. As we worked on the initial stages of this project, we…

    2025 · tinytpu.com

  20. 20

    Calculate the GPU memory you need for LLM inference

    2025

  21. 21DI

    Hi HN! I’m so excited to show my another open-source project here. It is a PoC project. Distributed Inference is a project to demonstrate an approach to designing cross-language and distributed pipeline in deep learning/machine learning domain, using WebRTC and Redis Streams. This project consists of multiple services, which are written in Go, Python, and TypeScript, running on Docker. It allows setting up multiple inference services in multiple host machines, in a distributed manner. It does RPC-like calls and service discovery via my other open-source projects, go-inventa and…

    2023 · github.com

  22. 22

    Pool compute to run powerful open models

    Apr 2026

  23. 23OG

    Your phone has a GPU more powerful than most 2018 laptops. Right now it sits idle while you pay monthly subscriptions to run AI on someone else's server, sending your conversations, your photos, your voice to companies whose privacy policy you've never read. Off Grid is an open-source app that puts that hardware to work. Text generation, image generation, vision AI, voice transcription — all running on your phone, all offline, nothing ever uploaded. That means you can use AI on a flight with no wifi. In a country with internet censorship. In a hospital where cloud services are a compliance…

    Feb 2026 · github.com

  24. 24
    Modelbit125

    Heroku for Data Science, from the founders of Periscope Data

    2023

Ranked by how close each launch is in meaning, then by votes. Refine with a description →