nowfound

Alternatives

Products that do what Inferencer does

Run and deeply control local artificial intelligence models

  1. 1IR

    Private inference app that lets you see the token entropy, explore and change the token probabilities. Just released on macOS, iOS version next then other platforms. Here's a demo of it in action running DeepSeek Terminus: https://youtu.be/kts098EL2PQ Would love to hear any feedback or feature requests from the community.

    Sep 2025 · inferencer.com

  2. 2

    AI models that run on an inference cloud optimized for speed

    May 2026 · generalcompute.com

  3. 3

    Fast multimodal-native inference at scale

    Dec 2025 · gmicloud.ai

  4. 4WM

    We wrote our inference engine on Rust, it is faster than llama cpp in all of the use cases. Your feedback is very welcomed. Written from scratch with idea that you can add support of any kernel and platform.

    2025 · github.com

  5. 5PI

    Deploying vision models is time consuming and tedious. Setting up dependencies. Fixing conflicts. Configuring TRT acceleration. Flashing (and re-flashing) NVIDIA Jetsons. A streamlined, developer-friendly solution for inference is needed. We, the Roboflow team, have been hard at work open sourcing Inference, an open source vision deployment solution. Our solution is designed with developers in mind, offering a HTTP-based interface. Run models on your hardware without having to write architecture-specific inference code. Here's a demo showing how to go from a model to GPU inference on a video…

    2023 · github.com

  6. 6
    ZeroGPU309

    The compute efficient layer for AI inference

    Jun 2026 · zerogpu.ai

  7. 7AB

    I built AutoThink, a technique that makes local LLMs reason more efficiently by adaptively allocating computational resources based on query complexity. The core idea: instead of giving every query the same "thinking time," classify queries as HIGH or LOW complexity and allocate thinking tokens accordingly. Complex reasoning gets 70-90% of tokens, simple queries get 20-40%. I also implemented steering vectors derived from Pivotal Token Search (originally from Microsoft's Phi-4 paper) that guide the model's reasoning patterns during generation. These vectors encourage behaviors like numerical…

    2025

  8. 8
    Groq®237

    Hyperfast LLM running on custom built GPUs

    2024

  9. 9
    local.ai104

    Free, local & offline AI with zero technical setup

    2023

  10. 10

    Next generation analytics, infer gives analysts superpowers

    2023

  11. 11

    No-code AI Lab: Train models, access datasets, run inference

    Feb 2026 · neuro-block.com

  12. 12

    The 1T Parameters Open-Source Thinking Model - SOTA on HLE

    Nov 2025

  13. 13
    InstaVM109

    Instant computers for AI agents

    May 2026 · instavm.io

  14. 14

    Calculate the GPU memory you need for LLM inference

    2025

  15. 15IB

    We wanted to do something very challenging to prove to ourselves that we can do anything we put our mind to. The reasoning for why we chose to build a toy TPU specifically is fairly simple: - Building a chip for ML workloads seemed cool - There was no well-documented open source repo for an ML accelerator that performed both inference and training None of us have real professional experience in hardware design, which, in a way, made the TPU even more appealing since we weren't able to estimate exactly how difficult it would be. As we worked on the initial stages of this project, we…

    2025 · tinytpu.com

  16. 16

    Serve Any AI Model, Faster & Cheaper

    Mar 2026 · ionrouter.io

  17. 17
    Sage2

    Local AI Inference Engine

    Jun 2026 · conifer.build

  18. 18AH

    autoresearch@home is a collaborative research collective where AI agents share GPU resources to collectively improve a language model. Think SETI@home, but for model training. How it works: Agents read the current best result, propose a hypothesis, modify train.py, run the experiment on your GPU, and publish results back. When an agent beats the current best validation loss, that becomes the new baseline for every other agent. Agents learn from great runs and failures, since we're using Ensue as the collective memory layer. This project extends Karpathy's autoresearch by adding the missing…

    Mar 2026 · ensue-network.ai

  19. 19GA

    2021 · inferrd.com

  20. 20NT

    Hello HackerNews! I’m excited to share what we’ve been working on at nCompass Technologies: an AI inference* platform that gives you a scalable and reliable API to access any open-source AI model — with no rate limits. We don't have rate limits as optimizations we made to our AI model serving software enable us to support a high number of concurrent requests without degrading quality of service for you as a user. If you’re thinking, well aren’t there a bunch of these already? So were we when we started nCompass. When using other APIs, we found that they weren’t reliable enough to be able to…

    2024 · ncompass.tech

  21. 21

    Pool compute to run powerful open models

    Apr 2026 · anarchai.org

  22. 22

    Keep your OpenClaw agents running. Free beta, no code change

    Apr 2026 · openinfer.io

  23. 23

    The Windows AI runtime built for real inference

    Jul 2026 · lizard-llm.qendryx.com

  24. 24VI

    Most inference UIs that I've come across pretty much just give us a chat-like interface to toy around with models in a single visual conversation thread. Given the fact that we are limited to seeing only one output at a time, it's kind of hard to compare outputs from different models, adjustments made to the prompting, and sampler settings. But even when keeping the generation parameters the same (e.g., to test for reliability in the output) and just going for multiple passes, there is no easy way to have a side-by-side comparison to keep track of the outputs from the multiple "rounds". I…

    2024 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →