nowfound

Alternatives

Products that do what SearchAI Inference Server does

Run Private LLMs on CPUs.

  1. 1

    Fast multimodal-native inference at scale

    Dec 2025

  2. 2

    Open-source monitoring for machine learning models

    2021

  3. 3

    Calculate the GPU memory you need for LLM inference

    2025

  4. 4
    AskCodi230

    Custom LLMs, without training. Use via openai compatible api

    Nov 2025

  5. 5
    LM Studio209

    Discover, download, and run local LLMs (incl. DeepSeek R1)

    2025

  6. 6

    Advanced Visual Reasoning & Agentic Tool Use

    2025

  7. 7
    RunInfra156

    Describe the AI model you need and get an optimized AI

    Jul 2026

  8. 8

    Build better AI by using on-device data

    2020

  9. 9

    Use open-source AI in your Node.js apps, up to 67x faster

    2023

  10. 10

    APIs for building AI chat and search

    Feb 2026

  11. 11

    Free MCP for security AI: live BGP, DNS, threat graph

    May 2026

  12. 12
    NobodyWho106

    Run AI models on any device

    17d ago · github.com

  13. 13

    One API for all documents your AI agents need

    Mar 2026

  14. 14

    Pool compute to run powerful open models

    Apr 2026

  15. 15

    An on-device AI Agent that runs on your phone, open & secure

    Aug 2026 · openminis.app

  16. 16
    Lekh AI80

    Private, offline AI for iPhone, and iPad

    Jan 2026

  17. 17EN
  18. 18S1

    I wanted to build an inference provider for proprietary AI models, but I did not have a huge GPU farm. I started experimenting with Serverless AI inference, but found out that coldstarts were huge. I went deep into the research and put together an engine that loads large models from SSD to VRAM up to ten times faster than alternatives. It works with vLLM, and transformers, and more coming soon. With this project you can hot-swap entire large models (32B) on demand. Its great for: Serverless AI Inference Robotics On Prem deployments Local Agents And Its open source. Let me know if anyone…

    Nov 2025 · github.com

  19. 19PA

    Hello Hacker News! I am Bertrand from Pruna AI. With my associates, John, Rayan, and Stephan, we are fellow researchers in AI efficiency and reliability coming from TUM. We are building an optimization engine that combines compression methods (e.g. quantization, pruning, compilation, batching…) in the aim of saving compute power when running AI models. This optimization engine take one base model as input and returns a compressed model as output. It aims to help for two things: - Make various AI models faster and/or smaller for various hardware (because they can require significant…

    2024

  20. 20FS

    Hi everyone! I've been loving building with AI, and over the past few years I've been leaning more and more into Typescript (and bun). My team at inference.net is constantly trying to get more leverage out of AI and find ways to setup our codebase to be able to increase the level of correctness that our AI is able to write code at. This starter repo is a very opinionated way to lay out a repo to lean into AI heavily. It leverages Cloudflare Workers as a deployment target for the API (my goal is to never have to deploy an API on a AWS/Azure/GCP server ever again unless I get to a…

    2025 · abeahmed.com

  21. 21MI

    2022 · max.io

  22. 22CM

    Hey HN, I've been building AutoAgents, an AI agent framework in Rust. Today I'm sharing a feature I haven't seen done well elsewhere: composable middleware layers for LLM inference pipelines. The problem Every agent framework lets you swap LLM providers. Almost none of them give you a structured way to enforce safety, caching, or data sanitization in the inference path itself. You end up with guardrails as application-level if-statements, caching bolted on as a separate service, and PII handling as a "we'll add it later" TODO that never ships. This gets worse with local models. Cloud APIs…

    Mar 2026 · github.com

  23. 23PA

    We built PrivateClaw because the hosted OpenClaw platforms on the market today require you to trust them with plaintext. PrivateClaw removes that requirement at the hardware layer. PrivateClaw runs AI agents inside Trusted Execution Environments (TEEs), backed by AMD’s SEV-SNP standard. This means that your data is encrypted at the hardware level, enforced by the AMD Secure Processor outside the host OS trust boundary. PrivateClaw comes with inference that also runs inside TEEs, which means your prompts and completions are private as well. How it works: Each user gets a dedicated CVM…

    Apr 2026 · privateclaw.dev

  24. 245L

    We've built InferX, a specialized runtime environment that fundamentally changes how LLMs are served. The core problem we solve is the latency bottleneck in AI inference, especially with large models. Current systems waste resources or suffer from painfully slow cold starts. InferX's AI-native architecture, with its "snapshot" technology, enables: * *Sub-2s cold starts:* Spin up models instantly. * *High density:* Serve more LLMs on the same GPUs. * *Optimal efficiency:* Maximize GPU utilization. This isn't just another API; it's a new execution layer designed from the ground up for the…

    2025 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →