nowfound

Alternatives

Products that do what Mesh LLM does

Pool compute to run powerful open models

  1. 1

    Calculate the GPU memory you need for LLM inference

    2025

  2. 2

    The easiest way to use cloud GPUs

    2025

  3. 3

    Run many models side by side and fuse the best answer

    Apr 2026

  4. 4
    Banana235

    Serverless GPUs for Machine Learning inference

    2022

  5. 5
    RunInfra156

    Describe the AI model you need and get an optimized AI

    Jul 2026

  6. 6
    Groq®237

    Hyperfast LLM running on custom built GPUs

    2024

  7. 7

    Simulate AWS, GCP & DigitalOcean without paying the bill

    Jun 2026

  8. 8

    The turn key OpenClaw solution with unlimited LLM tokens

    Mar 2026

  9. 9
    Reefy75

    Turn any PC into a private AI machine

    May 2026

  10. 10
    Aqueduct107

    The easiest way to run open source LLMs

    2023

  11. 11

    Self-host AI/ML with the world's cheapest GPU cloud

    2025

  12. 12

    Swarm Agents That Turn Slow PyTorch Into Fast GPU Kernels

    Jan 2026

  13. 13

    A new SOTA for compact open models on the edge

    May 2026

  14. 14

    Skip the setup and run OpenClaw & Hermes, fully managed

    17d ago · cloudways.com

  15. 15S1

    I wanted to build an inference provider for proprietary AI models, but I did not have a huge GPU farm. I started experimenting with Serverless AI inference, but found out that coldstarts were huge. I went deep into the research and put together an engine that loads large models from SSD to VRAM up to ten times faster than alternatives. It works with vLLM, and transformers, and more coming soon. With this project you can hot-swap entire large models (32B) on demand. Its great for: Serverless AI Inference Robotics On Prem deployments Local Agents And Its open source. Let me know if anyone…

    Nov 2025 · github.com

  16. 165L

    We've built InferX, a specialized runtime environment that fundamentally changes how LLMs are served. The core problem we solve is the latency bottleneck in AI inference, especially with large models. Current systems waste resources or suffer from painfully slow cold starts. InferX's AI-native architecture, with its "snapshot" technology, enables: * *Sub-2s cold starts:* Spin up models instantly. * *High density:* Serve more LLMs on the same GPUs. * *Optimal efficiency:* Maximize GPU utilization. This isn't just another API; it's a new execution layer designed from the ground up for the…

    2025 · github.com

  17. 17SO

    A simple calculator that estimates how many concurrent requests your GPU can handle for a given LLM, with shareable results.

    2025 · selfhostllm.org

  18. 18RA

    Hi there, looking for feedback on my new project "Featherless.AI" The idea is to allow users to run all the models on hugging face instantly. Via the OpenAI API compatible endpoint. Why? Because its a real chore to download models and spin up GPUs, especially if you want to test multiple models. Not to mention GPUs cost multiple dollars an hour to rent. And if we want more people to use open source AI, we got to make it easier for them to try and play with all of them. So what if instead of spinning up dedicated GPUs per model (which is what every provider is doing) We can startup a LLM…

    2024 · featherless.ai

  19. 19CM

    Hey HN, I've been building AutoAgents, an AI agent framework in Rust. Today I'm sharing a feature I haven't seen done well elsewhere: composable middleware layers for LLM inference pipelines. The problem Every agent framework lets you swap LLM providers. Almost none of them give you a structured way to enforce safety, caching, or data sanitization in the inference path itself. You end up with guardrails as application-level if-statements, caching bolted on as a separate service, and PII handling as a "we'll add it later" TODO that never ships. This gets worse with local models. Cloud APIs…

    Mar 2026 · github.com

  20. 20

    Turn idle GPUs into cash. Get affordable AI for everyone.

    Nov 2025

  21. 21CT

    I had been looking to try <500M parameter language models but you wouldn't find an API to try them anywhere, so I built this cloudflare hosted static website that hosts weights and built an inference runtime for these models that uses WebGPU and runs inference from your browser. These are only so useful in a multi-turn conversation but it's still interesting to see what you can pack in a <250mb model. I tried using ONNX versions earlier, but there were too many quirks of using them with language models and the TPS wasn't too impressive. Inspired by svenflow&#x2F;webgpu-gemma, I put my codex…

    May 2026 · chonklm.com

  22. 22QS
  23. 23PA

    Hello Hacker News! I am Bertrand from Pruna AI. With my associates, John, Rayan, and Stephan, we are fellow researchers in AI efficiency and reliability coming from TUM. We are building an optimization engine that combines compression methods (e.g. quantization, pruning, compilation, batching…) in the aim of saving compute power when running AI models. This optimization engine take one base model as input and returns a compressed model as output. It aims to help for two things: - Make various AI models faster and&#x2F;or smaller for various hardware (because they can require significant…

    2024

  24. 24SR

    Hi HN! Sipp is an open-source AI inference library for running local models in browsers with up to 3x faster decode speeds than alternative libraries. My background is in HCI (human-computer interaction) and graphics programming. Me along with my co-founder have been experimenting and thinking a lot about what the next user experience will look like when tokens are commodified to the point of being essentially “free.” A motivation for us was to try to move beyond the chat app and information retrieval use cases that are dominant now, and figure out how AI could instead act as a continuous…

    Jun 2026 · sipp.sh

Ranked by how close each launch is in meaning, then by votes. Refine with a description →