nowfound

Alternatives

Products that do what InferiaLLM does

The Operating System for LLMs in Production

  1. 1
    Inferless749

    Deploy any machine learning models in minutes

    2025 · inferless.com

  2. 2TV
  3. 3
    Groq®237

    Hyperfast LLM running on custom built GPUs

    2024

  4. 4
    Athina AI509

    Monitor LLMs and automatically detect hallucinations in prod

    2024

  5. 5RM
  6. 6

    Fast multimodal-native inference at scale

    Dec 2025 · gmicloud.ai

  7. 7

    Calculate the GPU memory you need for LLM inference

    2025

  8. 8

    Curated resources related to deploying LLMs into production

    2023

  9. 9

    An independent receipt for every LLM API call

    23d ago · github.com

  10. 10TL
  11. 11

    Benchmarks local LLM engines on your hardware

    15d ago · github.com

  12. 12
    NVLM 1.0200

    Open frontier-class multimodal LLMs

    2024

  13. 13
    QWQ-Max126

    New LLM by Alibaba excelling in reasoning w/ "thinking mode"

    2025

  14. 14

    Fast LLMs for low-latency and high-performance workflows

    Jun 2026 · jetbrains.com

  15. 15
    InternVL3135

    Open MLLMs excelling in vision, reasoning & long context

    2025

  16. 16PL

    https://github.com/elijah-potter/ofc

    2025 · elijahpotter.dev

  17. 17LA

    G'day, HN! I'm one of the maintainers of `llm`. I've been working alongside a trusty group of contributors to bring this project to life, and we're now at a point where we're ready to share it with the world. Large language models (LLMs) are taking the computing world by storm due to their emergent abilities that allow them to perform a wide variety of tasks, including translation, summarization, code generation, and even some degree of reasoning. However, the ecosystem around LLMs is still in its infancy, and it can be difficult to get started with these models. `llm` is a one-stop shop for…

    2023 · github.com

  18. 18LT

    This is my take on the common "use llms to generate shell commands" utility. Emphasis is placed on good CLI UX, simplicity, and flexibility. `llm2sh` supports multiple LLM providers and lets LLMs generate multi-command sequences to handle complex tasks. There is also limited support for commands requiring `sudo` and other basic input. I recommend using Groq llama3-70b for day-to-day use. The ultra-low latency is a game-changer - its near-instant responses helps `llm2sh` integrate seamlessly into day-to-day tasks without breaking you out of the 'zone'. For more advanced tasks, swapping to…

    2024 · github.com

  19. 19
    Taylor AI118

    Fine-tune open source LLMs in minutes

    2023

  20. 20LS

    Hi, I was a corporate lawyer for many years working with a lot of financial services and insurance companies. In practicing law, I noticed there was a lot of repetition in the tasks I was working on even as a highly paid attorney that could be automated. I wanted to solve the problem of dealing with a lot information and data in a practical way, using AI. This motivated me to start AI Bloks/LLMWare with my husband, who had a deep background in software and is a very early adopter of AI. We have been on this journey with our open source project LLMWare for the past 4 months, producing a…

    2024 · github.com

  21. 21

    Pool compute to run powerful open models

    Apr 2026 · anarchai.org

  22. 22

    Free, gamified roadmaps for LLM engineering: an Inference Engineering path (KV caches, CUDA kernels, production vLLM serving) and a Model Training path (pretraining on a budget, scaling laws, SFT/DPO/GRPO) — 185 tasks with auto-verified milestones instead of a paper certificate.

    14d ago · inferquest.org

  23. 23FT

    2024 · github.com

  24. 24R5

    Hi HN, I built OpenGraviton, an open-source AI inference engine that pushes the limits of running extremely large LLMs on consumer hardware. By combining 1.58-bit ternary quantization, dynamic sparsity with Top-K pruning and MoE routing, and mmap-based layer streaming, OpenGraviton can run models far larger than your system RAM—even on a Mac Mini. Early benchmarks: TinyLlama-1.1B drops from ~2GB (FP16) to ~0.24GB with ternary quantization. At 140B scale, models that normally require ~280GB fit within ~35GB packed. Optimized for Apple Silicon with Metal + C++ tensor unpacking, plus…

    Mar 2026 · github.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →