nowfound

Alternatives

Products that do what NVIDIA Nemotron 3 Ultra does

The first open frontier model built for agents

  1. 1

    Powers faster, efficient reasoning for long-running agents

    Jun 2026 · developer.nvidia.com

  2. 2

    Open hybrid Mamba-Transformer MoE for agentic reasoning

    Mar 2026 · developer.nvidia.com

  3. 3
    Mistral 3415

    A family of frontier open-source multimodal models

    Dec 2025

  4. 4MO

    I wanted to share our new speech to text model, and the library to use them effectively. We're a small startup (six people, sub-$100k monthly GPU budget) so I'm proud of the work the team has done to create streaming STT models with lower word-error rates than OpenAI's largest Whisper model. Admittedly Large v3 is a couple of years old, but we're near the top the HF OpenASR leaderboard, even up against Nvidia's Parakeet family. Anyway, I'd love to get feedback on the models and software, and hear about what people might build with it.

    Feb 2026 · github.com

  5. 5

    Platform for measuring and training AI agents

    2016

  6. 6

    High performance in a 24b open-source model

    2025

  7. 7

    gpt-oss-120b and gpt-oss-20b open-weight language models

    2025

  8. 8
    ZeroGPU309

    The compute efficient layer for AI inference

    Jun 2026 · zerogpu.ai

  9. 98F

    Hi HN! I'm just sharing a project I've been working on during the LLM Efficiency Challenge - you can now finetune Llama with QLoRA 5x faster than Huggingface's original implementation on your own local GPU. Some highlights: 1. Manual autograd engine - hand derived backprop steps. 2. QLoRA / LoRA 80% faster, 50% less memory. 3. All kernels written in OpenAI's Triton language. 4. 0% loss in accuracy - no approximation methods - all exact. 5. No change of hardware necessary. Supports NVIDIA GPUs since 2018+. CUDA 7.5+. 6. Flash Attention support via Xformers. 7. Supports 4bit and 16bit…

    2023 · github.com

  10. 10

    Hi HN, we built an open source model gateway. It's a single place to manage our own self hosted, frontier, and open source models in one place. It’s is rust native, built for concurrency, and implements all the config quirks across models and providers (streaming formats, tool calls, model parameters, rate limits, and different error behavior). The gateway adds under 1 ms for BYOK requests and under 2 ms when Experiential supplies the provider key. It has every major inference provider, and 1000+ models refreshed daily via a codex agent that opens a PR. Compared to other similar projects…

    10d ago · github.com

  11. 11

    Massive local model speedup on Apple Silicon with MLX

    Apr 2026 · ollama.com

  12. 12

    The open-source era of 1M context intelligence

    Apr 2026 · huggingface.co

  13. 13
    GLM-4.5298

    Unifying agentic capabilities in one open model

    2025

  14. 14

    Host LLMs across devices sharing GPU to make your AI go brrr

    Oct 2025

  15. 15

    One unified API for all AI models like Gemini, GPT-4, DALL-E

    2023

  16. 16
    Qwen3.5307

    The 397B native multimodal agent with 17B active params

    Feb 2026

  17. 17

    The easiest way to access frontier AI models.

    Aug 2026 · tokenharbor.ai

  18. 18

    The first open model to beat Sonnet made for productivity

    Feb 2026

  19. 19TL

    Hey HN, we wanted to share our repo where we fine-tuned Llama 3.1 on Google TPUs. We’re building AI infra to fine-tune and serve LLMs on non-NVIDIA GPUs (TPUs, Trainium, AMD GPUs). The problem: Right now, 90% of LLM workloads run on NVIDIA GPUs, but there are equally powerful and more cost-effective alternatives out there. For example, training and serving Llama 3.1 on Google TPUs is about 30% cheaper than NVIDIA GPUs. But developer tooling for non-NVIDIA chipsets is lacking. We felt this pain ourselves. We initially tried using PyTorch XLA to train Llama 3.1 on TPUs, but it was rough: xla…

    2024 · github.com

  20. 20
    GLM-5154

    Open-weights model for long-horizon agentic engineering

    Feb 2026

  21. 21
    RunInfra156

    Describe the AI model you need and get an optimized AI

    Jul 2026 · runinfra.ai

  22. 22WM

    Try it out! https://glhf.chat/ Hey HN! We’ve been working for the past few months on a website to let you easily run (almost) any open-source LLM on autoscaling GPU clusters. It’s free for now while we figure out how to price it, but we expect to be cheaper than most GPU offerings since we can run the models multi-tenant. Unlike Together AI, Fireworks, etc, we’ll run any model that the open-source vLLM project supports: we don’t have a hardcoded list. If you want a specific model or finetune, you don’t have to ask us for it: you can just paste the Hugging Face link in and…

    2024 · glhf.chat

  23. 23DO

    Demo of agent based model on GPU with CUDA and OpenGL (Windows/Linux) Agent instances on GPU memory Uses SSBO for instanced objects (with GLSL 450 shaders) CUDA OpenGL interops Renders with GLFW3 window manager Dynamic camera views in OpenGL (pan,zoom with mouse) Libraries installed using vcpkg (https://github.com/KienTTran/ABMGPU)

    2023 · github.com

  24. 24OS

    Hi I'm Dan from Elodin, making an open source real-time capable flight software simulation. For AI Grand Prix contestants, the wait for the Round 1 virtual qualifier simulation has been grueling. If you’re competing, check out our simulation harness to tide you over, built to match the published competition constraints and message format. It runs against real Betaflight, which we learned requires at least 1000 sensor samples per second to run real-time correctly. The competition warranted introducing a new feature to generate the camera sensor directly in the simulation loop. Typically…

    May 2026 · elodin.systems

Ranked by how close each launch is in meaning, then by votes. Refine with a description →