nowfound

Life & fun · February 21, 2026

L3

Llama 3.1 70B on a single RTX 3090 via NVMe-to-GPU bypassing the CPU

Hi everyone, I'm kinda involved in some retrogaming and with some experiments I ran into the following question: "It would be possible to run transformer models bypassing the cpu/ram, connecting the gpu to the nvme?" This is the result of that question itself and some weekend vibecoding (it has the linked library repository in the readme as well), it seems to work, even on consumer gpus, it should work better on professional ones tho

Visit github.comAlternativestop 1% of February 2026

In plain words

This project enables running the Llama 3.1 70B language model on a single RTX 3090 GPU by directly accessing NVMe storage, bypassing CPU and RAM bottlenecks. Designed for experimenters and hobbyists interested in alternative GPU architectures, it streams model data directly from NVMe to GPU memory. The implementation works on consumer graphics cards, though performance should improve on professional hardware. The code is available on GitHub with supporting library documentation.

written from the facts on this page · September 2026

Does the same job

all alternatives →
  • TL
    Tune LLaMa3.1 on Google Cloud TPUs2024 · github.com · ▲189

    Hey HN, we wanted to share our repo where we fine-tuned Llama 3.1 on Google TPUs. We’re building AI infra to fine-tune and serve LLMs on non-NVIDIA GPUs (TPUs, Trainium, AMD GPUs). The problem: Right now, 90% of LLM workloads run on NVIDIA GPUs, but there are equally powerful and more cost-effective alternatives out there. For example, training and serving Llama 3.1 on Google TPUs is about 30% cheaper than NVIDIA GPUs. But developer tooling for non-NVIDIA chipsets is lacking. We felt this pain ourselves. We initially tried using PyTorch XLA to train Llama 3.1 on TPUs, but it was rough: xla…

  • 8F
    80% faster, 50% less memory, 0% loss of accuracy Llama finetuning2023 · github.com · ▲385

    Hi HN! I'm just sharing a project I've been working on during the LLM Efficiency Challenge - you can now finetune Llama with QLoRA 5x faster than Huggingface's original implementation on your own local GPU. Some highlights: 1. Manual autograd engine - hand derived backprop steps. 2. QLoRA / LoRA 80% faster, 50% less memory. 3. All kernels written in OpenAI's Triton language. 4. 0% loss in accuracy - no approximation methods - all exact. 5. No change of hardware necessary. Supports NVIDIA GPUs since 2018+. CUDA 7.5+. 6. Flash Attention support via Xformers. 7. Supports 4bit and 16bit…

  • FL
    Finetune LLaMA-7B on commodity GPUs using your own text2023 · github.com · ▲449

    I've been playing around with https://github.com/zphang/minimal-llama/ and https://github.com/tloen/alpaca-lora/blob/main/finetune.py, and wanted to create a simple UI where you can just paste text, tweak the parameters, and finetune the model quickly using a modern GPU. To prepare the data, simply separate your text with two blank lines. There's an inference tab, so you can test how the tuned model behaves. This is my first foray into the world of LLM finetuning, Python, Torch, Transformers, LoRA, PEFT, and Gradio. Enjoy!

  • IV
    I've built a locally running Perplexity clone2024 · github.com · ▲669

    The video demo runs a 7b Model on a normal gaming GPU. I think it already works quite well (accounting for the limited hardware power). :)

  • WM
    We made glhf.chat – run almost any open-source LLM, including 405B2024 · glhf.chat · ▲161

    Try it out! https://glhf.chat/ Hey HN! We’ve been working for the past few months on a website to let you easily run (almost) any open-source LLM on autoscaling GPU clusters. It’s free for now while we figure out how to price it, but we expect to be cheaper than most GPU offerings since we can run the models multi-tenant. Unlike Together AI, Fireworks, etc, we’ll run any model that the open-source vLLM project supports: we don’t have a hardcoded list. If you want a specific model or finetune, you don’t have to ask us for it: you can just paste the Hugging Face link in and…

  • RL
    Run Llama3.1 405B on a 8GB VRAM challenge [video]2024 · youtube.com · ▲19

    How to run Llama3.1 405B on a 8GB VRAM

More life & fun this month

the category →
  • TL

    Life & fun · 10d ago · louisabraham.github.io

  • Articos385

    Launch with confidence, not gut instinct Discussion | Link

    Life & fun · 12d ago · producthunt.com

  • Nex351

    Claude Cowork for high-volume GTM workflows Discussion | Link

    Life & fun · 3d ago · producthunt.com

  • Photosynthesis fires two of your iPhone

    Life & fun · 29d ago · photosynthesis.camera

  • Your multimedia mentor that takes you from mid to great Discussion | Link

    Life & fun · 11d ago · producthunt.com

  • SoloUno310

    Take control of hair pulling, nail biting & skin picking

    Life & fun · 28d ago · solouno.io

Launched alongside, February 2026

the whole month →
  • Rork Max1,430

    Best AI for iOS apps. Website that replaces Xcode

    Life & fun · Feb 2026 · rork.com

  • happycapy1,367

    The agent-native computer, for the rest of us

    AI · Feb 2026 · happycapy.ai

  • SuperX902

    All-in-one growth OS for serious 𝕏 creators

    AI · Feb 2026 · superx.so

  • KiloClaw871

    Hosted OpenClaw. No Mac mini required.

    Dev tools · Feb 2026 · kilo.ai

  • Talk it out and feel better

    AI · Feb 2026 · lovon.app

  • Claude’s most advanced model for agentic tasks

    AI · Feb 2026 · anthropic.com