Work · alternatives · 2026
24 alternatives to Agentic CUDA Kernel Optimizer
A langgraph based workflow with a C++ CUDA suite to optimize CUDA kernels - bertaye/agentic-cuda-optimizer
Agentic CUDA Kernel Optimizer is a langgraph-based workflow paired with a C++ CUDA suite designed to optimize CUDA kernels. It targets developers and engineers working with GPU acceleration who need to improve kernel… Below are 24 products that do a similar job, ranked by how close each is in meaning and then by launch-day votes. Prices are shown for the 1 we have checked on their own sites, 1 of them free or with a free tier.
- 1

- 2

Demo of agent based model on GPU with CUDA and OpenGL (Windows/Linux) Agent instances on GPU memory Uses SSBO for instanced objects (with GLSL 450 shaders) CUDA OpenGL interops Renders with GLFW3 window manager Dynamic camera views in OpenGL (pan,zoom with mouse) Libraries installed using vcpkg (https://github.com/KienTTran/ABMGPU)
2023 · github.com · its alternatives →
- 3

CUDA is NVIDIA's language for GPU programming, allowing you to mix write CPU and GPU code in C++ in one file. By chaining a few projects that compile CUDA to OpenCL, then Vulkan, then WebGPU, you can experiment with this GPGPU language on any hardware.
2025 · hipscript.lights0123.com · its alternatives →
- 4
Cursor 2.0▲957Our first coding model and new interface for agents
Oct 2025 · cursor.com · its alternatives →
- 5
Forge Agent▲113Swarm Agents That Turn Slow PyTorch Into Fast GPU Kernels
Jan 2026 · rightnowai.co · its alternatives →
- 6
RightNow CLI▲121I locked in for 18 hours and built Claude Code for CUDA. It is completely open source! It writes CUDA kernels, debugs memory issues, and optimizes for your specific GPU. It is a fully agentic AI with tool calling built specifically for the CUDA toolkit This is the CLI version of RightNow AI code editor, our GPU-native code editor with a built-in emulator, visual profiling, and remote GPU access. The CLI is open source, lightweight, and fast. It is designed for anyone who wants AI help with CUDA without setup I used Python because it is the most common language, so anyone can build on top of…
Oct 2025 · github.com · its alternatives →
- 7

- 8VA
2013 · zhehaomao.com · its alternatives →
- 9

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM - jmaczan/tiny-vllm
May 2026 · github.com · its alternatives →
- 10

Back in the old days, people used to do general-purpose GPU programming by using shaders like GLSL. This is what inspired NVIDIA (and other companies) to eventually create CUDA (and friends). This is an implementation of GPT-2 using WebGL and shaders. Enjoy!
2025 · github.com · its alternatives →
- 11

- 12
GitNexus (Akon Labs)free tier, from $29/mo▲151The open-source kernel for coding agents
Aug 2026 · akonlabs.com · its alternatives →
- 13

A pipeline that translates Rust GPU code into formal Coq models, as a foundation for memory model proofs - neelsomani/vericuda
Oct 2025 · github.com · its alternatives →
- 14AG
This is a vector index I built that supports insertion and k-nearest neighbors (k-NN) querying, optimized for GPUs. It operates entirely in CUDA and can process queries on half a billion vectors in under 200 milliseconds. The codebase is structured as a standalone library with an HTTP API for remote access. It’s intended for high-performance search tasks—think similarity search, AI model retrieval, or reinforcement learning replay buffers. The codebase is located at https://github.com/rodlaf/BinaryGPUIndex.
2025 · rlafuente.com · its alternatives →
- 15

Hi HN, Over the past few months, I've been building `dsc`, a tensor library from scratch in C++/CUDA. My main focus has been on getting the basics right, prioritizing a clean API, simplicity, and clear observability for running small LLMs locally. The key features are: - C++ core with CUDA support written from scratch. - A familiar, PyTorch-like Python API. - Runs real models: it's complete enough to load a model like Qwen from HuggingFace and run inference on both CUDA and CPU with a single line change[1]. - Simple, built-in observability for both Python and C++. Next on the roadmap is…
2025 · github.com · its alternatives →
- 16

An on-premises, bare-metal solution for deploying GPU-powered applications in containers - emergingstack/es-dev-stack
2016 · github.com · its alternatives →
- 17

Minimal kernel to make any AI coding agent stateful. Clone, point your agent, go. - oguzbilgic/agent-kernel
Mar 2026 · github.com · its alternatives →
- 18

After the incredible response to our launch of the first online CUDA playground, we have just shipped something we think all you GPU programming and ML enthusiasts will love. Introducing LeetGPU Challenges--the place to compete on writing the fastest CUDA kernels. We have problems like matrix multiplication, agent simulation, multi-head self-attention, with more dropping every couple of days! We have a lot of really cool things coming up, including support for PyTorch, TensorFlow, JAX, TinyGrad; Multi-GPU programs; H100, V100, A100 GPU options Give it a try and let us know what you think!
2025 · leetgpu.com · its alternatives →
- 19

Multi-GPU reinforcement learning using Deep Q-Network in TensorFlow for OpenAI Gym - vishar0/dist-dqn
2016 · github.com · its alternatives →
- 20

I built this mostly because I love the intersection of game AI, high-performance computing, and poker. I’d love for anyone interested in game theory or CUDA optimization to tear it apart, test the accuracy, and give me feedback. Happy to answer any questions about the algorithms, the transition from CPU to GPU, or poker AI in general!
Jul 2026 · bupticybee.github.io · its alternatives →
- 21AH
2015 · pskel.github.io · its alternatives →
- 22

Hey, recently I took inspiration from llama.cpp, ollama, and many other similar tools that enable inference of LLMs locally, and I just finished building a Llama inference engine for the 8B model in CUDA C. I recently wanted to explore my newly founded interest in CUDA programming and my passion for machine learning. This project only makes use of the native CUDA runtime api and cuda_fp16. The inference takes place in fp16, so it requires around 17-18GB of VRAM (~16GB for model params and some more for intermediary caches). It doesn’t use cuBLAS or any similar libraries since I wanted to be…
2025 · github.com · its alternatives →
- 23

Hi, I’m Nabeel. In August I released RunMat as an open-source runtime for MATLAB code that was already much faster than GNU Octave on the workloads I tried. https://news.ycombinator.com/item?id=44972919 Since then, I’ve taken it further with RunMat Accelerate: the runtime now automatically fuses operations and routes work between CPU and GPU. You write MATLAB-style code, and RunMat runs your computation across CPUs and GPUs for speed. No CUDA, no kernel code. Under the hood, it builds a graph of your array math, fuses long chains into a few kernels, keeps data on the GPU when…
Dec 2025 · github.com · its alternatives →
- 24

By attaching virtual GPUs through a new QEMU-based data plane, Kernel now offers GPU-accelerated cloud browsers that dramatically speed up agents on sites with canvas-heavy applications or more generally sites that rely on WebGL.
Mar 2026 · kernel.sh · its alternatives →
Also compare
Ranked by how close each launch is in meaning, then by votes. Prices were read from each product’s own site when checked and can change. Refine with your own description →