
AIONBD
Deploy vector search on branch sites, IoT devices ...
What it does
AIONBD, Vector Database for the Edge Deploy AI vector search where cloud won't work. 12x faster than Qdrant. Single binary. Production-ready. — BENCHMARKS (Fashion-MNIST 784-dim, top-k=10) AIONBD: 767 QPS, p95: 1.4ms ⚡ Qdrant: 65 QPS, p95: 21ms Same hardware. Same test.
Does the same job
all alternatives →
Actian VectorAI DBApr 2026 · actian.com · ▲203The portable vector database for AI agents beyond the cloud

- IRI replaced vector databases with Git for AI memory (PoC)2025 · github.com · ▲198
Hey HN! I built a proof-of-concept for AI memory using Git instead of vector databases. The insight: Git already solved versioned document management. Why are we building complex vector stores when we could just use markdown files with Git's built-in diff/blame/history? How it works: Memories stored as markdown files in a Git repo Each conversation = one commit git diff shows how understanding evolves over time BM25 for search (no embeddings needed) LLMs generate search queries from conversation context Example: Ask "how has my project evolved?" and it uses git diff to show actual…
- RARun an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhoneAug 2026 · github.com · ▲312

- AGA GPU-accelerated binary vector index2025 · rlafuente.com · ▲65
This is a vector index I built that supports insertion and k-nearest neighbors (k-NN) querying, optimized for GPUs. It operates entirely in CUDA and can process queries on half a billion vectors in under 200 milliseconds. The codebase is structured as a standalone library with an HTTP API for remote access. It’s intended for high-performance search tasks—think similarity search, AI model retrieval, or reinforcement learning replay buffers. The codebase is located at https://github.com/rodlaf/BinaryGPUIndex.
More ai this month
the category →
I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions about the model, training, Core ML, or the many things that didn't work.
AI · 17d ago · simedw.com
Astute▲585Automate your B2B brand going viral, with new media creators
AI · 19d ago · company-app.joinastute.com


Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, small robots and microcontrollers. We got really great feedback here, and have now incorporated the suggestions to release Needle 2. The whole model is a single 14MB binary that runs a full session in 28MB of RAM; 45m parameters at 2bit compression. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, sits between 400-1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges…
AI · 27d ago · cactuscompute.com

