nowfound

Alternatives

Products that do what Shimmy – 5MB privacy-first, local alternative to Ollama (680MB) does

  1. 1

    Massive local model speedup on Apple Silicon with MLX

    Apr 2026 · ollama.com

  2. 2

    The easiest way to chat with local AI

    2025

  3. 3

    Let every AI remember the same you.

    Jul 2026 · memmy.bot

  4. 4IP

    The stack: two agents on separate boxes. The public one (nullclaw) is a 678 KB Zig binary using ~1 MB RAM, connected to an Ergo IRC server. Visitors talk to it via a gamja web client embedded in my site. The private one (ironclaw) handles email and scheduling, reachable only over Tailscale via Google's A2A protocol. Tiered inference: Haiku 4.5 for conversation (sub-second, cheap), Sonnet 4.6 for tool use (only when needed). Hard cap at $2/day. A2A passthrough: the private-side agent borrows the gateway's own inference pipeline, so there's one API key and one billing relationship…

    Mar 2026 · georgelarson.me

  5. 5TO
  6. 6

    Your AI, fully offline with Zero data collection & 100% free

    May 2026 · lumichats.com

  7. 7SA

    Hello all :) I made this BBS for fun and thought you all would enjoy it. It's not perfect, but it's been a fun exercise! see it: https://asciinema.org/a/Emg6SWrXMV6cehfQxrw1GRu75 try it: ssh -p 2223 shhhbb.com host it: https://github.com/donuts-are-good/shhhbb/releases/latest why? Every year I challenge myself in some new way, this year it is to push one project per week. You might recognize my static site generator [0] or my releaser for go [1] from previous posts as one of these weekly projects. If you want to join me in doing this, it's…

    2023 · donuts-are-good.github.io

  8. 8LT

    hey guys. the other day i was migrating hosting providers and i just needed something not too heavy and convenient to spin up my backups for awhile and realised there is almost nothing out there. kimchi hasn't been updated for years and cockpit is heavy. so here's something i came up with in a couple hours because of a sudden urge, nothing fancy just basic creation with cloud init, lifecycle management and image/storage, but it's modern-ish and it compiles to a 8.4mb binary inclusive of the embedded web UI, CLI and API, and only dep is libvirt.

    2025 · github.com

  9. 9WM

    We wrote our inference engine on Rust, it is faster than llama cpp in all of the use cases. Your feedback is very welcomed. Written from scratch with idea that you can add support of any kernel and platform.

    2025 · github.com

  10. 10
    Cai179

    Press ⌥C on anything to run smart actions, locally

    Apr 2026 · getcai.app

  11. 11PF
  12. 12EL
  13. 13CA

    ChunkHound’s goal is simple: local-first codebase intelligence that helps you pull deep, core-dev-level insights on demand, generate always-up-to-date docs, and scale from small repos to enterprise monorepos — while staying free + open source and provider-agnostic (VoyageAI / OpenAI / Qwen3, Anthropic / OpenAI / Gemini / Grok, and more). I’d love your feedback — and if you have, thank you for being part of the journey!

    Jan 2026 · github.com

  14. 14

    An ultra-fast, single-binary MCP server written in Rust as a lightweight alternative to Node.js/Python. - StamManif/mcp-stama

    24d ago · github.com

  15. 15

    The first pure-Rust GGUF inference engine. No C. No Python.

    May 2026 · github.com

  16. 16SA

    Hi HN, I’ve been working on Shibuya, a next-generation Web Application Firewall (WAF) built from the ground up in Rust. I wanted to build a WAF that didn't just rely on legacy regex signatures but could understand intent and perform at line-rate using modern kernel features. What makes Shibuya different: Multi-Layer Pipeline: It integrates a high-performance proxy (built on Pingora) with rate limiting, bot detection, and threat intelligence. eBPF Kernel Filtering: For volumetric attacks, Shibuya can drop malicious packets at the kernel level using XDP before they consume userspace resources.…

    Feb 2026 · ghostklan.com

  17. 17
    SHIM2

    Prevent data leaks and slash your LLM bill by 40%

    Jan 2026 · getshim.tech

  18. 18TF

    I’d originally launched my app: Private LLM[1][2] on HN around 10 months ago, with a single RedPajama Chat 3B model. The app has come a long way since then. About a month ago, I added support for 4-bit OmniQuant quantized Mixtral 8x7B Instruct model, and it seems to outperform Q4 models at inference speed and Q8 models at text generation quality, while consuming only about 24GB of RAM[3] at 8k context length. The trick is: a) to use a better quantization algorithm and b) to use unquantized embeddings and the MoE gates (the overhead is quite small). Other notable features include many more…

    2024

  19. 19

    A playful leaderboard for Claude Code usage

    Jul 2026 · shibaita.ai

  20. 20

    The Confidential AI Gateway

    Dec 2025 · orgn.com

  21. 21BA

    Hi HN, I'm excited to share Bodhi App, a tool designed to simplify running open-source Large Language Models (LLMs) locally on your laptops. While we currently support M2 Macs, we plan to support other platforms as our community grows. # Problem To use LLMs, you typically need to purchase a subscription from providers like OpenAI or Anthropic, or use OpenAI API credits with compatible Chat UIs. These options can not only burden you financially, but also raise data security and privacy concerns. Many laptops are capable of running powerful open-source LLMs, but for non-tech users, setting…

    2024 · github.com

  22. 22DP

    Hey HN! Dinoki is a 6MB native AI assistant with pixel pets that live on your desktop. Built for power users who want privacy, performance – and a bit of fun. See it in action: https://www.youtube.com/watch?v=V_pv2BU0CZA What makes it different: - 6MB native app (SwiftUI for macOS, WPF for Windows) – starts instantly - Zero telemetry – your API keys and data stay local - Pixel art companions with physics-based behaviors - Flexible AI support – OpenAI, Anthropic, OpenRouter (300+ models), or Ollama (offline) Why pixel pets in an AI assistant? Beyond being fun, pixel art is…

    2025

  23. 23

    POC: Private PDF AI using only your browser with WebGPU

    Nov 2025

  24. 24

    1-click custom AI workflows from your clipboard

    May 2026 · tbidesign.com

Ranked by how close each launch is in meaning, then by votes. Refine with a description →