Dev tools · alternatives · 2026

24 alternatives to Cachinator
Open source express middleware for Rate limiting & caching
Below are 24 products that do a similar job, ranked by how close each is in meaning and then by launch-day votes.
- 1

- 2

- 3
ReliAPI▲87Stop losing money on failed OpenAI and Anthropic API calls.
Dec 2025 · kikuai-lab.github.io · its alternatives →
- 4

- 5
Aproxymade▲10Smart REST API Monitoring & Caching. Zero Code Changes.
Jun 2026 · aproxymade.com · its alternatives →
- 6AC
Multi-tier exact-match cache for AI agents backed by Valkey or Redis. LLM responses, tool results, and session state behind one connection. Framework adapters for LangChain, LangGraph, and Vercel AI SDK. OpenTelemetry and Prometheus built in. No modules required - works on vanilla Valkey 7+ and Redis 6.2+. Shipped v0.1.0 yesterday, v0.2.0 today with cluster mode. Streaming support coming next. Existing options locked you into one tier (LangChain = LLM only, LangGraph = state only) or one framework. This solves both. npm:…
Apr 2026 · its alternatives →
- 7GL
2019 · github.com · its alternatives →
- 8

- 9PF
2025 · github.com · its alternatives →
- 10
MakeHub.ai▲118LLM Provider arbitrage to get the best performance for the $
2025 · makehub.ai · its alternatives →
- 11PP
Hey all, I recently published this Go package, and would like to show off as well as get feedback!
2024 · github.com · its alternatives →
- 12

Connect AI agents to browser through raw CDP
Apr 2026 · openbrowser.me · its alternatives →
- 13

- 14
BossHogg▲65Agent-first CLI for PostHog analytics and feature flags
May 2026 · github.com · its alternatives →
- 15

Single-agent LLMs suck at long-running complex tasks. We’ve open-sourced a multi-agent orchestrator that we’ve been using to handle long-running LLM tasks. We found that single LLM agents tend to stall, loop, or generate non-compiling code, so we built a harness for agents to coordinate over shared context while work is in progress. How it works: 1. Orchestrator agent that manages task decomposition 2. Sub-agents for parallel work 3. Subscriptions to task state and progress 4. Real-time sharing of intermediate discoveries between agents We tested this on a Putnam-level math problem, but the…
Feb 2026 · github.com · its alternatives →
- 16

I built OpenSwarm because I wanted an autonomous “AI dev team” that can actually plug into my real workflow instead of running toy tasks. OpenSwarm orchestrates multiple Claude Code CLI instances as agents to work on real Linear issues. It: • pulls issues from Linear and runs a Worker/Reviewer/Test/Documenter pipeline • uses LanceDB + multilingual-e5 embeddings for long‑term memory and context reuse • builds a simple code knowledge graph for impact analysis • exposes everything through a Discord bot (status, dispatch, scheduling, logs) • can auto‑iterate on existing PRs and…
Feb 2026 · github.com · its alternatives →
- 17RL
Generative AI applications pose a unique challenge in production. They are computationally intensive and orders of magnitude slower than traditional data-intensive applications. Scaling these applications is further complicated by expensive hardware requirements and GPU shortages. Consequently, developers are scrambling to implement home-grown caching and rate-limiting solutions, which are error-prone and difficult to get right. FluxNinja Aperture delivers a production-grade experience with a purpose-built load management platform that provides rate & concurrency limiting, caching, and…
2024 · fluxninja.com · its alternatives →
- 18

High-performance cache policies and supporting data structures for Rust systems, with optional metrics and benchmarks. - OxidizeLabs/cachekit
Jan 2026 · github.com · its alternatives →
- 19RC
Hello HN! We're building a caching solution for LLMs (ChatGPT, Claude). By combining cutting-edge approaches, such as edge computing, prompt compression, vectorization, and others - it can reduce your AI bills by up to 10x and significantly lower response times. Key Features: - cost efficiency: our system stores frequent queries, reducing the number of upstream (paid) API calls - fast responses: with various nodes globally, we reduce latency by serving data from the nearest location - scalability: designed to handle increasing loads and data sizes without degrading performance. The cache…
2024 · edgematic.dev · its alternatives →
- 20RC
2016 · github.com · its alternatives →
- 21

A Python Library for Efficient LLM Query Caching
2024 · github.com · its alternatives →
- 22RA
While building my AI-powered dating app, I couldn't find a memory backend that was affordable,accessible and efficient. I built Redcache-ai to meet this need. Redcache-ai is also available as a Python package. Happy to receive feedback and answer questions. Note: I am not a native English speaker. Apologies for the typos and grammatical errors.
2024 · github.com · its alternatives →
- 23

- 24

Raymond here from Butter.dev, an LLM response cache built as a chat-completions proxy. Today we're launching a key feature for the platform: the ability to generalize on dynamic, templated inputs. Caching at the HTTP request level has the obvious problem of generalizability. Nearly no request is identical, due to templated variables (like names) and metadata (like timestamps), so exact-match cache lookups rarely hit. We solve this at Butter by using LLMs to detect dynamic content in requests and derive their inter-relationships, allowing the cache entry to be stored as a template + variables…
Jan 2026 · blog.butter.dev · its alternatives →
Also compare
Ranked by how close each launch is in meaning, then by votes. Prices were read from each product’s own site when checked and can change. Refine with your own description →