Alternatives
Products that do what Skill Bench does
Automated evaluation for Claude Code skills.
- 1

Store, review, and share your Claude Code sessions
Mar 2026 · bench.silverstream.ai
- 2

Multi-agent review catching bugs early in AI-generated code
Mar 2026 · claude.com
- 3

- 4

- 5

Ready-made AI analytics skills for your business data
Jun 2026 · databox.com
- 6

- 7

Contribute to alpbahadur/interns-review-plugin development by creating an account on GitHub.
1d ago · github.com
- 8PS
I got tired of playwright-mcp eating through Claude's 200K token limit, so I built this using the new Claude Skills system. Built it with Claude Code itself. Instead of sending accessibility tree snapshots on every action, Claude just writes Playwright code and runs it. You get back screenshots and console output. That's it. 314 lines of instructions vs a persistent MCP server. Full API docs only load if Claude needs them. Same browser automation, way less overhead. Works as a Claude Code plugin or manual install. Token limit issue:…
Oct 2025 · github.com
- 9SF
May 2026 · github.com
- 10

- 11

- 12AP
Mar 2026 · lab.puga.com.br
- 13FO
2025 · github.com
- 14RL
Jun 2026 · github.com
- 15IM
At my work they provided a single Claude subscription for everyone on the team. To be honest I like kiro better as it provides a way better SDD management. But the company can't provide it and I can't afford it yet. Turns out I had the skill creator skill in my claude instance so I made use of it to create this Skill. I made it fully by using Claude but I wanted to make it open source, so I asked it to help me make tests and preparations for it, even a CI to run python tests. Well, we got this results with it: - Phase 2A: 67 static assertions (Python script, runs in CI) - Phase 2B: 15…
May 2026 · github.com
- 16
Claude Fable 5.1▲142Claude’s most advanced models for coding and knowledge work
5d ago · anthropic.com
- 17

- 18

- 19HW
A bunch of companies that I spoke to had their own claude & codex OTel dashboards that showed spend + seats per month. However, none of the dashboards actually analyzed how the engineers worked with the tools and if there were any areas for improvement! That's why I created https://www.promptster.ai. Managers get aggregate level view of code quality and how that ties with team workflows (nothing on a per-engineer level). While engineers get personalized coaching on how they can save tokens while keeping output high. We also have a tool built for individuals to test their local…
Jul 2026
- 20AS
May 2026 · github.com
- 21AT
I recently built a small open-source tool to benchmark different LLM API endpoints — including OpenAI, Claude, and self-hosted models (like llama.cpp). It runs a configurable number of test requests and reports two key metrics: • First-token latency (ms): How long it takes for the first token to appear • Output speed (tokens/sec): Overall output fluency Demo: https://llmapitest.com/ Code: https://github.com/qjr87/llm-api-test The goal is to provide a simple, visual, and reproducible way to evaluate performance across different LLM providers, including…
2025 · llmapitest.com
- 22CR
Jan 2026 · github.com
- 23WO
Apr 2026 · npmjs.com
- 24

Ranked by how close each launch is in meaning, then by votes. Refine with a description →