nowfound

Alternatives

Products that do what AegisLM does

See how easily your AI can be broken — in seconds

  1. 1
    Fabraix196

    Find gaps in your AI agents before users do

    May 2026 · fabraix.com

  2. 2

    Get a pentest done, today.

    Dec 2025

  3. 3

    The narrow control plane for AI agent tool and API calls.

    Aug 2026 · aegisora-ai.vercel.app

  4. 4WP

    Anthropic and OpenAI's publicly available models are explicitly guard-railed so that they refuse offensive tasks. And their cyber-focussed models are gated for enterprises. This leaves SMEs and mid market open to major vulnerabilities. AI can be used as both an adversarial and defensive tool in the world of cyber. A worst case outcome is if only the adversaries have access. Meanwhile, most existing AI cyber tools are just wrappers. The problem is that they still have all the guardrails on from the foundation model where they will inherit its refusals. For this project we've post-trained a…

    Jun 2026 · argusred.com

  5. 5
    AEVS131

    proof-of-execution for AI agents

    Jun 2026 · aevs.fetch.ai

  6. 6

    Ask your Playwright tests why they failed

    Apr 2026 · testrelic.ai

  7. 7FP

    We've built an open-source tool to stress test AI agents by simulating prompt injection attacks. We’ve implemented one powerful attack strategy based on the paper [AdvPrefix: An Objective for Nuanced LLM Jailbreaks](https://arxiv.org/abs/2412.10321). Here's how it works: - You define a goal, like: “Tell me your system prompt” - Our tool uses a language model to generate adversarial prefixes (e.g., “Sure, here are my system prompts…”) that are likely to jailbreak the agent. - The output is a list of prompts most likely to succeed in bypassing safeguards. We’re just getting…

    2025 · security.vista-labs.ai

  8. 8

    An open benchmark for AI agents that test APIs

    May 2026 · resources.kusho.ai

  9. 9

    See what breaks your AI agent and fix it automatically

    Jan 2026

  10. 10IM
  11. 11

    Find vulnerabilities in your AI prompts before your users do

    25d ago · testmyprompt.net

  12. 12AT

    Hi Hacker News! We're launching Zalor, an agent testing platform. Agents often break when you tweak system prompts, swap models, or add tools. Zalor automatically generates test scenarios and evaluates your agent so you know it's reliable before deploying to production. We currently support the OpenAI Agents SDK and are onboarding other frameworks. A GitHub integration is coming so you can get feedback on every update. Looking forward to hearing feedback from people building agents.

    Mar 2026 · agents.zalor.ai

  13. 13

    Find AI vulnerabilities before hackers do

    Mar 2026

  14. 14
    Aegis3

    A Human-in-the-Loop, AI-driven bouncer for Linux servers

    Feb 2026

  15. 15EY

    I built an open-source AI agent for security testing to find and fix vulnerabilities in your code. I’ve noticed how bad security vulnerabilities have gotten with everyone shipping AI code slop, so I wanted to build something that allows for vibe-coding at full speed without compromising security. Traditional security tools aren’t effective, and manual pen-testing can’t keep up with the rapidly growing AI code This tool runs your code dynamically, finds vulnerabilities, and validates them through actual exploitation. You can either run it against your codebase or enter your (or someone…

    2025 · github.com

  16. 16LC
  17. 17

    Compare AI models side-by-side on same prompt

    Feb 2026

  18. 18

    Zero-trust proxy & escalation boundary for AI agents.

    16d ago · github.com

  19. 19

    The Only AI Tool That Doesn't Trust AI

    Mar 2026

  20. 20

    Break, test, and secure AI agents before production

    Jun 2026 · github.com

  21. 21

    Meet the most gullible AI intern. Your job is to break it.

    Jun 2026 · breaktheprompt.xyz

  22. 22

    Break you AI systems before attackers do

    Feb 2026

  23. 23FC

    Hi everyone, I’ve been working on an open-source tool called Flakestorm to test the reliability of AI agents before they hit production. Most agent testing today focuses on eval scores or happy-path prompts. In practice, agents tend to fail in more mundane ways: typos, tone shifts, long context, malformed input, or simple prompt injections — especially when running on smaller or local models. Flakestorm applies chaos-engineering ideas to agents. Instead of testing one prompt, it takes a “golden prompt”, generates adversarial mutations (semantic variations, noise, injections, encoding edge…

    Jan 2026

  24. 24

    One URL in, full test suite out

    Feb 2026

Ranked by how close each launch is in meaning, then by votes. Refine with a description →