ElevenAgents by ElevenLabsScale conversations without scaling your team
Promoted
Maker
π
Hey Product Hunt! π
I built Mutant because I kept seeing AI evaluations become static pretty quickly ,you write a dataset, run it, and eventually stop discovering new failure modes.
Mutant has two core pieces:
β’ Behavioral mutations β take an existing evaluation dataset and generate targeted variations to expand evaluation coverage.
β’ Red-team agents β let autonomous agents attack LLMs, RAG systems, and AI agents to uncover vulnerabilities and unexpected failure modes.
The goal is simple: make it easier to continuously find the ways AI systems can fail before your users do.
Would love to hear what kinds of failure modes or attacks youβre currently struggling to evaluate. π