verifier.org

Gray Swan AI

runs a public attack competition that finds real exploits — and describes its actual product in adjectives.

not published#frontier-labs-commission#researchers-wanting-to-c

a genuinely novel model for finding novel attacks, wrapped around a commercial offer you cannot evaluate from outside.

the arena is the interesting part: a crowdsourced attack competition with open enrolment and cash prizes, feeding discovered exploits back into gray swan's defences. it claims three million attack attempts and says most arena-discovered exploits remain unpublished. as a discovery mechanism for genuinely new jailbreaks that's more likely to work than an internal team, and it names google deepmind, openai, anthropic and meta as customers or partners.

it also publishes system cards for evaluations of frontier models including claude and gpt, and maintains a research hub it says holds 150-plus curated tools and datasets.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
not published
our verdict

a genuinely novel model for finding novel attacks, wrapped around a commercial offer you cannot evaluate from outside.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more ai red-teaming tools

NeMo Guardrails

nvidia's toolkit for constraining llm behaviour, used alongside testing

#guardrails#open-source

Giskard

open-source scanner for llm vulnerabilities and quality regressions

#open-source#scanner

ModelScan

scans model files for unsafe serialisation before you load them

#supply-chain#open-source

HiddenLayer

we checked thisnot published

fifty disclosed cves and thirty patents, across the broadest claimed scope here.

#enterprises-wanting-mode