PRE-LAUNCH · early access opening soon

Security scores for every AI agent.

Taint is the independent security evaluation layer for AI agents — we measure how secure an agent actually is, before it ships and continuously after, and publish it. We're building this now.

// scoring agents across the Claude & OpenAI marketplaces, coming soon · taint.ai
// illustrative sample — not real agents or scores yet
// prompt injection // tool misuse // data exfiltration // privilege escalation // multi-step trajectory attacks // behavioral drift
// how it works

Point it at an agent.
Get back the truth.

The dangerous failures don't live in a single prompt — an agent reads poisoned input, drifts off-task, escalates, and exfiltrates, each step plausible alone. Existing tools check one prompt or one tool call and miss the sequence. Owl Watch scores the whole run.

01 / connect

Plug in the agent

Any agent from the marketplaces, or your own, connected through our harness.

02 / attack

Run the battery

Adaptive adversarial attacks probe the full trajectory — offline in CI/CD, or online against live agents.

03 / score

Read the score

A security score, reproducible failing traces, and drift tracking that updates as the agent changes.

// two products, one standard

Free to see. Paid to secure.

FREE · PUBLIC

The leaderboard

Independent scores for marketplace agents — the neutral reference no guardrail vendor can credibly give you. This is what launches first.

  • One standardized battery across every agent
  • Reproducible traces, not just a number
  • Re-scored continuously as agents change
ENTERPRISE

The eval platform

Your agents, under continuous evaluation — before they ship and while they run.

  • Offline: full batteries in CI/CD on every change
  • Online: live agents probed for drift & new attacks
  • Attacks that evolve as your agents evolve
DESIGN TARGET
<2%

The number everyone else hides

Anyone can claim a high catch rate. What decides whether a control survives in production is its false-interrupt rate — how often it breaks legitimate work. We're building every score to report it, not just catch rate — so it's a real number you can hold us to, not a marketing line.

// early access

Get on the list.

Nothing's live yet — we're building the leaderboard and the eval platform now. Leave your email and we'll bring you in as it opens, starting with the free leaderboard.

No spam. Just a note when the leaderboard opens.