broke to built. Say hello
Menu
All games and apps

Company Bench

Open original

Can your AI agent hold a job? Open-source benchmark for AI agent trustworthiness, not capability: 29 chairs, 7 departments, 241 deterministic checks, 78 planted traps (prompt injection, dirty data, irreversible actions). Scored by code, no LLM judge. Outputs trust level L0-L3

Source code agent-benchmark · agent-evaluation · agent-safety · agentic-ai · ai-agents

Runs in your browser, served from our own domain. No account, no install. If it misbehaves in this frame, open it on its own.