StatusSign in to track progress Evaluate Final Answers with Agent-Aware Normalization
Score final answers while rejecting leaked scratchpad, tool, or evaluator content.
DifficultyEasy
Evals85%15mStatusSign in to track progress Score Agent Behavior with a Rubric
Score answer quality, tool use, grounding, recovery, and safety as separate rubric dimensions.
DifficultyMedium
Evals53%25mStatusSign in to track progress Detect Prompt Injection in Model OutputPro
Identify outputs that attempt to override system or task instructions.
DifficultyHard
Evals35%35mStatusSign in to track progress Compare Expected vs Actual Tool TracesPro
Evaluate agent trajectories against an expected trace order.
DifficultyHard
Evals38%35mStatusSign in to track progress Aggregate Stochastic Agent Evals Across SeedsPro
Aggregate seeded eval runs to expose pass rates, instability, and stochastic behavior.
DifficultyMedium
Evals55%20mStatusSign in to track progress Explain Agent Eval Failures Without Leaking TestsPro
Group hidden eval failures into actionable feedback without exposing secret cases.
DifficultyMedium
Evals51%25mStatusSign in to track progress Evaluate an Agent System Design with a RubricPro
Score a proposed agent architecture against a comprehensive rubric covering orchestration, tools, memory, safety, evals, and observability.
DifficultyHard
Evals30%40mStatusSign in to track progress Implement an Agent Final Review DiagnosticPro
Inspect an agent config, execution trace, and eval results; classify failures; prioritise root causes; and produce a targeted review plan.
DifficultyHard
Evals25%45m