How it works
Context and Brief
The candidate is placed into a scenario specifically designed for their role, industry, and seniority band. They are given raw context, constraints, and a deliverable objective.
Prompting and Delegation
The candidate must direct an AI assistant to complete the task. We evaluate Agentic Judgment (the quality of the delegation) and Prompt Quality. Seeded errors are confirmed delivered per session. Where a seeded error never surfaced, the candidate is not scored on it.
Evaluation and Catching Errors
Every simulation responds with deliberate, seeded errors tailored to the industry (e.g., a compliance violation for a financial scenario). The candidate must read, evaluate, and identify these errors. Critical Evaluation measures whether they caught these flaws or accepted them as fact.
Judgment and Override
Upon spotting an error, the candidate must actively override and correct it. We capture the complete timeline, immutable work revisions, and final submission to measure Output Quality based on unverified AI-originated substance versus verified candidate corrections.
The Final Submission
Once submitted, the evidence is captured on an append-only timeline. A controlled outcomes-grading process uses temperature-0 model calls with a median-of-three consensus on sampled graders to produce a final AI Collaboration Index.