What we measure
All nine scored dimensions
01 Prompt Quality
Evaluates the structure, context injection, and constraint setting of initial and subsequent prompts.
02 Critical Evaluation
Evaluates whether the candidate accepted flawed AI reasoning or actively identified it. A non-compensatory floor applies at every band—a strong overall profile cannot mask a failure here.
03 Judgment & Override
Measures whether a candidate's decision to trust AI output is well-placed, or if they produce AI-laundered errors (outputs that look reviewed but weren't).
04 Output Quality
The base of Output Quality is unverified AI-originated substance, measured by span attribution from the session timeline rather than textual overlap. Verified AI-originated content carries zero deduction at every band.
05 AI Efficiency
Measures whether the candidate uses AI to make useful progress without unnecessary prompting, repetition, or unmanaged rework.
06 Ethics Compliance
Evaluates adherence to ethical guidelines and safety boundaries during AI collaboration, ensuring generated content does not violate core policies.
07 Bias & Fairness
Assesses whether the candidate identifies and mitigates biased, unrepresentative, or exclusionary AI outputs when they occur.
08 Transparency & Accountability
Measures whether the candidate correctly attributes AI-generated work, cites sources when appropriate, and maintains an auditable trail of decisions.
09 Agentic Judgment
Measures the quality of the delegation as written, not the agent's response to it. It carries a fixed, published-internally weight in the composite.
Severity scaling by altitude
Severity floors scale with altitude. Tactical and execution accuracy is evaluated heavily at IC and Lead bands, whereas strategic alignment, compliance, and risk prevention are evaluated at Manager, Director, and C-suite bands. An executive who swallows flawed AI reasoning is an organizational risk multiplier, and seniority raises the verification bar rather than lowering it.
See pricing plans