Measure how
people actually work
with AI
Proofency measures how candidates and employees direct and verify AI in real work. A live work simulation (averaging 30 minutes) produces an AI Collaboration Index (0–100) across nine dimensions, with a stated precision band.
Your AI ROI depends on who you hire.
The adoption number you report to your board is not the number that produces returns.
Proofency measures how people actually work with AI, so hiring teams can separate routine adoption from demonstrated proficiency.
of employees at the studied organisation used AI regularly.
met the study's definition of highly sophisticated users.
Frequency of use was not a reliable productivity signal. Junior employees used AI more often, substantially for personal tasks.
The most sophisticated users were frequently above manager level—one reason to calibrate the proficiency bar by seniority rather than assume junior staff are more fluent.
Source: Harvard Business Review, 2026, reporting on a KPMG and UT Austin analysis of approximately 1.4 million workplace AI interactions.
The Cognitive Fingerprint
Four questions a hiring manager actually asks, answered by a descriptive profile of how the candidate collaborates with AI.
What a typical 30-minute session looks like
Not a quiz. Not a questionnaire. Your candidate works on an actual task — in their role, in their industry, in their language — with a live AI assistant. Completion times vary by role, averaging 30 minutes. Every simulation includes deliberate AI errors seeded into the responses. We observe how they prompt, evaluate, and override.
- Realistic constraintsIndustry-specific scenarios that match real job demands, preventing generic responses.
- Verified flaw deliverySeeded errors are confirmed delivered per session. Where a seeded error never surfaced, the candidate is not scored on it.
Every scenario is built for the role, the seniority, and the work
A marketing coordinator is tested on execution. An analyst is tested on data interpretation and AI error detection. A director is tested on strategy and root-cause thinking.
The six cognitive archetypes
Observed working styles that describe how a candidate tends to direct and verify AI.
What the AI Collaboration Index measures
Six of the nine scored dimensions are summarized here. The Cognitive Fingerprint is a separate profile, not a score.
Built to be defended, not just delivered
We do not score what the session did not deliver. Where a signal was not observed, we report it as not measured. Everything else is controlled, auditable, and defensibly engineered for enterprise hiring.
- Temperature-0 scoringRemoves sampling variance from the model call. Residual grader disagreement is measured rather than assumed away, which is why every score ships with its precision band.
- Median-of-three consensusEach sampled grader runs three times and the median is taken. The same submission scores the same way twice.
- Authored severity tiersHow serious a finding is comes from a signed criteria catalog, not from the model's judgment. The grader reports what it found; it does not decide what it is worth.
- Non-compensatory floorA floor applies at every band. A strong overall profile cannot mask a failure in Critical Evaluation.
{
"dimension": "critical_evaluation",
"score": 84,
"precision_band": "±3",
"methodology": {
"temperature": 0.0,
"graders_sampled": 3,
"resolution": "median"
},
"flags": [
"caught_compliance_error",
"override_accepted"
]
}Where the assessment applies
The assessment is currently calibrated across core enterprise functions and applies to every band from Individual Contributor to C-suite.
Severity floors scale with altitude: an executive who swallows flawed AI reasoning is an organizational risk multiplier. Seniority raises the verification bar rather than lowering it.
Your team is already using AI. Find out how well.
For workforce development, the assessment is just the baseline. Proofency emits an ordered session list targeting specific skill gaps, providing a personalized path to reach the proficiency bar for their exact role.
View the Employee Development frameworkCommon questions
What is an AI proficiency assessment?
An AI proficiency assessment measures how candidates actually work with AI in real work scenarios, rather than relying on self-reported skills or quizzes. It evaluates their judgment, critical thinking, and prompt quality under realistic conditions. It sits in the same category as skills assessment and pre-hire testing platforms. It is not an AI detection tool, and not a self-assessment questionnaire.
What is the AI Collaboration Index?
The AI Collaboration Index is a 0–100 score across nine weighted dimensions, representing a candidate's ability to direct and verify AI output. It is reported with a ±3 precision band. The separate Cognitive Fingerprint describes collaboration style but is not a scored dimension.
How is Proofency different from an AI skills quiz?
Proofency is a live work simulation where candidates complete a real deliverable using an AI assistant. It observes demonstrated behavior, such as catching deliberate AI errors, rather than asking multiple-choice questions about AI tools.
Can a candidate game the assessment?
The assessment evaluates a complete evidence trail rather than a single answer. Seeded errors are confirmed delivered per session; where an error never surfaced, the candidate is not scored on it. Evidence capture is separated from scoring and based on verifiable work revisions and an append-only timeline.
Which roles and seniority levels are supported?
Currently, we support roles like Mechanical Engineer and Data Analyst. Scoring applies across multiple bands from Individual Contributor (IC) and Lead up to Manager, Director, and C-suite levels, with a non-compensatory Critical Evaluation floor at every band.
How accurate is the score?
Scoring uses temperature-0 generation, a median-of-three consensus on sampled graders, and authored severity tiers. We report the score with a ±3 precision band, and where a signal was not observed, we report it as not measured rather than guessing.