The AI readiness standard for hiring

Measure how
people actually work
with AI

Proofency measures how candidates and employees direct and verify AI in real work. A live work simulation (averaging 30 minutes) produces an AI Collaboration Index (0–100) across nine dimensions, with a stated precision band.

Sample Report
AI Collaboration Index
AI-Fluent
74
±3
Composite Score
Prompt Quality88
Critical Evaluation84
Judgment & Override76
Output Quality75
AI Efficiency72
Ethics Compliance74
Bias & Fairness— Not measured this session
Transparency & Accountability55
Agentic Judgment68
The proficiency gap

Your AI ROI depends on who you hire.

The adoption number you report to your board is not the number that produces returns.

Proofency measures how people actually work with AI, so hiring teams can separate routine adoption from demonstrated proficiency.

~90%

of employees at the studied organisation used AI regularly.

~5%

met the study's definition of highly sophisticated users.

Frequency of use was not a reliable productivity signal. Junior employees used AI more often, substantially for personal tasks.

The most sophisticated users were frequently above manager level—one reason to calibrate the proficiency bar by seniority rather than assume junior staff are more fluent.

Source: Harvard Business Review, 2026, reporting on a KPMG and UT Austin analysis of approximately 1.4 million workplace AI interactions.

The Cognitive Fingerprint

Four questions a hiring manager actually asks, answered by a descriptive profile of how the candidate collaborates with AI.

Q1
Can I trust the work they produce with AI?
Override Pattern
Shows whether a candidate actually reads and evaluates AI output before using it, or lets errors pass through unchecked.
Q2
What are the specific risks of this hire in my industry?
Error Taxonomy Profile
Maps which types of errors the candidate catches or misses across logical, factual, compliance, strategic, and framing problems.
Q3
How much will I need to manage their AI-assisted work?
Trust Calibration
Shows whether decisions to trust or challenge AI output were well placed. Reported when the session delivered enough signal to support it.
Q4
Between these three finalists, who do I hire?
Cognitive Archetype
Describes the shape of how someone works with AI—not whether they are good, but what kind of AI collaborator they are.

What a typical 30-minute session looks like

Not a quiz. Not a questionnaire. Your candidate works on an actual task — in their role, in their industry, in their language — with a live AI assistant. Completion times vary by role, averaging 30 minutes. Every simulation includes deliberate AI errors seeded into the responses. We observe how they prompt, evaluate, and override.

  • Realistic constraints
    Industry-specific scenarios that match real job demands, preventing generic responses.
  • Verified flaw delivery
    Seeded errors are confirmed delivered per session. Where a seeded error never surfaced, the candidate is not scored on it.
Explore the assessment process
01
Context
Given a raw, incomplete brief relevant to their specific role.
02
Delegation
Prompting the assistant to generate the initial output.
03
Evaluation
Spotting the deliberate flaws seeded into the AI response.
04
Override
Correcting issues and finalizing the submission.

Every scenario is built for the role, the seniority, and the work

A marketing coordinator is tested on execution. An analyst is tested on data interpretation and AI error detection. A director is tested on strategy and root-cause thinking.

Role and seniority adapted
Scenarios are scoped to the exact function and altitude of the candidate.
Deliberate AI errors to catch
The AI assistant will confidently present flawed reasoning to test evaluation.
Industry and language context
Scenarios carry industry and language context for relevance, ensuring candidates work with familiar domain concepts.

The six cognitive archetypes

Observed working styles that describe how a candidate tends to direct and verify AI.

Amplifier
Fast, high-leverage, and efficient. Extracts substantial value from AI with minimal prompts.
Skeptic
Questions AI output systematically and checks compliance, logic, and factual claims.
Architect
Thinks in systems and automation, designing structured AI-assisted workflows.
Explorer
Learns and adapts during the session, changing approach as new evidence appears.
Executor
Works consistently through defined tasks with a steady, repeatable collaboration pattern.
Editor
Actively rewrites and restructures AI output through frequent, substantive edits.

What the AI Collaboration Index measures

Six of the nine scored dimensions are summarized here. The Cognitive Fingerprint is a separate profile, not a score.

Critical Evaluation
Did they catch the deliberate flaws, or accept them as fact? A non-compensatory floor applies at every band.
Output Quality
Measured by unverified AI-originated substance vs independently verified candidate work.
Agentic Judgment
The quality of the delegation as written, carrying a fixed, published-internally weight in the composite.
Prompt Quality
Clarity, constraint setting, and contextual injection in their interactions.
Judgment & Override
How appropriately they corrected or discarded bad AI output.
AI Efficiency
How effectively they use AI to make useful progress without unnecessary prompting or unmanaged rework.

Built to be defended, not just delivered

We do not score what the session did not deliver. Where a signal was not observed, we report it as not measured. Everything else is controlled, auditable, and defensibly engineered for enterprise hiring.

  • Temperature-0 scoring
    Removes sampling variance from the model call. Residual grader disagreement is measured rather than assumed away, which is why every score ships with its precision band.
  • Median-of-three consensus
    Each sampled grader runs three times and the median is taken. The same submission scores the same way twice.
  • Authored severity tiers
    How serious a finding is comes from a signed criteria catalog, not from the model's judgment. The grader reports what it found; it does not decide what it is worth.
  • Non-compensatory floor
    A floor applies at every band. A strong overall profile cannot mask a failure in Critical Evaluation.
{
  "dimension": "critical_evaluation",
   "score": 84,
   "precision_band": "±3",
  "methodology": {
    "temperature": 0.0,
    "graders_sampled": 3,
    "resolution": "median"
  },
  "flags": [
    "caught_compliance_error",
    "override_accepted"
  ]
}

Where the assessment applies

The assessment is currently calibrated across core enterprise functions and applies to every band from Individual Contributor to C-suite.

Mechanical Engineer
Data Analyst
Marketing Coordinator
Operations Lead
Strategic Director
C-Suite

Severity floors scale with altitude: an executive who swallows flawed AI reasoning is an organizational risk multiplier. Seniority raises the verification bar rather than lowering it.

Your team is already using AI. Find out how well.

For workforce development, the assessment is just the baseline. Proofency emits an ordered session list targeting specific skill gaps, providing a personalized path to reach the proficiency bar for their exact role.

View the Employee Development framework

Common questions

What is an AI proficiency assessment?

An AI proficiency assessment measures how candidates actually work with AI in real work scenarios, rather than relying on self-reported skills or quizzes. It evaluates their judgment, critical thinking, and prompt quality under realistic conditions. It sits in the same category as skills assessment and pre-hire testing platforms. It is not an AI detection tool, and not a self-assessment questionnaire.

What is the AI Collaboration Index?

The AI Collaboration Index is a 0–100 score across nine weighted dimensions, representing a candidate's ability to direct and verify AI output. It is reported with a ±3 precision band. The separate Cognitive Fingerprint describes collaboration style but is not a scored dimension.

How is Proofency different from an AI skills quiz?

Proofency is a live work simulation where candidates complete a real deliverable using an AI assistant. It observes demonstrated behavior, such as catching deliberate AI errors, rather than asking multiple-choice questions about AI tools.

Can a candidate game the assessment?

The assessment evaluates a complete evidence trail rather than a single answer. Seeded errors are confirmed delivered per session; where an error never surfaced, the candidate is not scored on it. Evidence capture is separated from scoring and based on verifiable work revisions and an append-only timeline.

Which roles and seniority levels are supported?

Currently, we support roles like Mechanical Engineer and Data Analyst. Scoring applies across multiple bands from Individual Contributor (IC) and Lead up to Manager, Director, and C-suite levels, with a non-compensatory Critical Evaluation floor at every band.

How accurate is the score?

Scoring uses temperature-0 generation, a median-of-three consensus on sampled graders, and authored severity tiers. We report the score with a ±3 precision band, and where a signal was not observed, we report it as not measured rather than guessing.

Stop guessing who can work with AI.