Should AI Usage Be a Performance Review Metric?
Meta grades engineers on AI-generated code, and Qorvo ties bonuses to AI adoption. Here's how AI usage is becoming a review metric, and how to measure it fairly.
Meta’s performance tracker, Checkpoint, now logs how many lines of code its engineers generate with AI versus how much they write by hand, feeding that into the broader impact assessment behind each employee’s rating. Land in the top of Meta’s four performance tiers, an assessment where AI-driven impact is now an explicit input, and you’re eligible for a 200% bonus multiplier. Meta has been careful to call this an “impact-evidence starter,” not a straight reward for AI usage, but the direction is clear: AI usage isn’t just a productivity hack anymore. For a growing list of employers, it’s something that feeds into what you get scored on.
This guide covers who’s already doing it, why it’s controversial, and how to build AI-usage metrics into reviews without punishing people for being cautious.
What “AI Usage as a Metric” Actually Means
AI-usage metrics measure whether and how employees use AI tools in their day-to-day work, then factor that into ratings, bonuses, or both. This differs from measuring output quality. A company using AI usage as a metric might track lines of AI-generated code, hours saved through AI-assisted drafting, or how many workflows an employee has automated, separate from whether the resulting work was good.
The distinction matters because usage and impact aren’t the same thing. An employee can run every task through an AI tool and still produce mediocre work, or barely touch AI and still be a top performer. Companies that conflate the two risk rewarding activity instead of results.
Who’s Already Doing It
Meta made the biggest and most public move. Starting in 2026, the company is tying performance ratings and bonuses to “AI-driven impact,” with Checkpoint aggregating more than 200 data points per software engineer, including AI-generated code volume, error rates, and bugs tied to an engineer’s work. Select engineering teams have been asked to generate over 75% of committed code using AI tools by mid-2026.
Microsoft told managers in mid-2025 that AI use was no longer optional, and Google’s CEO delivered a similar message to staff around the same time. Beyond big tech, some companies are formalizing the practice in writing: Qorvo disclosed a standalone performance objective in its fiscal 2025 executive incentive plan, weighted at 20%, focused specifically on AI tool exploration and deployment to cut costs and improve efficiency.
The throughline: these aren’t pilot programs. They’re showing up in the same performance cycles that determine raises and promotions.
Why HR Leaders Are Considering This
HR is the natural owner of this shift: they already set review criteria, train managers on consistent evaluation, and calibrate ratings across teams. When a new signal like AI usage enters the review process, HR decides whether it’s measured well, or not measured at all.
The logic is straightforward from a leadership seat. If AI tools genuinely make people faster and more effective, then AI fluency becomes a real skill gap, similar to how spreadsheet literacy separated performers a generation ago. Rewarding early adopters is one way to close that gap company wide instead of waiting for it to close on its own.
There’s also a signaling function. When a CEO says AI use will affect ratings, it moves AI adoption from “optional nice to have” to “how we work here” faster than any training program could. Meta framed Checkpoint explicitly as an “impact-evidence starter,” not an activity tracker, which is the distinction leadership teams are trying to thread.
The Fairness Problem
Most managers aren’t ready to evaluate AI usage as a review criterion fairly. As Forbes reported in May 2026, performance review season is changing as companies begin measuring AI fluency at work, and many managers lack a consistent framework for judging it. Without clear criteria, AI-usage scoring turns into a popularity contest for whoever talks about AI the most in their self-review.
Harvard Law School’s Forum on Corporate Governance has tracked a related trend at the executive level: companies increasingly disclosing AI-related objectives in incentive plans, often without standardized definitions of what “good” AI usage looks like. The same ambiguity that shows up in a proxy statement shows up in a manager’s rating rubric.
Harvard Business Review argues that performance management needs entirely new metrics for the AI era, not just old metrics with an AI checkbox added. Usage counts and adoption percentages are easy to measure and easy to game. Impact is harder to measure and the thing that actually matters.
How to Measure AI Usage Fairly
The fairest AI-usage metrics separate adoption from impact and score them individually rather than blending them into one number. Track how often and how broadly someone uses AI tools as one axis, and whether that usage actually produced better or faster work as a second, independent axis.
- Frequency and breadth of use: how often someone uses AI tools and across how many task types, not just whether they used one once.
- Time or effort saved: a concrete, attributable outcome (faster turnaround, fewer manual steps) rather than a raw usage count.
- Quality of the output: the work still has to hold up. AI-assisted output that requires heavy rework doesn’t count as a win.
- Judgment in applying AI: knowing when not to use AI (sensitive client communication, novel problems without training data) is itself a skill worth recognizing.
Our guide on measuring AI adoption with employee surveys has a ready-to-use question set for the frequency and breadth dimensions if you want to baseline before adding anything to formal ratings.
Building It Into Reviews Without Penalizing Skeptics
The fastest way to make this backfire is to treat AI hesitancy as a red flag on its own. Some employees are cautious because their role touches regulated data, client trust, or safety-critical decisions, and that caution is often correct.
A few guardrails help:
- Separate “uses AI well” from “uses AI often.” Write these as distinct criteria in your rubric so managers can’t collapse them into one score.
- Give employees a way to explain low usage. A self-review prompt like “where did you choose not to use AI, and why” surfaces judgment instead of penalizing restraint.
- Calibrate across managers before finalizing ratings. AI-usage criteria are new enough that individual managers will interpret them inconsistently without a calibration pass.
Windmill’s calibration feature flags rating discrepancies and manager patterns automatically, which is exactly the kind of check a new, unproven metric like this needs before it reaches a paycheck.
Where This Is Headed
AI-usage metrics are unlikely to stay confined to Meta-scale engineering orgs. As AI assistants become embedded in how Slack, Jira, and other everyday tools get used, “did you use the AI tools available to you” becomes as answerable as “did you hit your sales number.” The organizations getting ahead of this are building the measurement framework now, before it’s forced on them by a rushed rollout during review season.
If you’re weighing whether to formalize AI usage in your own review cycle, start with the survey data, not the rating rubric. Understanding actual adoption patterns first keeps the metric honest.
Frequently Asked Questions
Should companies use AI usage as a performance review metric?
Companies can use AI usage as one input among several, but it works best as a secondary signal alongside impact and quality, not a standalone score. Treating usage and impact as the same thing risks rewarding activity over results.
Which companies tie AI usage to performance reviews?
Meta ties bonuses and ratings to "AI-driven impact" through its Checkpoint tracker, which logs AI-generated code volume alongside 200+ other data points. Microsoft and Google have both told employees AI use is expected, and companies like Qorvo have disclosed standalone AI-adoption performance objectives.
How do you measure AI usage fairly in performance reviews?
Measure frequency and breadth of AI use, time or effort saved, and output quality as separate criteria rather than one blended score. Give employees room to explain low usage, since caution around regulated or sensitive work is often a legitimate judgment call, not a skill gap.
Is it unfair to grade employees on AI adoption?
It can be, if usage is measured without accounting for role differences or without a clear rubric, since managers currently lack consistent frameworks for judging AI fluency. Calibrating ratings across managers before they're finalized helps catch inconsistent scoring before it affects pay.
What's the difference between AI usage and AI impact as a metric?
AI usage measures whether and how often someone uses AI tools; AI impact measures whether that usage actually improved the work or saved meaningful time. Reviews that conflate the two risk rewarding employees who use AI a lot but produce output that needs heavy rework.