Toward a First-Party Benchmark: How Often Do Candidates Use AI Overlays?
Every published number on AI-assisted interview cheating so far comes from a survey, someone asking a manager whether they suspect it happened, or asking a candidate to admit it happened. Those are useful, but they're proxies. The number we actually want, how often live AI overlays are detected during real interviews, has to come from operational detection data, not self-report. We don't have a defensible version of that number to publish yet, and we're not going to publish an estimate dressed up as one.
This post is about what it would take to do that responsibly, and an invitation to hiring teams who want to help build it.
Why self-report numbers aren't the same question
The existing published figures we cite in our statistics roundup, a manager's suspicion rate, a candidate's admission rate, are real and useful for understanding how the problem feels from either side of the table. But suspicion isn't detection, and admission undercounts by definition; nobody who successfully concealed AI assistance is in the numerator of a self-report survey. A benchmark built from actual, correlated detection events during monitored interviews measures something closer to the real rate, for the population of interviews it covers.
What a credible benchmark would require
- Enough volume to be meaningful. A benchmark drawn from a small number of interviews isn't representative of anything beyond itself.
- Consistent detection methodology across the sample. Comparing sessions scored by different signal sets or thresholds produces noise, not a rate.
- A defined confirmation standard. A raw flag rate isn't the same as a confirmed-after-human-review rate. Reporting one as the other would overstate the finding.
- Full anonymization. No candidate, company, or interview should be identifiable in aggregate published data.
- Transparency about what's excluded. Role type, geography, and interview format all affect the number; a responsible report segments by these instead of collapsing everything into one headline figure.
Where we are, honestly
We would rather take longer and publish a number we can defend line by line, including what it doesn't cover, than rush out a headline figure that collapses everything into one statistic. When we have a sample that meets the bar above, we'll publish the methodology alongside the result, not just the number.
In the meantime, the most useful thing hiring teams can do is start building their own internal benchmark. Our guide to internal benchmarking covers what to track even before an industry-wide number exists.
Key takeaways
- Suspicion and self-report surveys measure sentiment, not the actual detection rate.
- A credible operational benchmark needs volume, consistent methodology, a defined confirmation standard, and full anonymization.
- We haven't published an aggregate detection rate because we don't yet have a sample that meets that bar.
- Build your own internal benchmark now, it's more actionable than waiting for an industry-wide figure.
If you want to help build this
If your team runs enough monitored interviews to contribute meaningfully to an anonymized, aggregate industry benchmark, we'd like to talk. Reach out and we'll share what a responsible data-sharing structure looks like before any commitment.
Help build the real number
We're working toward a defensible, first-party benchmark on live AI-assisted interviewing. Reach out if you want to be part of it.