Home / Blog / Toward a first-party benchmark
Research

Toward a First-Party Benchmark: How Often Do Candidates Use AI Overlays?

Every published number on AI-assisted interview cheating so far comes from a survey: someone asking a manager whether they suspect it happened, or asking a candidate to admit it happened. Those are useful, but they are proxies. The number we actually want, how often live AI overlays are detected during real interviews, has to come from instrumented sessions, and nobody has produced it properly.

Short answer

Survey numbers measure suspicion and admission, which bound the problem from opposite directions without measuring it. A real benchmark needs instrumented sessions, a published methodology, a disclosed denominator, and independent publication.

We have not published our own detection rate, because our sessions come from organisations that already suspected a problem. That rate would be higher than the market's and we could not say by how much. Publishing it would be useful to us and dishonest.

Why self-report is a different question

The two available survey methods do not fail through poor execution. They measure genuinely different quantities from the one people cite them for.

MethodActually measuresDirection of errorWhy
Ask managersSuspicion, ever, across a career.OverstatesSuspicion is cheap, unfalsifiable, and inflated by the expectation that AI use is now normal.
Ask candidatesWillingness to admit misconduct.UnderstatesSelf-reported wrongdoing is systematically low, by an unknown factor.
Ask employers about discovered casesFraud that was eventually caught.UnderstatesUndetected cases are, by definition, absent.
Instrument the sessionsDetection rate.MeasurableNobody neutral has run it at scale.

Note that three of the four rows have a known direction of error, which is genuinely useful information. The 6% admission figure is a defensible floor and the 59% suspicion figure is a defensible ceiling. What nobody can do is claim a point estimate between them.

We know the floor and the ceiling. Everything published as a rate is somebody choosing a point between them and not saying so.
What we can defend, and what we cannot 6% candidates admitted defensible floor 59% managers suspected defensible ceiling true rate: unmeasured every published "rate" is a point chosen in here Bounding a quantity is a real result. Presenting a point inside the bounds as a measurement is not.
The published surveys bound the problem. Nothing published so far measures it.

What a credible benchmark requires

  1. A published methodologyReproducible by someone who disagrees with the conclusion. This is the property that separates research from marketing.
  2. A stated definition of a detectionDoes a flag that a reviewer cleared count? Does a single signal count, or only a correlated finding? These choices move the number by multiples.
  3. A disclosed denominatorPer session, per candidate, or per hiring loop are three very different rates from identical events.
  4. Coverage across many organisationsOne vendor's customers are self-selected. So are one large employer's candidates.
  5. Confirmed separated from flaggedConflating them inflates everything downstream and is the single most common flaw.
  6. Independent publicationReleased by a party that does not profit from the number being large.
6Properties a citable benchmark needs
0Published benchmarks that currently have all six
0Detection rates we have published, for the reason below

Where we are, honestly

We are in a position to measure detections, and our data does not meet the bar above. It is worth being specific about why, because the same limitation applies to every vendor in this category.

  • Our population is self-selected. Organisations instrument interviews because they already suspect a problem, and they instrument the round they are most worried about. Both push the observed rate up.
  • We cannot quantify the selection effect. We do not know the market rate, so we cannot say whether our sessions run twice as hot as average or ten times.
  • Our denominator is not the market's. Our sessions skew toward remote technical rounds at companies hiring for high-compensation roles, which is exactly the highest-risk slice.
  • Definitions are ours. We would be choosing what counts as a detection, and we benefit from choosing generously.

A rate from that data would be commercially useful and epistemically dishonest. So we have not published one, and we would encourage buyers to ask any vendor who has published a figure which of the four points above they have addressed.

Key takeaways

  • Survey figures measure suspicion and admission, which bound the problem rather than measuring it.
  • 6% is a defensible floor and 59% a defensible ceiling. Any point estimate between them is a choice, not a finding.
  • A citable benchmark needs published methods, a stated detection definition, a disclosed denominator, multi-organisation coverage, confirmed-versus-flagged separation, and independent publication.
  • We have not published a detection rate because our population is self-selected and we cannot quantify by how much.
  • Your own funnel distribution is more actionable than any global average would be.

Why vendor rates deserve scepticism

Including ours, if we ever publish one. The structural problem is that a vendor controls every input to the number and benefits from it being large.

The vendor choosesEffect on the headline rate
What counts as a detectionCounting single signals rather than confirmed findings can multiply it.
The denominatorPer loop rather than per session inflates it several times over.
Whether cleared flags countIncluding them can double or triple the figure.
Which customers are includedSelecting high-risk accounts raises it arbitrarily.
When to publishPublishing after a bad quarter for the industry is a choice too.

None of this makes a published figure false. It makes it unverifiable, which matters because a business case anchored to an unverifiable number is one a sceptical finance partner can dismantle in a single question. Build the case on cost instead: see the cost of a bad hire.

The better alternative: your own data

The honest answer to "what is the rate" is that you should measure yours, and that it will be more useful than any industry figure even if one existed.

  • Elevated-session rate by role. Where your exposure actually concentrates.
  • Rate by interview type. Which formats produce signals, which is often a question-design finding rather than a fraud finding.
  • Rate by interviewer. Variation here means your process is subjective somewhere.
  • Proportion cleared after review. Your false-positive proxy, and the number that determines whether reviewers keep reading reports.
  • Signal clusters that recur. Tells you which specific method you are actually facing.

A global average would compress all of that into one number that tells you nothing about what to change. Your own distribution points at the fix.

How this could actually get built

The measurement is not the hard part. The agreement is.

  1. Agree a definition of a detectionAcross several vendors or several large employers. This is where the effort dies, and it has to come first.
  2. Fix a denominatorPer monitored live session is the most defensible, because it is the unit that was actually observed.
  3. Contribute aggregated counts, not customer dataNo candidate-level information needs to leave any organisation for this to work.
  4. Disclose the populationSector, role seniority and region, so readers can judge selection effects rather than guess at them.
  5. Publish through a neutral partyAn industry body or academic group, with the methodology released alongside the figure.

We would participate in that and cannot construct it alone, since a benchmark built from one vendor's data is the exact thing this article argues against. If you run hiring at scale and would find this useful, we are interested in the conversation.

Frequently asked questions

Why are survey numbers not a benchmark for AI interview cheating?

Because they measure a different quantity. Asking a manager whether they suspect AI use measures suspicion, which is cheap, unfalsifiable and inflated by the general expectation that AI use is now common. Asking a candidate to admit it measures willingness to disclose misconduct, which understates by an unknown factor.

The number people actually want, how often AI overlays are detected in real interviews, requires instrumented sessions and has not been measured at scale by anyone neutral.

What would a credible interview fraud benchmark require?

Six properties: a published methodology anyone can reproduce, a stated definition of what counts as a detection, a disclosed denominator so the rate has a unit, coverage across many organisations rather than one vendor's customers, confirmed findings separated from cleared flags, and independent publication rather than release by a party selling the remedy.

Missing any one of those turns the number into marketing.

Why has InterviewWatch not published its own detection rate?

Because our data does not meet the bar we just described. Our sessions come from a self-selected population: organisations that already suspected they had a problem and chose to instrument the round they were most worried about.

A rate from that population would be higher than the market rate and we could not say by how much. Publishing it would be commercially useful and epistemically dishonest, so we have not.

Would a vendor-published detection rate be trustworthy?

Treat it with heavy scepticism, ours included. The vendor defines what counts as a detection, chooses the denominator, decides whether cleared flags are counted, and profits from a higher number.

None of that makes a published figure false, but it makes it unverifiable, and unverifiable numbers should not anchor a business case that a sceptical CFO will attack.

What is the alternative to an industry benchmark?

Your own funnel data, which is more useful anyway. Elevated-session rate by role, interview type and interviewer, with the proportion cleared after review, measured over a quarter.

That reflects your sourcing, your question design and your candidate population, which vary enormously between companies. A single global average would not tell you what to change; your own distribution does.

How could a proper benchmark actually get built?

It needs several vendors or several large employers to agree a shared definition of a detection, contribute aggregated counts with disclosed denominators, and publish through a neutral party such as an industry body or academic group.

The hard part is not the measurement, it is the agreement on definitions and the willingness to publish numbers that may be commercially inconvenient. We would participate in that and cannot construct it alone.

References

  1. Interview fraud statistics 2026, for the sourced survey figures and their limits.
  2. 2026 mid-year read, on what has and has not changed.
  3. The cost of a bad hire, for a business case that does not need a rate.
  4. InterviewWatch methodology, for how we define a finding internally.

Help build the real number

If you run hiring at scale and want a benchmark that would survive scrutiny, we are interested in the conversation.

Get in touchTry now