Integrity Score Explained: Designing a Fair Interview Score
A pile of raw signals is not a decision. To be useful to a recruiter, dozens of events have to collapse into one number with a clear meaning, and that number has to be explainable, hard to game, and never the final word. Here is how to design a score that earns trust rather than merely producing one.
Short answer
An interview integrity score is a triage aid, not a verdict. Start every session at 100, deduct only for confirmed findings, weight each deduction by severity rather than counting events, and make every deducted point traceable to a specific signal and timestamp.
Report severity and confidence as separate values. Severity asks how serious an event would be if connected to the interview; confidence asks how well the evidence supports that connection. Blending them into one number destroys the distinction reviewers most need.
On this page
- Start from 100 and deduct with evidence
- Severity, not a flat count
- Severity and confidence are different axes
- Making the score hard to game
- Score bands should map to actions
- Explainability over precision theatre
- Handling false positives
- What never belongs in the score
- Calibrating after launch
- Review rubric for hiring teams
- Frequently asked questions
Start from 100 and deduct with evidence
The cleanest mental model is a trust budget. Every interview starts at 100. Confirmed findings deduct points in proportion to their severity, and every deduction is tied to the specific signal and timestamp that caused it. A score of 18 should be readable as a short list: this happened, then this, then this, and here is the evidence for each.
The alternative, building a score upward from accumulated suspicion, sounds equivalent and is not. Additive models make the absence of evidence look like partial evidence, and they make it hard to answer the question a candidate will actually ask, which is "what did I lose points for?"
Severity, not a flat count
Not all signals are equal, and a score that counts them equally will be wrong in both directions at once. A confirmed capture-excluded overlay or an active remote-control session is close to disqualifying on its own. A single focus change is informational. Weighting by severity, and letting informational signals annotate the timeline without lowering the number, keeps the score honest and avoids punishing harmless behaviour.
| Tier | Example findings | Effect on score | Default action |
|---|---|---|---|
| Critical | Active remote-control session with injected input; capture-excluded answer overlay aligned to questions. | Dominates via severity cap. | Escalate for review before any decision. |
| High | Assistant process plus question-aligned paste; capture device registered only for the session. | Large deduction. | Reviewer must read the timeline. |
| Medium | Repeated focus departures aligned to questions with no corroborating signal. | Moderate deduction. | Reviewer attention if combined with others. |
| Low | A display hot-plugged mid-session; an unrecognised background process. | Small deduction. | Context only. |
| Informational | Two displays present throughout; a paste during setup. | None. | Annotates the timeline. |
The informational row is the one teams most often get wrong. If ordinary behaviour moves the number, reviewers learn that a score of 88 means nothing in particular, and the whole scale loses its meaning within a month.
Severity and confidence are different axes
This is the single most useful structural decision in integrity scoring, and it is frequently skipped. Severity asks how risky an event would be if it were connected to the interview. Confidence asks how strongly the available evidence supports that connection. They are independent, and collapsing them loses information reviewers depend on.
| Low confidence | High confidence | |
|---|---|---|
| High severity | Remote-desktop client installed but never active. Serious if used; no evidence it was. | Injected input during the coding task. Act on this. |
| Low severity | An unidentified background process with no activity. Note and move on. | A brief focus change during setup. Certainly happened, means nothing. |
Presenting both axes lets a reviewer clear the top-left cell in seconds, which is where most review time is otherwise wasted. Presenting one blended number forces them to re-derive the distinction from the raw timeline every time.
Making the score hard to game
- Correlate, do not enumerate. A score driven by intersecting signals is far harder to defeat than one tied to a list of individual checks that can be neutralised one at a time.
- Do not publish the playbook. Surface findings with evidence without handing anyone an exact threshold to tune against. Publishing the bands internally is right; publishing the weights is not.
- Apply severity caps. One confirmed critical signal should dominate, so stacking trivial noise cannot dilute a serious finding, and neither can a long clean stretch.
- Do not let duration launder severity. A two-hour session should not score better than a forty-minute one simply because the same finding is a smaller fraction of it.
Score bands should map to actions
A score is easier to use when each band has a default action attached. Publishing the bands internally stops recruiters inventing their own thresholds, which is what causes two candidates in identical situations to be treated differently because they interviewed with different teams.
- 75 to 100, low riskProceed normally. No reviewer action, and the report is retained for audit only.
- 40 to 74, needs reviewA named reviewer reads the timeline before the loop advances. Interviewer clarification may resolve it.
- 0 to 39, critical riskThe process pauses pending review, and escalation goes to recruiting operations, legal or security depending on the role.
Note that even the bottom band pauses rather than rejects. A score should never be wired directly to an outcome, both because the evidence rarely supports it and because automated adverse decisions attract regulatory attention in a growing number of jurisdictions.
Explainability over precision theatre
A score that says 73 instead of 72 is not useful unless the reviewer can understand why. In hiring, the goal is not mathematical drama; it is a fair, repeatable summary of evidence. Each score movement should map back to a human-readable event: a remote-control process became active, a capture-excluded overlay appeared, a display changed mid-session, or a large paste followed a question with no intervening work.
This is also why integrity scoring should avoid black-box behavioural claims. A candidate should never lose points because they looked away, paused, spoke with an accent, typed slowly or seemed nervous. The score should be grounded in technical events around the interview environment, and it should always preserve the timeline that produced it. Pair it with a signed, tamper-evident report and the result is something a team can defend to a candidate, to HR, or in an appeal.
Handling false positives
Fair systems assume false positives will happen and design the review path for them rather than treating them as edge cases. A remote desktop client may be installed but inactive. A second monitor may be part of a normal workstation. A paste may be explicitly allowed in an open-book exercise.
The score should therefore separate presence, activity, timing and policy context, so a reviewer can clear a benign case in under a minute. If clearing a false positive takes as long as investigating a real finding, reviewers will start clearing everything by default, and the system stops working.
Key takeaways
- Start at 100 and deduct with evidence; every point traces to a signal and a timestamp.
- Weight by severity, and let informational signals annotate rather than penalise.
- Report severity and confidence as separate axes, never blended into one number.
- Correlate signals, cap severity, and avoid publishing exact thresholds.
- Map bands to actions internally so candidates get consistent treatment across teams.
- The score focuses a human decision. It never makes one.
What never belongs in the score
Do not include facial expression, eye movement or gaze tracking, accent, personality inference, camera background, room appearance, or subjective confidence judgements. These are noisy, bias-prone and usually unnecessary, because technical signals answer the question directly and with far less collateral.
A practical test: if a signal cannot be explained to the candidate in plain language, and defended as relevant to a policy they saw before the session, it does not belong in the score or in the review. Anything that fails that test is generating legal exposure without generating evidence.
Calibrating after launch
The first version of an integrity score should be treated as a calibrated workflow, not a finished truth machine. Review the first several dozen monitored sessions by hand and compare the score against the human decision. Pay particular attention to sessions landing near the review threshold: those are where small weighting mistakes create the most operational friction.
- Track cleared signalsIf reviewers clear the same signal repeatedly, its weight is wrong or its benign explanation is too common. Fix the model, not the reviewers.
- Check for buried severityAre genuinely serious sessions landing in the middle band? That usually means the severity cap is too weak.
- Compare across teamsIf two teams treat the same band differently, the bands are not published clearly enough internally.
- Review together, not in isolationRecruiting operations, legal, security and hiring managers should calibrate jointly. A score tuned by one function will optimise for that function's risk only.
Over time the best score becomes less surprising. Recruiters know what high risk means, hiring managers know where to look in the report, and candidates are evaluated against disclosed rules rather than subjective suspicion.
Review rubric for hiring teams
- Does the signal violate a policy the candidate saw before the session began?
- Did multiple independent signals cluster around the same answer or task?
- Is there a benign explanation that fits the timeline at least as well?
- Are severity and confidence both high, or only one of them?
- Would a clarifying question or a second interview resolve this fairly?
- Can the final decision be explained to a recruiter, a hiring manager, a legal reviewer and the candidate?
Frequently asked questions
What is an interview integrity score?
An interview integrity score is a single number, usually on a 0 to 100 scale, summarising how much technical evidence of policy-violating behaviour was observed during a monitored interview.
It is a triage aid rather than a verdict: its job is to tell a reviewer which sessions need attention and to point them at the specific events that produced the number. A score with no attached evidence trail is not usable, because nobody can act on a bare figure they cannot interrogate.
What does an integrity report actually mean?
An integrity report describes what technically happened around an interview: which signals fired, at what timestamps, how severe each was, and how confident the system is that each is connected to the interview.
It does not mean a candidate cheated. It means these events occurred, and a human should decide what they imply. A good report separates severity from confidence, so a reviewer can tell the difference between a serious event that might be unrelated and a minor event that certainly happened.
How should an integrity score be calculated?
Start every interview at 100 and deduct only for confirmed findings, weighting each deduction by severity rather than counting events flatly. Let informational signals annotate the timeline without moving the number.
Apply a severity cap so one confirmed critical finding dominates and cannot be diluted by stacking trivial noise. Every deducted point must trace back to a specific signal and timestamp, or the score cannot be explained and therefore cannot be defended.
What is the difference between severity and confidence in an integrity score?
Severity asks how serious the event would be if it were genuinely connected to the interview. Confidence asks how strongly the available evidence supports that connection.
A remote-desktop process that is installed but never active is high severity and low confidence. A brief focus change during setup is low severity and high confidence. Reporting them as a single blended number destroys the distinction reviewers most need in order to clear benign cases quickly.
What should never be included in an interview integrity score?
Facial expression, eye movement, gaze tracking, accent, personality inference, camera background, room appearance and any subjective confidence judgement. These are noisy, bias-prone and largely unnecessary given that technical signals answer the question directly.
A useful rule: if a signal cannot be explained to the candidate in plain language and defended as relevant to a disclosed policy, it does not belong in the score.
How do you stop candidates gaming an integrity score?
Three things do most of the work. Correlate rather than enumerate, so the score depends on intersecting signals instead of individual checks that can be defeated one at a time. Do not publish exact thresholds, since surfacing findings with evidence does not require handing anyone a number to tune against.
Apply severity caps so one confirmed critical signal dominates and cannot be washed out by accumulating harmless activity.
References
- InterviewWatch scoring methodology, for the production weighting model.
- Tamper-evident signed reports, on making a score defensible after the fact.
- How to detect AI assistance in remote interviews, on promoting observations to findings.
- Consent-first interview monitoring, on the disclosed policy a score must map to.
A score your team can defend
InterviewWatch reports severity and confidence separately, traces every deduction to a timestamped signal, and never makes the decision for you.