The assumption underpinning most assessment security is that if you control the browser and watch the screen, you can see what the candidate sees. That assumption is now false, and understanding precisely why is necessary to respond sensibly.
The mechanism
A current-generation real-time assistance tool works roughly like this:
- Audio capture. The application taps the system audio loopback — what the speakers are playing — plus the microphone. It hears the interviewer's question without needing access to the meeting application.
- Transcription and generation. Audio is transcribed and sent to a language model with context the candidate configured in advance: the job description, their résumé, the target company, sometimes a full solution library.
- Overlay rendering. The answer appears in a window drawn above all other applications, positioned near the camera so the candidate's gaze stays plausible.
- Capture exclusion. This is the critical step. Operating systems provide APIs that mark a window as excluded from screen capture. The window is visible to the person sitting there and absent from any screen share, recording, or proctoring screenshot.
Nothing here is a hack or an exploit. Every step uses documented operating-system functionality intended for legitimate purposes — accessibility overlays, presentation tools, and privacy features.
Why the usual defences do not see it
Lockdown browsers control the browser. The assistance is a separate native process. A lockdown browser has no visibility into other applications on the machine and no authority over them.
Screen sharing and recording receive the composited frame minus excluded windows. The proctor and the recording see a clean desktop. This is working as designed, from the operating system's perspective.
Browser-based proctoring extensions are confined to the browser sandbox. They cannot enumerate native processes.
Gaze tracking is genuinely useful signal, but the overlay is deliberately placed near the camera, and candidates practise. Gaze alone produces too many false accusations to act on.
The uncomfortable summary: the entire category of browser-level assessment security was designed for a threat model where cheating meant opening another tab. That threat model is obsolete.
The other techniques
Remote access takeover. A confederate connects to the candidate's machine and drives the assessment. The video shows the right person; someone else is typing.
Virtual machines. The assessment runs inside a VM while assistance runs on the host, entirely outside the assessed environment.
Second-voice coaching. Someone off-camera supplies answers. Detectable in audio analysis, invisible on video.
Deepfaked candidates. Real-time face replacement so the person interviewing is not the person who will show up. This has moved from theoretical to documented, and is the reason identity verification at hiring has become a live concern.
Straightforward second device. A phone or tablet running the same tooling, outside the camera frame. Low-tech and effective.
What detection actually requires
Since the assistance executes on the desktop, detection must too.
Operating-system level process and window inspection — enumerating what is actually running and which windows are marked capture-excluded. A window hidden from screen capture is a strong, specific signal with very few legitimate explanations during an assessment.
Remote-access and VM detection — identifying the environment rather than the content.
Post-session analysis — deepfake and synthetic-content detection across video and audio, second-voice detection, attentiveness and secondary-device signals, and cohort-level outlier ranking that flags a candidate whose pattern is statistically unlike their peers.
Defensible evidence. For a hiring or academic-integrity decision to survive challenge, "the model thought it was suspicious" is not sufficient. You want a binary determination where one is available — was a named restricted program running, yes or no, with process identity and timestamps — and clearly-labelled probabilistic signals where it is not.
This is the problem ScreenComply.AI addresses: a desktop agent operating below the overlay layer, plus browser-based post-session analysis, producing verdicts where the question is binary and signals where it is not.
Frequently asked questions
Can you detect AI use just by watching the candidate?
Not reliably. Behavioural cues — gaze drift, unnatural pauses, answers that are too structured — are suggestive and generate far too many false positives to support a decision. Candidates also practise specifically to suppress these cues. Behaviour is useful as corroborating signal alongside technical detection, not as a basis for a determination.
Do lockdown browsers still have a purpose?
Yes, for a narrower purpose than they are often sold for. They prevent casual tab-switching, copy-paste, and access to local files, which still matters. They cannot address native-application assistance, and should not be presented to stakeholders as if they do.
Is detecting overlay windows an invasion of privacy?
It depends entirely on scope and disclosure. Enumerating which windows are marked capture-excluded during a disclosed assessment window is far less invasive than capturing screen content or files. The defensible posture is signal-layer detection only, no content capture, a clearly bounded session, and explicit prior disclosure to the candidate.
Should we just allow AI in assessments instead?
For many roles this is the right answer — design a task where skilled AI use is part of the competency and assess judgement rather than recall. But it has to be a deliberate decision with a task designed for it. Leaving an unassisted-skills assessment in place while knowing it is being assisted measures nothing.
How common is this?
Tooling is commercially marketed, inexpensive, and explicitly advertised as undetectable, with large user bases. Any organization running remote technical assessments at scale should assume a meaningful share of candidates have access to it, and that the share is growing.
