The submission pack — Nextgen Recruiter

Built by Noorflows (noorflows.com) for Nextgen Recruiter, an agency placing candidates with client companies. Working system, running on demo data. Every candidate, employer, client and document is invented.

What it does

Before a candidate goes to a client, the system reads the documents sent with the application and produces a pack the agency hands over alongside them. The pack always says three things: what was checked and came back clear, what was found with the evidence attached, and what could not be checked from the documents at all.

The third part is the point. No system can promise that a fabricated application never reaches a client. What an agency can have is a record of what it checked before it sent one, and that record is what survives the conversation afterwards. So the limits are printed on every pack, including the clean ones.

The one idea: read the document twice

A CV is two documents wearing one filename. There is the text layer a machine extracts, and there is the page as it renders, which is what a person actually sees. Almost always they say the same thing. When they do not, that gap is the finding — text in white on white, at one point, positioned off the edge of the page, or built from zero-width characters. A parser reads it; a reader never does.

The check is deliberately not a list of tricks. Written against white-on-white and one-point type it would be out of date the day somebody invents a fifth trick. The property tested is that a reader cannot see the text, which is true of every trick including the ones not yet invented.

The idea is not ours and we say so. It comes from a study of roughly 200,000 real applications by researchers at Duke, UNC and Berkeley, who built the same pair of detectors and found each catches cases the other misses: on 62,029 documents the two agreed on 517 and disagreed on 361. They also found that general-purpose tools do badly at this, because a CV is long and the hidden part is short. About 1% of the documents they read carried something hidden, and they describe that as a conservative lower bound rather than a measurement.

Worth saying about that study, since we hold everybody else to it: the 200,000 applications were supplied by hireEZ, a recruiting-software company, which also co-authored. That does not make the finding wrong, and the method is published, but a corpus from a vendor in the same market is a fact a reader is entitled to weigh.

The measured number

Applications in the set
8, written by us
Carrying a planted defect
5
Built to look wrong and be right
3
Planted defects found
5 of 5
Traps wrongly flagged
0 of 3
Things that cannot be checked from documents at all
4, on every pack

Precision and recall both read 1.00, and that is close to meaningless without the next sentence: the applications were written by us, so we knew the answers before we started. A figure measured against a fixture we built is a statement about our checks, not about the world.

It is deliberately not a fraud-detection rate. Nobody knows how many fabricated applications sit in a real postbag, so any figure of the form "we catch this share of fakes" would be invented.

What it refuses to do

It does not score, rank or recommend anybody
There is no score in the pack and no field to put one in. This is a legal position as much as a design one: New York City's Local Law 144 applies to a tool whose output is the only or most significant criterion in an employment decision, or an override to a human decision. A tool that reports what two readings of a document disagree about, and ranks nobody, sits outside that.
It never infers anything from who somebody is
The name is held, because an agency putting somebody forward needs it, and it is used to address them and nothing else. Nationality, photograph, accent, age, gender and address are not held anywhere in the system, so there is nothing to sort by even by accident, and a test walks the source to keep it that way. Illinois already outlaws the tidiest version of this by banning zip codes as a proxy for a protected class.
It reports documents, not people
A document can contain text a reader cannot see. Only a person can decide what that means about the candidate, and often the answer is nothing much.
It decides nothing
Every finding goes to whoever is putting the candidate forward, with the evidence attached and the option to disagree.

What cannot be checked from documents at all

Whether the person who wrote these documents is the person who will do the job
Nothing in a file can establish that. It is settled by speaking to them, and by the identity checks the client runs at offer stage.
Whether the employers named actually employed them
That is a reference, and a reference is a phone call to a human being. We have not made it and cannot make it from the file.
Whether the qualifications named were awarded
That is a check against the awarding body. It is separate work with its own cost and it has not been done here.
Whether the work described was actually theirs
A CV describes a team's output in the first person as a matter of convention. Only a conversation separates the two.

The account a rejected applicant is owed

Colorado's rewritten AI law, in force on 1 January 2027, requires that within thirty days of a decision that went against somebody they receive a plain-language description of the decision and of the software's part in it, along with a route to correct wrong information and to ask for human review. The system generates that account, and a test asserts it contains no jargon.

On the European rules, since every guide still says the old thing

Recruitment sits in Annex III of the EU AI Act. The high-risk obligations that attach to it were deferred from 2 August 2026 to 2 December 2027 by the digital omnibus. The Article 50 transparency obligations were not deferred. Deferred is not removed, and this build is designed for the regime rather than the deadline.


Demo data throughout. No real application has been copied into this system, including for testing. The interactive version · How this handles your data