The screen is an assessment
I spent my early career in educational measurement. In that field, before an instrument is used to decide something about a person, you are required to argue for it.
Messick wrote the consequences of use into the definition rather than appending them to it. Validity, in his words, is “an overall evaluative judgment of the degree to which empirical evidence and theoretical rationales support the adequacy and appropriateness of interpretations and actions on the basis of test scores or other modes of assessment” (Messick, “Validity of psychological assessment,” American Psychologist, 50(9), 1995, p. 741).
Two phrases carry the weight here. Actions, which puts what a score is used to do inside the definition. And other modes of assessment, which is Messick extending validity to means beyond tests.
A résumé screen sits inside that definition. It produces information about a person, and that information supports a decision. Whether the information is a number, a rank, a knockout rule or a recruiter’s read does not change the position it occupies.
Two of his threat categories transfer without modification.
Construct underrepresentation. A prior question has to be asked first. What is the screen measuring? Present capability? Evidence of specific prior experience? Probability of performing the role well? Or whether the person is worth assessing further? Those are four different constructs and a requisition rarely says which one is meant — the question this library keeps asking of every system that scores something.
Whichever one is meant, the screen samples keyword presence, title adjacency, and continuity of employment, inside one document written in a specific genre. That is a narrow sample of any of the four constructs.
Construct-irrelevant variance. Résumé-writing skill. Formatting compatibility. Prestige of prior employers and institutions. Name effects. All of these vary across candidates. None of them is the construct, and none of them is detected.
Kane supplies the apparatus. He reframed validation as an argument you write down: state how the score will be interpreted and used, then evaluate each inference in the chain (Kane, “Validating the interpretations and uses of test scores,” Journal of Educational Measurement, 50(1), 2013). Apply that to a screen and four inferences need backing, which is the same chain I ran against a spending decision, pointed at a person.
The document becomes a match score. The score stands in for the candidate’s typical work. The score predicts performance on the job. The score justifies the decision.
There is a large literature showing that résumé-derived variables correlate with later outcomes (Schmidt & Hunter, “The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings,” Psychological Bulletin, 124(2), 1998). That literature does not finish this argument, and pointing at it is the usual way of avoiding the question. Kane’s question is narrower. Does this screen, applied to this requisition, support each of those four steps, for this population, at this threshold, for this decision. A general result about predictors is not a validity argument for a particular instrument in a particular use.
That particular argument is what is rarely written down anywhere a reader can see it — including at the standard the field itself names (Society for Industrial and Organizational Psychology, Principles for the Validation and Use of Personnel Selection Procedures, 5th ed., 2018).
The requirement has no time index
A requisition says what a role requires. It does not say when.
There is nothing wrong with requiring one capability on the first day and expecting another to develop over the first year. Organizations do this deliberately and they are right to. Some things have to be there on arrival because the work will not wait. Some things are built by the work itself.
Almost no requisition marks the difference. The list arrives undifferentiated, and a screen reads all of it as conditions to be met at the moment of measurement.
The clearest evidence that the list is not a day-one list is that it routinely holds things no outside candidate could have on day one. Familiarity with the internal data model. Experience with the student information system this particular organization runs. Everyone involved knows those get learned after hire. They sit in the same undifferentiated list as the requirements that genuinely cannot wait.
So the screen is not failing at something it set out to do. It measures current capability across every listed item, which is exactly what it claims to measure. The difficulty is that current capability across every listed item is not the quantity the decision needs. The decision needs a judgment about performance over the life of the role, and part of that performance rests on capabilities the organization intends to build after hire. The screen is a sound measure of the wrong quantity.
The missing time index also breaks the third step in Kane’s chain. That step claims the score predicts performance on the job. Without a time index nobody has said which job — the one on the first day, or the one twelve months later. Those are different targets and the screen is scored against neither of them in particular.
It is the same document I have taken apart from the employer’s side, where two roles get fused into one posting and the rubric ends up screening for the wrong one.
What comes back is one bit
A rejection is a binary outcome produced by a system with many causes. At least seven distinct states sit behind it. The role was cancelled. An internal candidate existed. The budget froze. The requirements changed. The screen misread the document. A stronger candidate applied. The person could not do the work.
The claim is narrower than saying the other six have nothing to do with them. Some of them partly do. It is about what the signal licenses. Nothing in the outcome warrants an inference about their capability, and that inference is the one the outcome invites.
It is also narrower in a second way. Some rejections do name a cause. The role was cancelled, or the budget froze, and the message says so. Where that happens the problem described here does not arise, because the candidate knows which state applied. This section is about the rejection that names nothing. Nobody currently counts how often that happens, which is why the third remedy below asks for the number.
The mapping from those states to the outcome is many-to-one. The inverse problem is therefore underdetermined. That is the property that matters, and it holds whatever notation is used to describe it.
The information-theoretic version is the short way to say the same thing. A binary outcome carries at most one bit. Seven states cannot be recovered from one bit, because seven states need close to three. And one bit is a generous ceiling here. As rejections compound, each one arrives with the same missing information as the last. The signal does not deepen by being repeated. But its reading does.
The recipient is left to supply the cause themselves. Of the seven, only one supports the conclusion a rejection most readily invites: I could not do the work.
Repeat that across a job search and a second consequence appears. They cannot learn. A near miss and a mismatch produce the same outcome, so there is nothing to correct against. The rational response to an uninformative signal is to send more applications rather than better ones. That raises the volume employers cite when they explain why screening had to be automated in the first place.
Whether that loop runs at scale is an empirical question and I am not asserting it here. The point is narrower. Messick put the consequences of use inside the definition of validity rather than beside it.
He was explicit that this was not an addition to validity. In his words, it is “simply not the case that values are being added to validity in this unified view.” The unified view “exposes the inherent value aspects of score meaning and outcome to open examination” rather than treating them “as an adjunct or supplement” (same paper, p. 748).
These are consequences of use. They fall on the person being assessed, and no part of the hiring apparatus is chartered to measure them.
A reason travels when there is a score behind it
The channel is not always this thin, and the exception is worth reading closely.
An instructional designer wrote publicly in August 2026 about a rejection she received after an AI-conducted interview. The email offered her the option of detailed feedback and she took it. It gave specific examples of what she had said against what the employer was looking for, and it included positive points about her responses. Her own summary was that she had never received that much feedback from any interview she had done. She drew the comparison herself. What she was measuring it against was the canned line that begins “after carefully reviewing your resume.”
So a reason can travel. Look at what made that possible. She had been scored. An interview produces a record of responses with dimensions attached to it, so something exists to send back.
A résumé screen may well hold plenty behind its decision. Knockout results, ranking scores, extracted fields, recruiter notes, disposition codes. The difference is narrower than it first looks. Both systems hold a record. Only one holds a record built to be interpreted, and so only one holds a record that converts into a reason.
An operations leader wrote publicly the same month about her own search. Two hundred applications, ten replies, two offers on the table by the end of the cycle. She took the one from the employer whose people had been candid throughout — where she stood, how many others were in the pool, when a decision was coming. Her own summary was that the volume of applications an employer received was not the point. There were still ways to be respectful of the people who took the time to apply, and one of the two employers had used them.
Two different mechanisms, both returning something the candidate could act on. In one case a machine returned a reason. In the other, humans returned candor about the process. Neither turns on who spoke. Both turn on whether anyone did.
Two posts are observations and not a rate, and I am treating them as such. Together they cut against reading this note as an argument about automation. In one case the automated step was the one that explained itself. In the other, the humans were.
The error nobody inside can see
Criterion validation runs on people who were hired.
That sample is restricted by construction. A person screened out generates no performance record, no tenure data, and no manager rating. Whether that person would have done the work well is recorded nowhere, and it is recorded nowhere because the screen decided they would not be given the chance.
Note what that rules out. The employer cannot label its own rejections as errors, because labelling them would require the observation the screen prevented. A bad hire eventually becomes observable. A rejected candidate who would have succeeded does not. One of the two errors can be managed and the other cannot be seen.
So the harm is invisible to the party positioned to fix it. The instrument censors the evidence needed to validate one class of its own decisions. Ordinary post-hire performance data cannot reveal it, because the screen decides who enters that data.
Range restriction is a known problem, not a novel one (Sackett & Yang, “Correction for range restriction: An expanded typology,” Journal of Applied Psychology, 85(1), 2000). Corrections exist. They require observing performance on rejected candidates, which is exactly the observation the screen prevents.
There are ways around this. Audit sampling. Following rejected candidates who were hired elsewhere. Quasi-experimental designs that admit a few people below the threshold. Each of them works. None of them is ordinary practice, and none of them is what a validity coefficient computed on hires is doing.
The seam
Every organization runs on contracts between systems. Most of them are written down. Data contracts between teams. Service agreements with vendors. Reporting definitions between functions. When those contracts are missing, the systems drift apart, and someone eventually pays to reconcile them.
The seam between an organization and the labor market is one where the contract was never written, and where the counterparty does not see the terms. It is not alone in that. Credit scoring, insurance underwriting and admissions all have the same shape. What makes this seam worth pulling on is that every organization has one, and almost none of them treats it as a contract at all.
An unwritten contract does not stay empty. A market grows in the gap and sells the candidate a guess at the terms. It used to arrive as a document someone else wrote for them. Now it arrives as a tool that rewrites their own document against the job description. A job description is a guess at the criteria. Nobody in that transaction has seen the definition being matched against — not the vendor, not the candidate, not the employer reading the result. A model writes the résumé and a model reads it, and neither one has been told what the job requires.
Three things would help here. This note has named three failures in three places, and each one needs its own repair. A single remedy for all three would be a remedy for none of them.
Instrument the construct. A signed construct definition is what a written contract looks like on that seam. Every requirement carries three attributes. What capability. What would count as evidence of it. When it has to be true, on the first day or built inside the first year. One statement sits above the list: what we would expect to observe in twelve months if this role worked.
Hiring manager and HR both sign it. Not as ceremony. A contract with no signatory is a note to self, and a note to self is what a requisition already is.
This is a version stamp rather than a freeze. Requisitions should change, and organizations that cannot reshape roles die of it. The stamp says what we are measuring against as of today, and it records that people measured against the previous version were measured against the previous version. Software teams do not argue that specifications should never change. They argue that changes should be recorded. Nobody calls that bureaucratic.
Versioning fixes construct drift. It does nothing about the one bit.
Return the reason. An AI screen generates a great deal of information about why one application ranked below another. That is what it is built to do. Match scores against stated requirements. Position in the pool. Which criteria were met and which were not.
The historical defence of the one-line rejection was cost. At two thousand applications, writing individual feedback was not possible, and nobody pretended otherwise. That defence has expired. The assessment work is already done and the text is close to free to produce. What is missing is any obligation to send it.
So the ask is small. If you are going to screen with a model, return what the model produced. Not the sentence about moving forward with other candidates. The thing behind it.
Two conditions on that, and they are not decoration.
The return has to be structured. Free text needs review on every candidate, which is the cost that was there before. A structured return is reviewed once, at the template. Three fields carry most of it. Which stated requirement was not met. Which stage the application reached. Where it sat against the criteria the requisition named.
And the return has to come second. Feedback trains behaviour. If the screen is reading résumé-writing skill and title adjacency, then telling candidates what it read teaches them to optimise those, and every future application converges on whatever the instrument rewards. The signal degrades. A reason is worth returning only when the instrument is measuring something worth optimising toward. Version the construct first. Returning a reason drawn from an unspecified construct broadcasts noise with the employer’s authority behind it.
There is one serious objection and it is not the cost one. Counsel advises against stated reasons because a stated reason is evidence. That advice assumes withholding reduces exposure. Withholding does not remove the record. The screen’s output already sits in the applicant tracking system, already discoverable, and already unreadable by the only person it describes. Withholding moves the asymmetry rather than closing it. I am not a lawyer, and this objection deserves a better answer than the one I can give it.
Publish the numbers. Enforcement is publication. Three numbers per search. Applications received. Decisions issued with no stated cause. Applications that received no response of any kind.
Stated cause needs defining, or the second number is trivial to game. A message saying we have decided to move forward with other candidates states no cause. It restates the outcome. A stated cause names something about the application: a requirement not met, a stage not reached, a comparison lost against named criteria. Anything that names none of those counts toward the second number, including every well-written template.
One case sits awkwardly under that definition and is worth naming. Some rejections do state a cause on the requisition side. The role was cancelled, the budget froze, the search is on hold. That is a real answer for the candidate, and it settles which of the seven applied. It still counts toward the second number.
The second number does not measure whether the candidate learned anything. It measures whether the employer said anything about the screen. A cancelled requisition says nothing about the screen, so it counts as a decision issued with no stated cause, like any other. Counting it as a stated cause instead would let an employer lower the number without ever saying what the screen did.
Be clear about what these three do not do. They say nothing about whether the screen is valid. An employer could publish two thousand applications, nineteen hundred and forty causeless rejections and no silences at all, and still be running an instrument that measures nothing.
These numbers are not the third remedy. They are the audit that tells you whether the first two happened. The first asks what the instrument was meant to measure. The second asks what the person was told. The numbers ask whether either was done.
The third number carries the most weight, because it is the one that separates two situations a candidate cannot tell apart. Silence is what happens when a requisition receives two thousand applications, and it is also what happens when nobody bothers. Those are different, and the difference is only legible if the denominator is published beside the silence count. As things stand nobody can check it. Not the candidate, not a regulator, and on the evidence of what organizations measure about their own searches, not the employer either.
None of the three requires new instrumentation. Any applicant tracking system worth the license fee already holds all three, and the third is derivable from a timestamp. They go unreported because nothing obliges anyone to report them, and no candidate is in a position to ask.
The first organization to publish them will look worse than every organization that publishes nothing, because a number always looks worse than an absence. That is the whole reason nobody goes first.
Until those numbers are visible, silence stays uninterpretable. Until the construct is versioned, the employer cannot say what the screen was meant to measure. Neither side can answer a question about the same decision, and the decision gets made anyway.
Written August 2026 for the Analytic Bytes Library. Standing practice: no employer, client or vendor is named in this piece.
Questions, pushback, or a problem that looks like this one? Write to chai@analyticbytes.systems.