←  Library
Essay 15

Validity is the layer beneath responsible AI.

Why responsible-AI rubrics in mental health are necessary, and what they need underneath.

Chaitanya Ramineni, PhDAugust 14, 202622 min read
Cover illustration for Validity is the layer beneath responsible AI.

The rubrics, and what they assume

A round of frameworks for governing AI in mental health has appeared in the last fifteen months. Lyra Health’s Polaris Principles articulated the industry-side posture. The National Academy of Medicine published its AI Code of Conduct for Health and Medicine in May 2025 with six commitments including monitoring performance. The American Psychological Association (APA) released ethical guidance for AI in health service psychology in 2025 and followed with a 2026 health advisory whose plain-language judgment about consumer AI in mental health was blunt: the chatbots and wellness apps lack the testing and safety measures needed for quality mental health support. VERA-MH (Validation of Ethical and Responsible AI in Mental Health) appeared as an open-source AI safety evaluation built around conversation simulation and clinician judging. ASL-MH (AI Safety Levels for Mental Health) proposed AI safety levels for mental health applications. The WHO convened an expert workshop on responsible AI for mental health in January 2026 with explicit recommendations about embedding mental health into AI impact assessments. Trovane, the Australian company behind the Lumara platform, published a thirteen-question handbook for evaluating any AI mental-health tool before deployment, written so that a clinical lead or a foundation program officer can carry it into a procurement conversation. The rubrics are necessary. This essay credits the work being done and names the layer the rubrics do not yet include.

Of that set, the handbook is the most operationally detailed, written as a scoring instrument rather than a statement of principles, so it is worth walking through. Each of its thirteen questions is scored on three tiers: Basic, Better, Best Practice. Five of the thirteen are gate questions, where a Basic score means the buyer does not proceed until the gap is resolved, and more than one Basic anywhere means the buyer does not deploy at all. The authors also publish an appendix naming where their own product falls below their own bar, which is rarer in vendor-authored guidance than it should be.

Before the handbook was released, Trovane previewed it publicly and pulled five of the thirteen forward as the ones to ask in a first vendor meeting. Those five carry most of the weight for the argument here:

  1. 01Where does AI inference happen — what is the data-residency and privacy architecture, and what flows where?
  2. 02Who is the named clinician accountable for crisis detection, and what does accountability actually mean?
  3. 03How does crisis detection perform across demographic groups?
  4. 04What happens after deployment — monitoring, drift, the post-launch discipline that distinguishes a deployment from a launch?
  5. 05Can we speak with the clinical lead directly?

These questions map almost exactly to the responsible-AI cluster appearing across the NIST AI Risk Management Framework (RMF), the EU AI Act risk tiers, the NAM Code of Conduct, APA’s ethical guidance, the WHO workshop outputs, and the industry frameworks circulating alongside them. They are the right floor for responsible AI in mental health. The institution that can answer them is operating at a much higher standard than the institution that has not asked them at all.

The preview post made one claim about those five that is worth sitting with: a responsible vendor answers all of them without hesitation. That is true, and it is also the reason the floor needs something underneath it. A vendor can answer all five cleanly, with inference on shore, a named clinician holding veto, stratified detection results, a monitoring regime with an owner, and a clinical lead who will take the call, and still never have been asked what the detector detects.

The question the rubric does not ask

There is a question the rubrics do not yet include, and it decides whether everything else they ask earns its place: does the system measure the mental-health construct it claims to?

This library has asked that question before — of higher-ed AI in What is this system actually measuring?, of impact investing in The valid dollar. Each time, the responsible-AI floor turns out to depend on an unexamined measurement claim. Mental-health AI is the third and hardest application, because the constructs the systems claim to measure are the ones the field itself has been actively contesting for over a decade.

Validity is the layer beneath the responsible-AI frameworks
The frameworks are necessary. They assume — but do not check — the discipline beneath.

The responsible-AI rubric as currently practiced assumes the measurement target is correct and asks how the model behaves around that target. Is the inference happening in a secure place. Is someone accountable for the outcome. Does the performance hold across populations. Is the system being monitored for drift. Is there a human you can talk to. Each is necessary. A crisis-detection model that is accountable to a named clinician, runs on data residing in the right jurisdiction, performs evenly across demographic groups, is monitored for drift, and connects to a reachable clinical lead is still a problem if what it is detecting is not crisis but a proxy for crisis the field has not validated yet. The rubric governs the deployment. Validity governs the claim the deployment is making. That distinction is not mine. Messick, 1995 settled it thirty years ago: validity is a property of the interpretation and use of a number, consequences included, not of the instrument that produced it. A tool cannot be valid. Only a claim made with it can be.

The Trovane handbook shows the pattern clearly, and it shows it because it is the most detailed of the set rather than the least. Its crisis-detection question asks whether detection accuracy has been audited across demographic groups: age, gender, First Nations status, users from culturally and linguistically diverse backgrounds, users with English as an additional language, neurodivergent users. That is a fairness question and a good one. It also assumes the construct is settled. A detector can only be tested for even performance across groups once there is an account of what it detects. None of the thirteen rows asks what construct the system is measuring, or what evidence supports the claim that it measures it. The closest any row comes is the audit question, whose top tier lists outcome tracking as two words alongside drift detection, with nothing behind it about which outcome or against what. What carries the safety claim across all thirteen is clinical veto: a named, credentialed clinician holds documented authority to block any product decision that conflicts with clinical safety. Their handbook draws a sharp line between clinically informed, meaning a clinician was consulted at some point, and clinically governed, meaning a named clinician holds ongoing accountability and can say no. The distinction is real, and most of the market fails it. Clinical veto establishes who is accountable for a decision. Evidence that the decision is measuring what it claims is a different thing, and no row asks for it.

The same omission appears in work that cleared peer review at the top of the field’s literature, which makes it the more telling case. In August 2026, a team from the Max Planck UCL Centre for Computational Psychiatry, Oxford, and the UK AI Security Institute published a framework in Nature Medicine for auditing how AI chatbots behave in mental-health conversations. A model plays a user with a psychiatric vulnerability, talks to a target chatbot for up to ten turns, and a judge model scores every turn against a thirty-nine-dimension rubric, thirteen dimensions of which were selected in advance as psychiatrically relevant and carry the analysis. Eight hundred and ten conversations across nine chatbots produced more than ninety thousand turn-level ratings. The paper calls itself a clinically validated framework and it earns much of that title, on the strength of its validation battery: judge checked against judge, scores tested for stability when the rubric is reworded, a check for whether models grade their own output more leniently, and twenty-seven licensed clinicians producing four hundred and eighty-eight independent ratings.

Every one of those checks validates the rater. The closest any of them comes to the target is a known-groups test, where the judge separated high- from low-risk conversations at a median AUC of 0.98. The Methods describe where those conversations came from: a separate model, Grok-4, was given the written descriptions of the judge’s own dimensions and asked to generate conversations matching them. The stimulus was written from the rubric and then scored against the rubric. What 0.98 establishes is that the rubric’s own definitions are legible to a machine reading them. The known in known-groups means known to the rubric, not known to a clinician. The construct being scored is called concerning behavior, and the paper’s stated reason for using it is its high face validity, which means it looks right to people who know the area. The criterion check asks whether the model agrees with clinicians, and the clinicians did not apply the judge’s rubric — they rated each response on a single 1-to-10 slider, from safe and appropriate to clearly harmful, working from a standardized instruction sheet. Agreement there establishes that two instruments track each other on the same construct, not that the construct names a real thing. Those clinicians agree with each other at an intraclass correlation of 0.31, so the human criterion the model is validated against is itself unstable. The words construct validity, measurement, and psychometric do not appear in the article or in its supplementary material, including the table that holds the definitions of the thirteen dimensions.

That is the distinction this essay is asking for. Validating the instrument is not validating the target. A rubric can be reliable, reproducible, stable under paraphrase and agreed on by clinicians, and still be scoring something the field has not settled. A procurement handbook written by operators and a framework written by computational psychiatrists arrive at the same omission from opposite directions, which is reasonable evidence that the omission belongs to the field rather than to either set of authors.

Why validity has to hold in mental health

The point applies in every AI domain. It applies more sharply in mental health because every claim a mental-health AI system makes is a claim about a construct that cannot be directly observed. Depression is not measurable the way blood glucose is measurable. Suicide risk is not measurable the way blood pressure is measurable. Wellness has no laboratory test. Every mental-health AI tool that claims to detect, predict, score, flag, or assess one of these constructs is making a measurement claim about something inherently latent. The measurement is an inference from a proxy. The validity question — whether the proxy represents the construct — is upstream of every other governance question in the rubric.

The field has been candid about this for fifteen years. In 2010, the National Institute of Mental Health published its Research Domain Criteria (RDoC) framework with a one-sentence diagnosis of why a new framework was needed: the weakness of the Diagnostic and Statistical Manual of Mental Disorders (DSM) is its lack of validity. The standard mental-health diagnoses, unlike definitions of ischemic heart disease or lymphoma or AIDS, are based on consensus about clusters of clinical symptoms, not an objective laboratory measure.

Twelve years later, Thomas Insel, who led NIMH during the period he is critiquing, wrote in Healing that the research had largely failed to help Americans struggling with mental illness despite billions of dollars directed into it. The Lancet Commission on Global Mental Health and Sustainable Development in 2018 reframed the field at a global-public-good scale, with a projected sixteen-trillion-dollar cost by 2030 measured against a construct whose foundations remained partly underdetermined. The WHO has made the same point.

None of these are fringe positions. They are the field’s most authoritative voices saying, on the record, that the underlying construct on which mental health AI tools now make their measurement claims is in active contestation.

Not for lack of response. RDoC has generated fifteen years of dimensional-construct research, HiTOP has proposed an empirically-driven alternative taxonomy, DSM-5-TR has folded in dimensional criteria alongside categorical, and precision psychiatry and digital phenotyping are actively building next-generation measurement infrastructure. Those are extensions of the measurement framework around the constructs. They are not resolutions of the constructs themselves, or of construct drift as those constructs are applied — the level at which consumer-scale AI mental-health tools now make their claims.

That is the condition on which AI is being deployed. The proxy that an AI system has learned to associate with a mental-health outcome may be the proxy the field has been using as a stand-in for a construct it has not yet fully adjudicated. Responsible-AI rubrics that screen for privacy, accountability, equity, monitoring, and transparency do not catch this condition. They presuppose it has already been resolved. In most cases it has not.

The cost calculus is the second reason. Validity is owed in every domain. What changes across domains is how much evidence you should require before acting on the number, and mental health is where that requirement runs highest. The cost of measuring the wrong thing in administrative AI shows up as a wrong scholarship decision, a wrong advising flag, a wrong staffing recommendation. The cost of measuring the wrong thing in mental health AI shows up as a wrong crisis-detection call for an adolescent who needed it, a missed signal in someone with intent, a chatbot response that reinforced isolation when connection would have helped, an AI companion that taught a vulnerable user a pattern of relating that diminished rather than supported their self-efficacy and locus of control.

The most vulnerable users — adolescents, people in active crisis, those without access to traditional clinical support — are the users whose construct presentations are most heterogeneous and for whom the validity case is thinnest. They bear the cost of the validity gap. Adolescents, at least, are already there in numbers. In a nationally representative survey of US adolescents and young adults conducted in November 2025, 19.2 percent reported using an AI chatbot for mental-health advice; the same team’s survey a year earlier, on slightly different wording, found 13.1 percent. Of those who used one, 42.8 percent did so at least monthly and 91.7 percent rated the advice somewhat or very helpful. Girls were about twice as likely as boys to have used one at all (adjusted odds ratio 2.10), and among users, Black respondents were 5.45 times more likely than White respondents to be monthly users, though on an interval wide enough to hold almost anything, 1.44 to 20.66. Most of all: 63.3 percent had told no one (McBain et al., JAMA Pediatrics, June 2026).

One finding in that survey runs against the intuition. Use was higher, not lower, among young people who had spoken with a physician about their mental health in the previous six months — an adjusted odds ratio of 1.89. Chatbot use is not only filling a gap where care is absent. For some of these young people it is running alongside care, and the survey did not ask whether the clinician was told.

The canonical mechanism is documented. Obermeyer and colleagues in Science in 2019 examined a risk-prediction algorithm applied to millions of patients and found it was not predicting illness. It was predicting next year’s healthcare spending and using that as a stand-in for who was sick. Spending tracks the care a person receives, not the illness they have, and less is spent on Black patients at the same level of illness. So two equally sick patients generated different predicted costs, and the algorithm read the cheaper one as healthier. Changing the prediction target from cost to illness would have raised the share of Black patients in the highest-risk group from 14.1 percent to 26.7 percent. The predictors performed about as well as each other on every outcome tested, including each other’s, so the choice of label moved who got flagged far more than it moved accuracy. None of this is visible to a fairness audit that checks whether performance holds at a given risk score, because performance did hold at a given risk score. The score was the wrong quantity.

The mental-health version runs on a different proxy in the same way. Engagement is what most consumer tools have to hand: sessions, messages, retention. A user who comes back monthly, rates the advice helpful, and has told nobody they are using it scores well on every one of those measures. Nothing in the measures separates that user from one who has substituted the tool for telling someone. The survey’s own authors make the narrow version of this point, cautioning that perceived helpfulness may be picking up a chatbot’s tendency toward sycophancy and flattery rather than the quality of the advice. Obermeyer’s algorithm could not distinguish low spending from low need, because spending measures contact with a system rather than need for it. An engagement metric cannot distinguish heavy use from benefit, for the same reason. And here the misreading has no outside check, because 63.3 percent of these users have told no one at all.

Four questions the rubric depends on

So what does the responsible-AI rubric need underneath it. What is needed is not another principle or another commitment added to the list of six or eight or ten, but a prior layer, asked first, before any of the rubric’s other questions can be evaluated. Four questions constitute the validity layer that the rubric depends on.

A vertical stack of four numbered cards, Q1 through Q4, listing the validity questions the responsible-AI floor depends on: name the construct; produce the validation evidence and say what use it was built for; name where the case is thinnest and what policy applies there, with a decision owner and a start date; specify a drift-monitoring cadence and what it triggers. A footer notes that three of the four are questions about use rather than about the instrument, citing Kane 1992 and Messick 1995.
The four validity questions the responsible-AI floor depends on.

What construct is the system claiming to measure. Name it explicitly — depression, suicide risk, anxiety, crisis, engagement, well-being. Whatever the system claims, the institution adopting it has to be able to write down, in one sentence, what construct the system is producing scores or flags about. If the answer is fuzzy, the rubric cannot help. The construct has to be named before anything else can be evaluated against it.

What is the validation evidence and what it covers. Sample sizes, populations, instruments, comparators, time horizons. Published or proprietary. Independent or vendor-conducted. The institution has to be able to read the validation case the same way it reads any other evidence-based health-tech procurement: with the question of what the evidence says the system measures, what populations the evidence covers, and what use cases the evidence was built for. The rubric assumes this case exists. Often it does not exist at the resolution the deployment requires.

Where is the validation case thinnest, and what does the institution do about that. Every validation case has gaps. The validity layer requires the institution to name the gaps before the deployment, not after a regulator asks. Which populations is the system under-validated against. Which use cases does the evidence not yet cover. What is the institution’s policy for deployment in the under-validated zones. Disclosure to users in those zones. Restricted use. Human-in-the-loop. The answers vary; the question is the same. Naming the thin parts of the validity case before deployment is what distinguishes responsible deployment from responsive deployment.

Matthew Krome, who co-founded Trovane and built the safety architecture underneath its products, pushed on this question when he read a draft of this essay. His point is that naming a thin spot does not reliably produce a policy. Sometimes it sits open as an unresolved decision, waiting on a named person with the authority to make the call. He is right, and his own handbook is the evidence. Its author-disclosure appendix names continuous drift monitoring with named owners as documented but not yet fully automated, and no mitigation is attached to that gap. It is simply open. The interval between naming a gap and deciding what to do about it deserves to be measured on its own rather than treated as delay before an answer arrives. So the third question needs a fifth element: name the decision owner, and name the date the gap was opened. One caution, since this is now a measure and measures invite gaming. The clock has to start when the gap became knowable, at the model change or the population shift, not when someone got around to writing it down. An institution that starts the clock at disclosure can improve its record by naming things late.

What is the monitoring cadence for drift, and not just the kinds of drift the modern responsible-AI literature usually names? Construct drift, where what the field calls “depression” or “wellbeing” or “engagement” itself shifts relative to the validated proxy. Population drift, where the people using the system shift in composition. User-effect drift, the longitudinal-validity dimension the recent discourse has been naming — whether the system’s effect on users holds up over weeks, months, and years of repeated interaction, whether the cognitive offloading or excessive reassurance-seeking or shifts in self-efficacy that point-in-time evaluation cannot see are being measured at all.

Longitudinal experimental evidence from Oxford, Stanford, and the UK AI Security Institute (Ibrahim et al., 2026) documents the pattern under controlled conditions. Across five preregistered studies with 3,075 participants, including a three-week census-representative arm, sycophantic AI delivered the in-the-moment support people associate with close relationships, while the people receiving it reported lower satisfaction with their real-world social interactions than people given a neutral AI, and showed none of the downstream benefits — intellectual humility, feeling understood by other people — that human support provides. Two cautions the paper is clear about and a reader should carry with the finding. The satisfaction gap was small, 5.51 against 5.70 on a seven-point scale, on a preregistered exploratory measure. And the absent benefits are absent as null results rather than as measured declines. What the authors describe is a perceptual shift rather than a behavioural one; time spent with other people did not move. So this is the third drift mode observed in a laboratory over three weeks rather than in deployment over years, which is the argument rather than an objection to it. Nobody is monitoring for it in the wild, so nobody knows what it does there. The validity layer requires that monitoring address all three drift modes and that the institution name what it does when any of them surface.

The Trovane handbook names three drift modes too, and they are a different three. Model drift is where the underlying model gets updated and the system’s responses shift, including how it handles distress. Conversational drift is where small prompt and content changes accumulate until the tone and scope of the tool no longer match its clinical design. Detection-accuracy regression is where performance degrades for one user group while the aggregate numbers still look fine. Those three watch the instrument. The three above watch the claim the number is making: what it means, who it applies to, and what it does to the people it is applied to. Between them there are six modes in two families, and the responsible-AI literature currently covers one family well and the other not at all. A monitoring regime that runs the first three and none of the second three will catch a model update and miss the fifteen-year argument the field is still having about what the model was built to detect.

Six drift modes in two families. The instrument family — model drift, conversational drift, detection-accuracy regression — is what responsible-AI monitoring already covers. The claim family — construct drift, population drift, user-effect drift — is what the validity layer adds.
Six drift modes, two families. Responsible-AI monitoring covers the first family. The second is what the validity layer adds.

These four questions sit underneath the responsible-AI rubric. They do not replace the rubric; they earn it. A mental-health AI tool whose adopter can answer all four — names the construct, knows the validation evidence and its gaps, has a plan for the under-validated zones, has a monitoring cadence for all three drift modes — is one a clinician can stand behind, a foundation can fund, a regulator can audit. A tool whose adopter cannot answer them is being deployed against a construct claim no one has examined closely enough to defend, regardless of what the responsible-AI rubric says about privacy, accountability, or monitoring of model behavior. The other questions rest on a foundation that has not been checked.

Calibration, not exemption

A real objection from technology leaders is that requiring all four validity questions before any deployment paralyzes adoption. Eduardo Bunge frames the strongest version of that case: roughly 80 percent of the 1.2 billion people worldwide with a mental disorder never receive treatment, and even the field’s best-evidenced conditions cap remission near 51 percent (Parente Health, July 2026). The cost of not adopting is real. The objection has weight, and the answer is calibration, not exemption. Some deployments are low-stakes enough that the validity layer can be carried by post-deployment monitoring rather than pre-deployment certainty. A wellness-tracking feature with a corrective off-ramp is not in the same risk class as a crisis-detection model that escalates to a clinician. Others, including clinical decision support, child safety, crisis detection, and anything that triggers action without a human in the loop, are not negotiable. The discipline is to decide which deployment is which before the deployment rather than after a failure. The four questions do not block all deployment. They ask the institution to know, for each one, which deployments it is making with the validity work done and which it is making on a promise to do that work in flight.

This is offered as the prior discipline the responsible-AI rubrics assume, and as recognition that the rubrics are doing the right work at the right altitude. The Trovane handbook’s thirteen questions are the right responsible-AI floor. So are the NAM Code of Conduct’s six commitments, the APA’s ethical guidance, the WHO workshop’s three recommendations, the Polaris Principles, the VERA-MH evaluation work, the ASL-MH framework, and the Nature Medicine auditing framework, which shows how carefully an evaluation instrument can be checked before anyone relies on its scores. All of these are necessary. None of them is sufficient if the validity question is not answered first.

The instrument itself is fragile

The VERA-MH benchmark itself has already shown what validity fragility looks like in practice. Spring Health, which released VERA-MH in October 2025, quietly switched its recommended judge from GPT-4o to GPT-5.4 in June 2026 — as Bryan Jester documented in July 2026. The switch happened in a repository update without a peer-reviewed validation of the new judge as more clinician-aligned than the original. Under the new judge, Jester reports, frontier general-purpose models that had been scoring in the low-to-mid 60s dropped into the low-50s on the same conversations. Those figures are from his own re-run rather than a published benchmark table, and I have not reproduced them. What does not depend on the size of the drop is that the “safety” every public number was reported against had, six weeks earlier, quietly become a different measure. That is the validity gap operating in real time: the instrument changed, and what had appeared to be measured no longer was. It illustrates the concern this essay names — not because the judge switch was wrong, but because the framework being adopted at the responsible-AI floor depends on a measurement instrument that can change without notice and without independent validation.

There is an architectural answer to that specific failure, and it does not require resolving the construct question first. The rule is that model output alone never triggers a decision. Every escalation has to carry at least one hard, checkable signal behind it, and the response at the moment of disclosure is written in advance by a clinician rather than generated on the spot. Trovane publishes this as the governance layer underneath its platform, which is the clearest statement of the pattern I have found, though the account is the company’s own and I have not seen it independently audited. The construct question stays open under that design. What closes is the path by which an unvalidated model becomes the scorer without notice, and that is worth naming, because it is a thing an institution can write into a contract now rather than waiting for the field to settle what depression is.

Some validity work is starting to happen. In February 2026, Spring Health published a human-validation study of four VERA-MH judges — Claude Sonnet 4.5, GPT-4o, GPT-5.2, and Gemini 2.5 Flash — tested against licensed clinicians. The clinician-to-clinician inter-rater reliability baseline was 0.77; the four judges landed at 0.82, 0.81, 0.78, and 0.77, on confidence intervals overlapping each other and the clinician baseline. That is the exact discipline this essay argues has to sit under the responsible-AI framework: check whether the measurement instrument agrees with the human standard it claims to approximate, on the population it will judge. It is doable but not yet done systematically.

And the judge that replaced these four in the recommended slot, GPT-5.4, has not been part of a human-validation study of comparable rigor. Spring Health’s VERA-MH GitHub README claims an internal inter-rater analysis supporting the switch, but the data and methodology are not publicly available for independent review.

The discipline already exists

The discipline that would answer it already exists. Measurement validity is the oldest discipline in measurement science. It is the discipline higher education has been working with for a century in assessment, and the discipline ETS has been engaging with for decades in automated scoring. That lineage is mapped in Before it was called AI evaluation across eight bridges from that literature to modern AI eval. It is also the discipline the FDA applies through its post-market surveillance frameworks for diagnostics. The field of mental health has been engaged with construct-validity questions at the underlying-construct level since the RDoC initiative was launched by NIMH in 2010, explicitly because the DSM’s weakness was its lack of validity — its diagnostic categories did not map cleanly to the underlying biology and behavior they claimed to describe. The discipline the responsible-AI rubrics need underneath them is not new. The task is to apply it to the specific systems now being deployed at consumer scale against populations whose mental health was already underserved, and to name validity as the central question the rubrics inherit at the point of deployment.

The institutions that handle this moment well will be the ones that ask the validity question before the responsible-AI question, for the users whose mental health depends on the answer.

Written August 2026 for the Analytic Bytes Library as the third piece in the measurement-validity arc, following What is this system actually measuring? (higher education) and The valid dollar (impact investing). The rubric material is drawn from Trovane’s handbook for evaluating AI mental-health tools (v1.2, 2026), including its author-disclosure appendix, and from the published description of the governance layer underneath the Lumara platform, which is the company’s own account and is cited as such. Matthew Krome, co-founder of Trovane, read a draft and contributed the governance-latency argument in the third question; the clock-start caution attached to it is mine. Trovane was known as Mental Health Chat until 13 August 2026; the handbook, the preview post, and the platform described here are the same work under the current name. Krome read the final draft and confirmed he is comfortable being named.

Four studies carry an argument here about how a number was produced: McBain et al., JAMA Pediatrics, June 2026 (DOI 10.1001/jamapediatrics.2026.2015), a RAND-led nationally representative survey; Ibrahim et al., 2026 (arXiv 2605.07912); Obermeyer et al., Science, 2019; and Weilnhammer et al., Nature Medicine, 7 August 2026. The VERA-MH human-validation study is cited as the February 2026 preprint, arXiv 2602.05088, which is the version this essay worked from and the version the Jester analysis responds to. The Bryan Jester analysis is a single secondary account and is attributed as such in the text rather than treated as an independent result. Other citations — Parente Health (July 2026), the NIMH RDoC framework (2010), the APA, NAM, WHO, Polaris, and ASL-MH documents — are used for their stated positions rather than for how a number was produced.

Since first publishing on 14 August 2026, four substantial revisions have been made and are named here. On Nature Medicine, the original said none of the paper’s checks validates the target, which was too strong — the known-groups test, where the judge separated high- from low-risk conversations at a median AUC of 0.98, is target-side evidence, though the Methods say those conversations were generated by a separate model, Grok-4, prompted with the judge’s own dimensions, so the stimulus was written from the rubric and then scored against it; separately, the essay said the clinicians in the criterion check applied the judge’s rubric, but they did not — they rated each response on a single 1-to-10 slider from safe and appropriate to clearly harmful. On the youth-use figures, the demographic paragraph opened with “roughly one in eight US adolescents,” which is the same team’s 2024 figure on different survey wording; the 2025 figure this paper reports is 19.2 percent, the Black-versus-White odds ratio of 5.45 ran without its interval (1.44 to 20.66) and without its condition (it holds among young people who use a chatbot, not among all of them), the girls-twice-as-likely ratio belongs to any use rather than monthly use, and use was higher, not lower, among young people who had recently spoken with a physician — two paragraphs were rebuilt around that finding. On Obermeyer, the essay said changing the prediction target from cost to illness would have raised the share of Black patients flagged from 17.7 to 46.5 percent, but those figures belong to a different operation in that paper — a counterfactual patient swap that requires already knowing which patients are sick; the label change itself is Table 2, and it runs from 14.1 to 26.7 percent. On the six drift modes, the second family was originally called the target family and is now the claim family — target still names a property of the measurement, so both families ended up describing the instrument, while claim moves the second family out to interpretation and use, which is where Messick locates validity and where those three modes actually do their work; the first version of the figure circulated on LinkedIn under the earlier label. Smaller fixes to framing and sourcing are kept in the working file.

Analytic Bytes
From fragmented to decision-ready.

Questions, pushback, or a problem that looks like this one? Write to chai@analyticbytes.systems.