The missing discipline in responsible AI discourse — brought to the practice-operating scale.
The AI governance conversation talks a lot about principles (fairness, transparency, accountability, safety), controls (documentation, oversight, human-in-the-loop / AI-in-the-loop (HITL/AITL — see AB audit for the signal-informed proxy formulation of the load-carrier axis)), and risk categories. What's largely missing: measurement-validity discipline. Do the controls measure what they claim to? Is what's being audited well-defined? Does the evidence support the conclusion drawn?
Kane's 4-inference validity framework — mature in educational / psychological testing since the 1990s — is basically absent from responsible AI frameworks (NIST AI RMF, EU AI Act, ISO/IEC 42001). AB brings that tradition to AI-native operations at a scope existing frameworks don't traditionally address: solo-practitioner and boutique consulting scale.
The proof-of-concept: the AB 90-day audit — and the record of how it audited itself in real-time.
Not a checklist. A record of the discipline moves that surfaced during one audit's execution — visible, dated, available for peer review.
The audit did to itself what it audits others for. That IS the demonstration.