AB Governance Maturity Scorecard
Six dimensions of self-observation discipline. Today → Phase 2 (Q4 '26) → Phase 3 (Q1 '27). Testable, updated each 90-day audit.
v1.1 · Phase 1 baseline · scored Jul 31, 2026 · next refresh Nov 1, 2026 · source: ab_90_day_audit_apr_jul_2026.docx §7
▶ Operator hypothesis
This audit is expected to score
weakest on governance right now. Governance/measurement discipline is the USP, but descriptive instrumentation is thin. Reflection cadence just formalized. Dispatch prescriptive-not-descriptive. Scorecards are recent. Phase 2 is designed to move the scores. Failure to mature is data.
Today · Phase 1
MODERATE
USP anchors H · instrumentation LM
Phase 2 target · Q4 '26
MOD–HIGH
Tests whether cadence held
Phase 3 target · Q1 '27
HIGH
Client-transferable
Score:
HHIGH
MHMOD-HIGH
MMODERATE
LMLOW-MOD
LLOW
Reflection cadence discipline
Do monthly + quarterly cycles actually hold?
LM
→
MH
→
H
Just formed
Now Just formalized Jul 2026. Machinery exists (scheduled tasks); no full quarterly cycle held yet. P2 3 monthly + 1 quarterly cycle held Aug–Oct. P3 2 quarterly + 4+ monthly cycles complete.
Scorecard maturity
PDDS + Decision + DM scorecards at published quality
M
→
MH
→
H
Priority build
Now PDDS Governance Scorecard published (external, NIST-aligned). AB Decision Scorecards + DM Dashboard exist but not yet at published quality. P2 All 3 at scorecard visual style + PDDS-level construct discipline. P3 Client-transferable, deployable in engagements.
Validity spine
Kane 4-inference framework · NIST AI RMF · reliability + construct discipline
H
→
H
→
H
USP anchor
Now Demonstrated in PDDS eval/VALIDATION.md · Decision Validity essay · classification rulebook. Not aspirational. P2 & P3 Held.
Instrumentation depth
Transcript persistence · dispatch logging · token measurement
LM
→
MH
→
H
On track
Now Corrected Jul 31: extract v2 now 30/93 sessions (was 23). Session-start dispatch logging not yet in place. P2 Transcript path verified durable · dispatch log accruing · monthly re-extract cadence held. P3 All measurable dimensions have descriptive data.
Descriptive vs prescriptive
Do we report what actually happened, or only what we recommend should happen?
L
→
M
→
H
Priority build
Now Dispatch rulebook says what SHOULD happen; audit doesn't yet report what DID. Same for reflection cadence. P2 At least one dimension where actual (measured) is reported alongside prescribed. Dispatch is the first candidate. P3 Every prescriptive claim has a descriptive counterpart.
Semantic modeling
Categories, load-carrier axis, delegatable %, dispatch guidance — defensible definitions
H
→
H
→
H
USP anchor
Now Classification rulebook v2 (11 categories; Practice Ops kept as a query tag, not a claimed lane), reflex-tested boundary calls. Not vibes. P2 Held; boundary calls re-audited quarterly. P3 Kit-ready — rulebook transferable to client audits.