AB Governance Craft Log

How the AB 90-day audit audited itself in real-time.

v1.0 · Jul 30-31, 2026 close · sourced from ab_90_day_audit_apr_jul_2026.docx §9 · refreshed each 90-day audit cycle

▶ THE POSITIONING

Measurement-validity discipline for AI governance — the missing discipline in responsible AI discourse.

The AI governance conversation focuses on principles (fairness, transparency, accountability), controls (documentation, oversight, human-in-the-loop / AI-in-the-loop (HITL/AITL — see AB audit for the signal-informed proxy formulation of the load-carrier axis)), and risk categories. What's largely missing: the measurement-validity discipline that asks whether the controls measure what they claim to. Kane's 4-inference framework — mature in educational / psychological testing since the 1990s — is basically absent from responsible AI frameworks (NIST AI RMF, EU AI Act, ISO/IEC 42001). This log brings that tradition to AI-native operations at a scope existing frameworks don't traditionally address: solo-practitioner and boutique consulting scale.

▶ TOP-LINE FINDING
23 discipline moves + 7 emergent principles across 5 phases. Methodology documentation done inline, not deferred to future work.
The audit's own USP is measurement/governance discipline. This log is how that discipline held under stress during the audit's own execution — visible, dated, available for peer review. Section 7's honesty note said audit-of-audit was Phase 3+ future work; this log revises that. Done, not aspirational.

Executive summary

What this is. The methodology narrative of the AB 90-day audit's own execution — not the findings the audit produced, but the discipline moves the audit applied to ITSELF during the Jul 30-31 close. Verification, definition-sharpening, instrument-questioning, scorer-correction. 23 moves catalogued, 7 emergent principles named.

The Kane 4-inference frame applied to the audit itself. Kane's validity framework (Scoring → Generalization → Extrapolation → Implication) is what AB claims to bring to client work. This log shows the same discipline running on the audit's own execution. Each of the 5 phases maps to one inference layer — with the audit's mid-flight corrections as the empirical evidence.

Why it matters. Traditional governance craft lived in the practitioner's head; only the finished memo surfaced. AI-collaborative work captures the scaffolding by default — transcripts, corrections logs, mid-flight amendments. This log makes that scaffolding observable: the demonstration of methodology-documentation capability, not the aspiration.

Who this is for. Executive evaluators (hiring committees for governance/CDO roles, BD partners assessing engagement fit): scan the top of this log for the story shape. Deep peer reviewers (methodology practitioners, audit specialists): expand the move cards below to interrogate specific discipline choices.

The 5 phases at a glance

Each phase corresponds to one aspect of methodological discipline. Click a chip to jump to that phase.

Phase A
Definitional discipline
Definitions, categories, boundary calls, weights — the vocabulary the audit runs on
9 moves
Phase B
Corpus & instrumentation
Data provenance, tooling verification, operator-scorer alignment — the ground the audit measures from
5 moves
Phase C
Data refresh
Currency discipline, freshness cycles, and findings that emerged from the refresh itself
3 moves
Phase D
Quality standard
Visual scanability, rigor bar, defensibility — the presentation-and-craft discipline
3 moves
Phase E
Dissemination
Public safety, narrative correction, library-compatibility — the discipline of shipping honestly
3 moves

Phase A · Definitional discipline (9 moves)

The vocabulary the audit runs on. Every quantitative claim depends on prior definitional choices — what counts as a session, a category, a hub, a weight. When definitions are loose, downstream findings are un-defensible.

A-1"What IS a session?" formalized
Triggered by
Operator challenge to loose usage of "session" in early conversation.
Discipline surfaced
Definitional precision at the unit-of-analysis layer. If the atomic unit isn't defined, no downstream counts are meaningful.
Resolution
Formalized: "One Claude session = one row. A session is a named conversation, whether spawned manually, by scheduled task, or by sub-agent. Each session gets exactly one category — no session is split."
Teaching
Definitions of atomic units must be codified BEFORE quantitative work. Retrofit is possible but expensive.
A-2Session ≈ work-package heuristic with hub exceptions
Triggered by
Operator pushback on whether one session always equals one work package.
Discipline surfaced
Recognition that a heuristic isn't universal law. Formalizing the exception makes the heuristic defensible.
Resolution
Formalized: "Most sessions follow one new session per new research question / idea / exploration. Session count ≈ work-package count under that principle. Hub sessions violate this principle."
Teaching
Every heuristic needs its exception class named. "Usually true" without "except when" is un-auditable.
A-3Hub session taxonomy (single-category vs cross-category)
Triggered by
Operator surfaced that some sessions host many work packages (Job Apps hub, speaking hub, Gather publications).
Discipline surfaced
Taxonomic specificity. Not all hubs are the same shape; different exception patterns have different downstream implications.
Resolution
Formalized: single-category hubs (many packages, all one category — under-count work volume but don't distort categorical distribution) vs cross-category hubs (many packages across categories — distort BOTH count and distribution). Gather flagged as the only current cross-category hub.
Teaching
Exception patterns often have sub-taxonomies. Name them explicitly or the taxonomy leaks.
A-4Cross-category hub attribution convention
Triggered by
Rulebook needed a rule for how to assign a hub session's category when the hub crosses multiple categories.
Discipline surfaced
Convention-setting discipline. When natural attribution is ambiguous, define the convention and disclose it.
Resolution
Formalized: "Assign to the dominant intent category as of formation rather than re-classifying continuously. Note the hub role explicitly so readers can mentally re-weight adjacent categories upward." Also: "Library and Brand are structurally under-counted because much of that work lives inside the Gather cross-category hub."
Teaching
When attribution is judgment, name the judgment and its cost. The cost becomes a caveat, not a hidden distortion.
A-5PDDS thread confirmation
Triggered by
Scorer uncertainty about which session was the PDDS thread. Operator confirmed: "Public data charts examples" IS the PDDS thread.
Discipline surfaced
Scorer-operator alignment on referential specifics. Don't assume; confirm.
Resolution
Confirmed and locked. Related later correction: Philip Steigman panel design work moved to the PDDS thread — the two threads are related but distinct.
Teaching
Scorer working from labels alone will misidentify referents. Operator confirmation on referential specifics is a one-time cheap discipline.
A-6Session-by-session classification pass with operator-reviewed boundary-call resolution
Triggered by
Operator asked for transparent classification of all 92 sessions before quantitative analysis proceeded.
Discipline surfaced
Show-your-work per unit of analysis. Making the session→category mapping public and challengeable BEFORE aggregation.
Resolution
Produced session_classification_check.md (draft with ⚠ boundary calls flagged: 7 ambiguous sessions) → operator adjudication → session_classification_v2_resolved.docx (all 92 mapped, all boundary calls resolved).
Teaching
Quantitative analysis loses credibility if the underlying classification is opaque. Always publish the mapping. Cost is small; credibility payoff is large.
A-7Category 11 (Client Engagement) added mid-audit
Triggered by
Two sessions (Graham Windham one-pager, Philip Steigman) didn't fit existing 10 categories cleanly.
Discipline surfaced
Rulebook extension discipline. New categories added when the existing taxonomy fails to accommodate real data — with clear distinction from adjacent categories.
Resolution
Category 11 added: "Client Engagement — soft-prospect / advisory work for a specific client / partner org. Distinct from Job Apps (which targets a role for you inside an org): Client Engagement targets a contract / engagement / advisory relationship between AB and an org."
Teaching
Taxonomy fixes should preserve boundary discipline, not blur it. When adding, be explicit about what makes the new category distinct from adjacent ones.
A-8Foundation Layer / Category 12 (AI-Native PracticeOps) as a foundation layer, not a surface category
Triggered by
Editorial standards, voice specs, context onboarding kept surfacing as "invisible work" the taxonomy didn't capture.
Discipline surfaced
Foundation-vs-category distinction. Not everything the audit measures belongs on the same taxonomic layer. Some things are foundation layers, not surface entries.
Resolution
Coined AI-Native PracticeOps as a Foundation Layer / Category 12. Not a flat category alongside the 11 surface categories — the underlying layer. Parallel to DevOps in software engineering.
Teaching
New concepts sometimes need new taxonomic layers, not new same-layer entries. Force-fitting to existing structure hides the point.
Later revision (Aug 2026)
Foundation Layer was retired as a claimed library arc after further review. Practice Ops now lives as a Knowledge Repo query tag only, not as a distinct discipline AB Library is developing. The DevOps parallel is preserved as a descriptive frame in the Auditing field note, without the library-lane commitment. Historical move preserved here for methodology-record honesty.
A-9Weights table definition + provenance disclosure gap
Triggered by
Audit's task allocation view used weighted-proxy volume shares. Weights were mechanical (per-category × session count). Operator asked what the weights actually were.
Discipline surfaced
Provenance discipline for underlying parameters. Weighted methods need weight origins disclosed — literature, SME panel, or empirical calibration.
Resolution
Weights table printed (Strategic 5.0, Automation 0.36, Web/Tech 2.67, etc.). Provenance surfaced honestly: weights are practitioner judgment, NOT literature, NOT SME panel, NOT empirical calibration. Defensible for self-audit; NOT defensible for external audit kit product as-is.
Teaching
When method parameters are judgment-based, that has to be disclosed at the moment the method is used. Undisclosed judgment inside quantitative methods is the most defensible-looking form of hidden bias.

Phase B · Corpus & instrumentation discipline (5 moves)

The ground the audit measures from. Corpus scope, tooling verification, scorer-operator alignment — what makes the audit's evidence base defensible before any interpretation happens.

B-1Corpus scope verification via Corpus Sources table
Triggered by
Operator challenge: "Are we still updating the corpus and rerunning all the data analysis for the artifacts to match the window of May 1 → July 31?"
Discipline surfaced
Provenance discipline at the corpus level. Every artifact must map to a specific source + date window, disclosed transparently.
Resolution
Corpus Sources table added to Method section: 6 rows, one per artifact, each with corpus + source + date window + notes/gaps. Made corpus-artifact ambiguity impossible to hide.
Teaching
"Which data underlies which artifact?" should be answerable by pointing at one table. If it takes prose to explain, the audit is opaque.
B-2Time-window discipline — operating window vs audit-run day
Triggered by
Operator caught the conflation of operating window (May 1 → Jul 31) with audit-run day.
Discipline surfaced
Temporal specificity. Different types of analysis have different appropriate closing dates. Confusing them creates false precision.
Resolution
Formalized: operating window = May 1 → Jul 30, 2026. Audit-run day = Jul 30-31. Time-series artifacts share operating-window closing date. Snapshot artifacts carry their own stamps.
Teaching
Time semantics matter. "When does the window close" and "when is the audit run" are distinct questions and should be answered distinctly.
B-3Instrument tooling verification — the extract-path detective sweep
Triggered by
Operator challenge: "Are you sure you looked in Documents/Claude for transcripts? Are the transcripts really purged?"
Discipline surfaced
Verify-the-instrument discipline. Before accepting a data pipeline's output, verify the pipeline is pointing at the right source.
Resolution
Detective sweep via read_transcript MCP disproved the "macOS purged transcripts" story. Root cause: extract_token_usage.py globbed only /var/folders/*/T/claude-hostloop-plugins/, missing the persistent ~/Library/Application Support/Claude/local-agent-mode-sessions/ location. v2 script fix landed same day, re-extract recovered 30 sessions / 23.23B tokens (was 23 / 9.37B).
Teaching
"The instrument said X" is not the same as "X is true." Verify the instrument before defending its output. The audit almost shipped a wrong finding because I trusted the pipeline's silence.
B-4Scorer boundary corrections (Princeton, Philip Steigman) → Rule 1
Triggered by
Two operator corrections mid-close: Princeton mis-labeled as Client Engagement (was Job Apps); Philip Steigman mis-labeled as Website & Tech (was Client Engagement).
Discipline surfaced
Pattern-detection across scorer errors. Two errors in 24 hours surfaced a common failure mode: scorer inferring category from visible last-turn output rather than originating intent.
Resolution
Rule 1 added to classification rulebook: "Outputs migrate; intent doesn't. Categorize the originating session by its originating intent, not by its residual output. Practical test: ask 'what was this session started to do?' before 'what does its last turn contain?'"
Teaching
Scorer errors often have common root causes. When operator catches two of the same pattern, it's a rulebook update, not two isolated fixes.
B-5Definitional discipline as operator-scorer alignment discipline
Triggered by
Operator observation surfaced during weights conversation: "definitional discipline is becoming key for you and me to align."
Discipline surfaced
Meta-level naming of a discipline the audit had been practicing implicitly. Ambiguity in definitions drives most scorer errors — requires operator-scorer co-definition BEFORE analysis, then mutual accountability when either drifts.
Resolution
Recognized as a methodology principle in its own right, not a side-effect. Named as a Phase B move because it's about the underlying vocabulary the audit runs on.
Teaching
The most defensible-looking quantitative work depends on operator-scorer alignment about what the quantities MEAN. Skip this step and the numbers are technically-correct but strategically-empty.

Phase C · Data refresh + analysis-derived findings (3 moves)

Currency discipline (making sure data matches the window claim) and the discipline of naming findings that emerge from the refresh itself.

C-1Data freshness — DM curve rescore
Triggered by
Operator observation: "any over-time analysis should close at the same date. The DM curve stops scoring Claude threads at May 14 but the audit window closes Jul 30."
Discipline surfaced
Freshness discipline. If the window is Jul 30, the data supporting time-series claims must extend to Jul 30. Freshness gaps must be either closed or explicitly disclosed.
Resolution
14 Claude threads from May 15 → Jul 30 scored on 6-dim rubric. DM corpus: 102 → 116 threads. Curve extends to Jul 30. Pattern held.
Teaching
Freshness gaps are the most under-discussed audit risk. "We just haven't updated it" isn't a defense — it's the finding.
C-2Metadata-only vs full data refresh discipline
Triggered by
Scorer initially did a metadata-only pass on task-allocation-view (updated stats + framing, kept old chart data). Operator pushed back: "full data currency required."
Discipline surfaced
Refresh depth discipline. "Refreshed" can mean many things. Partial refresh creates the illusion of currency without the substance.
Resolution
Task-allocation-view fully rebuilt: session-count shares recomputed across 93 sessions, token shares recomputed across 30 measured / 23.23B tokens, chart axes rescaled.
Teaching
Partial refresh is worse than no refresh — it creates false currency. When time is tight, either fully refresh or explicitly stamp "partial refresh, X pending."
C-3Analysis-derived finding surfaced during correction (Finding 4.6 · Automation optimization)
Triggered by
During Section 4 correction (Automation flipped from 1.4% to 57.4% of tokens), operator surfaced: "one implication of this audit is how to bring down token usage for recurring automation."
Discipline surfaced
Discipline of naming findings that emerge from the refresh. When corrected data changes the story, interrogate whether new findings surface.
Resolution
Finding 4.6 added: "Automation is 57% of measured tokens — a win AND a Phase 2 cost-optimization opportunity." Six concrete optimization levers named.
Teaching
Findings can emerge from methodology corrections, not just from initial analysis. The audit's own correction cycle is a source of new evidence.

Phase D · Quality standard discipline (3 moves)

Visual scanability, rigor bar, defensibility. Presentation-and-craft discipline that determines whether the audit's findings are legible to peer review at all.

D-1Scorecard density critique + reframe: defensibility ≠ visual density
Triggered by
Operator critique: "the AB Governance Maturity Scorecard, PDDS Governance Scorecard, and AB Decision Scorecards are dense and boring, except the Speaking Opportunity Scorecard which is engaging."
Discipline surfaced
Reframe discipline. When a quality standard is criticized, ask whether the underlying assumption is wrong before defending the artifact.
Resolution
Reframe adopted: DEFENSIBILITY ≠ VISUAL DENSITY. Kane-honest content can live inside engaging design. Scorecard-visual grammar can carry the same construct discipline as dense grey tables.
Teaching
The quality bar for measurement artifacts is TWO dimensions (defensibility + scanability), not one. Optimizing only for defensibility produces artifacts nobody wants to read.
D-2Governance + Decision scorecards redesigned in scorecard visual style
Triggered by
Direct consequence of D-1 — the reframe demanded execution.
Discipline surfaced
Execute-the-reframe discipline. Naming a quality bar shift without acting on it is aspirational, not real.
Resolution
AB Governance Maturity Scorecard fully rebuilt in scorecard visual style. AB Decision Scorecards got a top "at-a-glance" summary strip. PDDS Governance Scorecard redesign deferred to a dedicated pass (external, higher stakes).
Teaching
Reframe + immediate execution is what makes a methodology shift real. Deferred execution reverts to the old standard by default.
D-3Quality bar revision in Critical 5 (scorecard visual + PDDS construct discipline)
Triggered by
Critical 5 item #4 originally said "match PDDS Governance Scorecard as the quality bar." But PDDS itself was flagged as dense.
Discipline surfaced
Recursive quality-bar discipline. When the benchmark shares the flaw the standard is trying to fix, the benchmark is wrong.
Resolution
Critical 5 item #4 revised: quality bar is now "scorecard visual style + PDDS-level construct discipline."
Teaching
Benchmarks must be audited too. "Match X" is only a good instruction if X itself represents best practice on all the dimensions that matter.

Phase E · Dissemination discipline (3 moves)

Public safety, narrative correction, library-compatibility. Discipline of shipping honestly — including how the audit's own outputs reach the world.

E-1Instrument re-registration for public use (DM curve pinned via Cowork-approved libraries only)
Triggered by
DM curve HTML used external CDN scripts not on Cowork's approved list. Pinning required rewriting to comply.
Discipline surfaced
Platform-compatibility discipline. When moving an artifact from private to public/pinned surface, verify it survives the new platform's constraints.
Resolution
Rewrote JS: Chart.js bumped to 4.5.0, dropped annotation plugin + date-fns adapter, converted x-axis to linear with epoch ms + custom tick formatter. All 116 threads preserved. Registered as pinned artifact.
Teaching
Public-facing surfaces have their own constraints. Retrofitting a private artifact for public use is often more work than expected. Design for the public target from the start where feasible.
E-2Public labeling review — "Job Apps & Prep" → "Opportunity Evaluation" for public chart
Triggered by
Operator caught: "I don't want to put job search & prep in public space chart" when reviewing the correction chart draft for LinkedIn.
Discipline surfaced
Public-safety discipline. Internal classification labels may broadcast unwanted signals when exposed publicly.
Resolution
Neutral public label adopted: "Opportunity Evaluation." Preserves truth without labeling any as job search. Internal classification rulebook still uses "Job Applications & Prep."
Teaching
Public / internal labeling should be decoupled. Internal labels optimize for scorer precision; public labels optimize for reader appropriateness.
E-3Corrected-narrative chart built before public dissemination
Triggered by
Operator flagged an earlier LinkedIn post used pre-refresh estimated data.
Discipline surfaced
Correction-as-content discipline. When public artifacts need correction, one option is silent redaction; a stronger option is making the correction visible as part of the practice's own methodology story.
Resolution
Built two chart iterations: (1) before/after Jul 31 correction, then (2) three-metrics view (raw count + weighted proxy + real tokens) which surfaced the operator's insight that the ratio between weighted-volume and token-share IS the human-effort / LLM-effort split — an emergent methodology finding.
Teaching
Making corrections visible as methodology-in-action is stronger than silent correction. It turns a stumble into a demonstration.

The 7 emergent meta-principles

Heuristics that surfaced during the Jul 30-31 close. Not preset — emerged from the discipline moves themselves. Each now belongs in the practice's rulebook going forward.

Principle 1
Rule 1 — Outputs migrate; intent doesn't.
When a session's output has been moved to a hub or follow-on session, categorize by originating intent, not residual output. Practical test: "what was this session started to do?" before "what does its last turn contain?"
Principle 2
Defensibility ≠ visual density.
Kane-honest content can live inside engaging design. Measurement discipline as a USP is not license for dense measurement-y presentation. Quality bar is TWO dimensions (defensibility + scanability), not one.
Principle 3
Two-window discipline.
Operating window (when the work happens) and audit-run day (when the audit is written) are distinct temporal semantics. Time-series artifacts share operating-window closing date; snapshot artifacts carry their own stamps.
Principle 4
Hub-session taxonomy.
Not all sessions are atomic work packages. Single-category hubs and cross-category hubs each have distinct downstream effects. Naming the exception classes preserves the heuristic; leaving them un-named breaks it.
Principle 5
Methodology documentation can be done inline.
The audit-of-audits doesn't have to be a future Phase 3+ item. When the audit's own execution is observable, the meta-methodology can be captured in real-time — turning "future work" into "already demonstrated."
Principle 6
Definitional discipline is operator-scorer alignment.
Ambiguity in definitions drives most scorer errors. Requires operator + scorer to co-define terms before analysis, then hold each other accountable. Skip this and quantitative work is technically-correct but strategically-empty.
Principle 7
Three-lens principle for practice measurement.
Raw session count (throughput) · weighted volume (attention estimate) · real tokens (cost measurement) — each answers a different question. Where they align, the story holds. Where they diverge, one method needs review. Ratio between weighted-volume and token-share operationalizes the human-effort / LLM-effort split (load-carrier axis, previously judgment-only).

Prior art and anchor frameworks

What this log adopts, what it stays adjacent to, and what it deliberately doesn't compete with — because the scope isn't the same.

✓ ADOPTED · CITED THROUGHOUT

Kane's 4-inference validity framework — the audit's spine. Applied to the audit's own execution (see Kane mapping below): Scoring → Generalization → Extrapolation → Implication. Peer-reviewed since the 1990s.

NIST AI RMF — anchor for the PDDS Governance Scorecard (the audit's external-facing product, which IS at org / product scale where NIST applies).

~ ADJACENT TRADITIONS

Messick construct validity · AERA/APA/NCME Standards for Educational and Psychological Testing · broader measurement-science tradition that informs the practitioner's disciplinary background but isn't formally cited in the audit itself.

× DELIBERATELY NOT ADOPTED (WITH REASON)
NIST AI RMF at the practice level
Scope mismatch. NIST targets organizations deploying AI to serve customers. This audit is a solo practitioner's own operating practice. NIST is cited where scope fits (PDDS); not force-fit where it doesn't.
COBIT / COSO / SOC 2
Wrong domain. COBIT is IT governance; COSO is financial internal controls; SOC 2 is security / availability audit. None target AI-collaborative practice methodology.
Balanced Scorecard / OKRs / CMMI
Different object of measurement. These measure business performance or software-org process maturity. This audit measures governance maturity of the practice itself — a distinct object.
Model Cards (Mitchell et al.) / Datasheets for Datasets (Gebru et al.)
Different object. Model Cards are for MODELS; Datasheets are for DATASETS. Adjacent methodology tradition (transparency-oriented reporting) but not applicable to practitioner-scale governance.
Huawei AI-Native Enterprise / PwC AI-Native Enterprise / SAFe AI-Native
All target ENTERPRISE scale. The AI-Native family emerging in 2026 doesn't address practice-level scope (solo / boutique / fractional). AB's practitioner-scale framing (session classification, three-lens measurement, HITL/AITL mapping, dispatch guidance) targets that unaddressed scope.
▶ WHAT THIS AUDIT CONTRIBUTES

Practitioner-scale AI governance instruments (session classification, three-lens measurement, HITL/AITL mapping, dispatch guidance) at practice-operating scale — a scope existing AI-Native frameworks don't cover. Behaves like DevOps in software engineering: an operating discipline the other task categories depend on.

Measurement-validity discipline applied inline, with the audit's own execution as the empirical demonstration (this log). Not a new framework — an application of established validity theory to a new object at a new scale.

Methodology documentation done inline, not deferred as Phase 3+ future work — and evidenced with 23 mid-flight discipline moves + 7 emergent principles catalogued above.

What this log proves

"The audit did to itself what it audits others for. That IS the demonstration."

The Kane 4-inference framework applied to the audit itself. Kane's validity framework was cited throughout the audit as the discipline AB claims to bring to client work. Each phase of this log maps to one Kane inference layer:

Link 1 · SCORING
Did we measure it right?
Verifying atomic units, categories, weights — the vocabulary the audit runs on.
→ Phase A · 9 moves
Link 2 · GENERALIZATION
Real, or just noise?
Verifying data pipeline, catching extract-path error, establishing Rule 1 for scorer boundary corrections.
→ Phase B · 5 moves
Link 3 · EXTRAPOLATION
Reflects what matters?
Closing freshness gaps, surfacing new findings that emerged from the correction cycle itself.
→ Phase C · 3 moves
Link 4 · IMPLICATION
Worth acting on, for whom?
Making outputs legible for peer review + public dissemination — the audit's findings actually reach their audiences.
→ Phases D + E · 6 moves

This is what makes AI-native governance work observable. Traditional governance craft lived in the practitioner's head; only the finished memo surfaced. The scaffolding was invisible. In AI-collaborative work, that scaffolding is captured by default: transcripts, corrections logs, mid-flight amendments. This log surfaces what already exists. Not new evidence manufactured for external audience — the practice's audit trail as demonstration of the practice's discipline.

Downstream: this log is the foundation for the AB Governance Portfolio (near-term, Aug: public-facing packaging positioning AB's governance craft as evidence for hiring evaluators and BD partners) and the AI-Native Practice Audit Kit (Q4 productization: the transferable methodology other practices can adopt). Both use this log as their reference case.