How AI-work is typically measured, and what this audit does differently
Most reporting on AI use at the practitioner level is testimonial. “It saved me hours.” “It writes my drafts.” The claims are common, but the evidence base is usually one anecdote.
Where the reporting gets quantitative, it typically stops at usage counts: sessions, prompts, tokens, screen-time. Usage is easy to measure. It is also a weak claim. High usage tells you the practice ran. It does not tell you whether the practice measured what it claimed to measure. It also does not tell you whether AI carries the load or only assists it.
This audit is different in three ways. It defines the atomic unit of analysis (one session = one named conversation, one category per session) before counting anything. It measures the same work through three lenses that answer different questions. And it applies measurement-validity discipline (Kane-framed inference, construct-definition rigor, scorer-operator alignment) to its own findings before publishing them.
Method: three lenses and the HITL/AITL split
The audit runs on three measurement lenses over the same 90-day corpus (n ≈ 93 sessions across ten task categories).
Session count. How often AI is engaged, per category. The cheapest measure, and the most misleading if used alone.
Message volume. How much conversation happens inside those sessions. A closer proxy for cognitive load.
Token usage. What the model actually computed. The real system-cost signal, measurable from API records.
The three lenses tell three different stories about the same practice, and the divergences carry the diagnostic weight. Recurring automation is fifty-five percent of sessions and fifty-seven percent of tokens after the Jul 31 v2 extract correction. That is a rare category where session count and cost move together. Website and technical development is a small share of sessions but a large share of per-session token weight. Strategic and meta-reflection has the fewest sessions and the highest cognitive load per session. Session count is a bad proxy for time, and token usage is a bad proxy for engagement frequency. Both are needed.
Cutting across the three lenses is a fourth distinction the audit uses to describe the shape of AI engagement per category. This is the naming contribution: HITL / AITL used as a paired framework.
HITL (Human-in-the-Loop). AI drives; the human oversees. Recurring automation fits here — the AI carries almost the full workflow, and the human sets it up, monitors, and corrects when it drifts.
AITL (AI-in-the-Loop). The operator drives; AI assists. Strategic thinking, product/framework work, and long-form writing sit here. The operator carries the work; AI is a thinking partner, drafts scaffolding, or synthesizes.
The two frames are standard AI-ops vocabulary. The move that matters here is using them as a signal-informed proxy for the load-carrier axis. HITL and AITL name who or what carries the primary work. That axis was previously described in prose (AI leads versus AI assists) without an operational way to name it per category. Placement per category was informed by token distribution, message volume, and user-turn patterns per session — categories where AI carried the output with few user turns landed HITL; categories carried by high-turn operator iteration landed AITL. No shared-middle option, because a category that reads shared on the surface typically has one side actually carrying the decision when you look closely. (The signals informed operator judgment on band placement; formalizing them into a per-session assignment rule with reliability data is a next-cycle instrument.)
Findings: what ninety days showed
Six things surfaced across the audit that a testimonial would have missed.
1. The working pattern is dialogic, not delegative. The dialogue-maturity curve spans 116 scored threads across the twelve-month arc — 88 GPT and 28 Claude — with all 28 Claude sessions inside the 90-day window. The mean per-session pattern is iterative back-and-forth. Session shape shifted upward on the 6-metric scorecard between the GPT-era portion (substantive threads cluster 5–7 composite; the tactical resume and job-app volume pulls the aggregate lower) and the Claude-era portion (mean composite 7.06, most sessions clustering 7.0–8.5, top of the corpus at 8.83). The gain concentrates in three of the six rubric dimensions: voice ownership, meta-awareness, and generative reframing.
2. The distribution of the practice’s attention has three visible modes. Instrumenting the practice — automation, meta-work, research — took roughly forty-five percent of estimated message volume. Positioning the practice — job search, market-identity work — took another fifteen percent. Producing the practice’s actual output — product, library, brand, talks, website — took about thirty-eight percent. That mix reads as a practice-in-formation. Naming which stage you are in is the discipline.
3. Session count and token usage disagree about where work happens, and both are right. The dumbbell view on the same ten categories shows recurring automation aligned across the two lenses after the v2 extract correction. Website and technical dev over-indexes on tokens (small session count, heavy per-session weight). Research, job applications, and product under-index (many short sessions, low token weight per session). Dispatch decisions (hire priorities, cost management) should default to the token lens when cost matters and the session lens when frequency matters. Aggregating them into a single “how much AI” number destroys the signal.
4. Dispatch discipline exists at three nested levels, and it is currently prescriptive rather than measured. Cross-tool dispatch (Claude vs GPT vs Gemini). Within-tool model dispatch (Haiku vs Sonnet vs Opus). Within-model effort dial (low vs medium vs high vs max). The rulebook says what should happen at each level. The audit has no per-session record of what actually happened. Closing the prescription-to-measurement gap is the highest-leverage single change for the next audit window.
5. The infrastructure has five active layers. Data, tooling, workflows, artifacts, and governance, each with named components, connected by five workflow loops. Two components are honestly dormant (Card Maker, Content HQ).PDDS is a working, public example of measurement-validity discipline expressed in code and policy. It runs a two-component LLM-as-judge setup (narrator and judge separated, deterministic code-veto on grounding claims, advisory-only judge scoring, with anti-Goodhart guardrails on both sides) and a Kane-framed validation approach; the current agreement number (Cohen’s κ = 1.00 on a small co-evolved 12-item gold set) is suggestive not proof, and the eval spec names the residual limits (single model, no held-out set, no variance runs, correlated LLM training bases).
6. The governance dimension is at a moderate baseline, with the two distinctive-discipline anchors scoring HIGH. Self-scored on the seven-dimension governance scorecard (self-assessment, not third-party audit). Validity spine (Kane four-inference framework, NIST AI RMF alignment, reliability + construct-validity discipline) scores HIGH. Construct-definition rigor (session unit, category rulebook, hub taxonomy, weights table, dispatch guidance) scores HIGH. Reflection cadence and instrumentation depth score LOW-to-MODERATE. These are the dimensions the audit itself just formalized. Descriptive-vs-prescriptive (rulebook exists, per-session measurement doesn’t) scores LOW. This is the expected baseline of a practice that just finished its first ninety days.
Handling the N=1
A field note built on one practice invites the obvious objection: N=1 is not generalizable. That objection is correct and worth naming up front.
Two responses.
The value here is not “here is what all AI-native practices look like.” The value is “here is an instrument other practices can apply to their own operations.” Session-unit definition, three-lens measurement, HITL/AITL mapping, dispatch levels, Kane-framed validity discipline — these travel. The findings themselves do not generalize. N=1 is a limitation on findings and a reasonable starting point for building instruments. Ethnographic tradition, single-subject research designs, and autoethnography in the social sciences all operate under the same trade.
The audit’s distinctive discipline is measurement validity. Publishing an unaudited claim about that discipline would be self-refuting. The audit documented its own methodology before publishing its findings on AB — twenty-three discipline moves and seven emergent principles logged as they happened during the Jul 30–31 close (self-selected and self-catalogued; Section 9 of the internal document). That work is methodology documentation, done inline, dated, and available for peer review. Replication is the invitation this piece extends.
The contribution: practitioner AI governance as a distinct discipline
Two conversations about AI governance dominate the current discourse. Data governance addresses asset stewardship: who owns, curates, and secures the data AI runs on. Enterprise AI governance addresses policy and organization: how a company writes its AI-use policy, sets up its review committee, and audits its vendor models.
Practitioner AI governance is a third slice, and it is under-built as a discipline. It addresses how one operator governs their own AI-native practice: the session-unit definitions, the model-and-effort dispatch, the load-carrier calibration, the validity discipline applied to the operator’s own outputs. It runs whether or not an organization has written a policy or stood up a data governance team. Every operator working with AI is already doing practitioner AI governance, badly or well. Almost none are doing it with measurement discipline attached.
The claim this audit stands on is that measurement-validity discipline for AI governance, at practitioner scale, is a distinct discipline worth naming. This audit argues it sits as a third slice alongside data governance and enterprise AI governance, with its own scope and instruments. The discipline has its own instruments: session classification, three-lens measurement, HITL/AITL mapping, dispatch guidance. It runs on its own validity spine: Kane-framed inference, construct-definition rigor, scorer-operator alignment. It produces its own artifacts, including the Governance Craft Log, the Executive Card, and scorecards.
Practitioner AI governance behaves like DevOps in software engineering: an operating discipline that the other task categories depend on, rather than a category of work alongside them. The instruments a practitioner uses to run this discipline — session classification, three-lens measurement, HITL/AITL mapping, dispatch guidance — apply across whatever categories the practice happens to work in.
From operator governance to agent governance
Practitioner AI governance and agent AI governance are complementary layers, not competing frames. Both are early in their formation. Both need measurement-validity discipline for the claims each makes to hold up.
The frameworks introduced here should transfer to the agent context with implementation adjustments (the transfer is proposed here, not tested).
HITL/AITL as the load-carrier axis. Agents are the HITL case, by definition: the AI drives the workflow, and the human oversees. The measurable question at the agent level is where within an agent’s execution path human oversight actually engages, and how the boundary is calibrated.
Kane’s four-inference validity discipline. The four inferences are scoring, generalization, extrapolation, and implication. The stack applies unchanged to agent evaluation. Scoring: did the agent’s action correspond to the intended output? Generalization: does the action hold across the agent’s task family? Extrapolation: does it hold in production contexts the eval did not reach? Implication: does the decision the agent supports actually produce the intended outcome?
Gate discipline. The checkpoints where a human review is required in the practitioner setting become the escalation and human-in-loop boundaries in agent systems. The taxonomy is the same, though the stakes are higher. This is the argument Actions, Not Answers made for agentic deployments: the free human checkpoint has to be designed back in.
Task and session classification. The unit-of-analysis discipline (what counts as a session, what counts as a category, how boundary calls get resolved) is directly transferable to agent-run task decomposition.
Two pieces do not transfer cleanly. Voice-map compliance works for an operator; agent output does not have a voice to comply with in the same sense. Scale changes: procedural governance at the operator level has to become programmatic governance at the agent level, and the failure modes shift accordingly.
The reason practitioner-level discipline matters for agent governance is architectural. Operator-level and agent-level practice share the same instruments; the operator level is where those instruments are calibrated first. Practitioner-level discipline is where instruments get built and tested before they run inside systems without a human at the keyboard.
Links out
The audit produces artifacts, not just findings. Four came out of the audit itself:
Governance Craft Log — the 23 discipline moves catalogued during the audit’s own execution, phase by phase.
Executive Card — the one-page synthesis of the audit and its methodology anchors.
AB Governance Maturity Scorecard — the seven dimensions with current baseline and next-90 targets.
Three Lenses — six task categories read through session count, weighted-proxy volume, and token share, mapped to HITL and AITL bands.
What’s next
Two immediate follow-ups. A governance scorecard post that surfaces the seven-dimension baseline in one page, for operators who want to score their own practices against it. And a Q4 audit refresh that closes the biggest instrumentation gap: the descriptive-to-measured shift for per-session dispatch (which model, which effort, actually used, per session).
The instrument this piece describes is v1. Later versions of the instrument will improve as other operators run it against their own practices and report what breaks. Send what you find.
Written August 2026 for the Analytic Bytes Library. The public report from AB’s own 90-day audit — three measurement lenses (session count, message volume, token usage) applied to n ≈ 93 sessions across ten task categories, mapped against the HITL/AITL load-carrier axis and Kane-framed validity discipline. Draws on the Governance Craft Log, Executive Card, and Maturity Scorecard produced during the audit.
Questions, pushback, or a problem that looks like this one? Write to chai@analyticbytes.systems.