In 2005 I was a doctoral student on a field team visiting small teacher-education programs across the country to help them make the case for accreditation. My job was data auditor. The team would visit a program in a small town, one at a time, and the observations and interviews took precedence; those built the case. The data was the audit layer underneath. I would sit in a back room with my small laptop, paper, pencil, and calculator, and I would meet the faculty member who had collected or logged the program’s evaluation data.
The data was incomplete. Rows missing. Codes inconsistent. Not because the faculty member was inattentive. Because holding the data was on top of the day job. The person in front of me was a teacher-educator who had absorbed the responsibility for the program’s data because someone had to, and they were the one who was capable with numbers or interested in them or both.
That was 2005. I would spend the next twenty years watching the same shape take different forms in different sectors, and I would only recently name what I was actually watching.
The pattern
Data infrastructure work in mission-driven organizations gets absorbed by whoever in the organization is capable with numbers or interested in them. A math teacher building professional-development slides. A writing faculty member logging automated-scoring results. A clinician maintaining a caseload tracker. A program coordinator running the grantee spreadsheet. A Director of Instructional Technology fielding data requests because they have system access.
The result is an absorption pattern. When there is no funded infrastructure role, the work goes to whoever can carry it, and it stays there. Never specialized. Never advocated for. Never scaled. Until the person carrying it leaves, and the informal knowledge walks out with them.
The visible symptom is spreadsheet sprawl. The invisible cost is that the work never gets to specialization. And the person carrying it spends their delivery time on something that was supposed to be someone else’s job.
Three buckets
Data work in any organization decomposes into three kinds of work, and the failure mode is trying to collapse them into one role.
Bucket 1 — data entry. Recording what happened. Documenting a session. Logging attendance. Reporting a grantee touchpoint. This is legitimately part of every delivery role and has always been. Clinicians do it. Teachers do it. Program officers do it. This bucket stays distributed across delivery.
Bucket 2 — interpretation and action. Reading a dashboard. Acting on a flag. Adjusting practice based on a signal. Also part of every delivery role, but this bucket requires two things the delivery role often does not have: data literacy to interpret the signal correctly, and decision-ready data to act on. When the delivery role is asked to interpret an ambiguous or unreconciled signal, they either default to their own read (drift) or hoard their own local data snapshot (sprawl).
Bucket 3 — curation, integration, semantic definition, quality assurance, governance. Maintaining the definitions. Reconciling how “attendance” or “session” or “enrolled” or “at-risk” is defined across systems. Building the pipeline. Catching data quality issues. Owning the semantic layer. This bucket is specialized capacity. It is not part of delivery. It requires a dedicated, funded role.
The three buckets exist in every organization. The question is where each lives and how each is funded.
The arc
The failure mode is when bucket 3 has no funded home and gets absorbed by bucket 1 or 2. I watched this take different shapes across four sectors.
In the 2005 accreditation work, bucket 3 was absorbed by teacher-education faculty. Their day job was preparing teachers. Their absorbed job was holding the evaluation data. The data was incomplete because there was no one whose primary responsibility was making it complete.
In 2014 I worked with a public university writing program on automated scoring for student writing samples and remedial placement. The data was being collected, logged, tracked, and analyzed by two writing-program faculty who had reached for the technology because they wanted to see if it could serve their classroom. The scoring worked. The infrastructure to sustain it stayed with the two faculty who had built it, on top of teaching.
In 2023 I stepped into a K-8 charter network. Before the data role existed, the reporting was being done by a math teacher who built the professional-development slides. Fifty Google Sheets carried the semantic layer over an unconsolidated legacy stack. The core measure, student persistence, was independently recalculated eight or more times across sheets. Each version was correct under its own definition. None of them agreed with the others. A Director of Instructional Technology managed the operational systems (student information system, learning management, assessment platform) and had become the de facto data provider because he had access. The role played both ways: users routed around the org chart to reach the person with system access; the person with system access held that access as leverage. Half-configured systems. Under-utilized fields. Users hoarding needed snapshots in their own local sheets. Before long, everyone was arguing about the single source of truth.
At a behavioral-health agency running Certified Community Behavioral Health Clinic services, a measurement-based-care pilot ran across two waves with thirty or more patients, producing a forty percent lift in engagement after reminder automation was added. The clinicians who participated were the ones who had elected in; they believed measurement-based care could make a difference. They had to be provided time by administration to attend pilot meetings, learn dashboard access, and provide feedback. Even the officially provided time came out of billing hours. In the long run, the discipline could have recouped that time through more efficient visits and better outcomes evidence. In the short run, the delivery role absorbed the bucket-2 capacity-building on top of caseload.
At a national mission-driven foundation, the infrastructure was funded. Fivetran ingestion. Snowflake warehouse. A dbt-modeled semantic layer over four to six instrument versions of the same construct across years of survey data. Digital products managed the client-facing surfaces. DevOps managed the pipeline. A research-data-steward managed instrument documentation. Each role solved a different problem. None of them had the mandate or the disciplinary depth to design the semantic layer, the reconciliation of what “engaged participant” meant across instrument versions, the crosswalks between survey generations, the decisions about what could and could not be compared across years of program data. The infrastructure existed. The semantic layer stayed unowned. Bucket 3 was distributed across three roles, none of whom was the semantic-layer architect the work required.
Data literacy versus decision-ready data
These are not substitutes. They serve different buckets, and mission-driven organizations regularly try to use one to compensate for the absence of the other.
Data literacy is what the delivery role (bucket 2) needs to interpret and act. It is coaching. Professional development. Sense-making capacity. A clinician who understands what a no-show pattern means, in context. A teacher who understands what an assessment score does and does not tell them. This is training and reinforcement work, and it never stops.
Decision-ready data is what the central function (bucket 3) delivers. The semantic layer, the reconciled records, the dashboard curated for a specific recurring call. Without this, the delivery role’s data literacy does not help. They are interpreting bad signal well.
You need both. They come from different capacities. The mission-driven failure is training clinicians and teachers to be data literate as if that would substitute for the semantic-layer role that was never funded. It does not. Data literacy layered on unreconciled data produces confident wrong interpretation, which is worse than no interpretation.
The architecture
Centralization and decentralization are not opposing poles that need to be balanced. They are properties of specific buckets.
Centralize bucket 3. Semantic definitions, integration, quality assurance, governance. Owned by a specialized function that does not do delivery. Its output is decision-ready data.
Decentralize bucket 2. Interpretation and action stay with delivery roles who understand the operational context. A central data function cannot decide what an at-risk flag means for a specific student in a specific school — whether it warrants calling the family, escalating to the counselor, adjusting the intervention plan, or watching closely for another two weeks. That is contextual judgment, and it belongs to the delivery role.
Distribute bucket 1. Entry stays where it has always been, part of every delivery role.
This architecture is a three-layer design that respects what each bucket requires and what each role can carry.
The mission-driven failure is trying to run all three buckets out of one role, the math teacher or the writing faculty member or the Director of Instructional Technology or the research data steward, because the specialized bucket 3 role has no funded home. It works until it doesn’t. When it doesn’t, the delivery role’s day job suffers, the data quality erodes, and users start hoarding their own snapshots because they no longer trust the central number.
What could work
The enterprise sector has been working on a version of this problem for a decade and has produced partial answers. Federated data governance and data mesh architectures locate a central function that owns the platform, the standards, and the semantic layer, while domain teams own the data products in their operational area. Data contracts formalize the interface between the two. A Chief Data Officer or Chief Data and Analytics Officer holds the accountability. What transfers to mission-driven contexts is the three-bucket separation, the data-contracts vocabulary, and the case for a specialized semantic-layer role. What does not transfer is the staffing scale. Most mission-driven organizations cannot fund a CDO with a team underneath, cannot afford the tooling that data mesh implementations run on, and do not have domain teams large enough to own the products the enterprise model assumes.
If the enterprise pattern does not port over intact, the mission-driven sector has been working on models that fit its constraints. Sector-level backbone organizations that hold shared definitions and shared infrastructure (Ed-Fi in K–12, CCBHC collaboratives in behavioral health, state longitudinal data systems, Candid in philanthropy) are the most established. The emerging operating model is a fractional bucket-3 advisory arrangement, where a specialized data architect works across several mission-driven organizations at a portion of an FTE each. Foundation-funded direct capacity grants have a mixed track record but have not been retired. Frontier-lab shared-infrastructure partnerships of the kind Anthropic and others are building are the newest and potentially the most durable, if they survive champion turnover. Coalition-owned semantic layers, held jointly by a group of peer organizations, are under-tried and worth trying.
The discipline that connects these models is customer and supplier. Tom Redman and Angela Saitta have argued in the CDO Magazine version of this that every delivery role in the organization is both a customer of upstream data and a supplier of downstream data, and that the specialized data function’s job is to architect the interfaces between them. That framing is what a bucket-3 role does in practice. Without it, data literacy programs end up teaching abstract concepts instead of the specific customer-supplier discipline the delivery role’s workflow requires. With it, data literacy has something concrete to attach to.
One last thing worth saying. Mission-driven work does not have to be back-breaking, low-pay, over-work. The absorbed-data-role pattern is one specific mechanism of the broader burnout the sector already knows it has. When the specialized bucket-3 role is unfunded, the delivery capacity pays the cost, and the person who absorbed the work pays it with their time and their day-job attention. The path forward is funding the specialized role at whichever level the sector can sustain: shared, fractional, foundation-supported, or lab-partnered. All of those are better than the absorption model that has been running by default.
Close
The boundary the sector needs to draw is between data entry, data interpretation, and data infrastructure. All three are real work. All three need funded capacity. Only two of them belong in delivery roles.
Every mission-driven organization I have worked with has held all three in the same person, because that person was capable and there was no one else. That model is the informal structure keeping mission-driven data operations running. It is also the reason the infrastructure never gets built. When the person absorbing the work leaves, or burns out, or gets pulled onto something else, the informal structure collapses. There is nothing underneath.
The specialized bucket 3 role is the one that does not exist in most mission-driven org charts. It is also the one the sector cannot function without. Naming it, funding it, and hiring for it is the piece of decision-systems architecture that mission-driven organizations keep skipping. The math teacher, the writing faculty member, the clinician, the program coordinator will keep absorbing the work in the meantime.
That was 2005 in a small town. Twenty-one years later, in most places, the pattern still holds.
Written July 2026 for the Analytic Bytes Library. A field-note argument drawn from twenty years of watching the same absorption pattern take different forms across four sectors: teacher-education accreditation, a public university writing program, a K-8 charter network, and a behavioral-health CCBHC, plus a national mission-driven foundation where the infrastructure was funded and the semantic-layer role still stayed unowned. The three-bucket separation and the customer-supplier framing borrow from Tom Redman and Angela Saitta’s CDO Magazine writing on the specialized data function.
Questions, pushback, or a problem that looks like this one? Write to chai@analyticbytes.systems.