Reach is what falls out for free
Every measurement stack has a gravity, and its gravity is reach. Count the people served, the dollars deployed, the workshops run, the grants closed. These numbers are not wrong and they are not useless — a funder who cannot say how many children a literacy program touched has a real problem. But ask why they are always the numbers you have. They fall out of the system for free. They are byproducts of transactions the organization was already logging for other reasons: the grant was disbursed, so the dollars are counted; the session happened, so attendance is counted. Reach is the exhaust of operations.
Outcomes are not exhaust. Nobody’s operational system logs “and it worked” as a side effect of cutting a check. Whether a program moved the thing it exists to move has to be constructed — the outcome defined, placed on a scale, and made comparable to the next program that defines success in its own words. Reach is counted. Evidence is built. The reach trap is mistaking the first for the second, and then blaming the reporting when the second never shows up.
Why the CRM can’t save you
This is where most teams go looking for the fix, and where it fails. The instinct is familiar: we have a CRM, it holds all our grantee data, surely the answer is a better view on top of it. But a CRM is a system of record for relationships and transactions — contacts, grants, touchpoints, disbursements, pipeline stage. It is engineered, correctly, to answer who did we fund, what did we do, what is the status. That is a real and necessary job. It is just not the job of telling you whether any of it worked.
The honest version of the claim is not “CRMs are bad.” You can bolt an outcome field onto any CRM. The trouble is that comparability is not a storage feature; it is a governance feature. Add a free-text “impact” field and, across forty grantees, you get forty differently-worded fields that do not read across. The seam you were trying to close reopens, now with a database licence attached. The CRM was never the wrong tool. Treating it as your evidence layer is the category error. Reach falls out of it for free; comparability does not exist unless you build the second system deliberately.
None of this says measurement people have ignored use. Utilization-focused evaluation has argued for decades that an evaluation nobody acts on has failed; the conviction here is inherited, not invented. What changes is where it lives. Shared-measurement frameworks reach comparability by making every grantee report the same indicators up front; metric catalogs like IRIS+ do it by prescribing a common dictionary to pick from. Both work when a portfolio shares one goal or one asset class. A grant book rarely does. Forty grantees carry forty theories of change, so the spine runs the other way: let each keep its own language, then place it on one scale afterward. The results-framework world already tried the alternative, and the OECD’s review of results-based management landed on a familiar pattern: the data gets collected, but it rarely reaches the decision. The obvious read is a culture problem, teams that never build the habit of looking. The deeper one is structural: the data was never shaped to be read as a decision, and the second system is what shapes it.
What the rubric actually is
The second system has two moving parts, and the first is the one people skip. A rubric is not a survey and it is not a KPI list. It is a shared outcome scale — an explicit, small, ordered set of what “good” looks like, the same scale for every grantee in the book. Four points, say: emerging, building, established, sustained. That much is intuitive. The part that does the real work is the second piece: a semantic layer that maps each grantee’s own language onto that scale.
The scale is not a universal yardstick for impact. It is decision-specific: for this book, at this decision point, it asks a single question, what stronger evidence of progress looks like, and it is honest only along that axis. Compare grantees doing genuinely different work on a scale built for a different decision and you get false equivalence, not comparability.
Grantee A says “kids reading at grade level.” Grantee B says “literacy gains held two years.” Grantee C says “fewer kids need remediation.” Left alone, those are three incomparable sentences. Run them through the semantic layer and they become placements on one scale — B at sustained, C at established, A at building. Different words in; one comparable placement out. Without the semantic layer, a rubric is just a fourth dialect nobody speaks. With it, forty theories of change finally line up on a single axis.
That mapping is the part people imagine is automatic, and it is not. Funder and grantee co-author it: they agree, in advance, what evidence justifies each placement. The layer does not dissolve the judgment about whether “held two years” outranks “reading at grade level.” It forces that judgment to be made once, in the open, and then reused, instead of relitigated grantee by grantee every cycle.
What the spine actually is
The rubric tells you where one grantee lands. The spine is what makes the whole book a book. It is the shared record shape — the same handful of columns for every grantee, every cycle: the outcome tier from the rubric, the leading signal you are watching, the confidence you have in the evidence behind it. That consistency is not clerical tidiness. The shared columns are the structure. They are what a CRM’s per-grantee custom fields can never be. One of those columns earns a definition the others assume. Confidence is not a mood. It names how much weight the evidence behind a placement can bear: how strong it is, how recent, whether it is comparable to the grantee in the next row. Keep it vague and one officer’s “high” is another’s “medium.” Define it and the column carries actual weight.
Keep the columns identical and the book reads across in one glance: same columns, one axis, comparability. You can sort, compare, and rank forty grantees on the same terms, and the portfolio question — where does the next dollar go — becomes answerable instead of rhetorical. Fragmented in, decision-ready out. Drop the columns, let each grantee report in its own shape, and you are back to forty PDFs that agree on nothing. The spine is unglamorous by design. It is a schema, not a dashboard. But it is the actual decision architecture for MEL, and it is the thing the reach trap leaves out.
One spine, three decisions
The reason this is worth building once is that the same spine serves three different seats without being rebuilt. The vocabulary changes; the architecture does not. A program officer reads a single grantee’s row and asks whether it is working. The spine gives them the leading signal before the post-mortem, not just the status the CRM already showed. A portfolio manager reads down the whole column and asks where the next dollar goes; comparability lets that view resolve into invest, scale, or stop instead of collapsing to a single green number. A CSR lead in India reads the same structure board-facing, against Schedule VII and the two-percent mandate, and asks whether to renew, scale, or exit. Their ERP already gives them spend and compliance. The spine gives them the impact half that spend alone can never defend.
Three readers, three questions, one record shape. That is the whole argument for treating the spine as architecture rather than a report: you build it once and it pays out at every altitude.
Where to start
You do not start by ripping out the CRM. You start by admitting it is a system of record for activity and will never be more than that, and then you build the small second system beside it. Name the four-point scale before you argue about metrics. Write down the semantic layer — the mapping from each grantee’s language to the scale — because that is the part everyone wants to skip and the part that makes the rest work. Fix the three or four columns that will be identical for every grantee, forever. That is a spine. It is less software than people fear and more discipline than they hope. But it is the difference between a stack that hands you reach and a decision system that hands you decisions.
Reach is a number. The decision is the system you architect around it. Analytic Bytes helps funders, portfolios, and CSR teams build that evidence spine — the rubric, the semantic layer, and the shared record shape — and move from fragmented to decision-ready.
Written July 2026 for the Analytic Bytes Library. A field note on the category error underneath most MEL stacks — the CRM as system of record for activity, and the evidence spine as the separate, deliberately-built system of record for outcomes.
Questions, pushback, or a problem that looks like this one? Write to chai@analyticbytes.systems.