A response to “Dollars in Pockets: A New Way to Measure What Matters” — NextLadder Ventures & GitLab Foundation, July 2026.
Getting to a common unit was the hard part
For years the economic-mobility field measured in units that do not convert. The authors’ own diagnosis is the right one: a venture return describes what comes back to the investor, and a philanthropic report describes how many people an intervention touched. Neither answers the question a funder is holding: did this do more good per dollar than the alternative? Put a legal-aid tool, a benefits app and a jobs platform side by side, and neither unit lets you rank them. The evidence was rich and stubbornly incomparable.
It was not for lack of trying, and the authors are careful to say so: they credit a lineage that includes the Robin Hood Foundation, Bridgespan, and TPG’s Y Analytics. Alongside that work the field also carried shared vocabularies: the Global Impact Investing Network (GIIN)’s IRIS+ catalog of standardized impact metrics, and Social Return on Investment (SROI)’s older tradition of monetizing outcomes into a return ratio. But IRIS+ is a menu of indicators, not a single convertible number. And SROI, which does denominate its ratio in currency, leaves the valuation proxies behind that currency to be selected study by study — so one organization’s ratio is rarely comparable to another’s. Dollars in Pockets takes both steps together.
The authors draw the analogy themselves, and it is the right analogy. Global health once faced the same problem: vaccines, surgeries, and mental-health care all funded from one budget, with no way to rank them against each other. It answered by building the disability-adjusted life year (DALY) — a single unit expressing what a dollar buys in healthy life. Their claim is that money admits the same treatment. Fees avoided, wages earned, and benefits accessed are financially distinct events; denominate them in one unit and they become comparable, and the interventions producing them become comparable with them. That is the move the field needed. It is not a small thing to have made.
The analogy holds in a second way that matters more. The DALY arrived unsettled, and the argument about it ran for three decades. The disability weight (the number for how bad a given health state is) is elicited from surveys, not measured, and critics have argued the resulting figure has no clear meaning and overstates how much the world agrees (Nord, 2013). Two of the original parameters, age-weighting and time discounting, were contested for the better part of two decades and then dropped in the GBD 2010 revision. That they could be removed at all showed how much of the unit was normative choice made to look like measurement.
The contested history is part of what makes the analogy useful. Global health built the unit first and did the validity work afterward, slowly, under pressure from people who kept asking what the weights meant. Dollars in Pockets is at the front of that arc. The authors have done the hard part. The long part is still ahead of them, and it is the part this essay is about.
A magnitude is not a validity
A common unit raises a question it does not itself answer. Dollars in Pockets tells you how much. Whether you can stand behind that much is a second question, and it is the one that usually goes ungraded.
The arithmetic can be correct and the inference behind it still weak. When a fund sums its pathways, the weak one disappears into the total. The total reads healthy while part of it does work the evidence cannot support.
It matters exactly what the weak part damages. A soft housing-stability estimate does not make a hard fees-avoided estimate any softer, and each pathway keeps whatever standing its own evidence gives it. What the weak part damages is the composite claim — the sentence a fund says out loud about the whole figure. That sentence asserts every dollar in the total, so it carries only as much weight as the weakest dollar in it. Blending preserves the arithmetic and loses the warrant.
It also hides which pathway is the soft one, and that matters most the moment the number becomes a target. Goodhart’s law is the standing warning: once a fund optimizes toward its Dollars in Pockets ratio, optimization pressure will eventually find the softest pathway — the very one the blend has hidden. A metric that cannot show its weakest inference rewards inflating it.
The authors have an answer to this, and it is a fair one. They say the ratio is one input among several, weighed alongside team strength, strategic alignment, organizational effectiveness, execution and evidence risk, and expert judgment. I take them at their word. The difficulty is that this is a governance commitment held by the people who built the measure, not a property of the measure itself. Commitments of that kind travel less well than numbers do. Dollars in Pockets is designed to spread (that is the whole point of a common unit), and the funders who adopt it later will inherit the ratio without inheriting the restraint that currently surrounds it. If the gate is part of the number, later adopters get it whether they intend to or not. A norm alongside the number depends on them choosing to keep it.
Take the piece’s own worked example, Rasa Legal, whose tool helps people clear eligible criminal records. The authors assess it across three named pathways: legal fees avoided, wage gains that follow a cleared record, and savings from reduced housing instability. They draw on Rasa’s own data and outside research to estimate how often each benefit occurs, compare each against a researched counterfactual, then sum the three and divide by the investment required. The article publishes no Rasa-specific ratio, only the structure. The structure is enough to make the point.
Look at the three pathways side by side and they are not equally certain. Fees avoided is close to immediate and direct — the record is cleared, the fee is not paid. But “wage gains that follow a cleared record” and “savings from more stable housing” each depend on a chain of contribution holding: that the clearance causes the job, that the job causes the housing, and that none of it would have happened anyway. Those are the softer links. Sum the three into one figure and the near-certainty of pathway one underwrites the causal reach of pathways two and three. The composite cannot tell you which pathway carries the weakest inference. Adding does that to any set of numbers, whatever their quality, and Rasa Legal is not the problem here.
One caveat, and it matters. The three pathways above are named in the published article, but the complete methodology behind them is not yet public — the authors say they intend to publish it. So the question this essay raises cannot be settled from outside. Does that methodology already report per-pathway confidence, grading fees-avoided, wage-gain, and housing-stability separately rather than only in the blend? If it does, the point here narrows from “the composite hides the weak link” to “surface that per-pathway confidence next to the headline number, where a reader can see it.” Either way the discipline is the same, and the last section of this essay sets out what showing it would look like. The open question is whether it is already being done and not shown.
What Dollars in Pockets already gets right — and where the gap is
It would be easy, and wrong, to say the field just needs more rigor. Dollars in Pockets already does most of what a naive critic would demand, and the authors describe the practice in their own words: they “mark down optimistic projections, and weight every assumption by the strength of the evidence behind it.” Each modelled return is set against a researched counterfactual, and only the difference is counted. After the money moves, the estimate is checked again. NextLadder works with Opportunity Insights using independent methods, and GitLab Foundation resources its grantees to run their own measurement and evaluation, including direct feedback from the people served. The authors give these as examples rather than a complete validation design.
Two further choices deserve more credit than a critic would naturally give them. The first is the benchmark. NextLadder sets its ten-to-one target against direct cash transfers, on the reasoning that a dollar handed to someone is a dollar received, and that any more complicated intervention should have to clear that floor. A fund that picks a benchmark that unflattering has made one of the most disciplined choices in the framework. Most skeptics arrive expecting the opposite. The second is volume. GitLab Foundation reports applying the underlying ROI method across more than three hundred pre-investment impact models and two hundred grants. That is an empirical base, not a paper exercise.
The lineage named earlier has gone further than this, and it is worth being exact about how far. Y Analytics, TPG’s impact research arm, published its climate methodology in November 2025. It works pathway by pathway, and it grades the evidence behind each one. The risk rubric covers, in its own words, “rigor of evidence (i.e., RCT vs. observational study), geographical similarity, and number of linkages between company output and impact.” The result is a risk adjustment factor for each pathway.
Read that rubric again. Rigor of evidence is the distinction the three verbs below are built on. Number of linkages is the length of the chain. Y Analytics is already producing, pathway by pathway, most of the judgment this essay says a fund should publish. It then multiplies that judgment into the magnitude and reports one figure, the Carbon Yield.
That is a different problem from averaging, and the arithmetic is not the fault. Multiplying serial links is defensible. What the multiplication costs is the separation GRADE insists on, certainty rated apart from the strength of what is recommended. Fold a confidence grade into a magnitude and a reader can no longer tell a large number with thin evidence from a smaller number with strong evidence.
The authors also name a limit of their own. Under the heading of what they still cannot measure, they say Dollars in Pockets addresses economic success but not empowerment or dignity, and that they are still building yardsticks for those. That is an honest admission, and it is a different admission from the one this essay is pressing. Theirs is a limit of scope: which parts of a life the unit covers. Mine is a limit of inference: whether the number, inside the scope it already covers, supports the claim being made on it. A measure can have perfect scope and still fail the second test. Adding dignity to the unit would not close the gap named here. It would apply the same inference problem to more of a person’s life.
The rigor is real. The composite is where it breaks. All of that rigor still resolves into a single reported magnitude, and contribution (additionality, the “would this have happened anyway” question) is the one link a blended ratio is built to hide. A fund can do the work and still report a number that does not show it.
The gate: the weakest link, not the average
Testing solved a version of this problem long ago. Since Kane’s 1992 paper, high-stakes assessment has increasingly treated a decision as a chain of inferences to be examined link by link (Kane, 1992: validation as an argument, with evidence aimed at the weakest assumption), and has insisted that validity is a property of the use of a number, consequences included, not of the number itself (Messick, 1995: unified construct validity). Point that machinery at an impact figure and you get a small, hard discipline. Four questions, in operator language — a translation of the argument-based approach for impact capital, not Kane’s own four inferences restated:
- 01Measure it right — is the underlying quantity captured correctly?
- 02Real, not noise — would it hold up, or is it an artifact of a good month or a favorable sample?
- 03Reflects what matters — does the thing measured stand in for the outcome we care about?
- 04Ours, not adjacent — did our capital contribute to the effect, or did it ride alongside something that would have happened anyway?
Magnitude (is the effect big enough to act on) is deliberately not one of the four. It is a policy threshold, set by a fund’s own bar for what counts as worth doing. The four ask whether the number means what it says. How large a meaningful number has to be before it justifies capital is a separate decision, and mixing the two lets a fund answer the easy question in place of the hard one.
The scoring rule is what separates this from a maturity score. The verdict is the minimum of the four. One red link caps the claim no matter how strong the other three.
Kane does not give you that rule, and the borrowing is worth stating plainly. Argument-based validation tells you to aim evidence at the weakest assumption rather than pile it up where it is already strong. It does not prescribe an arithmetic. The minimum is the operating discipline proposed here, and it earns its place from what these four links are: they run in series. A quantity has to be captured correctly before its stability means anything. It has to be stable before construct relevance can be argued. All three have to hold before the contribution question is worth asking. Break one and every link after it has nothing to work with, so averaging lets the intact links pay for the broken one.
That series is what makes the minimum the right operator, and it is also what bounds it. The rule applies down a single pathway, where each inference depends on the one before it. It does not license reading one weak pathway as evidence against a strong one sitting beside it. A Dollars in Pockets figure sums several pathways: wages gained, fees avoided, time saved. Each has its own four links, so each carries its own minimum, and each keeps that grade. The composite claim is a separate sentence about the whole total, and it caps at the weakest pathway.
Averaging is not the only rival, and the strongest one deserves naming. If each link has some independent chance of holding, the product of the four is the joint chance the whole chain holds. That is an expected value, and it answers a real question: how likely is this claim to be true. The minimum answers a different one: what can this claim say out loud. It is a bound rather than an expectation, and a fund is entitled to say so.
The two rarely disagree by much. Multiply three strong links against one weak one and the strong links barely move the result, so the product lands just under the weak link’s own value. Minimum and product converge whenever the links that are not capping the claim are healthy, which is the case this essay is about. The difference is not in the number. A product gives a fund one more figure to report. A minimum gives it a word it has to defend.
Read the four as a chain rather than a scorecard and the asymmetry becomes obvious. The first three links are work a fund can do with its own data: define the measure, check it is not noise, argue it tracks what matters. Contribution is different in kind. It requires a comparison the fund’s own records cannot supply, and it is the link that has to carry the word produced. Averaging lets the first three pay for the fourth. The minimum does not allow that trade.
Medicine worked through a version of this two decades ago. GRADE, the system Cochrane, the WHO and many national guideline bodies use to rate evidence, does two things that matter here. It rates the certainty of the evidence separately from the strength of the recommendation (Guyatt et al., BMJ, 2008). And when it summarizes certainty across several outcomes, it generally takes the lowest of the ones judged critical (Guyatt et al., J Clin Epidemiol, 2013). The GRADE Working Group’s stated rationale is that the lowest rating among critical outcomes will generally provide an upper limit on certainty in the net benefit (Hultcrantz et al., J Clin Epidemiol, 2017).
So the minimum is a familiar move. A field with more at stake handles the same arithmetic problem the same way.
The analogy has a limit worth stating, because it weakens the rule proposed here. GRADE’s outcomes are entangled: mortality, hospitalization and adverse events all come from one intervention given to one set of patients, so weak evidence on one constrains what can be said about the balance. Impact pathways are separable. Whether legal fees were avoided for one group does not depend on whether housing improved for another. Taking the minimum across them is a harsher rule than GRADE’s, and a fund is entitled to say so.
The argument survives in a narrower form. Separability is a reason to keep the total out of the headline. But a published total still needs a label. Whoever repeats the total will attach a verb to it, and they will reach for the strongest verb on the page.
GRADE’s other half answers the question a program officer raises next. Low certainty does not forbid action. GRADE allows a strong recommendation on low-certainty evidence in five defined situations, though it treats these as rare, and audits find guideline panels reach for them more often than the guidance sanctions (Alexander et al., J Clin Epidemiol, 2016). Certainty governs what a fund may claim. It leaves the funding decision open. A pathway that earns only supports can still be worth funding, and both sentences can be said out loud without contradiction.
The responsible-AI frameworks a fund is likeliest to reach for do not close this gap either. The National Institute of Standards and Technology (NIST)’s AI Risk Management Framework (2023) names validity and reliability as the first characteristic of a trustworthy system, and the EU AI Act (2024) sorts systems into risk tiers with obligations scaled to the tier. Neither treats validity in the argument-based sense meant here. NIST does not define validation itself. It adopts the ISO 9000:2015 definition: confirmation that the requirements for a specific intended use have been fulfilled. That is a question about whether a system performs as specified. NIST is revising that framework, and no revised version has been published. Argument-based validity asks whether a reported figure can carry the decision resting on it. The same absence shows up in higher education, wherever a number is asked to justify a decision.
Validity tells you which verb the number earned
So far the four links have worked as a constraint. Here they start doing something a fund can use. Additionality is one link of the four, and the impact field already treats it as a question of its own: the Impact Management Project’s five dimensions (what, who, how much, contribution, risk) name contribution as a dimension in its own right, the one that asks whether the change would have happened anyway.
Three different counts are now in the essay, and they answer different questions. The five dimensions describe what a claim is about: they are the fields a complete impact statement has to fill in. The four links test whether the argument holds: they are inferences, and they are scored, not described. The three verb rows grade one link (contribution) by the strength of counterfactual standing behind it. Contribution appears in all three: fourth of the five dimensions, fourth link in the chain, and the thing the three rows grade. Most of what is hard about an impact number shows up there.
Dollars in Pockets already does counterfactual work at that dimension, which is what a strong verb is bought with. So the discipline does not slow the claim down. It tells a fund how much weight a given number can carry:
— No counterfactual behind the contribution link → the number supports, is consistent with, the outcome.
— No controlled comparison available, but rival explanations tested and defeated → contributed to is earned.
— A matched, quasi-experimental, or randomized comparison behind it → produced is earned. How far produced reaches depends on which of those it was; a matched comparison and a randomized one do not license the same sentence, and a fund should say which it has.
The middle row deserves more than a footnote. A controlled comparison is expensive, and the effects it is trying to detect often take years to compound. It also asks for a particular expertise: knowing how to set up a quasi-experiment, and knowing how to read what comes out of one. Most organizations do not hold that in-house, and cannot buy it at the scale their reporting asks for. So it is reasonable to expect that many funders and grantees are in the middle row. The cost of the method predicts it. It also means a great many impact claims will be settled in that row or not settled at all. That is the situation contribution analysis (Mayne, 2012) was built for: assemble the contribution story, then test it against the rival explanations that would make the outcome happen without you.
Mayne’s own logic is argument. Contribution analysis assembles a verified theory of change and accounts for the other factors at work. Mayne calls the result a warrant: a case a reasonable person would accept from the evidence and the reasoning. This essay has been borrowing that word.
Mayne does not run his method backwards to pick a verb. This essay does. Start with the strongest verb you would like to claim and let the rivals knock it down. Say the claim is that a record-clearing tool produced a wage gain. The rivals: the local labor market tightened that year, the person would have found a free legal-aid clinic within a few months anyway, the employers who hired them never ran a record check. Each rival you cannot defeat kills a stronger verb. Contributed to is what stands when you run out of rivals — the strongest verb none of them could defeat, which is a different thing from the strongest verb you can defend in a meeting.
“Here is our return, and here is how much of it we can stand behind, and where it is thin” is a stronger thing to say to a skeptical funder or board than a bigger number with no gate behind it. Validity is what survives the next question.
Keep the unit. Publish the warrant beside it.
The common unit is the achievement here, and nothing above argues for giving it up. Break the figure back into incommensurable pieces and the field loses what it spent years building. A funder is left where they started, holding a wage estimate and a fee estimate with no way to weigh one against the other.
The fix is smaller than that. A common unit makes dollars comparable. It does not make the evidence behind those dollars comparable. So report the composite, and report the warrant next to it, pathway by pathway.
Three of those four columns are things a fund already has. The dollars come from the model it ran. Where the pathway caps comes from the four questions. The verb comes from the evidence behind that link. This adds almost no new measurement. It asks a fund to show grading it already did.
There is a failure condition here, and it bounds the claim just made. Rashi Trivedi has spent a decade building monitoring and evaluation systems for governments, multilaterals and corporate foundations, and calls the advisory and assurance ecosystem around them the Framework Economy. She put the test this way: the table has to answer “What evidence allows you to call this produced, rather than supported?” If that takes minimal extra machinery, it is useful disclosure. If validating the warrant needs its own methodology, the table has become the next layer of that economy.
The check itself is not avoidable. Someone has to be able to test the verb, or the column is decoration. What the table can do is make the check cheap. The third column states where the pathway caps and on what evidence, so a reader who disagrees has something specific to argue with. That is reading, not auditing.
The table fails her test the day a fund needs a certified assessor to sign the warrant column. India’s regime shows how that happens. It regulates who assesses and stops short of what would make the inference sound. Empanelment for the warrant would be the same move, one layer up.
The gain is specific. A reader can see that most of the figure comes from a claim the fund can defend, and that a smaller share comes from a claim it cannot defend yet. That is more useful than a single number with a single verdict attached.
It also tells a fund where the next evaluation dollar goes, though not by pointing at the weakest row. A soft warrant on a small number may not be worth buying down. What matters is both at once: how large the number is, and how weak its warrant. That is the dollars column read against the claim column. In the table above, housing has the weakest verb and the smallest amount. Wage gain carries twice the dollars at a verb one step stronger. The table does not settle which to fund next, but it puts the tradeoff somewhere a board can argue about. One direction matters here. The fund publishes this about its own claim, not about its grantees. A gate a funder turns downward, asking applicants to arrive with counterfactual evidence, would sort organizations by evaluation budget rather than by what they do.
The table also shows where evaluation spending went nowhere. Under the minimum rule, only evidence at the link that caps a pathway can move the verb. Take a pathway capped at link four, contribution, because there is no counterfactual behind it. The fund then buys a better instrument, one that measures the outcome more precisely. That is link one, and link one was already strong. The cap is still link four, so the verb stays at supports. The fund paid to improve a link that was not holding the claim back. A dashboard of indicators does not show that. A table stating what each pathway’s claim rests on does.
Goodhart’s law has less room to work here. Suppose a fund publishes warrant per pathway. If it then grows its total by leaning on the softest one, the table shows a bigger number at a weaker verb. Anyone reading the table can see the trade.
GRADE has one more move worth borrowing, and it comes before the arithmetic. Not every outcome enters its summary rating. Only the outcomes a panel has judged critical to the decision are counted toward overall certainty; the rest are reported and set aside.
The same move raises a question a fund does not currently ask. Which of these pathways is the claim about?
If housing stability is central to what the fund is claiming, it counts, and it caps. If it is a secondary finding included because the data happened to exist, housing does not belong in the aggregate. The fund can still publish that number. It no longer caps a claim it was never central to. The declaration has to happen before the total is computed, and it has to be visible. That is what keeps it from working as an escape hatch: a fund that reclassifies its weakest pathway after seeing the cap has done something a reader can see.
For a small fund, the declaration does something else as well. Every pathway costs evaluation money and staff time, including the ones no decision depends on. A fund that declares its critical pathways first can see where that spending is worth it. It can also see where it has been paying for evidence behind a claim it never needed to make.
The blended row still costs a fund something. The dollars add without trouble. The claim does not add the same way, because one sentence about the whole figure asserts every part of it. A fund that says produced about a total containing an unexamined pathway has claimed more than its evidence supports. That is not an argument for reporting less. Publish the total and the three rows above it together, and the strong pathway keeps the verb it earned instead of losing it in the blend.
An invitation, taken up
The Dollars in Pockets authors asked the field to apply the measure to its own work and tell them where it breaks. This essay is one answer, offered in the same spirit and with real admiration for what they built. The metric is good, so what the field needs is not a better one. It needs a reflex: a fund asking “is that a valid dollar?” as automatically as it now asks “what is the ROI?”
The reflex is worth building now, while Dollars in Pockets is still a voluntary standard. The authors are already building for spread: they say they are developing AI-native tooling, in partnership with Anthropic, to make the measure easy to adopt. Tooling is how a measure becomes infrastructure, and a gate added to infrastructure afterward costs far more than one built in. Who runs the gate once the tooling spreads is a separate question. When the stakes are the mission argued that mission-driven organizations adopting AI rarely hold the evaluation capacity that discipline assumes, and that the capacity has to be built as a shared good rather than organization by organization. A validity gate on an impact number runs into the same constraint.
Mandating a measure does not make it valid. India’s regulated regime already runs the experiment. Rule 8(3) of the Companies (Corporate Social Responsibility Policy) Rules, as amended in 2021, requires large corporate spenders to commission independent impact assessment of their larger completed projects. India’s Social Stock Exchange, regulated by the Securities and Exchange Board of India (SEBI), requires social enterprises that have raised funds on the exchange to file an annual impact report assessed by an empanelled Social Impact Assessment Organisation. A September 2025 amendment tightened the form further, setting a floor of 67 percent of prior-year program expenditure. The same amendment allowed entities registered on the exchange without a listed security to file a self-reported report instead. SEBI consolidated the whole framework into a Master Circular in January 2026. Both are real obligations with real cost attached, and SEBI goes further than most regimes, with prescribed formats and empanelled assessors. But a regime can specify the form of an assessment without saying anything about the strength of the inference inside it. Rule 8(3) says only that the agency must be independent. It says nothing about what would make the resulting number sound. A common unit will spread faster than the discipline to audit it. The gate is cheapest to install before the number is everywhere.
The reach trap argued that a portfolio needs one comparable axis before anyone can say where the next dollar should go. Dollars in Pockets is that axis. Getting it does not settle what each dollar on it can claim. A magnitude, however rigorously built, is not yet a validity. The impact field needs a sounder path from the number to the claim, and it needs the weak link named early, while fixing it is still cheap.
Written August 2026 for the Analytic Bytes Library as a response to “Dollars in Pockets: A New Way to Measure What Matters” (NextLadder Ventures & GitLab Foundation, July 2026; contributing authors Kyle Nelson, Tamsin Chen, Rhett Dornbach-Bender, and Ellie Bertani), offered in the spirit of the authors’ own invitation to apply the measure and report where it breaks. Every description of Dollars in Pockets is drawn from that published announcement; the complete methodology had not been published at the time of writing. The Indian regulation text was read in a secondary brief quoting it verbatim, as SEBI’s attached PDF could not be opened directly. The failure condition in the warrant section was named by Rashi Trivedi in correspondence after publication and is used with her permission. The Y Analytics material is drawn from Evidence-based Impact in Climate (Y Analytics and TPG Rise Climate, November 2025), the firm’s own published account of its approach rather than an independent evaluation of it.
Questions, pushback, or a problem that looks like this one? Write to chai@analyticbytes.systems.