The Field Ran the Experiments. Nobody Built the Instrument.
Aug 17, 2026
Here’s what the data shows.
Moving to Opportunity (MTO) followed 4,604 households for ten to fifteen years, moving families from high-poverty neighborhoods to low-poverty ones and tracking adult earnings and employment.
No detectable effect on the adults.
The same experiment, reread through tax records years later, found that children whose families used a low-poverty voucher to move before they turned thirteen earned roughly 31 percent more in their mid-twenties. The effect surfaced only when researchers changed whose outcomes they measured, and when.
What you choose to measure decides what an experiment can see.
The Health Profession Opportunity Grants ran across dozens of sites with a randomized design. Family Rewards paid families for meeting health, education, and employment benchmarks. CareerAdvance paired parent workforce training with children’s early education, evaluated against a matched comparison group.
Null or mixed on adult employment and earnings.
Mixed means exactly that. At roughly three years, HPOG 2 moved credential receipt up 14.3 percentage points and health care employment up 6.9 points, while its 94 dollar quarter-twelve earnings effect went undetected. Real movement on intermediate measures, nothing the earnings table could call durable.
The Transitional Jobs Reentry Demonstration put justice-impacted workers into subsidized jobs. Employment rose while the subsidy lasted and fell back when it ended. At twenty-four months: no significant effect on unsubsidized employment or on the principal recidivism outcomes.
These are not small studies. They are well-funded, carefully executed, several of them randomized. The field ran the experiments it said it needed.
Before going further, what I am not claiming
These studies measured household outcomes.
Family Rewards tracked household income, poverty, material hardship, and savings. The Family Options Study measured five domains of family wellbeing. HPOG 2 reports monthly household income, poverty, emergency-funds capacity, food security, housing security, and public assistance receipt. MTO reports household income and public assistance.
Anyone who says the field never looked at households is wrong, and will be corrected in public.
What none of them built was a composite. Not one constructed a single cross-domain durability score and tracked it for the same household across twenty-four months.
That distinction is the whole argument. Everything below rests on it.
What the field concluded
That these interventions do not work.
That the evidence base for 2Gen programming is thin. That transitional employment produces temporary gains and nothing durable. That neighborhood mobility does not move adult economic outcomes.
That conclusion is what a decade of briefing books carried.
The wider literature echoes it. The largest meta-analysis of active labor market programs, 207 studies and 857 estimates, found effects near zero in the short run and more positive two to three years after completion.
The zeros got the headlines. The window where effects emerge is the one funding decisions never wait for.
The result that complicates it
The Center for Employment Opportunities trial does not fit the pattern, and the honest version of this argument has to say so.
CEO significantly reduced recidivism. MDRC reports relative reductions of 16 to 22 percent among participants who enrolled within three months after release, and 58.1 percent of the program group was ever incarcerated across the three-year follow-up against 65.0 percent of the control group.
The employment effects still faded. The recidivism effects were concentrated in the first year, and the cumulative difference across three years remained significant.
So the picture is not uniformly null. It is mixed in a specific way: the outcome that moved was the one measured furthest from the job placement.
Hold that thought.
Why a composite is a different instrument
A study reporting household income, housing security, and public assistance as three separate outcomes can tell you whether each one moved.
It cannot tell you whether the household became more stable.
A household whose income rose while its housing stability collapsed reads as a partial success in a table of separate outcomes. In a composite with the weakest domain visible, it reads as what it is: a household that is not holding.
The separate-outcomes design cannot detect the thing the intervention was built to produce, because durability is a property of the household as a system, not of any single measure taken alone.
That is why CEO’s recidivism finding matters. It was the outcome furthest from the placement metric, and it was the one that moved.
What this costs, and who is paying for it
Here is the ROI math, with its limits stated up front.
No primary source publishes a total evaluation expenditure for MTO, HPOG, TJRD, or the CEO trial. Not one. I searched; no consolidated figure has been published.
What is published is fragmentary. A single HPOG 2 task order to Abt was worth up to $7,635,864. TJRD’s program cost was $4,897 per participant in 2024 dollars. CEO’s gross program cost was $4,263 per program group member.
One more absence worth naming: no primary source publishes the marginal cost of extending a follow-up from twelve to twenty-four months. Anyone who claims the second year of measurement is too expensive is citing a figure nobody has published.
And this, which should stop you:
The Department of Labor’s total evaluation investment in FY 2022 was $23,180,000, or 0.16 percent of its $14.2 billion discretionary budget. Its evaluation office budget request has since fallen from $11.5 million for FY 2024 to $4.3 million for FY 2026.
So the field makes measurement decisions on an evidence base whose total cost nobody has published, funded by an evaluation budget that rounds to zero against program spending, and whose requests are shrinking.
The assumption is transparent: an evaluation line against a discretionary budget, different kinds of money. The point is not the ratio. It is that we do not know what we spent to learn what we think we know.
For funders
Ask what the outcome structure was, not only what the horizon was.
An evaluation reporting six household measures separately is not the same as one scoring the household. A table of domains cannot tell you whether households became more stable, only whether individual measures moved.
In the next 90 days: on one grant you already fund, require a composite household score at intake and month twelve. Not six separate metrics. One number, with the weakest domain visible.
For policymakers
The null results in your briefing books are real findings about the questions those studies asked.
They are not findings about household stability, because no study in that stack constructed a durability measure. Treating them as settled evidence against household-centered design reads a conclusion the research never reached.
In the next 90 days: before the next reauthorization cycle, ask whether any evaluation in your evidence base produced a cross-domain composite. If the answer is none, that is the finding.
For operators
You can have composite household data nobody else has within two years.
Five domains. Twenty sub-measures. One composite from 5 to 25, unweighted, with the weakest domain always visible. Scored at intake, six, twelve, eighteen, and twenty-four months.
In the next 30 days: score your current cohort at intake. That is your baseline, the thing the flagship evaluation literature is missing.
What comes next
I will be precise about the claim, because a hostile reader deserves it.
Household-level measurement has been done. Composite household stability measurement, tracked longitudinally as a primary outcome in a flagship evaluation, has not.
Practitioner instruments exist: self-sufficiency matrices in Snohomish County, Arizona, Colorado, and the Netherlands. None has been deployed as the primary outcome measure in a major randomized evaluation.
And these are not sketches. Colorado’s Family Support Assessment 2.0 covers fourteen domains with peer-reviewed interrater reliability between 0.79 and 0.96. The Dutch matrix has peer-reviewed psychometrics. What is missing is not the tool. It is the decision to make the tool the outcome.
The Padua trial comes closest. It randomized 427 participants, surveyed them at twelve and twenty-four months, the exact cadence I am arguing for, and printed a self-sufficiency matrix in its appendix. Then its impact analysis reported six separate domains and never combined them.
The trial carried the instrument to the door and left it there.
The field ran the experiments. It ran them honestly and at scale. It reported domains separately and asked whether the adults earned more.
Nobody built the instrument that would have answered the question.
Take the free Durability Index Self-Assessment. Twenty questions. Five domains. One score.
Until next time, keep building what they said couldn’t be built.
Khalil Osiris
Author & Founder, Khalil Osiris Consulting | Market Architect, 2Gen Economy Workforce Ecosystem | Fair-Chance Hiring · Household Stability · Workforce Durability | Publisher, The Durability Economy
Subscribe to The Durability Economy for workforce redesign, fair-chance hiring, and household stability. For leaders who measure what lasts.
Change the metrics. Change the outcomes.
Subscribe to The Durability Economy
Get the weekly briefing on workforce transformation, criminal justice reform, and the 2Gen Economy, delivered straight to your inbox. Evidence over ideology. Households over headcounts.
We hate SPAM. We will never sell your information, for any reason.