Get the Free Toolkit

The Field Graded Its Own Programs. Then It Filed the Results.

2gen economy fair chance hiring household stability workforce development Aug 03, 2026
Comparison of two five-domain instruments: the field's program instrument scores Program Leadership, Staff Characteristics, Offender Assessment, Treatment Characteristics, and Quality Assurance; the Durability Index scores Employment Retention, Housing Stability, Financial Resilience, Family Connection, and Justice-System Stability. One list describes a program. The other describes a household.

In October 1973, the United States Board of Parole adopted the Salient Factor Score as its primary decision instrument.

It scored the person. Prior convictions, age at first offense, employment history. A number came out, and that number helped decide whether someone went home.

Fifty-three years later, the instruments are more sophisticated and the object has never moved. The LSI-R scored the person. COMPAS scores the person. Every generation points the same direction: at the individual, measuring readiness, risk, and compliance.

Here is the question that took me years to ask properly. In all that time, who scored the program?

The Field Already Answered That Question

It did, and precision matters here, because the easy version of this argument is false.

The field does have program-quality instruments. The Correctional Program Assessment Inventory. The Correctional Program Checklist. The Standardized Program Evaluation Protocol. Trained assessors run site visits, interview staff, observe service delivery, and score programs against a rubric. Hundreds of programs have been assessed this way.

So the honest claim is not that nobody looked.

The honest claim is harder: they looked, they wrote down what they found, and the field built no mechanism that acted on it.

Five Domains, Five Domains

Look at what the field's premier program instrument actually measures.

The Correctional Program Checklist rates programs across five domains: Program Leadership and Development. Staff Characteristics. Offender Assessment. Treatment Characteristics. Quality Assurance.

Leadership. Staffing. Assessment. Treatment. Quality control.

Four of the five describe the program's own machinery. The fifth, the one that names a person, scores the program on how well it categorizes people. Even the participant-facing domain is about the program's sorting ability.

The revised Durability Index also carries five domains: Employment Retention. Housing Stability. Financial Resilience. Family Connection. Justice-System Stability.

Same count. Same architecture. Every single one describes a household.

One instrument asks whether a program follows the model. The other asks whether a family got more durable. A program can score high adherence on the first and leave every household exactly where it found it.

What Happened When the Field Looked

In 2006, Lowenkamp, Latessa, and Smith published a study of 38 Ohio halfway house programs serving people on parole in fiscal year 1999, matching 3,237 participants against 3,237 comparison cases on county, sex, and risk level.

In 28 of the 38 programs, 73 percent, the comparison group recidivated at rates equal to or lower than the people who went through the program.

That study is twenty years old and describes one state's cohort from one fiscal year. It does not prove programs fail. It proves something more specific about how this field operates.

The measurement existed. The finding was published in a peer-reviewed journal by respected researchers using the field's own instrument. And then what?

No program lost funding on those grounds. No agency restructured its contracts. The finding entered the literature and stayed there.

Four years earlier, three of the field's own leading scholars published "Beyond Correctional Quackery" in Federal Probation, arguing that corrections defaulted to tradition and common sense instead of research. That was twenty-four years ago. The diagnosis is not new. The consequence has never been built.

The Difference Between Measurement and Accountability

We measure individuals continuously and act immediately. A risk score changes a release date, a supervision level, a program assignment. That measurement has teeth.

We measure programs occasionally and act almost never. The score becomes a report, the report becomes a recommendation, the recommendation becomes a technical assistance plan, and the program continues.

Measurement without consequence is not accountability. It is documentation. Their score follows them. The program's score follows no one.

This Is Not My Idea

In 2022, the National Academies of Sciences, Engineering, and Medicine published a consensus study called The Limits of Recidivism: Measuring Success After Prison.

Its second recommendation reads: researchers should "develop and validate new measures to evaluate post-release success in multiple domains, including personal well-being, education, employment, housing, family and social supports, health, civic and community engagement, and legal involvement."

The committee credited its formerly incarcerated members with pushing that conclusion.

That is the National Academies saying, in print, that the field needs multi-domain trajectory measures and that the people who lived through the system were right about it. The Durability Index is one answer to a standing national recommendation. It is not one architect's grievance.

The Objection I Owe You Before You Raise It

If you have spent time in workforce evaluation, you know the problem with outcome instruments, and you should have my answer before you finish reading.

In 2002, Heckman, Heinrich, and Smith studied the performance standards attached to the Job Training Partnership Act. They found that systems built on short-term participant outcomes, with no counterfactual, create what the field calls cream skimming: programs select people who help hit the short-run numbers, people who in the authors' words "would have done well without it," rather than the people who would gain most.

Those short-run measures, they found, were weakly and sometimes perversely related to long-run impacts.

An instrument that scores households at month 18 and stops there reproduces that failure exactly. It rewards programs serving the most stable families and punishes those serving the hardest cases, which is the opposite of what this work is for.

So the design principle, stated in public where you can hold me to it: the Durability Index scores change from an intake baseline, not attainment at exit. A program serving households that begin in Crisis and move them to Fragile has produced more durability than a program serving households that arrive Stabilizing and stay there. The composite is read against where the household started, and any implementation that scores end-state alone is a misuse of the instrument.

For Funders

Ask every program in your portfolio for its program-quality assessment and its date. Many will have one. That is the point.

Then ask the second question, the one that has never had a good answer: what changed as a result of that score? If the answer is a technical assistance plan and nothing else, you are funding documentation.

Your ninety-day move: attach one consequence to one assessment in one contract renewal this quarter. Not a punishment. A remediation timeline, a scored rebid, a funding tier tied to household outcomes at month 12. One contract is enough to learn from.

For Policymakers

The National Academies asked federal agencies and foundations to fund multi-domain post-release success measures. That recommendation is four years old and largely unfunded.

Your ninety-day move: identify one state contract or grant program where success is currently defined by recidivism alone, and add a household-trajectory reporting requirement alongside it. Do not replace the recidivism measure. Pair it, so that the day arrives when the two numbers disagree and someone has to explain why.

For Operators

Pull your last program assessment, whatever instrument it used, and read the recommendations section.

Count how many were implemented. Count how many were funded. Count how many changed what a household experienced.

Then do what the assessment could not: pick ten households you served eighteen months ago and find where they are now across all five domains. That number is not in your grant report, and it is the only one that tells you whether the work worked.

What Changes When the Instrument Turns Around

We score people and act on it immediately. We score programs and act on it never.

We measure adherence to a model. We do not measure whether a family got more durable.

The instrument that scores the household turns measurement into accountability, because it produces a number a funder, a board, and a family all read the same way.

Fifty years of instruments have pointed at the person. It is time one pointed back.

Take the free Durability Index Self-Assessment. Twenty questions. Five domains. One score.

Take the Assessment →

Until next time, keep building what they said couldn't be built.

Khalil Osiris

Author & Founder, Khalil Osiris Consulting | Market Architect, 2Gen Economy Workforce Ecosystem | Fair-Chance Hiring · Household Stability · Workforce Durability | Publisher, The Durability Economy

Subscribe to The Durability Economy for workforce redesign, fair-chance hiring, and household stability. For leaders who measure what lasts.

Change the metrics. Change the outcomes.

Subscribe to The Durability Economy

Get the weekly briefing on workforce transformation, criminal justice reform, and the 2Gen Economy, delivered straight to your inbox. Evidence over ideology. Households over headcounts.

We hate SPAM. We will never sell your information, for any reason.

Ready to Build Household Stability?

Let’s discuss how the Durability Economy Architecture can transform your community.