Validating AI-Generated Officer Coaching for Community Corrections Supervision Visits
Cliff Hurt Johnston,
Ted Green,
Madeline Warren and
Valerie Meade
No m7d5g_v1, LawArchive from Center for Open Science
Abstract:
Meta-analytic evidence links sustained officer fidelity to core correctional practices with substantially lower participant recidivism; caseloads supervised by untrained officers recidivate roughly 38% more often than those supervised by trained officers (Chadwick et al., 2015), and fidelity after training — not the training event itself — is the moderator that makes the difference (Labrecque et al., 2023). Our companion study validated the measurement layer: a frontier large language model can grade recorded supervision visits against a standards-based rubric about as reliably as expert human auditors (Green, Warren, & Meade, 2026). This paper tests the practice-change layer: whether an LLM can turn those transcripts and scores into officer-facing coaching that an independent expert judges to be as good as coaching written by expert human reviewers. On the same 30-visit corpus, we assembled 120 coaching outputs — four per visit, from two expert human sources (authentic prior grader coaching notes) and two LLM sources (Claude Sonnet 4.6 and GPT-5.5), randomized within visit — and had a third expert, Valerie Meade, score every output on eight pre-specified 1–5 dimensions, rank the four outputs within each visit, and make a product-display decision. The review was blinded: source identities and source types were withheld until scoring was complete, and the reviewer’s own historical coaching was excluded from the packet. Decoded against the internal source key, the strongest LLM source (Claude Sonnet 4.6) was rated equal to or better than both directly compared human sources on the composite quality index (vs. one human source +0.60, Holm-adjusted p = 0.006; vs. the other +0.28, not significantly different and formally non-inferior at a 0.4 point margin, one-sided p = 0.0002), earned the best mean within-visit rank (2.07 of 4), the most first-place outputs (10 of 30), and the most display-ready outputs (13 of 30). Pooled by type, LLM coaching outscored human coaching on the visit-level composite (3.44 vs. 3.14; paired difference +0.30, 95% CI +0.10 to +0.52, p = 0.019). We report the result with its boundaries: one expert reviewer, two directly compared human sources, non-standardized human comparison text, a 30-visit corpus, and no claim of field behavior change or recidivism impact. Together with the grader validation, the finding supports the working thesis that agencies can responsibly act on validated visit scores with display-filtered LLM-generated coaching — the mechanism the training literature ties to lower recidivism — while the outcome link itself awaits field study.
Date: 2026-09-08
New Economics Papers: this item is included in nep-exp
References: Add references at CitEc
Citations:
Downloads: (external link)
https://osf.io/download/6aa0164588301b4487486271/
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:osf:lawarc:m7d5g_v1
DOI: 10.31228/osf.io/m7d5g_v1
Access Statistics for this paper
More papers in LawArchive from Center for Open Science
Bibliographic data for series maintained by OSF ().