SHLEX · 15 SEPTEMBER 2026 UTC

jeremiah_v1.
Four arms, twelve checkpoints.

c-1213 is jeremiah_v1, the operator-selected research incumbent: 1183/1280 shared diagnostic outcomes and 575/600 sealed acceptance outcomes, versus 267/600 for incumbent P0 and 305/600 for incumbent P1. The strict synthetic acceptance gate failed: credentials: 83/100, requires 85/100. Historical regressions remain tradeoffs; application release is pending.

The question

The previous continuation reached 240/240 synthetic transfer outcomes while older retention fell from 145/160 to 139/160. This experiment tests whether data composition, a lower learning rate or an earlier checkpoint can preserve both.

Matched experiment

All four full-SFT arms use Qwen2.5-Coder-1.5B-Instruct and start from the same September 9 incumbent with fresh optimizer state. Arms A/B use M0 replay; C/D use M1, with 8,000 new validated examples and 11,408 replay rows. The 8,000 new rows cover 7,290 unique semantic specifications, capped at four rows per specification. Both mixtures contain 19,408 rows. A/C use learning rate 1e-5; B/D use 3e-6. Each runs one epoch with microbatch 4, accumulation 4, seed 42, seeded shuffle, bf16 and maximum sequence length 1,792. Checkpoints at steps 303, 607 and 1,213 represent matched exposure.

One RTX 6000 Ada generated all diagnostics greedily. Incumbent P0 uses the current app context; incumbent P1 isolates the explicit environment-format context correction. The September 10 comparator and all candidates use P1. Commands execute through the Mac broker with independent fixtures and unchanged sandbox policy. Successful-empty tests also require a positive companion fixture.

Diagnostic results

Each cell shows executed outcomes · plan correctness. The 240 transfer, 160 retention, 280 wider and 600 development cases are selection diagnostics. They are not 1,280 independent acceptance cases.

Model/contextTransferRetentionWiderDevelopmentOverall outcomes
a-1213239/240 · 239/240153/160 · 143/160225/280 · 203/280392/600 · 368/6001009/1280
a-303235/240 · 235/240146/160 · 136/160218/280 · 193/280357/600 · 333/600956/1280
a-607239/240 · 239/240150/160 · 140/160227/280 · 205/280391/600 · 368/6001007/1280
b-1213211/240 · 210/240134/160 · 114/160216/280 · 192/280275/600 · 251/600836/1280
b-303196/240 · 196/240124/160 · 107/160213/280 · 190/280254/600 · 230/600787/1280
b-607206/240 · 206/240131/160 · 111/160215/280 · 191/280268/600 · 244/600820/1280
c-1213238/240 · 238/240132/160 · 114/160237/280 · 210/280576/600 · 576/6001183/1280
c-303237/240 · 237/240120/160 · 101/160235/280 · 213/280563/600 · 561/6001155/1280
c-607239/240 · 239/240139/160 · 124/160233/280 · 206/280572/600 · 572/6001183/1280
d-1213199/240 · 199/240128/160 · 111/160206/280 · 168/280490/600 · 466/6001023/1280
d-303193/240 · 193/240128/160 · 112/160178/280 · 138/280459/600 · 434/600958/1280
d-607199/240 · 199/240126/160 · 111/160204/280 · 165/280486/600 · 462/6001015/1280
incumbent-p049/240 · 40/240145/160 · 129/160229/280 · 213/280246/600 · 224/600669/1280
incumbent-p173/240 · 64/240139/160 · 127/160233/280 · 215/280263/600 · 238/600708/1280
sept10-p1240/240 · 240/240139/160 · 126/160238/280 · 218/280356/600 · 327/600973/1280

What changed under matched conditions

These are changes in executed outcomes at the final checkpoint. Positive numbers mean more passing cases. Lower-LR comparisons move from 1e-5 to 3e-6; mixture comparisons move from M0 to M1. Context-only compares incumbent P0 with P1. Each training arm has one seed, so these differences do not establish stability across repeated training runs.

ComparisonTransfer ΔRetention ΔWider ΔDevelopment Δ
lower lr m0-28-19-9-117
lower lr m1-39-4-31-86
new mixture lr 1e-5-1-21+12+184
new mixture lr 3e-6-12-6-10+215
context only+24-6+4+17

At the final checkpoint, lowering the learning rate from 1e-5 to 3e-6 changes M0 from 1009 to 836 outcomes and M1 from 1183 to 1023 outcomes, out of 1,280. This is a matched one-epoch comparison. It does not establish how the lower rate would perform with more training.

Selection and retention tradeoffs

The operator selected C final as the default jeremiah_v1, with D replacing it only if a D checkpoint strictly exceeds its total executed outcomes over the same 1,280 cases. A tie keeps C final. Better D checkpoints rank by total outcomes, development outcomes, total plans and then earlier step. Incumbents were evaluated on all four boards, including the new 600-case development board. The case counts define the weighting; they do not represent measured production traffic.

This choice was made after viewing A/B/C results. The original frozen rule required zero historical outcome or plan regressions and at least 38/40 passing both measures in each transfer family. Its results remain below as a separate audit. They are not the operator selection rule, and historical regressions are not erased by an overall improvement.

CheckpointRegressed historical casesWorst transfer familyEligible
a-12132439/40No
a-3033738/40No
a-6072439/40No
b-12135730/40No
b-3036723/40No
b-6076026/40No
c-12134539/40No
c-3034838/40No
c-6073839/40No
d-12137820/40No
d-30310820/40No
d-6078020/40No

C middle scores 1183/1280 and C final 1183/1280. Their retention scores are 139/160 and 132/160; development scores are 572/600 and 576/600. The operator explicitly preferred C final as the default. Equal totals do not mean identical capabilities.

Selected model: paired outcomes against incumbent P0

These counts compare the same requests case by case. Improvements are previously failed requests now passing; regressions are previously passing requests now failing. Aggregate gains do not cancel the practical cost of those regressions.

BoardImprovementsRegressionsNet change
transfer1890+189
retention518-13
wider2012+8
dev3300+330

Independent acceptance and limits

Acceptance opened: True. The selected nominee receives a paired historical MPS recheck followed by the sealed 600-case synthetic comparison. Both the original zero-regression rule and a historical check of positive paired gain plus 38/40 per transfer family are reported. Under the operator overall-performance policy, historical shortfalls are tradeoffs and do not withhold this single-nominee acceptance assessment. Synthetic acceptance requires 540/600 overall, at least 85/100 per stratum, positive paired gain and zero false acceptance on unsupported requests. The final C/D choice was locked before opening these 600 cases. Only that selected checkpoint and the two incumbent contexts were evaluated on acceptance; no second nominee is allowed. Independent human evaluation has not been collected.

historical on MPS

Model/contextOutcomesPlansFalse acceptanceFalse refusal
c-1213606/680560/680027
incumbent-p0426/680384/6803029
incumbent-p1449/680409/6802929

Selected model against incumbent P0: 212 improved outcomes and 32 regressions. Paired gain +26.47 percentage points; cluster-bootstrap 95% interval [+21.65, +31.79] points (451 clusters, 5,000 draws, seed 42).

filename-retention on MPS

Model/contextOutcomesPlansFalse acceptanceFalse refusal
c-121337/4034/4000
incumbent-p039/4038/4000
incumbent-p139/4037/4000

Selected model against incumbent P0: 0 improved outcomes and 2 regressions. Paired gain -5.00 percentage points; cluster-bootstrap 95% interval [-12.50, +0.00] points (40 clusters, 5,000 draws, seed 42).

These 40 cases are already part of the historical wider board; they are not additional independent evidence.

acceptance on MPS

Model/contextOutcomesPlansFalse acceptanceFalse refusal
c-1213575/600575/60000
incumbent-p0267/600242/60011104
incumbent-p1305/600275/600679

Selected model against incumbent P0: 312 improved outcomes and 4 regressions. Paired gain +51.33 percentage points; cluster-bootstrap 95% interval [+46.50, +56.00] points (300 clusters, 5,000 draws, seed 42).

Acceptance strata

Stratumc-1213incumbent-p0incumbent-p1
boolean99/10043/10052/100
capability100/10063/10074/100
credentials83/1005/10010/100
filename96/10088/10090/100
scope_format98/10044/10042/100
time_order99/10024/10037/100

Next steps

Use development diagnostics to design the next intervention. Retire the opened acceptance set from independent testing and author a fresh sealed set before the next experiment.

Credentials is the first follow-up priority: 15 of its 17 failed acceptance cases involve a changed ordering field. Use contrastive examples that preserve requested ordering while varying provider and credential wording, and separately exercise limit, all-results and file-type constraints. Field counts can overlap; this is an observed output-error pattern, not a claim about the model's internal reasoning.

For a subsequent training experiment, prioritize the measured regression patterns below. Collect contrastive examples that change one requested field at a time, such as file type, output mode or scope, and pair each with retention replay. Keep the selected checkpoint fixed as a baseline and test one intervention at a time at matched exposure. Treat the opened acceptance set as diagnostic evidence from now on. Freeze fresh development and acceptance data before comparing additional variants, and keep the new acceptance set sealed while tuning on development.

Field errors among selected-model outcome regressions against incumbent P0 (a case can contribute to multiple fields): {"all_omitted": 1, "invalid_document_plan": 6, "scope_changed": 2, "types_invented": 11, "types_omitted": 11, "wrong_output_mode": 7}.

The requested model name is jeremiah_v1; its status is research incumbent selected. The research-model selection does not deploy the model into the installed application. Mac replay ran concurrently with the fixed GPU queue. All twelve full checkpoint files were recovered to the Mini with recorded SHA-256 verification before GPU deletion. Laptop copies were verified before final aggregation. Estimated Ada compute cost: $7.43 (controller estimate; final provider billing may differ). Public evidence includes aggregates and hashes only; corpora, private prompts, raw predictions, fixtures and weights remain private.