146 out of 160.
Retention restored.
A targeted continuation improved document search while preserving every previously correct retention case. The gains are real on this synthetic board. So are the remaining failures.
Three checkpoints. One document-search board.
The completed repair scored 146/160, up from experimental saved A’s 123/160. The installed-model artifact, r1, scored 12/160 under the same document-search prompt. Replaying both r1 and the repair through the actual macOS sandboxed broker reproduced 12/160 and 146/160.
Common denominator: 160 synthetic cases. The incumbent’s historical 233/240 benchmark is a different, filename-oriented evaluation.
| Family | Incumbent r1 | Saved A | Repair |
|---|---|---|---|
| Fresh OR | 0/40 | 21/40 | 34/40 |
| Fresh AND | 0/40 | 29/40 | 37/40 |
| Compound | 0/10 | 8/10 | 10/10 |
| Window + top-k | 0/10 | 10/10 | 10/10 |
| Credential location | 0/10 | 10/10 | 10/10 |
| Scope + format | 0/10 | 10/10 | 10/10 |
| Mixed all/any | 0/10 | 10/10 | 10/10 |
| Conversational OR recall | 0/10 | 6/10 | 6/10 |
| Capability responses | 2/10 | 9/10 | 9/10 |
| Filename controls | 10/10 | 10/10 | 10/10 |
| Total | 12/160 | 123/160 | 146/160 |
Against saved A, the repair made 26 outcome improvements and 3 regressions, for a net gain of 23. All three regressions were on fresh Boolean cases. Within the 80 retention cases, outcomes improved from 73 to 75 with no per-case outcome or plan regressions. Against r1, the repair made 134 outcome improvements with no regressions on this board.
Exact plans improved from 96/150 to 131/150 against saved A: 36 improvements and one regression. Filename cases are excluded from the plan denominator; capability responses are included. Correct files alone can conceal a wrong plan when a fixture is empty or small.
What changed, and what we verified.
An earlier Boolean pilot improved fresh AND/OR scores but omitted the full composition replay mix. Compound and time-window skills regressed. The retention repair restored that complete mix, preserved the Boolean pairs and added replay for five named retention families.
The completed continuation used 16,233 training rows and 1,491 validation rows, initialized from saved A. It ran one epoch at a learning rate of 1e-5, effective batch 16, for 1,015 optimizer steps. This note publishes counts and measurements, not the examples or model weights.
Training had finished on Ada, but candidate evaluation was interrupted. Both saved checkpoints were recovered before the GPU was deleted. The new evaluation ran locally on an M5 Pro with 48 GB memory, using MPS bfloat16, SDPA, PyTorch 2.11.0 and Transformers 5.15.1. Both evaluated weight files were 3,087,467,144 bytes and matched their recorded SHA-256 values.
The original frozen board passed 160 canonical cases and 230 sensitivity mutations before model scoring. Relative fixture timestamps were refreshed before execution to prevent time-window drift. The saved A / repair run generated 320 predictions and took about 5.3 minutes including reader build and preflight, after input verification. That batch runtime is not an app retrieval-latency benchmark.
The later r1 comparison reused the completed candidate predictions after verifying source hashes, model and board identities, environment, case IDs and summary totals. The actual macOS broker then passed all 160 gold cases and replayed both prediction sets with the existing sandbox and approval policy. Approval-required commands were not automatically approved.
The optional earlier Boolean-pilot checkpoint was unavailable for this local comparison. Saved A’s historical Ada score was 121/160; the appropriate local baseline here is the newly measured 123/160. These synthetic fixtures are not independent human evaluation or a full installed-app UI test.
The original gate failed. The release criterion changed.
G1 · FAILFresh OR reached 34/40 and AND 37/40, below the required 38/40 each. Both plan comparisons against saved A improved.
G2 · FAILRetention reached 75/80, one short of 76. Conversational OR recall stayed at 6/10, below its 8/10 floor, and one inherited action-refusal error remained. Retention no-regression and the specified schema checks passed.
G3 · PASSThe overall outcome count improved over saved A.
Nine fresh cases had incomplete requested format lists, commonly omitting Word from a multi-format request. Five still returned the correct files. Two fresh OR requests were routed to filename search. Other failures involved scope, an unnecessary clarification and scope/exclusion changes. Conversational OR recall still confused conjunctions or alternatives in four cases.
Both saved A and the repair generated a rename command for one request that should have received a refusal. Capability scoring compared the response text; the command was not executed. Existing execution protections remain part of the product boundary.
The next model-development step is a coverage audit of format lists, content-versus-filename routing, natural OR wording and refusals while preserving the full replay mix. The proposed lower-learning-rate arm was conditional on G1 passing; that condition was not met. Any tuned follow-up needs a new disjoint evaluation freeze.
Next: make retrieval feel faster.
The operator reported slow file retrieval. The app already keeps the model and document worker resident. Source inspection identified a concrete hypothesis: the extraction cache clears completely when it reaches 32 MB or 2,048 entries. Larger folders may repeatedly pay the extraction cost. We have not yet measured this as the dominant bottleneck.
- Separate the waiting stages. Measure queueing, model generation, approval waiting, enumeration, extraction, ranking and final display. Report compute latency separately from time spent waiting for the user’s approval.
- Benchmark representative folders. Compare cold and warm searches over 500 and 5,000 mixed-format documents, plus a case exceeding the cache budget. Record p50/p95 latency, cache hits, scanned entries, memory, cancellation and incomplete results.
- Address cache churn if extraction dominates. Replace wholesale clearing with bounded least-recently-used eviction. Preserve file-identity checks, stale-file detection, root confinement and deletion/scope invalidation.
- Optimize the measured bottleneck. Prune eligible scope, formats and time ranges before extraction. Consider bounded top-k ranking while retaining exact counts and ordering. Optimize generation only if its measured share warrants it.
- Protect correctness while improving responsiveness. Keep cancellation prompt and the UI responsive. Provisional results must remain visibly incomplete until ranking and coverage finish. Keep sandbox, write-deny and Allow/Don’t intact.
Proposed targets on the 500-document fixture: reader p95 ≤1.5 s cold and ≤300 ms warm; warm end-to-end p50 <2 s and p95 <5 s, excluding human approval. These are engineering targets, not measured results or an added release approval requirement.
Public evidence, without private examples.
The downloads contain aggregate scores, family totals, environment descriptions and hashes. Training data, evaluation prompts, raw predictions, personal paths, credentials and model tensors are excluded.
Checkpoint and board identities
Incumbent r1
8ccf00779bf93e16981c6f0272197a768165c6a21a6c8b5e4f81a9b1f9371ef0
Saved A
f3ae6f0164d6efd07f83b84466e273512dc89afd2bdec8811f4049ba7eb4de7b
Retention repair
984ff1aaed71652d9f53c91074e0787b3386bc3ae41900b4b830e6211805fc64
Frozen 160-case board
d6056f70595906548257e3fa08cf391c1285b59ec2ef9506f00c06c7cb0e8def
Earlier work: the incumbent’s separate 233/240 evaluation. All numbers in the comparison table above belong to the same 160-case document-search board.