Skip to all papers
PorkiCoder Research

Every result.
Nothing buried.

Seventeen published studies and product notes, summarized side by side. Each result keeps its limits, full paper, and public evidence attached, including findings that were later superseded. Tab-namer sessions 1-11 and the Bash SFT series are separate.

17 published studies and notes 9 tab-title studies 5 Bash SFT notes 1 Shlex evaluation note 1 matched-control study 1 product architecture note
Complete library

All published results.

Newest first, with the headline result and claim boundary visible before you open a paper.

Terminology as of 16 August 2026. Hybrid A means only the 2+1 glue rule. The production tab-title stack is mid-GSG B-9500 + beam-4 + centroid v2 + Hybrid A glue; older papers used the name for a complete greedy-era system. SO-board and Terminal-board are separate leaderboards; do not mix their FLAN numbers. Bash SFT is a new series starting 21 August 2026, separate from the tab-title namer notes.

  1. Shlex · Best-fit search

    Shlex 0.9.0: nothing you type is required, and a small model that offers other words

    A file search that required every typed word found less the more you said. Ranking by how much of the request a document holds, a board written and blind-reviewed by another model, a bake-off of four sources of meaning, and a 1.5B model trained for under two dollars to offer other words for the same thing.

    Completed experimentOn a board nobody who built the fix wrote, the holdout half goes from 19 of 117 cases passed with Shlex 0.8.4 to 62 of 117 with 0.9.0, by executed and ranked tests; every older board holds.
  2. Shlex · Unseen board

    Jeremiah v4: rows for three string rules, measured on a 312-case unseen board

    560 reviewed rows for the string rules a small model kept getting wrong, one pass on a rented GPU, and a 312-case board of never-seen providers and topics with confidence intervals, failure clusters and next steps.

    Completed experimentv4 step 856 passes 284/312 (91%) of a 312-case board whose providers, topics, repos and nouns were never trained on (95% interval 87 to 94%); the model in the app passes 239/312 (77%) (72 to 81%). On the unseen v3 development cases 856 passes 88/96 (92%) against 76/96 (79%); canaries 2/2 (100%); zero false acceptances and zero unsafe executions on every board. The v2 development and regression boards are not yet scored, so no snapshot has passed the full gate set; jeremiah_v4 ships in Shlex 0.7.0 and 0.7.1 as an operator-authorized field test. The run cost 0.36 USD on a rented RTX 6000 Ada, created on its first attempt. Update, the same day: the plural and spelling rule this note proposed now runs in the app (Shlex 0.7.1), not in the model. Replaying the same cached model lines through it, v4 step 856 passes 304/312 (97%) of the unseen board (95 to 99%) and the model of the previous round passes 289/312 (93%), more than this run's training gave it; no case newly fails for any of the four models and there are still zero false acceptances and zero unsafe executions.
  3. Shlex · RunPod

    Jeremiah v3: more rows for the weak families, trained on a rented RunPod GPU

    1,200 new rows, one pass on a rented GPU, six models on the same executed tests, and the friction log from changing GPU providers.

    Completed experimentv3 step 825 passes 77/96 (80%) of the unseen v3 development cases (the model in the app: 56/96 (58%)) and 367/403 (91%) of the v2 development board (app model 358/403 (89%)); canaries 2/2 (100%); regression 1246/1330 (94%) against 1225/1330 (92%). No snapshot passed every gate; nothing is nominated. The run cost 0.38 USD on a rented RTX 6000 Ada; four failed launches added 0.19 USD.
  4. Shlex · Content-first

    Jeremiah v2: teaching Shlex to find documents by what they contain

    Two matched arms, six snapshots, every result an executed test on the real broker.

    Completed experimentrepair-750 answers 2/2 operator canaries (baseline 0/2) and 82/108 (76%) of the held-out content-first development cases (baseline 20/108 (19%), control 27/108 (25%)); regression 1225/1330 (92%) versus baseline 1208/1330 (91%). No snapshot passed every gate; nothing is nominated.
  5. Shlex · Retention matrix

    jeremiah_v1: Shlex’s four-arm retention matrix

    Matched data mixtures, learning rates and checkpoint timing, with same-Ada diagnostics and real Mac execution.

    Completed experimentc-1213 is jeremiah_v1, the operator-selected research incumbent: 1183/1280 shared diagnostic outcomes and 575/600 sealed acceptance outcomes, versus 267/600 for incumbent P0 and 305/600 for incumbent P1. The strict synthetic acceptance gate failed: credentials: 83/100, requires 85/100. Historical regressions remain tradeoffs; application release is pending.
  6. Shlex · Latest evaluation

    Shlex live repair: 240/240 transfer, but retention regresses

    Same-Ada predictions replayed through the real macOS broker: new transfer improves 49/240 to 240/240; older outcomes fall 145/160 to 139/160.

    Release gate failed15 regressed older cases · candidate not promoted

    Synthetic transfer gains do not establish general reliability. Full replay retained; GPU deleted after verified checkpoint recovery.

  7. Shlex · Latest evaluation

    Shlex retention repair: 146/160, with the old composition skills retained

    The completed local repair improves saved A’s 123/160 to 146/160. The r1 incumbent scores 12/160 on this document-search board; actual sandboxed broker replay reproduces the comparison.

    Latest result 146/160 outcomes · 131/150 exact plans · zero retention regressions

    Three fresh-case outcome regressions versus saved A. The original gate failed; the operator’s revised better-than-incumbent release criterion is met. This is synthetic document-search evidence, not the separate 240-case filename benchmark.

  8. Bash SFT · Note 05

    Bash SFT note 05: 233/240 release candidate on the human-like board

    A new independently authored 240-item set is the only selection number: 233/240 versus 230/240 for the prior champion and 25/240 for the previous production model. The targeted continuation is the Shlex release candidate.

    Release candidate 233/240. We ship run-shlex-v31r1.

    Unseen human-like hard set, Ada greedy generation, Darwin execution, frozen before generation on it. Older generator-matched boards are not used to pick the ship candidate. Product wiring is a later session.

    • 233/240 release candidate
    • 230/240 prior champion
    • 25/240 previous
    • no further training
  9. Bash SFT · Note 04

    Bash SFT note 04: live-style training dropped 92/98 to 82/98

    One continuation from the 92/98 checkpoint on 2,240 shorter prompts scored 82/98 on the frozen macOS set and 2/27 on a new live-style set. The model still invents a folder name instead of using a dot.

    Published result Do not ship. Keep last40 wd=0.1 at 92/98. Continuation 82/98 frozen, 2/27 live-style.

    Arm 0 reproduced 92/98. Envelope-only dropped the incumbent to 86/98. Training JSONL is unpublished. The Ada box was deleted after a hashed weight pull.

    • keep 92/98
    • continue 82/98
    • live-style 2/27
  10. Shlex · Candidate report

    Shlex live-query training: 79/80 synthetic, 4/13 human

    Faithful conversational SFT transformed the synthetic live-language score. A frozen canary from the real app showed that the gain did not generalize far enough.

    Release decision Do not ship. Synthetic live-style rose from 7/80 to 79/80, but human-query usefulness reached only 4/13 against a predeclared 8/13 gate.

    All 13 commands were policy-valid and none escaped the chosen folder. One false refusal and one secret-sensitive search still vetoed deployment.

    • 79/80 synthetic
    • 4/13 human
    • not shipped
  11. Bash SFT · Note 03

    Bash SFT note 03: file search reached 92/98, but more data did not help

    Two experiments answer different questions. A diverse training mix performed best at general Bash imitation. For executable file-search, the kept checkpoint passed 92 of 98 locked macOS tasks; more traces and repeated hard cases did not improve it.

    Published result The kept file-search checkpoint passed 92/98 tasks on one attempt each. The full general-Bash mix had the lowest validation loss, 0.7180.

    These are separate scoreboards. Test the checkpoint on real searches before renting more GPUs. The historical handoff canary remains 0/1, and the training JSONL is unpublished.

    • 92/98 locked
    • mix A 0.7180
    • fair baselines 90/8/4
  12. Bash SFT · Note 02

    Bash SFT note 02: 3,000 executed examples, four arms, and a fair-base correction

    Cumulative Session 2 report. We corrected an unfair base-model interface comparison with symmetric Markdown normalization and an identical contract-calibrated prompt. The large in-domain gap remained; external transfer did not.

    Published result Fair contract-calibrated comparison: stock 8/98 versus Arm 2 at 89/98 on the same locked tasks.

    The original strict 0/98 is retained only as a product-contract audit. External historical canary: 0/1 for all six tested models. No SHELLper comparison.

    • 3,000 executed rows
    • 89/98 fair-prompt SFT
    • 8/98 fair-prompt stock
    • 0/1 external
  13. Bash SFT · Note 01

    Bash SFT note 01: a 4.5-second train and a 1.8-point exact bump

    Research note, not a paper. First note in a new series. Tiny mix, destroyed the train GPU before scoring, asked for extra GPUs that would not have helped. Exact bash 12.2% to 14.0%. Text F1 is mostly the SFT wrapper.

    Published result One-seed overall point estimates, n=500: exact bash 12.2% to 14.0%, stem 50.4% to 56.8%, tool 86.4% to 91.6%, bash F1 41.5% to 43.8%.

    Not executable pass@1. Cloud n=9 stayed 0% exact. No significance test. Weights not published.

    • 137 train rows
    • exact +1.8 pp
    • CUDA RTX 6000
  14. Tab titles · Terminal-board

    Session 11: more epochs on the Aug-12 mix

    Continue Title-SFT FLAN on the leak-checked Aug-12 mix. flat1e4 is the ship. Hybrid A stays off. Next training move is more data, not another lr sweep.

    Published result Ship: Title-SFT FLAN raw flat1e4. Holdout 1000 7.59 / 85.7% versus Session-10 ship 7.38 / 83.5%. Holdout 1000b 7.55 / 84.5% versus 7.34 / 82.5%.

    35M distill shelved. Session-9 schema-argmax continue-SFT lost. Holdout 200 stays closed.

    • flat1e4 ships
    • holdout 7.59
    • no Hybrid A
  15. Tab titles · Terminal-board

    Session 10: every namer vs the clock

    Title-SFT FLAN raw is the ship candidate. Confirm-v1 labels beat P1. Size and 29 ms are accepted. 6t stays wired until the SFT worker exists.

    Published result Ship: Title-SFT FLAN raw, 6.55 / 71.3% on confirm-v1 335, 7.28 / 87.8% on ship-board 500. P1 6.41 / 6.79. Hybrid A on this SFT lost. +193 MB, 29.3 ms.

    Honorable mentions: P1, 6t, CPU picker, neural ranker. Holdout 200 stays closed.

    • SFT raw ships
    • confirm 6.55
    • no Hybrid A
  16. Tab titles · Terminal-board

    Session 9: Dense P1 + Hybrid A beats 6t

    First look at sealed Terminal-v2 (516 unused rows). Same Hybrid A glue on ten arms. The intent-pointer head is the new Terminal best versus shipped 6t. FLAN is still ahead.

    Published result P1 + Hybrid A 6.43 beats 6t 6.30 (+0.13). Loses to FLAN 7.02 (−0.60).

    3.1× fewer live parameters than title-tuned FLAN-T5-small (24.5M vs 77.0M). Keep-namer and free decode lose. Electron still loads 6t ONNX.

    • +0.13 vs 6t
    • −0.60 vs FLAN
    • 24.5M vs 77M
  17. Tab titles · Terminal-board

    176k Flash-Lite CE: quantity did not beat 6t

    175,642 Gemini 3.5 Flash-Lite titles on four 35M learning rates. After the production beam-4 + Hybrid A stack, the best continue still lost to 6t_argmax and to FLAN.

    Published result Ship stays 6t_argmax. Hottest CE inverted the judge.

    Packet 2, same 300 Terminal-board tasks: A 6.05 versus 6t 6.16 versus FLAN 6.62 (paired A−6t −0.112). Arm B reached CE 0.91 and still finished last. This run does not replace the ship weights.

    • 6.05 vs 6.16
    • −0.112 vs 6t
    • 175,642 rows
  18. Tab titles · Terminal-board

    Techniques log: what moved the namer, and what did not

    A single page of every Terminal-board arm after mid-GSG, uniform mean / useful-rate tables, bar charts, and 115 glued titles so you can read pin → r3 → 6t → FLAN yourself.

    Published result Ship 6t_argmax. min10 (train only on ≥10-word tasks) is tied with it at n=1000. Later hinge, glue, and ranker plays were flat or negative.

    6t_argmax is about +0.12 versus the pin on Terminal-board. FLAN + centroid v2 is still 7.11. Dead arms are listed so we do not rerun them.

    • +0.12 vs pin
    • 115 titles
    • 7.11 FLAN bar
  19. Tab titles · Terminal-board

    Session 4: we beat our own namer. FLAN is still ahead on terminal.

    Judged-SFT on 4,338 terminal tasks. Copying FLAN titles that can be spelled from the page nudged the pin; copying our own best-of-K guesses did not.

    Published result r3 beats the pin on Terminal-board in two blind packets. The SO-board near-miss is a different race.

    Confirmation packet: 5.67 versus 5.56, paired +0.12. Session 2.5 FLAN + centroid v2 on the same board is still 7.11. The 19/40 Codex usefulness gate stays sealed; the reusable 200-row SO-board audit did not improve.

    • +0.12 vs pin
    • 7.11 FLAN bar
    • SO-board flat
  20. Tab titles · Sealed result

    Four beams, no new weights

    The unchanged 35M B-9500 checkpoint used safe four-beam search to beat the title-tuned FLAN stack on a sealed 1,000-task test.

    Published result Beam 4 reproduced the FLAN win; the conjunctive final gate still did not clear.

    The paired lead was +0.752. The sealed n=40 audit missed its 60% usefulness bar; a later preregistered dev-200 audit estimated 40.0% usefulness (95% interval 33.5–46.9%).

    • 6.20 vs 5.45
    • +0.752 paired
    • 40.0% dev audit
  21. Tab titles · Development result

    The 35M that caught Hybrid A

    A scratch 35M checkpoint trained on 4.86 million body-to-title pairs was tested with the same glue as the earlier system.

    Published result B-9500 beat the older 35M stack and tied raw FLAN, but remained behind FLAN plus centroid.

    It scored 5.65 versus 5.10 for the older stack and 5.64 for plain FLAN. Later beam search superseded this stopping point without adding weights.

    • 5.65 mean
    • 534 / 110 / 356
    • 35M parameters
  22. Tab titles · First comparison

    Two recipes, one glue

    The first tab-title study compared a FLAN-plus-centroid Mac recipe with a smaller 35M-plus-centroid phone-class recipe on the same holdout.

    Published result Both device recipes beat the incumbent; later work superseded the original production choice.

    FLAN plus centroid reached 77% useful, while the older 35M stack reached 64% at half the namer weights. The full seven-way packet remains published.

    • 77% useful
    • 64% phone-class
    • n=1,000
  23. Model analysis · Matched controls

    The Sniff Test: A seed lottery, dissected

    A literary embedding initialization appeared to reduce repeated words in one 12.7M-parameter checkpoint, then met a deterministic 31-run control.

    Published result The eye-catching fixed-pair difference was real; the proposed literary mechanism was not established.

    Sequence order had no stable advantage. Ordered Verne averaged 14.75% repeated titles versus 13.79% for frequency-matched shuffled Verne.

    • 31 control runs
    • 907 tasks
    • 9 matched donors
Audit trail

Evidence attached to the claims.

Frozen or aggregate public artifacts for the studies that publish separate evidence bundles.

Bash SFT note 05

Aggregate corpus, independent-hard scores, hashes, and synthetic examples. No data rows, private traces, raw evaluations, or weight tensors.

Bash SFT note 04

Live-style continuation aggregates, mix counts, synthetic samples, and byte hashes. No training JSONL, private traces, fixtures, or weight tensors.

Bash SFT note 03

Mix-composition arms, locked GFR, harvest, last40, Kimi, and matched-prompt kept/stock/specialist aggregates with byte hashes. No training JSONL, private traces, fixtures, or weight tensors.

Bash SFT note 02

Sanitized aggregate data manifest, all four training recipes, corrected fair-base comparisons, paired bootstrap intervals, external canary summary, and byte hashes. No private traces, completions, fixtures, or weight tensors.

Bash SFT note 01

CUDA score of stock vs 137-row SFT on 500 Mini-trace holdout rows, train summary, mix manifest, per-row scores, and byte hashes. Gold commands are not published. No weight tensors.

Four beams, no new weights

Decoder and glue source, frozen tasks and outputs, blinded audits, score rows, environment pins, and byte manifests.

The Sniff Test

Aggregate control outcomes, the prospective continuation rule, recorded runtime, source revision, and checksums.

The Sniff Test artifacts publish aggregate outcomes and provenance hashes, not production task text, source identifiers, or row-level predictions. Its interactive paper remains readable without JavaScript.