Shlex · local file search · 2026-09-20

Shlex 0.9.0: nothing you type is required, and a small model that offers other words

Result. On a board nobody who built the fix wrote, the holdout half goes from 19 of 117 cases passed with Shlex 0.8.4 to 62 of 117 with 0.9.0, by executed and ranked tests; every older board holds. The search now orders documents by how much of what the person typed they hold, in the text, the file name and the folder names, and requires none of it. A 1.5B model trained on a rented GPU for 1.12 USD answers a second question, other words for the same thing, and never malformed: that takes the hardest family from nothing found to a wanted document in the first three results in 10 of 23 design cases, where a frontier model's wordings reach 12.
19 → 62of 117 holdout cases passed, 0.8.4 against 0.9.0
11 → 61of 120 design cases
817 → 824cases passed over every older board and the design half by the same app, shipped weights against new
108 → 0malformed answers to the second question, of 117
1.12 USDtwo training runs on one RTX 6000 Ada, pods deleted after the weights were verified at home

Two faults, in the user's words

Shlex is a desktop pet that finds files. A small local model turns a request into a plan, and a sandboxed reader walks the folder. Until 0.8.4 a document was a result only if it held every word of the plan, letter for letter. Its user put the two consequences plainly: "every word needs to be in the document to match is a terrible requirement, because 'job offer' will miss 'employment agreement' even though they are kind of the same thing", and "we need the model to understand that 'bovine' and 'cow' are the same thing".

The first fault means that saying more finds less. Each helpful word is one more thing every document must contain: a plural where the file says the singular, a year that is only in the folder's name, one slip of the finger, and the list is empty. The second means that matching is string lookup; the language model was only used to pick words. Neither had been chosen by the user, and the boards of the earlier notes could not see them: their expected answers were plans in the app's own convention, so they measured the every-word rule against itself.

A board nobody who built the fix wrote

The expected answer of a case is now documents a person would want, never a plan. A frontier model (Grok, through its command-line tool) wrote 292 cases in ten families: a small invented home folder of nine to twelve documents, a request as a person types it, the documents wanted, and the documents it would be plainly wrong to show. A second, blind reading that saw only request and folder had to name the same wanted documents; 54 cases where the two disagreed were dropped, 237 kept. A seeded split gave a design half of 120 and a holdout half of 117, hashed before any fix was scored. The builder read the design half only; the holdout was scored once, on the build that shipped.

The check is ranked: a case passes when the wanted documents fill the first places, in any order, and no "never" document is shown at all. Every request is answered by the model exactly as the app does, every later step is the app's own code through the release binary, and the line that results is executed by the real sandboxed reader on the case's folder. Token loss decides nothing anywhere in this note. The board is harsher than a real folder: its "never" documents mention the request's words in passing on purpose, and its "same meaning" documents avoid the request's key words in every form.

The rule

A document's place in the results is how much of what the person meant it holds. Nothing typed is required. More words can only reorder the results, never empty them. The app writes every plan of topics itself from the person's own words, one idea per word in either number; the model's plan only adds what words cannot say (a format, a folder that exists, what to leave out, newest first, how many, a time window asked for in words). What the person asked for explicitly stays strict: quoted text, exclusions, named formats, a real folder, stored-credential searches. The sandbox, the permission model and the command policy are untouched.

Results on the new board

Holdout half, scored once (117 cases)Shlex 0.8.4, shipped modelShlex 0.9.0, new model
All families19/11762/117
a long request whose words are spread over the file name, a folder name and the text0/109/10
plural typed, singular written (or the other way round)1/115/11
one slip of the finger in a long request0/93/9
two kinds of document joined by "and"1/126/12
a year that is only in a folder name, only in the file name, or only in the text1/127/12
letters and digits run together ("tax2026", "Invoice2031.pdf")1/97/9
the same thing in other words; the request's key words are nowhere in the wanted document0/221/22
one document holds everything, others hold part: it must come first9/1211/12
documents that hold one common word of the request and have nothing else to do with it2/116/11
Desktop, Documents, Downloads ...: a file under Music/ is not about music4/97/9
cases where nothing at all is shown737
cases where a "never" document is shown322
Design half (120 cases)0.8.40.9.0
All families11/12061/120
a long request whose words are spread over the file name, a folder name and the text0/1010/10
plural typed, singular written (or the other way round)1/118/11
one slip of the finger in a long request0/104/10
two kinds of document joined by "and"1/127/12
a year that is only in a folder name, only in the file name, or only in the text0/125/12
letters and digits run together ("tax2026", "Invoice2031.pdf")0/93/9
the same thing in other words; the request's key words are nowhere in the wanted document0/234/23
one document holds everything, others hold part: it must come first6/1312/13
documents that hold one common word of the request and have nothing else to do with it3/113/11
Desktop, Documents, Downloads ...: a file under Music/ is not about music0/95/9

0.8.4 shows more "never" documents than its pass count suggests it could only because it shows nothing at all in most cases. 0.9.0 shows something nearly always, and the strict "never" check then bites on documents that mention every word in passing. That is the larger part of what still fails, with misspelt key words, wanted scans whose name says nothing, and meaning.

A few invented requests of the design half

FamilyTyped0.8.40.9.0 shows first
documents that hold one common word of the request and have nothing else to do with itlab report for bailey2 shown, the wanted ones not firstLabs/Bailey/CBC Chemistry Report - Bailey.pdf
Labs/Bailey/urinalysis_report_bailey.docx
letters and digits run together ("tax2026", "Invoice2031.pdf")n400 interview 20260 shown, the wanted ones not firstDocuments/Citizenship/N400InterviewPrep.docx
Documents/Citizenship/N400InterviewNotice.pdf
one document holds everything, others hold partgirls under 13s winter training plan2 shown, the wanted ones not firstCoaching/Girls U13s/Winter Training Plan.docx
the same thing in other wordsthe kids report cards0 shown, the wanted ones not firstDocuments/Education/Emma_Fall_Term_Grades.pdf
Documents/Education/Owen_Academic_Marks.docx
plural typed, singular written (or the other way round)tax returns for affidavit6 shown, the wanted ones not firstTaxes/Tax Return 2022 - Rahul Sharma.pdf
Taxes/Tax Return 2021 - Rahul Sharma.pdf
Taxes/Tax Return 2023 - Rahul Sharma.pdf
a long request whose words are spread over the file name, a folder name and the textadult literacy tutor background check consent0 shown, the wanted ones not firstRiverRead/adult-literacy/Tutor Consent.pdf
compliance/adult-literacy/Tutor Background Form.docx
Desktop, Documents, Downloads ...download speed complaint1 shown, the wanted ones not firstDocuments/Broadband/Openreach download speed complaint.docx
Documents/Broadband/ISP_complaint_ref_8841.pdf
two kinds of document joined by "and"invoices and contracts for loftwell2 shown, the wanted ones not firstClients/Loftwell/Invoice_LW-191.pdf
Clients/Loftwell/Invoice_LW-184.pdf
Clients/Loftwell/MSA_Loftwell_signed.pdf

Every older board holds

The older boards expect exact sets, and best fit shows near matches below the documents that hold everything, so they are read gold first: every gold path returned, above every other; order among them free unless the person asked for newest first. Same run, same binary (SHA-256 43dd379cfef012bb…), the shipped weights against the new ones:

Boardjeremiah_v4 (shipped)jeremiah_v5
Unseen board of the previous note, 312 cases, as the app runs it309/312307/312
v3 development board102/109106/109
v4 development board56/5655/56
Operator canaries2/22/2
Refusal board, design half62/6462/64
Refusal board, holdout half46/4846/48
Format-word board52/5652/56
Format-word holdout35/3836/38
Format-word overlay7/76/7
File-name board, historical38/4038/40
File-name board, everyday53/5453/54

The new weights lose two cases of the 312-case unseen board, one v4 development case and one overlay case against the shipped ones, and gain on the v3 development board and a format-word board. The first training pass had lost five unseen cases: the second question's rows dilute the first a little, and a second pass over the same data recovered most of it.

Where should meaning come from? A bake-off

The reader's half is one small rule: a plan may carry other ways a document might say an idea, and an idea said another way counts at half the worth of the person's own word. Four sources were measured on the 23 "same meaning" cases of the design half, before any were built into the app.

SourceStrict passA wanted document in the first 3Wanted documents shownNothing shown
Best-fit search alone0/230/230/4915/23
A table of static word vectors in the app (potion-base-8M, nearest words)1/232/232/4913/23
A small sentence encoder over the documents (all-MiniLM-L6-v2, Python prototype)8/2318/2328/490/23
Other wordings written by a frontier model shown the request only: the ceiling for any model that writes them4/2312/2324/493/23
The shipped 1.5B model asked the second question untrained0/230/230/4915/23
The new 1.5B model, trained for the second question4/2310/2317/496/23

No discount between a quarter and full worth moved the ceiling: what limits a list of other wordings is that nobody can guess "Occupancy Remittance" from "rent for the workshop". Word vectors mostly return inflections. Document vectors are the stronger mechanism even with the smallest encoder and no tuning, and they are the next project: inference inside the sandboxed reader, a cache and first-run indexing are weeks of work, where a list of wordings was a day. The two add up.

Teaching a 1.5B model the second question

The model is asked a second, separate question with its own 657-byte context: for each idea of the request that a document might word differently, the person's own words first, then up to four other words or short phrases, as a JSON list of lists. The first question, its context, its 13,696 training rows and its boards stay byte for byte as they were, so retention is measured by the same gates. The app trusts nothing in the answer: it can only add wordings for ideas the person typed, container words ("files", "records") are refused whatever a model writes, and a malformed answer changes nothing. It costs about a third of a second per content search, and only a model trained for it is asked.

Grok wrote 3,900 requests with their other wordings across thirty themes, and reviewed them. Reading all of it before training changed four things. The strict second reading had struck "employment agreement" for "job offer", the user's own example, so a strike counts only against a single word, which matches widely when it is loose; a specific phrase costs a near match at most. Container words were removed as wordings. Misspelt words now get the word that was meant as their first wording (180 found with a dictionary): the authoring prompt had ruled "spelling variants" out, which left the one wording that helps most unwritten, and it is why the typo family moves. And every answer was put through the app's own code, so the model is taught exactly what reaches the reader. 3,673 rows were accepted (489 of them answer that there is nothing to add), 3,296 went into training beside the unchanged 13,696.

The first pass (1,062 steps, 0.40 USD) learned the format at once and was still gaining on both questions when it ended; the second run exposed every row twice under one schedule (2,124 steps, 0.72 USD) and its step 1593 is the model in 0.9.0 (SHA-256 1cca31f623f03d93…). Both pods were deleted by the controller only after the snapshots had been pulled and their hashes verified at home. Snapshots were picked by tests passed, never by loss.

What still fails, and what is next

  1. Documents that mention everything in passing. A document that holds every word is never cut, which is right for a person and costs cases on a board whose decoys are written that way. Better evidence of what a document is about than its name, its folders and its first line is the way forward, not a harder cut.
  2. Meaning beyond a list of wordings. Document vectors inside the reader, measured on the same family against the same ceiling.
  3. A misspelt key word. The model now offers the word that was meant; a spelling rule in the reader was left out on purpose, because a respelling rule risks false corrections on a small folder.
  4. The head noun. "checking account statements" ranks the account-opening note first: the last word of a noun phrase says what kind of document is wanted, and the scoring does not know that yet.