A process.exec step in rote captures the first 65,536 bytes of stdout and
discards the rest. The step exits 0. The runner reports the run completed. The linter cannot
see it. I found it because it was quietly corrupting my own play’s output — then measured
how much of the public registry stands in the same place.
Drawn to scale. The solid band is everything the play could read; the hatching is what the command wrote and the step threw away. No error was raised at any point.
A process.exec step captures at most 65,536 bytes into
body.stdout.text. It sets body.stdout.truncated: true alongside
bytes and preview_bytes, and writes the complete payload to
.rote/artifacts/processes/@N/stdout.json — a path a
steps_with_presentation body cannot open. The full data exists on disk and is
unreachable from the code that needs it.
The cap is exact, not approximate. Across 4,825 process observations recorded on this machine, 14 were truncated, and every one of them kept precisely 65,536 bytes.
A three-line play reproduces it on rote 0.80.0:
steps:
big:
type: process.exec
argv: [python3, -c, "print('x'*100000)"]
emitted by command : 100001 bytes
captured text len : 65536
truncated flag : true
step exit : { kind: "code", code: 0 }
run status : succeeded
It is undocumented. 65536 appears twice in
rote guidance typescript play-creation — once as browser.extract's
unrelated truncated field, once as an author-supplied maxBytes
argument. Neither describes this cap. Filed upstream as
rote-releases#10.
This is the part worth writing down. The failure is invisible at every gate that exists to catch failures:
| Gate | What it reports |
|---|---|
| the child process | exits 0 |
| the step | completed |
| the runner | N/N completed, 0 failed, 0 blocked |
| the run | exits 0 |
rote play lint | passes — fixtures are small, so it never reproduces |
It only appears on real input, at real scale, in production. And there are two failure modes, of which the second is far worse.
The prefix breaks the parser. JSON.parse throws on the cut
object. You get an error and blame the provider for a malformed response. Annoying, but at
least it is loud.
The prefix is still valid input. If the step emits newline-delimited
records — or the body splits on \n and filters — the truncated data parses
cleanly.
You get a confident, wrong answer, with no error anywhere. Not a crash. A number that is simply too low. That is the mode that ships.
Eight distinct steps across five published plays, recorded on one laptop over two days of ordinary use — a ninth cut belongs to the synthetic three-line repro I wrote to confirm the cap, and is excluded here. Each bar is drawn to the same scale as the one above it.
The last one is mine. review-gate answers a single question before you merge:
are there unresolved review threads or failing checks? Pointed at
DataDog/dd-trace-js#10068 —
a pull request carrying 917 checks — it returned a confident verdict computed from four
fifths of the evidence.
A user running the audit reported that it only ever saw steps that were cut — his step had stopped being cut and started timing out instead, bypassing every degrade path he had written. He was right, and my first explanation of my own bug was wrong. I said a killed step was invisible by construction: the run fails, never reaches the presentation plane, no record is written.
That is not what happens. Forty-eight failed runs on this machine wrote presentation records, and the timed-out one was among them:
run_20260905_084310.629_1 status: failed
step evaluate -> status: "failed"
output.diagnostic.response_id: 2
output.diagnostic.exit: { kind: "timed_out", timeout_ms: 30000 }
The rule, measured across all 896 step outcomes on disk: a completed step
files its observation under outcome.output.body; a failed step
files it under outcome.output.diagnostic. Never both, never swapped.
| Outcome | body | diagnostic | Steps |
|---|---|---|---|
| completed | yes | — | 787 |
| failed | — | yes | 55 |
| blocked | — | — | 50 |
| failed (no observation) | — | — | 4 |
A killed step is a failed step. My audit read bodies. So it reported zero timeouts on a machine that had one — and the record had been sitting there the whole time, step name attached. It was never structural. It was a place I had not looked, and I had called it a hole in the runtime.
Truncation and timeout are two outcomes of one condition: the payload exceeded what the step could carry. Only the first sets a flag anyone can read.
Having found it in my own code, the question was whether anyone else had. 302 packages pulled and statically analysed — 298 published, plus local working copies of my own four plays, which are excluded from every published figure so my work is not double-counted.
| Class | n | % |
|---|---|---|
| consumes stdout, never reads the flag | 221 | 73.2 |
| consumes stdout, reads the flag | 32 | 10.6 |
| no stdout read | 22 | 7.3 |
| no process steps at all | 16 | 5.3 |
| legacy no-steps body | 11 | 3.6 |
Of the 249 published packages that consume process stdout, 221 — 89% — never read the truncation flag. The 28 that do belong to seven owners.
My first denominator was wrong, and the way it was wrong is instructive.
rote registry play list has no list-all — it requires an owner, org or
community. So I enumerated by sweeping search queries: 30 queries surfaced 292 plays,
45 surfaced 532. It was tempting to call 532 “the registry.”
It is not. Per-owner listing is exhaustive per owner, so I listed all 122 owners directly and got 697. Every play search had found was in that set; 165 were not. Search had missed 24% of the registry — and no number of additional queries would have revealed that, because a search sweep cannot report what it failed to surface.
If you want to enumerate that registry: list owners, do not sweep queries.
The instinct is to ask for a bigger buffer. That is the wrong shape of fix — it moves the cliff without removing it. The rule that generalises: don’t move the payload, move the answer.
Aggregate where the rows still exist, inside the step, before stdout is ever written.
gh’s --jq and --template run in-process, so
filtering there costs nothing. Carry out a bounded number of detail rows plus an explicit
retained-vs-total count, so any list you print is a labelled lower bound rather than a
confidently wrong number.
Applied to the step that started this — filtering passing checks server-side instead of client-side:
A 99.8% reduction with no loss of the actual answer. The payload was never the point.
per_page=100 on a REST call is safe.gh pr view --json status rollup, and raw
git log -p or git ls-files piped into the body are not.body.stdout?.truncated === true — it is on the typed
ProcessExecStream the presentation SDK hands you.The finding changed five times, and the tool caused every correction. This is the part I would want a reader to take seriously, because each wrong answer was confident.
Wrong. All four were matching their own unrelated clamping logic that happened to use
the word truncated.
True across the 26 plays I had installed. False across 302.
Wrong again. The proximity heuristic — looking for truncated near
stdout — had a 29% false-positive rate, 12 of 41, matching things like
findings_truncated sitting beside stdout_bytes. A false
“this is handled” is the worst possible output for a tool like this, so the
detector now requires member access on the stream itself, with frontmatter and comments
stripped first.
The share of the registry actually missed is 165/697 = 24%. 31% was 165/532 — a different question, silently answered.
It reads about 84%. The script I wrote to verify the play walked only single-shape step outputs and skipped 3,328 fan-out items — so I trusted a throwaway checker over the audited tool without ever running one against the other. This one reached another person before I caught it.
Three of the five corrections came from other people running the play and reporting what it got wrong. Every one of those was a failure mode I had not personally hit — which is the argument for being wrong in public quickly rather than careful in private.
Using the tool at scale also found defects in the tool itself, including the one I like most: its own output reached 41 KB against the 64 KB cap at 302 plays. The truncation auditor was on course to truncate its own evidence. Compacting the index to positional rows brought it back to 29.5 KB.
The audit is published as a play, so the check is reproducible rather than a one-off script. It is read-only, needs no network and no API key, and reports demonstrated loss separately from untested exposure — the two are never merged.
rote play run https://play.modiqo.ai/aryangorde/truncation-audit root=~/.rote