Rote Playoffs · September 2026 · aryangorde

Cut at 65,536

A process.exec step in rote captures the first 65,536 bytes of stdout and discards the rest. The step exits 0. The runner reports the run completed. The linter cannot see it. I found it because it was quietly corrupting my own play’s output — then measured how much of the public registry stands in the same place.

repo-leak-doctor · history_check 95.1% never read
1,346,888 bytes produced 65,536 carried

Drawn to scale. The solid band is everything the play could read; the hatching is what the command wrote and the step threw away. No error was raised at any point.

What the cap actually is

A process.exec step captures at most 65,536 bytes into body.stdout.text. It sets body.stdout.truncated: true alongside bytes and preview_bytes, and writes the complete payload to .rote/artifacts/processes/@N/stdout.json — a path a steps_with_presentation body cannot open. The full data exists on disk and is unreachable from the code that needs it.

The cap is exact, not approximate. Across 4,825 process observations recorded on this machine, 14 were truncated, and every one of them kept precisely 65,536 bytes.

A three-line play reproduces it on rote 0.80.0:

steps:
  big:
    type: process.exec
    argv: [python3, -c, "print('x'*100000)"]

emitted by command : 100001 bytes
captured text len  : 65536
truncated flag     : true
step exit          : { kind: "code", code: 0 }
run status         : succeeded

It is undocumented. 65536 appears twice in rote guidance typescript play-creation — once as browser.extract's unrelated truncated field, once as an author-supplied maxBytes argument. Neither describes this cap. Filed upstream as rote-releases#10.

Why nothing reports it

This is the part worth writing down. The failure is invisible at every gate that exists to catch failures:

every signal a play author would check
GateWhat it reports
the child processexits 0
the stepcompleted
the runnerN/N completed, 0 failed, 0 blocked
the runexits 0
rote play lintpasses — fixtures are small, so it never reproduces

It only appears on real input, at real scale, in production. And there are two failure modes, of which the second is far worse.

The prefix breaks the parser. JSON.parse throws on the cut object. You get an error and blame the provider for a malformed response. Annoying, but at least it is loud.

The prefix is still valid input. If the step emits newline-delimited records — or the body splits on \n and filters — the truncated data parses cleanly.

You get a confident, wrong answer, with no error anywhere. Not a crash. A number that is simply too low. That is the mode that ships.

What it cut here

Eight distinct steps across five published plays, recorded on one laptop over two days of ordinary use — a ninth cut belongs to the synthetic three-line repro I wrote to confirm the cap, and is excluded here. Each bar is drawn to the same scale as the one above it.

repo-leak-doctor · tracked_files 88.9%
587,958 produced65,536 carried
pr-review-verdict · fetch_inline 85.4%
448,889 produced65,536 carried
tech-debt · scan_suppressions 57.9%
155,641 produced65,536 carried
repo-onboarding-brief · read_runner_scripts 55.2%
146,366 produced65,536 carried
pr-review-verdict · fetch_reviews 50.9%
133,553 produced65,536 carried
repo-onboarding-brief · read_secondary_ci 35.5%
101,672 produced65,536 carried
review-gate · merge_checks 20.9%
82,822 produced65,536 carried

The last one is mine. review-gate answers a single question before you merge: are there unresolved review threads or failing checks? Pointed at DataDog/dd-trace-js#10068 — a pull request carrying 917 checks — it returned a confident verdict computed from four fifths of the evidence.

The second outcome

A user running the audit reported that it only ever saw steps that were cut — his step had stopped being cut and started timing out instead, bypassing every degrade path he had written. He was right, and my first explanation of my own bug was wrong. I said a killed step was invisible by construction: the run fails, never reaches the presentation plane, no record is written.

That is not what happens. Forty-eight failed runs on this machine wrote presentation records, and the timed-out one was among them:

run_20260905_084310.629_1   status: failed
step evaluate -> status: "failed"
  output.diagnostic.response_id: 2
  output.diagnostic.exit: { kind: "timed_out", timeout_ms: 30000 }

The rule, measured across all 896 step outcomes on disk: a completed step files its observation under outcome.output.body; a failed step files it under outcome.output.diagnostic. Never both, never swapped.

where an observation is filed, by step outcome
OutcomebodydiagnosticSteps
completedyes—787
failed—yes55
blocked——50
failed (no observation)——4

A killed step is a failed step. My audit read bodies. So it reported zero timeouts on a machine that had one — and the record had been sitting there the whole time, step name attached. It was never structural. It was a place I had not looked, and I had called it a hole in the runtime.

Truncation and timeout are two outcomes of one condition: the payload exceeded what the step could carry. Only the first sets a flag anyone can read.

How much of the registry

Having found it in my own code, the question was whether anyone else had. 302 packages pulled and statically analysed — 298 published, plus local working copies of my own four plays, which are excluded from every published figure so my work is not double-counted.

302 installed packages, by how they treat process stdout
Classn%
consumes stdout, never reads the flag22173.2
consumes stdout, reads the flag3210.6
no stdout read227.3
no process steps at all165.3
legacy no-steps body113.6

Of the 249 published packages that consume process stdout, 221 — 89% — never read the truncation flag. The 28 that do belong to seven owners.

A methodology correction worth recording

My first denominator was wrong, and the way it was wrong is instructive. rote registry play list has no list-all — it requires an owner, org or community. So I enumerated by sweeping search queries: 30 queries surfaced 292 plays, 45 surfaced 532. It was tempting to call 532 “the registry.”

It is not. Per-owner listing is exhaustive per owner, so I listed all 122 owners directly and got 697. Every play search had found was in that set; 165 were not. Search had missed 24% of the registry — and no number of additional queries would have revealed that, because a search sweep cannot report what it failed to surface.

If you want to enumerate that registry: list owners, do not sweep queries.

The fix

The instinct is to ask for a bigger buffer. That is the wrong shape of fix — it moves the cliff without removing it. The rule that generalises: don’t move the payload, move the answer.

Aggregate where the rows still exist, inside the step, before stdout is ever written. gh’s --jq and --template run in-process, so filtering there costs nothing. Carry out a bounded number of detail rows plus an explicit retained-vs-total count, so any list you print is a labelled lower bound rather than a confidently wrong number.

Applied to the step that started this — filtering passing checks server-side instead of client-side:

before
82,822 bytes — truncated
after
173 bytes — clean, check count still exact

A 99.8% reduction with no loss of the actual answer. The payload was never the point.

  • per_page=100 on a REST call is safe.
  • An unpaginated GraphQL body, a gh pr view --json status rollup, and raw git log -p or git ls-files piped into the body are not.
  • Detect it with body.stdout?.truncated === true — it is on the typed ProcessExecStream the presentation SDK hands you.

What I got wrong

The finding changed five times, and the tool caused every correction. This is the part I would want a reader to take seriously, because each wrong answer was confident.

  1. “Four authors handle this.”

    Wrong. All four were matching their own unrelated clamping logic that happened to use the word truncated.

  2. “Only my play handles it.”

    True across the 26 plays I had installed. False across 302.

  3. “45 are guarded.”

    Wrong again. The proximity heuristic — looking for truncated near stdout — had a 29% false-positive rate, 12 of 41, matching things like findings_truncated sitting beside stdout_bytes. A false “this is handled” is the worst possible output for a tool like this, so the detector now requires member access on the stream itself, with frontmatter and comments stripped first.

  4. “Search found 532 of the registry, so listing finds 31% more.”

    The share of the registry actually missed is 165/697 = 24%. 31% was 165/532 — a different question, silently answered.

  5. “The audit reads 14% of the available evidence.”

    It reads about 84%. The script I wrote to verify the play walked only single-shape step outputs and skipped 3,328 fan-out items — so I trusted a throwaway checker over the audited tool without ever running one against the other. This one reached another person before I caught it.

Three of the five corrections came from other people running the play and reporting what it got wrong. Every one of those was a failure mode I had not personally hit — which is the argument for being wrong in public quickly rather than careful in private.

Using the tool at scale also found defects in the tool itself, including the one I like most: its own output reached 41 KB against the 64 KB cap at 302 plays. The truncation auditor was on course to truncate its own evidence. Compacting the index to positional rows brought it back to 29.5 KB.

Run it on your own machine

The audit is published as a play, so the check is reproducible rather than a one-off script. It is read-only, needs no network and no API key, and reports demonstrated loss separately from untested exposure — the two are never merged.

rote play run https://play.modiqo.ai/aryangorde/truncation-audit root=~/.rote
truncation-audit which of your plays silently drop data, and which already have
flake-finder which CI jobs are flaky, grouped by job across runs rather than by run
reviewer-finder who should review this diff, by breadth of prior contact, flagging bus-factor-1 files
review-gate unresolved threads and failing checks, before merge