agent-manager · mock · reader · feedback on #115

Sub-agents as step rows, inside the one expandable

Built, not just drawn: section 0 is the running app, on the branch behind PR #115. The rest is the mock the design was worked out in. The second disclosure is gone. A sub-agent is a step — the Agent call already sits in the step list — so the row it already produces becomes the live, expandable sub-agent row. Chronology comes for free, and there is no second list to keep in sync. Every row below carries real values: descriptions from the sidecars, durations, tool counts and tokens from the completion notifications of three real sessions.

0 · The point of the feature: what is running, under the spinner, collapsed

Nothing expanded. The exchange is closed, the pane's own working line says it is busy, and under it is one line per sub-agent still going — task, how long it has been running, and when it last wrote. This is the running app, shot on the branch.

The collapsed exchange with five running sub-agents and one stalled
Five spinning, one that has not written in 44 minutes and has stopped claiming to run. Eight were spawned; the two that finished are gone from the strip and counted in the line above.
decisionwhat it doeswhy
running onlya finished sub-agent leaves the strip; the count in the summary line is what persistsa strip that keeps them grows for the length of the turn and stops being a glance
leaves on a fadea row that has just finished holds its place for 6s with its ✓ and duration, then fadesrows vanishing under a live spinner is the shape that produced the API-log flicker; and only rows this client saw running linger, or every old completion flashes in on page load
capped at sixthen one line: "…and 6 more running — open the work to see them all"the-gatherer's busiest turns spawn 40, 36, 22 and 12 — see the correction below
15-minute silence rulethe mark goes inert and the row reads no write in 15mthe measured worst silence inside a run that finished normally is 601s; nine of release-video's 22 sub-agents have no outcome at all and would otherwise spin forever
the pane gates the wordif the pane is not running, nothing under it says "running" — the rows read no resulta killed pane leaves sub-agents that never finish and never write again

a correction to what I reported last round

I said the busiest real exchange spawned four. That is release-video. Counting the-gatherer's parent as well: 168 Agent calls across 219 exchanges, and its busiest four turns spawn 40, 36, 22 and 12. Twenty-two at once is a real turn, the operator was right, and the cap is not decoration. (Those older turns also predate the per-sub-agent transcript: 168 spawns, 38 sidecars on disk. Rows for them still count and still carry their outcome, and say no transcript kept instead of offering a triangle that would 404.)

Twelve running sub-agents, six shown and an overflow line
Twelve running: six rows and the overflow line. The line is the only thing that grows.

and the drill-down is in the same rows

A strip row opened in place
A row from the strip, opened where it is: the task it was given, then its own trace with its own summary line and report.
The work expanded, sub-agents among the tools
The work expanded: the same rows, in the order they were spawned, among the parent's own tools. One expandable, as asked.

Shot against a live pane whose transcript is a fixture in the real record shapes, so all six states appear at once; the sidecars, the child transcripts, the descriptions and the notification numbers are real, and the pane really is running. The turn's own summary line reads 11 steps · 2 tools · 8 sub-agents · 1h 13m · 9.9k tok — sub-agents counted as sub-agents, not also as tools.

1 · The structure, before and after

#115 put the count in the summary line as its own button, so the row had two disclosures. The operator's point is that a sub-agent is not a separate category of thing — it is one of the steps. So the count becomes plain text in the line, and the single fold opens a step list in which the Agent steps are now live rows.

#115 — two disclosures
turn 24/3812:43 PM
two things to open, and the sub-agents sit outside the work
this mock — one
turn 24/3812:43 PM
the count reads in the line; the fold opens everything, sub-agents included

The count stays a count, not a status summary: what each one is doing is one row away, and the row is where state belongs. "2 still going" would fit in the line too, but it then repeats what the rows say two pixels below.

2 · Running, mid-turn, among ordinary steps

release-video's last turn, as it stands right now: eleven of the parent's own steps, then three sub-agents spawned within ninety seconds of each other, then the parent going back to work while they run. Two have reported; the third was spawned at 14:28 and is still writing.

reader · release-video · turn 38, live
dont overshoot the growth and make the echo a bit more discrete or more continuous
turn 38/3814:19 PM
Bash ×3python3 - <<'PY' def peak(c1)…, sed -i 's/ c1 = 1.70…
Read/tmp/n3f.png
Write ×2/data/workspaces/release-video/film/demo/build_zoom.py
Bashcd /data/workspaces/release-video/film/demo; X264="aq-mode=3:ps…
sub-agent Legend beside arrows and agent focus 49m 33s154 calls212k tok
sub-agent Rapid fire title-then-UI restructure 47m 35s136 calls197k tok
sub-agent Capture group view in reader running 1h 10mwrote 2m ago
Bash ×2for f in approved/scene*.mp4; do …, M=/home/node/local/agent-state…
Read ×3…/rf/contact-rf3.png, /tmp/wlsheet.png, …/gfx5/coordination-focus…
thinkingTwo of the three are back. The group-view capture is the one I still…
working· Bash cd /data/workspaces/release-video/film
real turn, real ordering: the sub-agents sit between the parent's own steps because that is where they were spawned

Two facts on the running row, and neither is a verdict: 1h 10m is how long since it was spawned, wrote 2m ago is how long since its transcript last changed. Silence gets reported, never judged — measured gaps between consecutive records inside a live sub-agent reach 601s, so a rule that greys a row out after a minute would be wrong several times an hour.

Its own spinner sits next to the pane's working line, one indented and one not. Two things are working and both say so; the row's spinner is the same braille cell, in the same accent, so it reads as the same kind of statement.

3 · Several at once — and the number is smaller than it looks

In release-video they arrive in runs of two or three across 38 turns; per exchange the distribution is 4, 3, 3, 3, 3, 2, 2, 2, 2, 1, 1, 1. But that session is not the ceiling: the-gatherer spawns 40, 36, 22 and 12 in single exchanges (168 Agent calls across 219 turns). So both shapes are real — most turns spawn two or three, and a fan-out turn spawns dozens — which is why the collapsed strip is capped and the step list is not.

reader · release-video · turn 24, the busiest in the session
1. for the reader switch, the indicator can be even bigger. 2. also there can be an overlay of the two panes…
turn 24/3812:43 PM
Bashcd /data/workspaces/release-video/page && python3 - <<'PY' p='re…
sub-agent Recapture second-agent sequence 1h 13m169 calls331k tok
sub-agent Rapid-fire feature captures 1h 43m218 calls329k tok
sub-agent Add step legend to comms animation 14m 56s39 calls131k tok
Bash ×14ffmpeg -v error -y …, hf upload lvwerra/agent-artifacts …, du -sh sift
Read ×2/tmp/cm.png, /tmp/margins.png
sub-agent Reshoot scene 3 with clean sidebar 11m 48s37 calls121k tok
thinkingAll four are back. The sheet needs one more pass before I fold them…
four sub-agents, in the positions they were spawned, between the parent's own work

What this turn was under-reporting. Its own line says 53 tools. Its four sub-agents made 463 tool calls and spent 913k tokens between them — most of the turn's work, and none of it reachable from the reader before.

one thing this needs from the code

Consecutive calls to the same tool collapse into one row today — Read ×4 — and spawns arrive consecutively: in this very turn, three in a row at 12:43–12:44. Left as it is, they would render as a single Agent ×3 row with no state, no duration and nothing to open. Sub-agent steps have to be exempted from that grouping in stepsOf. Small change; it is the difference between this design and a row that says ×3.

when a turn spawns twenty

the-gatherer does, four times over. The proposal is the pattern the reader already uses for older turns: show the first few, then one line that opens the rest. The sub-agents stay chronological, and a pathological turn costs one row instead of twenty.

sub-agentGather: mirror-hunt blocked Zhihurunning4m 26s
sub-agentGather: SemiAnalysis speculationrunning3m 58s
sub-agentGather: DeepSeek-R1 / Kimi 1.5 long-form analysis2m 52s13 calls41.7k tok
the fallback, drawn with the-gatherer's real rows — that session's busiest turn spawned 40

4 · All finished, collapsed

The state the reader is in most of the time: nothing spins, nothing is indented, and the line says how much happened underneath it.

can you add a few more backgrounds: e.g. a mix of drift and pixel noise/flickering
turn 21/3809:31 AM

Both grounds are built — drift-plus-noise at two amplitudes and the intricate grid. The grid one took an hour and a half; most of it was re-rendering at 60fps after the first pass read as static.

closed, the sub-agents are a number in the line and nothing more

5 · One expanded: its task, then its trace

Opening a sub-agent row shows the two things the operator asked for. First the task it was given — the Agent call's own prompt, clamped. Then its trace, in the same rail the parent's steps use, because a sub-agent's transcript is an ordinary trace and the reader can already render one.

reader · release-video · "Reshoot scene 3 with clean sidebar", open
turn 24/3812:43 PM
Bash ×14ffmpeg -v error -y …, hf upload lvwerra/agent-artifacts …
sub-agent Reshoot scene 3 with clean sidebar 11m 48s37 calls121k tok
You are re-shooting one scene of the Agent Manager release film in /data/workspaces/release-video/film. A previous agent built this scene successfully; two cosmetic blemishes from the live Space got baked into the frames and I do not want to ship them. Everything else about its work was right and must be preserved. START by reading these files before you touch anything: scene3k.js, build_scene3work.py, work/scene3k.log, lib.js, tools/on_ground.js and ov/render_ov.js.
its trace115 records · 37 tools · 12 minopen in a pane ↗
Skill ×2environment, character
Bash ×9ls -la && echo "=== approved ===", cat scene3k.js, cat work/scene3k.log
Bashcurl -s -H "Authorization: Bearer $HF_TOKEN" "https://lvwerra-agen…
thinkingThe composer is still queued on the live Space, so the frame I need…
Bashcd …/film; for ep in /api/state /api/sessions; do curl -s …
replyI did not ship the rebuild — both fixes hit a hard stop. Here's where…
thinkingAll four are back. The sheet needs one more pass before I fold them…
real task text, real steps, real report — the sub-agent's own words are the last row of its rail

"completed" means it came back, not that it worked. This one's notification says completed; its report opens "I did not ship the rebuild — both fixes hit a hard stop." The row shows delivery, and the reason to open a row is that delivery and success are different things. Worth keeping the mark honest: ✓ is "handed something back", not "succeeded".

6 · Depth 2, which is where a design like this breaks

rl-llm-wiki ran three sub-agents; one of them ran five more. Inside an open row those five are one rail inside another, which holds. The question is what happens when you open one of the five: a third rail, inside a row already indented twice.

This session also shows an asymmetry worth knowing about before reading the rows: the parent pane recorded no completion for its three sub-agents, while the depth-1 agent recorded all five of its own. A notification lands in the transcript of whoever spawned the agent — so a missing outcome one level up says nothing about the level below.

a · nest it anyway
sub-agentReview remaining topic articlesno resultwrote 11:52, 16 Jul0.36 MB
sub-agentReview RLVR core articles2m 33s5 calls61.8k tok
Read ×4…/verifiable-rewards-and-reasoning/rlvr.md
Bashwc -w …/rlvr*.md
replyFour articles, grades A-, B+, B, B-. The RLVR core page is the strongest…
sub-agentReview phenomena + objectives articles2m 09s8 calls82.1k tok
sub-agentReview training-systems + preference-data2m 33s7 calls75.9k tok
three rails. Legible at this width; the text column has lost ~4.5em, and on a phone it stops working
b · one level nests, deeper levels navigate
sub-agentReview remaining topic articlesno resultwrote 11:52, 16 Jul
in this row:Review RLVR core articles·2m 33s · 5 calls · 61.8k tok← back to the five
Read ×4…/verifiable-rewards-and-reasoning/rlvr.md
Bashwc -w …/rlvr*.md
replyFour articles, grades A-, B+, B, B-. The RLVR core page is the strongest…
the row's body swaps to the deeper agent and names where you are; the indent stops at two

Recommendation: (b), but only past depth 2. One rail inside a row is worth having — it is how you see that a sub-agent fanned out at all, and it reads fine. A third is where the text column starts paying for the structure, and it is unusable at phone width, where the reader already narrows its gutter to 8px. Swapping the body and showing a crumb keeps the depth reachable without spending indentation on it, and it reuses an idea the reader already has in the card's full history ↗. Depth 3 has not been observed here — 7 of 70 sub-agents are depth 2, none deeper — so this is a guard, not a common path.

7 · The unhappy state, which must not look like "running forever"

Not a rare case: of release-video's 22 sub-agents, nine have no recorded outcome — the notification never arrived, because the pane was restarted or the message was lost to a compaction. Their transcripts are 4–12 MB each and then stop. The reader has to say that plainly.

reader · release-video · turn 14, from a pane that has since been restarted
i create a a `video-feedback.md` in this folder with my ideas of a scene. start working on it
turn 14/3810:20 AM
Bash ×3cat /data/workspaces/agent-manager/video-feedback.md, ls /home/node…
sub-agent Build Scene 1 intro animation no result wrote 10:49, 6 days ago4.1 MB
sub-agent Build the fold transition no result wrote 10:31, 6 days ago6.1 MB
sub-agent Build the browser overlay for Scene 4 no result wrote 10:37, 6 days ago9.6 MB
Bash ×26ffmpeg -v error -y …, cd /home/node/local/film2 && cat > scene2b.js
no spinner, no ✓, no claim: the mark is inert, the state is named, and the only time shown is the last write
  • The mark does not spin. A spinner is a claim about right now, and it is only earned while the pane's own process is alive. This pane has since been restarted, so nothing here is running.
  • It is not a failure either. That third transcript is 9.6 MB of real work; nothing says it failed, only that nothing came back. no result is the honest word — dashed, not red.
  • Still openable. The transcript is on disk, and its last message is usually the report the notification failed to deliver. This is the case where opening the row is worth the most.

failed and killed are different states and read differently — red ✗, not a dash. This session's notifications include four of each, so both are real.

8 · What this asks of the code, if the operator likes it

  • Drop the second disclosure. The count moves into the fold's own label; SubAgentsButton and SubAgentsList stop existing.
  • Exempt Agent from step grouping in stepsOf, so three consecutive spawns are three rows and not Agent ×3.
  • Teach the step row one new kind. StepRow gains an agent case: spinner-or-mark, the task as the detail line, the notification's facts on the right, and a body that is the task prompt plus the child trace #115 already fetches.
  • Keep the roster call. It resolves a spawn to its transcript when the launch receipt has scrolled out of the window, and supplies depth, size and last-write.
  • Stop nesting at two. Depth 3 swaps the body and shows a crumb instead of indenting again.

Nothing here needs a new data source: every value on every row above is already on disk, and #115's two endpoints already serve it.

where the numbers came from

panelsession · turnreal
2 · runningrelease-video, turn 38 (live)6 tools, 42m, 87.2k tok; two notifications (154/136 calls), one spawn at 14:28 still writing
3 · four at oncerelease-video, turn 2453 tools, 2h 39m, 1.06M tok; four notifications: 169/218/39/37 calls, 331k/329k/131k/121k tok
4 · collapsedrelease-video, turn 2111 tools, 1h 51m, 4.6M tok, two notifications
5 · expandedrelease-video, agent a441b23d…115 records, 37 tools, 708s; task text and report quoted from the transcript
6 · depth 2rl-llm-wiki, agent a99944597…five notifications inside the child (153/128/151/153/194s); depth-1 rows have none, and say so
7 · no resultrelease-video, turn 1481 tools, 58m, 1.33M tok; three spawns with no notification, last writes and sizes from stat
3 · fallbackthe-gatherer38 sub-agents in the session; the quoted row's notification is 172s, 13 calls, 41.7k tok