A case study in why “the automated tests are green” and “the bug is fixed” aren’t always the same claim — and what it took to close the gap between them.
The setup
The bug, as first reported, sounded simple: in toolcrib’s demo app, a collapsible bottom panel (“Live AI Event Bus Monitor”) has a Collapse button that’s supposed to shrink the panel down to exactly the height of its own toolbar. Instead, the toolbar was visibly cut off — clipped shorter than it should be.
User (new turn — the opening report, not an interjection): “in the demo, the bottom panel is ‘clipped’ when collapsed. see image. it should only collapse to the exact height of the toolbar.”
What actually happened over the next two rounds of investigation is a good illustration of a specific failure mode: an agent can verify a fix thoroughly, with real browser automation and real pixel measurements, and still be confidently wrong — because it’s asking the automation the wrong question. Getting from “my tests pass” to “the bug is actually gone” took a human physically dragging a window and noticing something no synthetic viewport sweep had thought to check.
A note on how to read the quotes below. Some of these are ordinary turns — the agent finishes a response, the user replies. Others are marked interjection — the user’s client surfaces these as arriving while the agent is still mid-tool-call, not after it stopped to report back. That distinction matters here specifically: several of the clarifications below didn’t wait for a summary to react to. They landed in the middle of an unrelated step, sometimes before the agent had even finished the previous piece of the investigation, which is part of why the second bug got found as fast as it did — the human side of this wasn’t reading reports and replying, it was watching the work happen and correcting course in real time.
Round 1: the panel doesn’t account for its own resize handle
Investigating
The demo’s Collapse button worked by picking a fixed percentage (MAIN_SPLITTER_MIN_SIZE = 5) and telling the <Splitter> component to jump to it. Reading Splitter.tsx‘s own code:
flex: `0 0 ${split}%` // top panel
flex: `1 1 ${100 - split}%` // bottom panel
Those two percentages summed to exactly 100% of the container — but the Splitter also renders a resize handle between the panels, a fixed 0.625rem (10px) strip that isn’t part of either percentage. Nothing reserved room for it.
First fix attempt (demo-only)
The first pass replaced the fixed 5% with a measured target: use useAdaptiveSize (a real ResizeObserver-backed hook) on both the whole Splitter and the toolbar itself, and compute the exact percentage needed to make the panel match the toolbar’s real height.
Tool calls that mattered here:
Writea temporary Playwright spec (investigate-auto-density-temp.spec.ts-style pattern) driving a real dev server, clicking Collapse, and reading backgetBoundingClientRect()on the panel vs. the toolbar.- First real measurement: panel came out to 18px against a 28px toolbar — a shortfall, but the interesting part was that this 18px number stayed suspiciously constant no matter what viewport height was tested.
Finding the real root cause
A shortfall that doesn’t scale with container size is the signature of a fixed pixel loss, not a percentage-math error. That pointed straight at the one fixed-size thing in the layout: the resize handle. Confirmed by inspecting the panel’s own computed flex style directly in a real browser — the math traced exactly to the handle’s un-reserved 10px.
The actual fix landed in Splitter.tsx, not the demo:
flex: `0 0 calc(${split}% - ${HALF_HANDLE_SIZE_REM}rem)`
flex: `1 1 calc(${100 - split}% - ${HALF_HANDLE_SIZE_REM}rem)`
— splitting the handle’s footprint evenly between both panels, and exporting SPLITTER_HANDLE_SIZE_REM so a consumer computing a pixel-exact split could account for it themselves.
Verification, thoroughly:
- Real Playwright measurement at three viewport heights (900px / 650px / 500px) — panel landed at exactly 28px in all three.
- Full
vitestsuite: 1317/1317 pass. - Full e2e chromium suite: 66/66 pass, including the whole-app
accessibility.spec.tsWCAG sweep. - Gemini’s PR review flagged a real-sounding concern (a negative
calc()could break theflexshorthand entirely at the split’s0/100boundary) — verified false via a directpage.evaluate()test of raw CSS behavior in Chromium: the browser clamps to0px, it doesn’t discard the shorthand. Replied on the PR thread with the evidence before merging. - Merged as PR #457.
This was a real, general defect in the shared Splitter component — not a demo-only quirk — so fixing it there meant every future Splitter consumer gets the correct behavior, not just this one demo panel.
At this point, every signal available said: fixed, verified, shipped.
An earlier decision that quietly made Round 2 possible: the commit hash in the header
Before any of this, in the same session, a much smaller exchange had already happened — and it wasn’t idle polish, it came directly out of a real mix-up: the agent had just reported a fix as live, the user checked, and what they were looking at didn’t match — because they were testing against the deployed GitHub Pages build while the fix so far only existed in local dev. Out of that confusion came the actual request:
User: “is it worth it putting a commit hash or anything in the top header?”
The first pass at answering undersold it — for a local dev server, git status/git log already answer “what am I looking at” in one command, so a header stamp there would mostly duplicate something already one command away. That framing missed the point the confusion had just demonstrated:
User: “negative, i go to the github pages version and test.”
GitHub Pages is the user’s primary interaction point with this project, not local dev — and a deployed static page has no local checkout to compare against; there’s no git status to run against https://escape-llc.github.io/toolcrib/. So it got built: vite.config.ts reads the real commit hash via git rev-parse --short=7 HEAD at build time and injects it as a literal string via Vite’s define; the demo header renders it as a small link straight to the commit on GitHub. (It wasn’t quite that simple in practice — the first version referenced the hash as a bare global, which broke a separate CI job that copies demo/App.tsx raw into a real Next.js project to smoke-test it; fixing that is its own small story, not this one.) Shipped as PR #451.
It was built to solve one specific, already-experienced confusion. It turned out to be the exact mechanism that made Round 2 tractable at all, a confusion of the very same shape but higher stakes:
- The user’s bug report —
"i am viewing demo at this commit https://github.com/escape-llc/toolcrib/commit/918a293"— was only that precise because the header made the commit hash something to literally read off the screen and paste, rather than something to guess at (“whatever’s live right now,” “probably the latest”). - Every verification script written afterward scraped that same header link (
page.locator('a[href*="github.com/escape-llc/toolcrib/commit/"]')) and logged the commit it found — so every measurement in this whole investigation carried its own built-in confirmation that agent and user were looking at the identical build, not two different deploys separated by a few minutes of GitHub Pages CDN propagation.
Neither side asked for this coordination mechanism explicitly going into Round 2 — it existed already, built for a smaller instance of the identical problem, and happened to be exactly the tool a “the fix isn’t actually live for you” vs. “the fix isn’t actually working” ambiguity needed. Infrastructure built to fix one real confusion paying for itself in a completely different situation later is its own small lesson.
Round 2: “still clipped” — and the automation says otherwise
Some time later, in the same session:
User (new turn): “i am viewing demo at this commit https://github.com/escape-llc/toolcrib/commit/918a293 and the bottom panel is still clipped.”
This is what makes it land so hard: the report comes back after the prior work had already been reported as done — a full, verified, merged fix — not in the middle of one. This is the moment the story gets interesting. The natural first move was to reproduce it exactly as reported — against the live deployed page, at that exact commit, not local dev:
- Confirmed 918a293 was the correct, deployed commit (
gh run list --workflow=deploy-demo.yml). - Wrote a fresh Playwright spec pointed at
https://escape-llc.github.io/toolcrib/, clicking Collapse and measuring the panel against the toolbar. - Swept four real viewport sizes (1400×900, 1280×720, 1920×1080, 1024×768).
Result: zero clipping, in every single case. clipped=false across the board, with the panel landing at exactly the toolbar’s height every time.
This is the part worth sitting with: the automated reproduction, run against the exact same URL the user was looking at, genuinely could not find the bug. That’s not a small disagreement — it’s two different, both-true-sounding claims: “the fix works” (verified) and “the user is still seeing it broken” (also, presumably, true). Something had to give.
Asking for what only the human could see
Rather than keep guessing blind, the next move was to ask — specifically, for the two most likely environmental explanations for a visual bug an automated test can’t reproduce:
Agent: “Could you check the browser’s zoom level… Could you do a hard refresh (Ctrl+Shift+R) to rule out a stale cached bundle?”
Notably, the RDP detail itself arrived before that question was even asked — as an interjection, surfaced mid-tool-call while the agent was still writing the viewport-sweep script, not offered in reply to anything:
User (interjection): “i am running in RDP if that matters.”
The zoom/hard-refresh question that followed was shaped by that detail. The user then ruled out both directly, as the reply to that question:
User: “Already confirmed 100% zoom” / “Already hard-refreshed, still clipped”
Then, again interjecting rather than waiting to be asked anything further, the user did something more useful than any follow-up question could have prompted: they attached a screenshot of exactly what they were seeing, live.
Reading the screenshot
The screenshot showed the toolbar’s buttons — “Export JSONL”, “Clear Log”, “Expand” — fully legible, icons intact, not visually squished mid-character the way the original bug had looked. It just sat right at the bottom edge of the captured image, with no visible breathing room below it. That was a genuinely ambiguous signal on its own: is the content actually cut off, or does the page just naturally end there in a short window?
The pivotal clue
This is the line that actually cracked it — the direct answer to one more clarifying question (“what’s the actual window height, or is something below that row getting cut off”) — but the content of the answer is a piece of empirical, hands-on testing only the human could have done, because it required physically manipulating the real browser window and watching what stayed constant:
User: “if i vary the size of chrome window by dragging, the splitter and bottom panel track exactly the same, and it is clipped consistently. it must be mismeasuring where the ‘bottom’ is.”
Read that again: the user independently rediscovered the exact diagnostic signature from Round 1 — a shortfall that doesn’t scale with window size, meaning it’s a fixed-pixel loss, not a percentage-math bug — and correctly reasoned from it to “something is mismeasuring.” That’s not a bug report anymore; that’s a root-cause hypothesis, arrived at by direct interaction with the running product in a way no synthetic Playwright sweep across four preset viewport sizes had happened to surface.
(A fifth viewport size might eventually have shown it too, if the bug were the same shape as Round 1’s — but it wasn’t proportional at all, so no viewport size would have. The automated sweep’s blind spot wasn’t “not enough viewports,” it was “comparing the wrong two elements” — see below.)
Finding the second bug
With that reframing — “trust that this is fixed-pixel, and Splitter itself is already verified correct, so look at what the demo is measuring” — the actual defect took one more targeted Playwright script to confirm precisely:
BEFORE collapse:
toolbar (role=toolbar) height=28
my ref div height=28
Card.Header (DIV.) height=45 padding=8px 16px
AFTER collapse:
panel height=28
Card.Header height=45
→ PANEL TOO SHORT (Card.Header clipped)
The demo’s own measurement code (eventLogToolbarRef) was wrapping the inner <Toolbar> element — 28px — instead of the outer <Card.Header> that actually contains it, padding and all — 45px. The panel was dutifully, correctly sizing itself to match its own (wrong) 28px target. It wasn’t broken math this time; it was measuring the wrong DOM node entirely.
This is exactly why Round 1’s automated live-page sweep reported clipped=false: that test compared the panel’s height against the [role="toolbar"] element specifically — the same wrong reference the buggy code itself was using. Two independently-written pieces of code (the fix and the test) happened to share the same blind spot, because both were built from the same mental model of “the toolbar is the thing that needs to fit.” The bug was in that shared assumption, not in either piece of code considered alone.
Right as the fix was landing, interjecting again — this arrived mid-tool-call, while the Card.Header investigation was still being written, not after any report back — the user supplied one more piece of precise framing that confirmed the diagnosis was pointed the right way:
User (interjection): “the splitter’s bottom panel is positioned ‘perfectly’ regardless of resize; it is the inner thing not calculating correctly.”
— distinguishing, correctly, between “the Splitter mechanism” (fine, already fixed, positioned exactly where told) and “the inner thing” (the demo’s own measurement target, still wrong). Then, also interjecting rather than waiting for the fix to land first:
User (interjection): “is this a toolkit issue or a usage issue?”
The honest answer, given right there mid-fix: usage. Splitter/Card/Toolbar were all behaving correctly; the demo was just pointing its own ruler at the wrong object.
The fix
Moved the measurement ref from wrapping <Toolbar> to wrapping <Card.Header> itself:
<div ref={eventLogToolbarRef}>
<Card.Header paddingMode="compact">
<Toolbar>...</Toolbar>
</Card.Header>
</div>
Verified again, the same way:
- Real Playwright measurement at three more viewport sizes (1400×900, 1024×768, 900×500) — panel landed at exactly
Card.Header‘s real 45px height in all three. - Full
vitestsuite: 1317/1317. Full e2e chromium suite: 66/66. - Filed issue #462, opened PR #463.
Closing the loop for the next person
The user asked for one more thing — again interjecting, this time while the full e2e suite was still running as part of verifying the fix, before any final report had gone out: make sure this lesson doesn’t stay buried in a commit message.
User (interjection): “we should probably augment the ai-docs with this tidbit”
A moment later, once the JSDoc tags were underway, a second interjection made sure the request wasn’t satisfied by a hand-written comment alone:
User (interjection): “also present this in the generated ai-docs as well”
The general shape of the mistake — when computing a pixel-exact “fit to this element” target, measure the outermost box whose padding/border actually needs to fit, not an inner child — got recorded in two places:
Splitter.tsx‘s ownSPLITTER_HANDLE_SIZE_REMcomment (real vendored source every consumer reads).- A proper
@manifestAntiPatternAvoid/@manifestAntiPatternInsteadJSDoc tag pair onSplitter‘s own component declaration — whichgenerate-manifest.js/generate-docs.jspick up automatically intoai-docs/component-manifest.jsonand the generated Anti-Patterns table inai-docs/CORE.md. Not hand-typed into the generated file (which the next regeneration would silently overwrite) — wired into the actual generation pipeline, so it ships to every consumer who runstoolcrib init/mergefrom here on.
What actually made the difference
Strip away the specific bug and what’s left is a fairly clean demonstration of where each side of a human+agent debugging session has a real, non-overlapping advantage:
- The agent’s advantage: systematic coverage. Four viewport sizes, checked in seconds, with exact pixel measurements and zero fatigue. When Round 1’s real bug was found, it was found because the shortfall was measured precisely enough (18px vs. 28px) to notice it was constant, not proportional — a distinction a human eyeballing a screenshot would likely miss.
- The human’s advantage: embodied interaction with the real, live thing. Dragging an actual window and watching what tracked together (splitter position) versus what stayed weirdly fixed (the clipping amount) surfaced a signal the agent’s own pre-planned viewport sweep never would have — not because the agent couldn’t have tested more viewports, but because more of the same kind of test would have kept confirming the same wrong conclusion. The fix wasn’t “test harder,” it was “test a different comparison” — and that reframing came from a person physically manipulating the product, not from a bigger automated matrix.
- The failure mode this avoided: an agent trusting its own green tests past the point where a human says the bug is still live. The tempting move after Round 1’s live-page sweep came back
clipped=falsefour times in a row would have been to conclude the user was looking at a stale cache or a rendering quirk and stop there. Asking two direct, falsifiable questions (zoom level, hard refresh) — and believing the answers when they ruled out the easy explanations — is what kept the investigation moving instead of stalling on “well, my tests say it’s fine.” - The unglamorous prerequisite: none of this works if the two sides can’t first agree on what they’re even looking at. The commit-hash header is what let a bug report carry an exact, checkable build identity instead of “the live site” — collapsing an entire class of false leads (stale cache, propagation delay, wrong deploy) before they could ever eat investigation time.
Leading the agent vs. following it
There are two distinct modes a human can be in relative to an agent’s work, and this whole story is really about the difference between them.
Following looks like: hand off a well-specified task, let the agent run, read the conclusion when it stops. Most of this session was following mode, and it’s the mode that makes an agent actually useful for throughput — the original bug report that opened Round 1 was a single, complete turn, sent and then left alone; what came back was a fully investigated root cause, a fix, a three-viewport measurement pass, a 1317-test unit suite, a 66-test e2e suite, a Gemini review verified line-by-line, and a merged PR — all without the human needing to watch any of it happen. Following mode is what let that entire arc complete in one continuous stretch instead of a dozen back-and-forths. It’s also, not coincidentally, exactly the mode in which the agent’s own conclusion — “fixed, verified, shipped” — turned out to be incomplete, not wrong exactly, just short of the whole truth. Nothing about following mode would have surfaced that on its own; the gap only became visible once someone acted on the conclusion in a context the agent’s own tests hadn’t covered.
Leading looks completely different: stay present while the agent works, and push new information in as it becomes available, without waiting for a stopping point. Round 2 was leading mode, densely so — count the interjections: the RDP detail, arriving mid-tool-call before any question had even been asked about it; the live screenshot, sent unprompted the moment it existed; “it is the inner thing not calculating correctly,” landing while the next investigation script was still being written; “is this a toolkit issue or a usage issue?”, asked mid-fix; both documentation requests, one arriving while a full test suite was still running. None of these waited for the agent to stop and report. Each one reached the investigation at the moment it was true, not at the next natural checkpoint — which meant each one had the chance to change what got investigated next, instead of only being able to critique what had already shipped.
The two modes aren’t in tension so much as complementary, and the value of each shows up specifically where the other one is weak. Following mode is cheap for the human and excellent at grinding through anything mechanical — four viewport sizes, a full regression suite, a Gemini finding that needs a page.evaluate() to actually check — precisely because nobody has to watch it happen. But following mode has a blind spot by construction: it can only ever react to a conclusion, and a conclusion that’s subtly wrong (verified against the wrong comparison, in this case) looks exactly like a correct one from the outside, right up until someone tests it for real. Leading mode is what catches that — but it isn’t free. It costs the human’s continuous attention, which is exactly why it isn’t the default mode for everything; using it selectively, at the moments where the agent’s own plan might be heading somewhere wrong, is what made this fast rather than either extreme (full autonomy risking a half-fixed bug shipping quietly, or full moment-to-moment steering burning attention on work automation already handles fine).
The fastest path through this whole story ran through both modes, used deliberately for what each is good at: follow while the work is mechanical, lead the instant the agent’s own account of “done” stops matching what’s actually being seen.
Appendix: tool-call trail (abbreviated)
| Phase | Tool calls |
|---|---|
| Commit-hash header (earlier, unrelated) | Edit vite.config.ts (define + git rev-parse) · Edit demo/App.tsx header link · Bash real npm run build, byte-check the injected hash against git rev-parse · Bash gh pr create (PR #451) |
| Round 1 investigation | Read Splitter.tsx/DataTable.tsx toolbar code · Write + PowerShell Playwright spec (local dev server) · Read screenshot |
| Round 1 root-cause | Edit Splitter.tsx flex-basis (calc()) · Bash re-run Playwright at 3 viewport heights · PowerShell full vitest + e2e chromium suites |
| Round 1 PR review | WebFetch GitHub Expressions docs (unrelated PR, same session) · Write temp spec testing raw calc() flex-basis in isolation · Bash gh pr comment posting the false-positive resolution |
| Round 2 initial repro | Bash gh run list --workflow=deploy-demo.yml (confirm deployed SHA) · Write + run Playwright spec against the live https://escape-llc.github.io/toolcrib/ URL at 4 viewport sizes, scraping the header’s own commit-hash link on every run to confirm agent and user were looking at the identical build |
| Round 2 clarification | AskUserQuestion (zoom level, hard refresh) · direct read of user-supplied screenshot |
| Round 2 root-cause | Write targeted Playwright spec comparing Card.Header vs. the toolbar ref · Read a cropped screenshot confirming the visual clip |
| Round 2 fix | Edit demo/App.tsx (move ref from <Toolbar> to <Card.Header>) · re-run Playwright at 3 more viewport sizes · full vitest + e2e chromium suites |
| Documentation | Edit Splitter.tsx comment + @manifestAntiPatternAvoid/Instead JSDoc tags · PowerShell npm run generate-manifest / generate-docs · Bash grep verifying the new row landed in ai-docs/CORE.md |
| Shipping | Bash gh issue create ×2 · Bash git checkout -b / commit / push ×2 · Bash gh pr create ×2 |