Lab · case study · 4 min read
Agent Review
An experiment in making design systems executable for coding agents. A React component system with its tokens and Storybook, the rules an agent reads before touching it, a controlled experiment where Claude Code added a feature inside it, the validation that caught drift, and the tool a person uses to review the consequences and send a precise correction back.
- Role
- Everything: the system, the rules, the product, the experiment and the review of it
- Built with
- React, TypeScript, Vite, Storybook 10, Vitest with axe, Playwright, Claude Code, the Storybook MCP server
- Size
- 13 components, 161 stories, 12 rules, 3 branches of one feature
- Status
- Live. The product, the Storybook, and the source and the record
The loop
-
01The system
- One token layer, thirteen components, every state a story
- Axe runs on every story as a test
- A lint that reports any value written where a token exists
-
02The rules
- Twelve rules an agent reads before changing UI
- Each names the command that checks it
- The Storybook readable by an agent over MCP
-
03The experiment
- One request, two real runs: with the rules, and without
- Both clean on every check; the differences were judgment
- A third branch seeded with the four drifts the checks exist for
-
04The review
- Each finding reproduces itself in a live preview
- A correction composed from findings, not typed from memory
- What the system learned, as a diff
The question
Coding agents can produce an interface quickly. The harder problem is whether what they produce inherits the decisions a team has already made: the components, the tokens, the states, the accessibility floor. I built the smallest complete version of the loop to find out, then ran one feature request through Claude Code twice, with the rules in its context and without, and a third time by hand with the four things the rules forbid written in on purpose.
The product
Agent Review is the tool a product team would use to inspect UI a coding agent has changed before it merges. A queue of changes, each with its validation in five lanes. A change opened to its findings, with the affected screen rendered from the branch beside them, before and after, at three widths. Selecting a finding puts the preview in the state that shows it, at the width that shows it, with the element outlined. Then a decision: accept, reject, or return to the agent with a correction.
The product and the screen it reviews are built from the same component system, as an internal tool and a product at one company would be. That is what lets a finding say “Button / secondary / compact already covers this” and open the story in one press. Try it; the three experiment branches are the first three rows of the queue.
The system
Thirteen components from one token layer in a plain CSS file, with no component kit underneath and no utility framework, because the component model is the point and the tokens should not hide behind utility names.
- Every value a token: spacing on a 4px scale, four type sizes, semantic color, two radii, one shadow, one focus ring, three durations.
- Every state a story, and not only a Default: a Button has ten, a Table nine, the review’s own components sixty-nine. 161 in all.
- Every story a test: Vitest renders each one in Chromium and runs axe on it, with violations set to fail. A token lint reports any literal written where a token exists. Playwright holds the screens to their baselines and every toolbar to its own box at 768.
Making the system legible to an agent
A design system is documentation people read. For an agent it has to be context it can act on, so the repository carries a CLAUDE.md a Claude Code session loads on entering, pointing at twelve rules in skills/ui-quality/SKILL.md. Each says what to do, how to tell whether you did it, and what happens if you did not. Three of them, as written:
- Never write a value where a token exists. How to check: the token lint. If no token fits, that is a finding about the token layer, not a reason to write a literal.
- Do not change a shared component to solve a local need. A change with a reason in writing is allowed, and the reason has a shape: what you measured, on which story, before and after.
- A new interaction pattern needs a person to look at it. Build it behind a story, describe it, and mark the change for review.
The rules end with how to describe a change, in seven sections, so a reviewer finds the shared changes and the new patterns where the rules say they will be. The Storybook also runs an MCP server, verified answering, so a logged-in agent can read a component’s props and run a story’s tests without reading the source.
The experiment
The request, verbatim, and nothing else about the feature: Add bulk actions to the customer table using the existing component system. A person should be able to select several customers and archive or export them. Each run was a fresh session in its own worktree with the repository as its only context; the record says exactly how, including why the Storybook MCP server, though connected, was not reached.
- Run 1, with the rules
- Clean on every check. Checkbox for the rows, the Table’s own selected-row composition, a second Toolbar for the selection, Button in three variants, EmptyState for a table archiving empties. Seven new stories with play functions; 92 of 92 pass with axe. One shared change, with a reason and a measurement. One new pattern, listed for review.
- Run 2, without the rules
- Also clean on every check, with the same components. It swapped the title toolbar out instead of adding a bar, flagged the same head-row growth the first run fixed rather than touch a shared file, and left the emptied table with a head and nothing under it.
- The seeded drift
- Not an agent’s run. The first run’s screen with four things done to it the rules forbid: a local button with its own stylesheet, the count in a literal ink, a bar that does not wrap at 768, and the shared Button’s padding changed so the two line up.
- What that says
- The components, the tokens and the stories did most of the work. Two runs, with and without the rules, both stayed inside the system. What the rules changed was judgment at the edges, and where a reviewer finds things.
What caught the drift
The seeded branch is where the mechanism shows. Each catch is a different instrument, and each is the check’s output as it was printed.
- The token lint: seventeen values off the token layer, with file and line.
padding: 10px 14px,border: 1px solid #d0d0d0,padding: 0 10pxin the shared Button, and fourteen more. - Axe, on the Selected story:
insufficient color contrast of 2.87 (foreground color: #7a8db8, background color: #e9effc). Expected contrast ratio of 4.5:1. The states with no story were not tested, which is what the rule about stories is for. - The 768 check:
toolbar “Selected customers” overflows its box (792 vs 734). - The visual baselines: four frames differ, one of them the Invoices screen, which the change never mentioned. Every compact Button is 4px wider: the shared change, seen where nobody was looking.
Human review
Three decisions, made in Agent Review the way a reviewer would make them. The first run: accepted, both open questions in its favour, its branch merged. The second: returned with two corrections, neither a defect a check can see. The seeded branch: returned with all five, two of them blocking, and Accept disabled while they stand, with the reason beside the button rather than greyed out silently.
What changed in the system
The review is only worth doing if the system is different afterwards. The diff after the three decisions:
- The rules. The bar of actions over a selection joins the list of common misses. The inline confirmation becomes the system’s second shape for a destructive action, with the first run’s story as its reference. Rule 9 says what a reason for a shared change looks like, and that “so the buttons line up” is not one.
- The component and the product. The Table’s control column fits a checkbox without growing the row. The customer table has bulk actions, seven stories reach its new states, and an emptied table renders the system’s EmptyState.
- The tests. The 768 check measures every toolbar’s own box; the visual tests start their own server rather than reuse whatever is on the port, after the first run’s visual pass measured a different checkout and passed.
What I learned
The most useful part of a design system, for an agent, turned out to be the part that runs. Two runs with and without twelve carefully written rules both stayed inside the system, because the components, the tokens and the stories were there as code and the agent read code. The rules earned their keep at the edges: the empty state nobody asked for, how to treat a shared file, where to write down a new pattern. A smaller job than I expected, and a more precise one. The checks also found more defects in my own work than in the agent’s, which is the same instrument pointed at the project; a system that measures itself is the only kind an agent should be trusted inside.
And the judgment did not go away; it moved. Not to approving or rejecting, which a check could do, but to the questions a check returns unanswered: is this shared change a defect exposed or a shortcut taken, is this pattern one the product wants, is this red frame the feature or the break. The product is built so that those are the only questions left on the screen when a person gets to it.