Writing · 5 min read
What a design system needs to explain to an agent
Components and tokens tell a coding agent what an interface is made of. To build the right thing, it also needs to know what each part means, how it should behave, where the usual answer holds and where it doesn’t. I’ve been working that out in Agent Review, a prototype I built for supervising a coding agent’s design work.
Key idea 1 of 5
An agent can already build from the parts
Before I designed anything for review, I ran a small experiment. I gave Claude Code the same request twice: add bulk actions to a customer table built from a React component system. One run read twelve written design rules first. The other had only the components, the tokens and the stories.
Both stayed inside the system. Each reused existing components and wrote every value as a token, and the checks found no defect in either version.
Design systems have always been good at this. On the SAP Fiori program, one style guide and component library was much of how many teams, building at different times, inherited the same decisions. An agent is one more builder reading the same files.
Stripe’s Katie Dill made the broader case in a talk at the Lenny and Friends Summit (opens in a new tab). With more people and more automation doing the building, there is no longer always a designer in the room to fill in what was never written down, so a system has to carry intent, not only consistency. My question is narrower: which parts of that intent can a design system actually carry?
Key idea 2 of 5
Following the rules can still leave the meaning open
Agent Review’s simulated run shows where the parts run out. An agent moves 36 hard-coded values in a set of billing screens onto design tokens, changing nothing anyone can see. Two of the values are the same red: a payment that failed, and an old plan struck through because it was replaced.
The system has a token for each meaning, and today both are the same red. So either name passes the token lint, the visual comparison and the accessibility check.
The trouble comes later. If the danger red gets stronger, every replaced value that borrowed it gets stronger too, and a routine plan change starts to look like an error. Two things can look the same and say different things, and a change to the look alone changes what both of them say.
The checks established
- A token was used, not a raw value
- The screen matches its baseline
- The text clears its contrast floor
No check could establish
- Which of the two tokens was meant
- Whether the two reds should mean the same thing
- What a later change to either one will do
A passing check establishes what that check measures, and nothing beyond it.
Key idea 3 of 5
The designer decides what each thing is for
The question isn’t which token to use. It’s whether these two reds should mean the same thing, and that is a product decision. I’d answer it as I would on any team: ask what each one tells the reader, then picture the next change. A failed payment needs someone to act. A replaced plan records what changed. They should stay separate, so only the failure gets louder.
In Agent Review the agent doesn’t guess. The limits on its job tell it to ask when the evidence for a value’s meaning conflicts, so it stops with one question, the evidence for each reading and a recommendation. A person answers.
The same gap shows up far from color. An illustration, not an Agent Review feature: a list of invoices comes back empty. The system’s empty state renders correctly whether nothing is overdue, a filter hid everything, or this person can’t see the account. The first is good news, the second needs a way to clear the filter, and the third needs to say who can grant access. “Use the empty state” is followed all three times. Only someone who knows why the list is empty can say what it should help a person do next.
Key idea 4 of 5
Keep the answer, but not always as a new rule
Once a decision is made, the reflex is to write a rule. In Agent Review an answer becomes one only when a person checks a box that starts unchecked, and the rule says where it applies and why, and can be revoked. The answer about the two reds settles three more waiting changes, and a matching case won’t ask again.
A rule is only one place to keep an answer, though. Reviewing the experiment’s three versions changed the system in four places, and its twelve rules stayed twelve:
- A better test. The layout check had passed a toolbar 58 pixels too wide, because a parent hid the overflow from what it measured. Now it measures each toolbar.
- A clearer rule. Changing a shared component already needed a reason in writing. The rule now says what a reason looks like: what was measured, where, before and after.
- A condition written down. The runs asked before archiving customers, and I kept that confirmation as a pattern. The rule for destructive actions now says when a confirmation fits and when an undo does.
- A pointer to what exists. The system already had the bar of actions a selection needs. The rules now name it, so an agent finds it before building another.
Most of what a review teaches belongs in something the system already has. Each new rule is one more thing an agent, and its reviewer, reads before every run.
Key idea 5 of 5
Some answers only show up when people use the product
Some decisions can’t be written down ahead of time, because nobody knows the answer yet. In the experiment, the visual comparison flagged four changed screens on every version and couldn’t tell the feature from a break. Whether the product wanted the new confirmation took a person reading it. Whether it slows down someone who archives customers all day would take watching them work.
Agent Review is at that stage too. It’s a working prototype that plays simulated runs, informed by one small experiment, and I haven’t tested it with anyone. Next I’d put it in front of designers and engineers supervising real agent work. Can they say what the agent did from its account alone? Do they catch a decision it got wrong? Can they set up guidance without spending more time reviewing?
A design system can carry what a team has decided. Noticing what it hasn’t decided is still the job.