Rollcall

Multi-service deploy console · working prototype · synthetic data

A failed deploy the developer can read in one place and put back in one press

Something broke in a project of six services. The screen answers the five questions in the order they come.

  • What’s live, never confused with what’s newest.
  • A failed build, explained on one card.
  • A rollback, previewed before it runs.
  • One press to put it back, one to undo.

Try the prototype How it was designed →

The console on a deploy that is live and failing: a heading reading “api is live, and failing since the deploy 12 min ago”, then a white card with three figures, 14% of api requests erroring against a smaller “was 0.2%”, a p95 of 2.8 s against 340 ms, and 41 worker jobs failed against 0. Under the figures, what changed, one commit, what depends on it in a sentence, and what to do, then a sage button reading “Roll back api and worker”. Beside the card, a panel headed “The last good deploy” with its commit, how long it was live and its error rate, and under it “What depends on it” as a row of chips, web ok and worker failing.
When the deploy succeeded and the errors started, the screen reads rather than refuses. The figures since the deploy beside the figures before it, the one commit, what depends on it, and the one action, asked for rather than taken. Full size ↗, live but failing
Problem
The screen a developer opens when a deploy breaks holds every fact and answers no question.
Hypothesis
Arrange the screen in the order the five questions are asked, and the developer stops reconstructing the story from the parts.
My role
Product strategy, interaction design, visual design, prototyping and the front-end build.
Artifact
A working browser prototype. No platform behind it: the services, the deploys, the logs and the figures are one object at the top of the script.
Context
Self-directed, on synthetic data, informed by shipping this site through its own checks and by the operations floors in the case studies. Not built for or with any platform.
Next
Put it in front of developers who have been on call. Then let it propose the rollback itself, and find out whether anyone still reads the figures.

Not a reproduction of any product, and not yet tested with a developer. The project, its services, commits, authors, logs and figures are invented. I ship through a branch, a commit, the checks and a static deploy every day; cloud infrastructure across many services is the part I have read about and not run.

The screens a developer opens most are the ones that open when something broke

A developer with six services gets a message that a deploy failed, or no message and a customer. The platform has every fact: which commit is serving, which was pushed, what the build said, what the health check saw, which service calls which. Shown as two lists and a log, those facts are a puzzle with no picture on the box, and the developer reassembles the story by hand at the moment they have the least attention to spare.

Failure is a first-class state, not an edge case.

An illustration of a developer at a desk late at night with two monitors, one showing a list of services and one a scrolling log, a phone lit up beside the keyboard and a mug gone cold.
A failed deploy lands in the evening, beside a log that says what happened in the machine’s words and a phone that says a customer noticed first.

The design question Can one screen answer the five questions a developer asks when something broke, in the order they ask them, without hiding the parts?

Assumptions behind the prototype

Assumptions, not observations: that the five questions come in one order for most developers, that six services is the point where stitching the picture together by hand starts to cost time, and that a developer would rather see the log’s last lines under the sentence naming the error than the whole log first. The order came from shipping this site and reading what the log said each time something failed; a developer who has carried a pager is as likely to reorder the questions as to confirm them. The console shows six services because six fit on a screen with room to read; at sixty the table becomes a search, which is a different design.

Try the 90-second prototype

One project, six services, five situations: the whole journey, from the screen a developer opens before anything is wrong, through a first deploy, a build that fails and a deploy that succeeds and errors, to a rollback. The screen changes shape each time, because an empty state, a refusal, a reading and a record are different kinds of answer.

The whole journey, in five situations
  1. Entryall live: what is serving, before anything is wrong
  2. Emptyfirst deploy: a job that has never run, and what it needs
  3. Failurea build failed: the refusal, the live deploy untouched
  4. Failurelive, but failing: the reading, a decision asked for
  5. Recoveryrolled back: in order, with a record and a way back

Three things to try

  1. Deploy the job that has never run, then run it once. On the failed build, find which commit touched the file the error names, then open the log.
  2. On the live-but-failing deploy, read the figures before and after, then press Roll back and read what would stay.
  3. Roll it back, read who decided and on what, then undo it. Open any service for its history.

What this is

  • Designed and built by me: plain HTML, CSS and JavaScript, no framework, no build step.
  • Responsive and keyboard accessible, with changes announced. At 390px the table keeps three columns and folds the rest into the row.
  • Synthetic data and a plain script, not a platform. Source: deploy-console.js, the project as one object at the top.
  • Built on the site’s own system, extended rather than fought: one token added, and the disclosure row, the device chrome and the motion vocabulary shared with the cockpit.

Pick a scenariothe screen below changes with it

Every service

Five decisions, and what each one looks like

  1. Live is a fact about traffic, not about time

    The deploy serving requests and the deploy most recently attempted are two different things, and no cell on the screen holds both. Every row carries the live commit with how long it has served, and the newest deploy with its state: the same one, a failed one, or one still building. When a build fails, the first sentence says nothing is down.

    Why I designed it this way
    • “Latest” is the word that causes the outage call. A list sorted newest-first shows a failed deploy at the top and a developer reads it as the one serving traffic. The word “live” is reserved for the deploy that is serving.
    • The first question is answered before it is asked. The project card counts what is serving, what is a datastore and what has never deployed, so “what is going on with my system” has one line of answer before any scrolling.
    • An age is a fact too. Six days live with 0.2% errors is the sentence the failing deploy is measured against, and it sits beside the failure rather than in a history tab.
    • The empty row says why it is empty, and leads somewhere. A cron job that has never deployed reads “never deployed” and, one press down, “nothing is wrong; nothing has happened yet.” The first-deploy situation is that row taken to its end: what the deploy needs, where each thing came from, the four steps in order, and a first run to see it succeed.
  2. Treat a refusal and a reading as different things

    A build the platform would not promote is a refusal: the platform’s own verdict, set on the caution ground with the caution edge, with the live deploy untouched under it. A deploy that succeeded and is erroring is a reading: there is nothing to refuse, so the card stays white, sets the figures since the deploy beside the figures before it, and asks for a decision. The two never share a channel.

    Why I designed it this way
    • A verdict and a doubt must not look alike. If a failed build and a rising error rate are painted the same, the reader learns that the color means nothing in particular. Caution ground is spent on one thing: a deploy the platform refused.
    • The figure carries its before. 14% means nothing until “was 0.2%” sits under it. The pair is the whole sentence, and the before is in the caution ink because it is the reading, not the verdict.
    • Correlation is stated as correlation. The errors started when the deploy did, and the commit calls a service outside the project. The card says both and lets the developer decide which one broke.
    • A refusal snaps; a reading arrives. The status line for a failed build appears at once on its ground; the reading rises in, like everything else the screen composed. The same rule the cockpit draws, applied to a second screen.
    1. A caution-outlined card on a pale caution ground, headed “a3f9c1e failed at the build, step 3 of 4”, with four labeled answers under it: what failed, with the TypeScript error line; what changed since it last built, two commits, one marked as touching src/billing/invoice.ts; what depends on it, web and worker, both still on the live deploy; and what to do, fix the build and push, there is nothing to roll back. A folded row offers the last nine lines of the build log and a quiet button reads “Retry the build”.

      A refusal The platform would not promote it. Caution ground, the live deploy untouched, and the one thing to do. Full size ↗, the refusal

    2. A white card with a dark edge, headed “Since the deploy”, with three figures: 14% of api requests erroring, was 0.2%; p95 2.8 s, was 340 ms; 41 worker jobs failed in 12 min, was 0. Under them, what changed, one commit; what depends on it, in a sentence naming worker as failing and web as unchanged; and what to do, roll api back to 7e1b2c9 with worker, with a note that the tax service it calls is outside the project and may be the thing that broke. A sage button reads “Roll back api and worker”.

      A reading The platform has nothing to refuse. White card, the figures beside their befores, and a decision asked for. Full size ↗, the reading

  3. Gather the failure; do not make the developer reconstruct it

    One card answers what failed, what changed since it last worked, what depends on it and what to do, in that order, each with its fact: the command and the line it stopped on, the commits since the last good deploy with the one touching the named file marked as a suspect, the dependents with their health, and a sentence that ends in an action. The log sits one step down, with its failing line marked again.

    Why I designed it this way
    • An error message is an interface. “Build failed” is a status; it says something happened, in the machine’s words, and ends the conversation. What happened, to what, why, and the one thing to do next is a sentence a person can act on.
    • A suspect is not a verdict. The commit that touches the file the error names is marked, in words, and the line under it says that is a suspect. The developer who wrote the other commit deserves not to be accused by a color.
    • Progressive disclosure is order, not hiding. The line is on the card because most readers need only the line. The log is one press down because the rest need to believe it. Nothing load-bearing is behind the fold.
    • Retry is offered, and honest about itself. Pressing it says a retry builds the same commit and fails on the same line. A button that quietly runs the same failure again is a button that teaches people to press twice.
  4. Roll back as a coordinated act, previewed, with a way back

    Rolling api back takes worker with it, because the two moved in one push and worker calls api, and the plan says so before anything moves: every service is a row, the two that change show their commits, and the four that stay show the reason they stay. One press, in dependency order, shown in that order. Undo takes the button’s place, inert for half a second. The record carries who decided, when, from what to what, and the figures it was decided on.

    Why I designed it this way
    • What stays is half of what makes it safe. A preview that lists only the two services changing leaves the developer wondering about the other four. The rows that say “unchanged” and why are the ones that let somebody press.
    • A preview is the confirmation. There is no dialog after the press, because the plan was the dialog, and it was read rather than dismissed. The way back sits where the action was.
    • The record is complete on purpose. The cockpit’s override saved the reason and admitted the rest was not written yet. Here the who, the when, the from, the to and the figures are all written with the deploy, so a decision can be reconstructed by someone who was not there.
    • The order is stated. api first, then worker, because worker calls api and a worker rolled back onto an api that has not moved yet is a second incident. The plan says the order and the status line repeats it.
  5. Keep every service in reach, and let nothing look like nothing

    The answer zone is assistance, not the only route. The project card draws who calls whom, and every node opens its row. Every service is a row, every column the answer read is in the table, every column sorts, every row opens on its history with the failures kept, the rows run comfortable or compact, and a row with no deploys says so and can be deployed from there. If the answer zone vanished tomorrow, this would still be a deploy console.

    Why I designed it this way
    • Density is a kindness when attention is short. A developer on call wants every service on one screen, not six cards with a lot of air in them. The compact end is the one a team would actually run.
    • A diagram for the project, words for the service. How six services connect is one drawing, and it answers “what is going on with my system” at a glance, with the failing line in the caution ink. What one service calls and what calls it is a row of words with its health beside each, because at that level the reader wants the one thing wrong, not the shape.
    • The history teaches. A deploy list with the failures kept in it, each with the line it failed on, is how a developer learns the shape of this project’s bad days.
    • A phone gets three columns and folds the rest. The service with its state under it, its health and the button; everything else one press down, inside the row it belongs to.
What one round of critique changed
  • The button says what the plan does. It read “Roll back api” under a sentence saying worker goes too; one rule now names everything the plan will take, on the button, in the plan and in the status line.
  • The plan leads with what changes. The two changing services sat under an “unchanged” row in project order; they sort to the top.
  • The rollback happens in order on screen. The plan promised api first, then worker, and the press landed in one frame; the building state the data carried is shown, step by step.
  • A row opens like a drawer. The detail appeared at full height while the rows under it were shoved down; the panel’s height is the animation now, and the rows follow its edge.
  • The screen stopped lecturing. Three lines of this page’s argument had been written into the product and came out; the words that remain are the product’s.

What I need to learn from developers

The prototype shows an arrangement. It does not show that the arrangement finds the fault faster, and the first session with somebody who has carried a pager would change it.

A service list and a deploy log

The facts as most consoles hold them, in two lists and a log.

compared with

The same facts, arranged as five answers

This screen: live against newest, the failure gathered, the plan previewed.

Over matched incidents: a build that fails on one line, a deploy that succeeds and errors, a rollback that takes a second service with it. Do developers name the fault and the action sooner, and do they open the log once the card has quoted the line?

  1. Is the order right for a team?

    The five questions came from one person shipping alone. On a team, the first question may be who pushed, and an experienced developer may want the log before the summary. If so, the log moves up and the summary becomes its caption.

  2. Does anyone read the before?

    The reading leans on “was 0.2%” being read under 14%. If it is skipped, the before moves into the headline; if that is skipped, the card is asking for attention the incident does not have, and the figure belongs on the row of the service it is about.

  3. Does the plan reassure, or delay?

    Six rows of “unchanged” are meant to make the press safe. At two in the morning they may be six rows between a developer and the fix. If so, the unchanged rows fold, and the two that change stand alone with one line saying the rest stay.

Nothing on this screen acts without a press. The version after this one would propose the rollback when the figures cross a line, hold the press for a person, and write down every proposal with the figures behind it. That is the line between a console and an agent, and its interface is this screen with one more column: who decided, and whether anyone checked.

The hard part of designing for a developer’s bad day is not showing the facts. The platform has all of them. It is deciding which question each fact answers, putting the answers in the order the questions come, and making sure nothing on the screen looks like nothing.