Design & User Experience

Do Not Rewrite an Untested UI Before Mapping Its Journeys

You inherited a codebase with no trustworthy UI tests, and the next request is to make an awkward screen easier to use. My position: QA should resist a broad visual cleanup, even if the code is ugly. Introduce the improvement through one observable interaction, establish what that interaction does today, and make the change reversible. Otherwise, a nicer screen can conceal a broken workflow.

The first UI change should create evidence before it creates components

Start with a user action that has a clear beginning and end: submitting a search, correcting a form error, or returning to a list after saving. Avoid choosing a whole page as the unit of change, because a page can contain several unrelated contracts. Write down the action, its HTTP response, the resulting URL, visible text, keyboard focus, and any persisted state. Those observations give QA something more precise than “it looks the same.”

Add UX without a rewrite in a legacy backend codebase gets the scope constraint right, but I would delay the first visible patch until QA has captured the chosen workflow, because otherwise nobody can separate a pre-existing defect from a regression. This is not a demand for comprehensive coverage. For a pilot, choose one workflow and, as a working sample size to adjust, five representative records: a normal record, an empty result, an invalid input, a record with long text, and one the current user cannot access.

Put those records behind repeatable test data rather than borrowing live accounts. A PostgreSQL seed script, a fixture endpoint restricted to the test environment, or a disposable database snapshot can work; the winning choice is the one the team can reset reliably. Record the fixture version alongside the test run, because a screenshot comparison is meaningless if its underlying record changed. For an inherited application, this data-control work often exposes more risk than the first CSS edit.

Run a browser test before altering the screen. The following Playwright test is an executable starting point for a stable fixture homepage; install @playwright/test, point BASE_URL at the test server, and create the initial image with npx playwright test –update-snapshots. Move the URL to the selected workflow once its fixture is ready.

const { test, expect } = require('@playwright/test');
test('fixture page retains its rendered baseline', async ({ page }) => {
  const response = await page.goto(process.env.BASE_URL || 'http://localhost:3000/');
  expect(response).not.toBeNull();
  expect(response.status()).toBeLessThan(400);
  await expect(page.locator('body')).toBeVisible();
  await expect(page).toHaveScreenshot('fixture-page.png', {
    animations: 'disabled'
  });
});

Playwright’s documented default test timeout is 30 seconds; treat that as a runner default, not evidence that a slow page is acceptable. Keep the screenshot only if you can make its data, viewport, fonts, and time-dependent content stable. If the existing page varies by session, assert a particular status message and action instead, because a noisy snapshot teaches the team to approve differences without reading them.

A server-rendered seam usually beats a new UI runtime for the first change

There are two plausible ways to improve an old interaction. Option A: enhance the existing server-rendered form or template. It wins when the server already owns validation, permissions, and navigation; it costs some care around legacy markup and CSS. Option B: mount a React island inside the page. It wins when the interaction truly needs substantial client-side state; it costs a build pipeline, a second rendering path, and tests for synchronization with server state. For an untested codebase, I would choose Option A first unless the interaction demonstrably requires Option B, because each additional state owner expands the regression surface.

Front-End Craftsmanship: Build Faster, Cleaner UI makes cleaner interface code an attractive goal, but I would not introduce a design system as the first QA-led change, because replacing markup across screens makes failures harder to attribute. Instead, place one small template partial around the selected form or result row. Keep its inputs explicit: record data, current validation errors, and permitted actions. Leave the existing route and database write in place.

Progressive enhancement gives that seam a useful test boundary. The form should still submit through its existing HTTP method when JavaScript is unavailable. If you add client-side validation or a loading state, check that it does not change what the server accepts: client checks improve feedback, while server checks remain authoritative because requests can bypass the browser. Preserve the response’s redirect behavior, including any HTTP 302 used after a successful POST, unless a product requirement explicitly changes it.

Before editing CSS, inspect the computed styles in Chrome DevTools and identify the selector that owns the behavior. A narrowly scoped class on the new partial is preferable to changing a global button rule, because the global rule may affect screens you cannot yet test. If the stylesheet already uses CSS cascade layers, keep the new rule in the appropriate layer; adding a layer solely for this patch can alter precedence in surprising ways. These are small decisions, but they keep the first deployment attributable to one workflow.

Preserve the old form in source control rather than maintaining two live forms behind a permanent flag. A temporary server-side feature flag can help with rollout, but two active implementations double the states QA must exercise. Give the flag an owner and removal date when it is created, because “temporary” paths tend to survive after their test fixtures disappear.

A matching screenshot does not prove the interaction still works

Visual regression is useful for catching layout drift, but its oracle is pixels. It will not tell you that keyboard focus vanished after an error, that a button’s accessible name changed, or that a saved record belongs to the wrong user. Test those as separate contracts. For a failed submission, assert the server’s error state, then press Tab through the recovery path and check where focus lands. For a successful submission, read the saved record from a fresh request rather than trusting a transient success banner.

Use axe-core 4.x, directly or through @axe-core/playwright, on the changed state. Treat its results as a detector, not certification, because automated rules cannot decide whether an error message actually helps a person recover. Check relevant WCAG 2.2 AA requirements manually as well. WCAG specifies a 4.5:1 minimum contrast ratio for ordinary text; that number is a standard’s threshold, whereas the legibility of the chosen error wording still needs human review. Prefer native HTML labels and buttons before adding ARIA attributes, because native controls already supply keyboard behavior that custom widgets must recreate.

Characterization tests need careful interpretation. Suppose a baseline run measures keyboard focus disappearing in 3 of 12 fixture cases. Do not encode “focus disappears” as the permanent expected result merely to make the suite green. Mark those cases as known failures, link them to the observed states, and specify the corrected behavior for the case being changed. That preserves the distinction between documenting the inheritance and endorsing it.

Test permissions through the real route, too. A hidden “Edit” button is not an authorization boundary, because a user can send the request directly. Use two fixture accounts with different access, verify the rendered affordance for each, and issue the disallowed request separately. If the application returns an established denial response, assert that contract without forcing a new status code into this UI task. This keeps the improvement from accidentally turning a presentation change into a security change.

Finally, run the same core checks with JavaScript disabled if the workflow previously worked without it. Playwright can create a browser context with javaScriptEnabled: false. That test is especially valuable for a server-rendered seam, because it reveals when an enhancement has quietly become a dependency. It need not cover every animation or convenience feature; it must cover the submit-and-recover path that users already had.

A rollout is testable only when its stop condition is observable

Release the altered workflow to a small cohort or test environment that uses production-like authentication and data shapes. Before enabling it, specify a rollback condition tied to that workflow: for example, choose a provisional threshold of 1% failed submissions over a defined observation window, then adjust it against the site’s own baseline. A percentage without a denominator, event definition, and window is not an actionable alarm.

Instrument the action rather than logging every click. OpenTelemetry spans can connect the browser request to the server handler if the application already uses tracing; otherwise, structured server logs keyed by route and outcome are enough for this pilot. Do not put form contents or personal data in event names, because debugging does not justify leaking user input. Compare the old and new paths using the same definition of success, such as “valid submission followed by a persisted change,” rather than counting a displayed toast as success.

Watch performance at the interaction, too. Google’s published “good” threshold for Interaction to Next Paint (INP) is at or below 200 milliseconds at the 75th percentile; use it as an external reference, not a claim that this legacy page already meets it. Browser data from the changed action is more informative than a single Lighthouse CI run, because a lab run may miss slow devices and delayed server responses. Keep Lighthouse CI as a repeatable smoke check if it is already in the pipeline.

When an alarm fires, the rollback should be mechanical: disable the server-side flag or redeploy the previous partial, then rerun the original fixture test. Do not ask QA to approve a hurried second rewrite, because that would discard the very baseline that makes the failure diagnosable. Once the cohort is stable, remove the old path and its flag in a separate change, with the characterization and accessibility checks still running.

The first ticket should freeze one risky state

Tomorrow, pick the form whose failure would most inconvenience a user and create its resettable fixture before changing its markup. Capture a valid submission and one recoverable error, including the resulting focus position. If those observations cannot be repeated locally, fix that test setup first. The first UI improvement should begin only when QA can tell whether it made that specific action better or broke it.