Theme studio

Design a theme, preview it live, then export it. Saved in this browser.

Quick picks #10b981
Generated scale
50100200300400500600700800900
Some text is below AA
Accessibility

Shipping WCAG evidence from unit tests

Nexera UI tags its tests with WCAG 2.2 success criteria and builds the conformance table from the tests that pass, so every claim points at a check.

NAccessibility team· 5 min read

Most component libraries describe their accessibility in prose: "fully accessible", "WCAG compliant", "keyboard friendly". Those sentences are hard to check and easy to let drift. In Nexera UI, the accessibility table is generated from the test run, and a success criterion only appears next to a component when a passing test supports it.

The problem with accessibility claims

A claim written by hand is true on the day someone writes it. After that, a refactor can drop a focus ring, a new variant can lose its label, or a prop can stop reaching the right element, and the claim stays the same.

We wanted the opposite arrangement. The tests already exercise keyboard paths, roles, names and live regions, so the evidence exists. What was missing was a link between each test and the criterion it demonstrates, and a report that reads that link instead of anyone's memory.

Tagging a test with its criteria

Every test that demonstrates a WCAG criterion starts its title with the criterion numbers in square brackets. Here is one from the Button tests:

app.tsx
it("[2.1.1] activates with a click, Enter and Space, once each", async () => {
  const onPress = vi.fn();
  const { user } = renderUI(<Button onPress={onPress}>Approve</Button>);
  await user.click(screen.getByRole("button"));
  expect(onPress).toHaveBeenCalledTimes(1);
  await user.click(document.body);
  await user.tab();
  expect(screen.getByRole("button")).toHaveFocus();
  await user.keyboard("{Enter}");
  expect(onPress).toHaveBeenCalledTimes(2);
  await user.keyboard(" ");
  expect(onPress).toHaveBeenCalledTimes(3);
});

The tag [2.1.1] says this test is evidence for Keyboard. The rule for contributors is to tag only what the assertions show. A test that presses Enter can claim 2.1.1. A test that only renders the button cannot.

Direct and supporting evidence

Some criteria cannot be exercised in jsdom. It has no layout and computes no CSS, so a unit test cannot measure a focus ring or a 24 px target. For those, a test can still check that the right classes or markup are present. We tag those tests with a tilde:

app.tsx
it("[~2.4.7 ~1.4.11] keeps the focus ring and halo classes after class merging", () => {
  // asserts the shared focus-ring classes survive a consumer className
});

The report keeps the two kinds apart. Direct evidence means a passing test exercises the behaviour. Supporting evidence means a passing test checks the classes or markup that implement it, and the browser run or manual testing has to confirm the result. The Button row of the table lists 2.1.1, 2.4.7 and 4.1.2 as direct, and 1.4.11, 1.4.12 and 2.5.8 as supporting.

How the report is built

pnpm --filter @nexera-ui/react a11y:report runs Vitest with the JSON reporter and then scripts/a11y-report.ts, which reads the results and applies a few strict rules:

  • A criterion is listed for a component only when at least one tagged test for it passed.
  • If any tagged test for that criterion and component failed, the criterion is dropped for that component.
  • Untagged tests are counted and reported as proving nothing.
  • A tag that is not a WCAG 2.2 criterion stops the script with an error, so a typo cannot invent evidence.

The script writes two files: a11y/conformance.json, which the docs site reads, and a11y/CONFORMANCE.md, a table by criterion and by component.

Three sources of evidence

Unit tests are one source. The report adds two more.

The browser run. A Playwright suite opens every story of the built Storybook and runs axe with the wcag2a, wcag2aa, wcag21a, wcag21aa and wcag22aa rule sets in Light, Dark and RTL at 1280 px, then checks reflow at 320 px. This is where colour contrast (1.4.3), target size (2.5.8) and reflow (1.4.10) are measured, because the jsdom axe helper switches off color-contrast. Disabled elements are exempt from text contrast under WCAG, so contrast results on disabled controls are ignored and everything else is reported. The report records the run only when it passed.

The token audit. The tokens package lists 96 colour pairs that must meet AA in both modes: 4.5:1 for text and 3:1 for control boundaries, focus indicators and graphics. Five pairs in the Figma tokens fail today, such as sidebar/text-muted on sidebar/active at 4.27:1. They are documented in KNOWN_CONTRAST_EXCEPTIONS, and a test asserts that the failing set equals that list exactly, so a new failure breaks the build and a fixed one forces the list to shrink.

What the evidence does not cover

The generated files open with the same caveat: this is evidence from automated checks, and it is not a legal conformance claim. Nothing in the tests or the design files claims full WCAG conformance.

Some things are outside what a component library can prove:

  • Screen reader testing with real assistive technology.
  • Your content: labels, error messages, headings and alt text.
  • Your colour choices when you override tokens or build a theme. createTheme() warns when a pair drops below AA, but you decide what to ship.

For this reason, most components document their consumer duties next to the API. OTPInput, for example, leaves the label, helper and error wording, and the verification itself, to you.

Reading the table

Open the Button page or the CONFORMANCE.md file in the repository and read it as a list of pointers. Every criterion next to a component names tests you can open, run and break. When we refactor a component, the table changes only if the evidence changes. That is the property we wanted: the claim and the check live in the same place.

Keep reading

All posts →

One email when we publish.

New posts and releases, about twice a month. Or follow the RSS feed.