top of page

Accessibility Testing You Can Run Today

Shawn West
Mar 23
10 min read

Updated: 2 days ago

The signup form scored 100 on Lighthouse. Green across the board. axe reported zero violations. By the numbers on the dashboard, it was accessible — and the team shipped it.


Then a support ticket arrived from a user who navigates with a keyboard, not a mouse. They tabbed into the email field, into the password field, and then focus disappeared — no visible ring, no highlight, nothing on screen to say where they were. They were pressing Tab into a void. When they guessed and hit Enter, the form rejected their password, and the only signal was the field border turning red. A sighted mouse user sees red and knows. A screen-reader user hears silence. The error was never announced, never associated with the input, never surfaced to anyone not looking at the exact pixels that changed color.


None of that showed up in the automated scan. All of it showed up in a ten-minute manual check: unplug the mouse, tab through the form, then turn on the screen reader and try to submit it wrong on purpose.


That gap — between a passing scan and a usable product — is the subject of this article. Automated accessibility tools are fast, cheap, and genuinely useful, but they inspect the structure of a page, not the experience of using it. They can tell you an input has no label. They cannot tell you that your focus indicator is invisible, that your error message is announced to no one, or that your custom dropdown traps a keyboard user with no way out. To catch those, you layer three checks that each see what the previous one is blind to, in the order you should run them.



Why a green score isn't a passing grade


Start with the honest limit of the tool you already have. Automated scanners evaluate rules checkable by reading the DOM: is there an alt attribute, does this input have an associated label, is this contrast ratio above threshold, is the heading order sane. Those are real problems worth catching automatically. But a large share of what makes a page usable — whether the alt text is meaningful, whether focus order matches reading order, whether a custom widget operates with a keyboard, whether a live update is announced — cannot be read off the static DOM. It has to be operated to be found.


How big is the blind spot? Precise figures get thrown around, but they vary by study, ruleset, and how you count a "criterion," so treat any single percentage with suspicion. The defensible, widely-cited framing: automated tooling reliably detects only a minority of WCAG success criteria — commonly cited as somewhere around a third — because most criteria require human judgment. Keyboard and screen-reader testing pick up a large additional share; a remaining slice needs genuine expertise (complex ARIA patterns, cognitive load, meaningful sequence). Deque, which builds axe, says as much itself: automated testing is a first pass, not a certificate. (See Sources.)


Do this week: read your last "accessible" page's automated report as a floor, not a ceiling. Write down the three things it can't tell you — is focus visible, is the flow keyboard-operable, are errors announced — and treat those as unchecked until you run the next two layers.


The developed example: one signup form, three checks


(Developed example — composite scenario, assembled from patterns we see repeatedly on real forms.)


Take one flow — a signup form with email, password, a "Show password" toggle built as a <div>, a country dropdown, and a Submit button — and run it through each layer in order.


Layer 1 — automated (axe / Lighthouse), 30 seconds. You run the scan. It flags two real issues: the country dropdown's trigger has no accessible name, and the helper text under the password field sits at a 3.9:1 contrast ratio against white — below the 4.5:1 minimum for normal text. You fix both. Re-run. Zero violations, Lighthouse 100. The scanner is now satisfied. It has passed the "Show password" toggle without comment, said nothing about the focus indicator, and never attempted to submit the form. It read the page; it didn't use it.


Layer 2 — keyboard-only, 3 minutes. You unplug the mouse and press Tab from the top. Email: fine. Password: fine. Now the "Show password" toggle — the one built as a <div> — Tab skips right past it. It isn't in the tab order at all, because a bare <div> isn't focusable. A keyboard user can never reveal their password. You keep tabbing. The country dropdown opens on Enter, but once open, Tab moves focus behind the open menu instead of through its options, and Escape does nothing: a focus trap. And through the whole run, you notice you're guessing where you are, because the focus outline was removed in CSS (outline: none) and never replaced. Three defects the green scan reported as clean.


Layer 3 — screen-reader spot-check, 5 minutes. You turn on NVDA (or VoiceOver) and submit the form wrong on purpose. The screen reader announces the email field and the password field correctly — the labels are real. But when validation fails, it says nothing. The error text appears visually in red beside the field, but it isn't associated with the input and isn't in a live region, so the screen reader never speaks it. The user is told, in effect, that submitting did nothing. Then you tab to Submit and the reader announces "button" — no name — because the label was applied with a background image and no text. Two more defects, both invisible to layers 1 and 2.


One form. The scan found 2 issues in 30 seconds; keyboard testing found 3 more in 3 minutes; the screen reader found 2 more in 5 minutes. Three lenses, seven defects — and six of the seven would have shipped on the strength of a perfect Lighthouse score.


What each check catches — and misses


The layers aren't ranked by quality; they're ranked by what part of the experience they can observe.


Check

Catches

Misses

Automated (axe / Lighthouse / Pa11y)

Missing alt, inputs without labels, low contrast in static text, missing lang, empty links/buttons, broken heading order, missing ARIA roles

Whether alt text is meaningful, whether focus is visible, keyboard operability, focus order, whether dynamic content is announced, whether errors reach the user

Keyboard-only

Elements not reachable by Tab, focus traps, invisible focus indicators, illogical focus order, modals that don't trap or don't close on Escape, mouse-only controls

Whether announcements are correct, whether names/roles/states are spoken, whether contrast is sufficient, whether alt text conveys meaning

Screen-reader spot-check

Unlabeled or mis-named controls, unannounced errors and status changes, headings not exposed as headings, decorative images that speak, wrong roles/states

Contrast (you're listening, not looking), subtle visual layout issues, and anything the flow you tested didn't touch

Contrast check

Text below 4.5:1 (normal) / 3:1 (large, 18pt+ or 14pt bold+), UI components below 3:1

Whether the low-contrast text even matters, focus visibility, everything non-visual


Read the "Misses" column as your to-do list: each layer's blind spot is the next layer's job. That is why order matters — and why no single tool is a pass.


Do this week: after your next automated run, cross off only the "Catches" cell you actually verified. The uncrossed cells are your remaining test plan.


The keyboard checklist you can run in three minutes


Keyboard testing is the highest-yield manual check and needs zero tooling. Unplug the mouse — don't just "avoid" it, remove the temptation — and run this on any flow:


  • Tab reaches everything. Every interactive element (links, buttons, inputs, custom widgets) is reachable with Tab / Shift+Tab. Anything you can click but can't tab to is a defect.

  • Focus is always visible. At every stop you can see, without hunting, where focus is. If you ever lose it, note the element — a removed outline with no replacement is the usual cause.

  • Focus order matches reading order. Focus moves top-to-bottom, left-to-right the way the content reads. Jumps that skip around or dive into a hidden menu are defects.

  • Operate, don't just reach. Enter activates links and buttons; Space toggles buttons and checkboxes; arrows move within radio groups, menus, and dropdowns. A control you can focus but not operate is still broken.

  • Modals behave. Opening a dialog moves focus into it, Tab cycles within it (a trap, the good kind), Escape closes it, and focus returns to the control that opened it.

  • No trap you can't escape. You can always Tab your way back out to the page. If you get stuck inside a widget, that's the worst keyboard defect there is.


Do this week: run these six checks on your single most important flow — signup, checkout, or login. Log every failure as a bug with the keystroke that triggered it. That list beats any scan report.


The escalation: how to know you've hit each ceiling


Accessibility testing is a discovery exercise, not a checkbox. You escalate through four layers, and each has a ceiling — a signal it has told you everything it can. Running them out of order wastes effort: you don't want a human hand-checking contrast a tool measures perfectly, or a tool vouching for an experience only a human can judge.


  1. Automated first. Cheapest, fastest, runs in CI on every commit. Ceiling reached when the violation count sits stable at zero and new failures are only regressions of rules you already fixed. Beyond that, the tool has no more to say about experience.

  2. Keyboard second. No tooling, three minutes, catches the largest class of defects the scan can't. Ceiling reached when you can complete the whole flow with the keyboard, focus is visible and logical throughout, and nothing traps you. You've now proven operability but not comprehension.

  3. Screen reader third. Ten minutes on the key flow tells you what a non-visual user experiences. Ceiling reached when every control announces a sensible name, role, and state, and every status change and error is spoken. Interpreting complex widgets reliably takes practice — the signal to escalate.

  4. Expert last. Complex ARIA patterns, cognitive-load review, meaningful-sequence judgment, audits against the full WCAG criteria. You've hit your own ceiling when you're guessing whether a custom pattern is "right" rather than observing that it works. That guess is the moment for specialist review or usability testing with disabled users — the only layer that validates the whole experience rather than its parts.


The discovery move underneath all four: the defects that hurt most are the ones the cheaper layer is structurally blind to. So the question driving the escalation isn't "did it pass?" — it's "what can this layer not see, and have I run the layer that can?" That is the same instinct behind why we test at all: you choose the check that can observe the failure you're worried about, not re-run the one that already can't.


Do this week: map your current testing to these four rungs and find the highest one you actually run. If you stop at rung 1, your next improvement isn't a better scanner — it's three minutes with a keyboard.


Wire the floor into CI so it never slips


The automated layer is the one you can make permanent. Run axe inside your existing Playwright suite so no new violations can merge while you burn down the existing ones:


import { test, expect } from '@playwright/test';
import AxeBuilder from '@axe-core/playwright';

test('signup page has no new accessibility violations', async ({ page }) => {
  await page.goto('/signup');

  const results = await new AxeBuilder({ page })
    // Scope to the ruleset you're committing to enforce.
    .withTags(['wcag2a', 'wcag2aa'])
    .analyze();

  expect(results.violations).toEqual([]);
});

The pattern that works on a real codebase: zero new violations allowed, existing violations tracked as debt. A greenfield page asserts toEqual([]). A legacy page with known issues gets a temporary allowlist (axe supports .disableRules([...]) or .exclude(selector)) with a ticket per exclusion, so the count only ever goes down. Gate it the way you'd gate any test — smoke scans on every PR, the fuller sweep on merge — the same staged rhythm that keeps an end-to-end suite from becoming the team's biggest pain point.


What CI cannot do is press Tab. Automated checks are your regression floor; they stop the two-out-of-seven class of defect from creeping back. The other five still need a human before release — which is why the manual layers belong in your definition of done, not your acceptance criteria: "keyboard-tested and screen-reader spot-checked" is a bar every story clears, not a feature-specific requirement someone remembers to write down.


Do this week: add the axe-in-Playwright test above to one route and wire it into your PR check at "no new violations." You'll have a permanent floor by Friday.


Where the ten-minute check ends and expertise begins


Two more layers deserve naming, because they're where manual testers stall most often.


Errors and status. The signup story turned on a visual-only error. The fix is mechanical — associate the message with its input (aria-describedby) and put status changes in a live region (aria-live="polite") so they're announced without stealing focus. The rule: never signal state with color or position alone. If the only way to know something changed is to see a pixel change color, a screen-reader user doesn't know. This is exactly the failure that exploratory testing — a discipline, not a hack surfaces, because you find it by trying to break the flow, not confirm the happy path.


ARIA, used sparingly. The strongest accessibility fix is usually deleting ARIA, not adding it. Native HTML carries semantics for free: a <button> is focusable, operable by Enter and Space, and announced as a button with no attributes at all. A <div role="button"> reproduces none of that until you hand-wire tabindex, keydown handlers, and state — and most re-implementations are subtly wrong. The "Show password" toggle in our example failed for exactly this reason. The rule: reach for a native element first; add ARIA only to fill a gap native HTML genuinely can't.


Do this week: grep your codebase for role="button" and role="link". Each hit is a candidate to replace with the native element — a fix that usually removes code and adds accessibility at once.


Start today


You don't need a specialist or a budget to move from "passed the scan" to "usable." You need order: wire the automated scan into CI as your floor, unplug the mouse and run the six-item keyboard checklist on your most important flow, then spend ten minutes with a screen reader submitting that flow wrong on purpose. Each layer catches what the last one structurally cannot. When you're guessing whether a complex widget is "right" rather than watching it work, that's your signal to bring in expert review — and, eventually, disabled users themselves, the only test that validates the whole experience.


Pick one flow. Run all three checks before your next release. You will find something the green score swore wasn't there.


Sources


  • W3C — Web Content Accessibility Guidelines (WCAG) 2.2. Success criteria, contrast minimums (1.4.3 / 1.4.11), and keyboard-operability requirements referenced throughout. https://www.w3.org/TR/WCAG22/

  • Deque Systems — axe-core documentation. The maker of axe on what automated testing covers and why it is a first pass rather than full coverage. https://github.com/dequelabs/axe-core and https://www.deque.com/axe/

  • WebAIM — screen reader and accessibility analyses. Widely-cited practitioner data on real-world accessibility issues; basis for the qualitative "automated tools catch a minority of criteria" framing. https://webaim.org/

  • MDN Web Docs — ARIA and accessibility. Guidance on native semantics vs. ARIA and correct use of aria-describedby / aria-live. https://developer.mozilla.org/en-US/docs/Web/Accessibility


Automated-detection share is presented qualitatively (a "minority," commonly cited as roughly a third) rather than as a precise figure, because reported numbers vary by study and ruleset.


Keep learning. This article is part of the Software Testing Foundations path in the ShiftQuality Learning Center.

bottom of page