AI Accessibility Auditor · Case study
An accessibility auditor that writes the fix
Paste a URL. The app runs axe-core in headless Chromium, keeps every WCAG violation, and asks Claude to explain each one and hand back corrected HTML. Scan again later and it tells you what got fixed.
- Role
- Product design · Full-stack build · Agent orchestration
- Type
- Full stack product
- Timeline
- 1 day · October 2026
- Stack
- Next.js 16 · axe-core · Playwright · Claude · Neon Postgres · Inngest



The problem
Audit reports say what's broken. Not how to fix it.
axe finds the violation and points at a selector. A developer who is new to accessibility still has to work out who it hurts and what to type.
image-alt
Images must have alternative text. It doesn't say what to write.
color-contrast
The ratio is too low. It doesn't say which colour to pick.
label
Form fields need a label. It doesn't show the markup.
Scan
Paste a URL. Watch it run.
The scan is queued, Chromium loads the page, axe runs, and the page updates itself until the results land.

The fix
Every issue explained, with the HTML to paste.
Claude gets one violation at a time: the rule, the failing element and axe’s failure summary. It has to answer in a fixed shape.
explanation
Two to four plain sentences: what's wrong and who it affects.
fixSummary
One sentence.
fixCode
The corrected element only. Valid HTML, no fences, no commentary.

Over time
Scan it again. See what changed.
Each site keeps its history. Compare two scans and every issue sorts into fixed, new or persisting, keyed by rule and selector.


Under the hood
One request, five steps, nothing blocking the browser.
The route only queues an event. An Inngest function does the work in retryable steps, so a slow page or a flaky model call never times out the request.
playwright-core + a serverless Chromium
Full Playwright blows Vercel's 250 MB limit.
Drizzle on Neon Postgres
Sites, scans, issues.
Structured output via Zod
The model's answer is parsed, never regexed.
Private URLs rejected
Localhost and private ranges never get scanned in production.
Duplicate guard
Submit a site that is already queued or running and you get that scan back.
PostHog, hostnames only
Never full URLs or issue HTML.
Does the fix work?
I didn't want to take the model's word for it.
An eval harness scans fixture pages, applies every suggested fix to the HTML, and scans again. Two numbers come out, and I read them as a floor and a ceiling.
Strict
Rule plus selector gone from the re-scan.
Lenient
Fewer hits per rule, so a rewritten element still counts.
Introduced
New keys after the fix. An upper bound on regressions.
The design system
Bold outlines, flat fills, and a palette that passes its own audit.
Hand-drawn doodles on cream paper, hard offset shadows instead of blur, and every text colour checked against WCAG AA before it went in. Status is never colour alone.



Fredoka and Nunito
2px ink outline, everywhere
Offset shadow, no blur
AA ratio listed per token
Badge = label + glyph
Reduced motion in MotionConfig and CSS
How I built it
A plan, four phases, and a design spec before a redesign.
I wrote the build plan and made the calls that mattered: the browser that fits on Vercel, the database, the queue, no auth for v1. Claude Code built each phase while I reviewed, fixed the deploy, and wrote the design spec the redesign was built from.
Phase 1
URL scan pipeline: axe-core, Inngest, Claude explanations.
Phase 2
Scan history, re-scan, and comparison.
Phase 3 and 4
Eval harness, analytics, deploy.
Redesign
Tokens and a spec first, then primitives, then every page.
What I'd do next
Three things, in order.
Show the model the page
It only sees the failing element, so heading order and landmark rules score low.
Run the eval in CI
The harness exists. It should block a prompt change that makes fixes worse.
Fix the whole file
The applier already rewrites HTML in the harness. Put it in the product.
See it