01 / 05
PageAudit

A UX & GEO page auditor that tells owners exactly what to change.

Paste an address. Two audits run at once: UX conversion and AI visibility. Every finding says what's wrong, why it matters and how to fix it.

Solo2026100+ hoursLive
Claude Code in CursorClaude APIPuppeteerRailwayGitHub CI
02 / 05

Name the element. Then name the fix.

Free tools give scores and generic advice, and AI can sound certain about things it made up. So PageAudit names the exact element and the fix, grades only what it can reproduce, and says when it couldn't read a page.

Generic advice (example)

"Improve your CTA."

PageAudit

No guarantee or risk-reversal statement present

What… no such statement exists anywhere in the rendered page text
WhyVisitors choosing a plumber face real financial and property risk. …
FixAdd a short guarantee statement … e.g. '[X]-year workmanship guarantee on all installations' … use only claims the business can stand behind

Top: the kind of advice generic tools give. Bottom: a real finding from a real audit, shortened where you see … (business anonymised).

03 / 05

Find the root cause. Error-proof the fix. Repeat.

I read every report as an owner would, traced each wrong finding to where it started, fixed it there, and wrote the fix down as a rule. The biggest single cause was the prompt itself.

50% → 0118 of 238 findings carried an invented percentage, because the prompt asked for one. After the fix: 0 of 62.
  1. Build
  2. Test on real pages
  3. Find the root cause
  4. Error-proof the fix
  5. Update the standard
  6. Run again
    ↺ back to 1
04 / 05

From grades to things you can check.

Re-running the same unchanged page gave 2 critical issues, then 2, then 1. So I removed every grade the product couldn't reproduce: structured data is graded against Google's required properties, and everything else is a plain count of findings.

My first design
Today
05 / 05

Designed, built and shipped. Now it needs its first users.

Live, and checked against 6 real business sites. User sessions with owners come next.

What I learned: trust in an AI product has to be designed. Sometimes that means taking features out: the scores went, and the page generator stays off until it's good enough.

01The Problem

Owners need the fix,
not a score

PageAudit is for people who run a business and rely on their website: a shop, a plumber, a freelancer. A score doesn't help them. What helps is knowing which part of the page is wrong and what to change, clearly enough to pass to whoever builds the site, or to fix it themselves that afternoon.

Generic audit tool
PageAudit
"Improve your CTA"
The 'Call Example' button gives no benefit or urgency signal. Change the label to say what the caller gets.
"Add schema markup"
Your business name reaches search engines with raw HTML code in it ("&"). The report shows the value as it is now, and as it should be.
A score that moves, with no explanation
Only results that can be reproduced: a count of findings, and a structured-data verdict checked against Google's own rules.
One audit, hoping it covers the right things
Two audits of the same page run side by side: UX conversion and AI visibility.
0 of 3
free tools I tried covered UX conversion or AI visibility. HubSpot Website Grader, PageSpeed Insights and Lighthouse all measured speed and technical SEO.
1
AI-visibility tool turned up in my search. It asked for a business email before running, and gave back a score rather than fixes.
What the free tools showed
HubSpot Website Grader results
HubSpot Website Grader
PageSpeed Insights results
PageSpeed Insights
Lighthouse results
Lighthouse
Framer AEO scanner results
Framer AEO scanner
Who this is for

An online shop owner. Visitors reach the product page and leave without buying, and nothing says why.

A local tradesperson. The phone should ring from the homepage, and the call button gives no reason to press it.

A self-employed professional. One service page does the selling, and they can't tell what's missing, or whether AI assistants can read it.

AI visibility (GEO): 10 conversations with business owners
8 / 10hadn't heard of AI visibility, or that a business could work on it
1 / 10had heard of it but didn't know where to start
1 / 10was already working on it
Informal chats with business owners I know, 2026
The design challenge

Owners act on what the report says. If one line is wrong, they stop trusting the rest.

What one page can't tell you

Whether AI assistants recommend a business depends on its whole website and how it's talked about elsewhere. A single page is the part an owner can change today, so the AI visibility audit covers that page only, checks its structured data against Google's own rules, and says so in the report.

02The Solution

Four rules behind
every screen

PageAudit is an automated pipeline: it loads the page in a real browser, runs fixed checks, then asks the AI only for what the checks can't answer. It runs two audits on one page, UX conversion and AI visibility, and turns them into a single report with a PDF. Every design decision in it follows one of these rules.

01
Specific enough to act on
Every finding names the element on the page, what's wrong with it and what to change. "Your CTA could be stronger" gets ignored. "The Call button names the company, so say what the caller gets" gets fixed.
02
Only show what can be reproduced
Re-running the same page gave 2 critical issues, then 2, then 1, with nothing changed. So the owner never sees a grade the product can't reproduce. They see how many findings there are, and a structured-data verdict checked against Google's own rules.
03
Say what couldn't be checked
If the page can't be read, because of a security check or a failed load, the report says so and grades nothing. A confident report about the wrong page is worse than no report.
04
Ship the smallest version that keeps the promise
Two audits, one report, a PDF. Built on the Claude API: free to use, about $0.18 per audit, with a daily spend cap. A "Register interest" button tests demand for a whole-site audit before any of it gets built.
PageAudit results page on desktop and phone: 10 improvements found, 3 to fix first; structured data graded Good Foundation
The real results page, from a real audit of a local plumber (business anonymised).
03How I built it

Build, test on real pages,
fix the cause, repeat.

AI output looks convincing even when it's wrong, and only real pages expose it. So I ran the project as a continuous improvement loop (plan, do, check, act): test on real pages, read every report as an owner would, find the root cause, error-proof the fix, and write it down as a standard.

The build, phase by phase
Phase 01
Build
Obstacle: the AI's numbers moved

What I did. Two audits and a page generator, with the AI giving each page a score.

What broke. The same page came back with 2 critical issues, then 2, then 1, with nothing changed. The AI also flagged delivery information as missing when it was on the page.

Scores taken away from the AI. Key facts are measured by a real browser before the AI sees the page.
Why

If you run the audit twice and get two different scores, why would you trust either? So I stopped letting the AI score anything. And if a browser can check something, like whether the delivery info is on the page, I check it first instead of asking the AI to guess.

False CriticalDelivery information wrongly flagged as missing
Correctly GoodThe same finding marked Good after page facts were measured
Why

Honestly, the generated pages weren't good enough to show a client. Next to the audit they'd have made the audit look worse too. So I hid the generator and shipped what worked.

Generated page, first runThe generator's first run: a blank white page
Why

When I traced the wrong findings back, most of them came from what the AI was given. Its reasoning was usually fine. So fixing the input fixed a whole group of problems at once.

4 / 4
pages where the wrong main button was detected
57%
of one page's text cut without warning
1
hotel told it had no address, while the address was on the page
Why

The AI wasn't lying on purpose. The prompt asked for a number, so it gave one. Changing what I ask for was the only fix that works on every page, including ones I'll never see.

Early build: numbersEarly build: an AI-visibility score of 31 with five sub-scores
Next: grade wordsNext build: Needs Work grades on both audits
Today: a count and a rule-based verdictResults today: 10 improvements found, Good Foundation
Why

A report is only as good as the page it read. Reaching the bottom of a page isn't the same as having all its text, and having a page isn't the same as having the right one. Decision 03 covers the second.

Refused honestly, at no costWe couldn't read this page: a security check was shown instead
What else I fixed along the way
  • Typing an address without https:// left the Run button dead, with no explanation.
  • Reloading the page lost the whole report.
  • A counter read "0%" to the tool and "97%" to real visitors.
Tools, and why
  • Claude API. Cost measured on every audit, about $0.18, with a daily spend cap so a busy day can't run up a bill.
  • Puppeteer. A real browser loads the page, so the tool reads what a visitor sees, after the page's scripts have run.
  • Railway. Chosen after comparing seven hosts on the same six questions. The deciding one: does a two-minute audit fit inside the request time limit?
  • GitHub. All 16 test suites run automatically on every push, so a broken change shows up before anyone sees it.
03How I built it

Build, test on real pages,
fix the cause, repeat.

AI output looks convincing even when it's wrong, and only real pages expose it. So I ran the project as a continuous improvement loop (plan, do, check, act): test on real pages, read every report as an owner would, find the root cause, error-proof the fix, and write it down as a standard.

Tap a phase to see what broke and what changed.

Obstacle: the AI's numbers moved

What I did. Two audits and a page generator, with the AI giving each page a score.

What broke. The same page came back with 2 critical issues, then 2, then 1, with nothing changed. The AI also flagged delivery information as missing when it was on the page.

Scores taken away from the AI. Key facts are measured by a real browser before the AI sees the page.
Why

If you run the audit twice and get two different scores, why would you trust either? So I stopped letting the AI score anything. And if a browser can check something, like whether the delivery info is on the page, I check it first instead of asking the AI to guess.

False CriticalDelivery wrongly flagged
Correctly GoodDelivery correctly Good
Obstacle: the third output wasn't good enough

What I did. The visual system, a URL-only entry, the results layout, and a generator meant to rewrite the page.

What broke. The generator's first runs gave a blank white page. Once fixed, the pages loaded but looked worse than the originals: a redrawn logo, a missing gallery.

Generator hidden. Launch with the two audits that keep their promise.
Why

Honestly, the generated pages weren't good enough to show a client. Next to the audit they'd have made the audit look worse too. So I hid the generator and shipped what worked.

Generated page, first runBlank generated page
Obstacle: the AI wasn't shown the real page

What I did. Ran the same page twice, on several real business sites.

What broke. The wrong main button was detected on 4 of 4 pages, more than half of one page's text was cut without warning, and a hotel was told it had no address.

The tool reads what a real browser shows, and nothing is cut without saying so.
Why

When I traced the wrong findings back, most of them came from what the AI was given. Its reasoning was usually fine. So fixing the input fixed a whole group of problems at once.

4 / 4
pages where the wrong main button was detected
57%
of one page's text cut without warning
1
hotel told it had no address, while the address was on the page
Obstacle: the AI filled gaps with guesses

What I did. Read every full report as an owner would.

What broke. Findings carried invented percentages, because the prompt asked for them. Scores bunched in the middle of the scale.

Scores became words and counts; predictions banned. Suggested copy uses only facts from the page. Decision 02 has the full story.
Why

The AI wasn't lying on purpose. The prompt asked for a number, so it gave one. Changing what I ask for was the only fix that works on every page, including ones I'll never see.

Early build: numbersEarly numeric scores
Next: grade wordsGrade words
Today: a count and a rule-based verdictResults today
Obstacle: a page that wasn't the page

What I did. Deployed it, added the PDF report, and made the layout responsive.

What broke. One long product page was only 61% read, and one audit described a security-check screen instead of the business's homepage.

Full-page capture, a check that it's the right page, and the scope frozen.
Why

A report is only as good as the page it read. Reaching the bottom of a page isn't the same as having all its text, and having a page isn't the same as having the right one. Decision 03 covers the second.

Refused honestly, at no costRefused page card
What else I fixed along the way
  • Typing an address without https:// left the Run button dead, with no explanation.
  • Reloading the page lost the whole report.
  • A counter read "0%" to the tool and "97%" to real visitors.
Tools, and why
  • Claude API. Cost measured on every audit, about $0.18, with a daily spend cap so a busy day can't run up a bill.
  • Puppeteer. A real browser loads the page, so the tool reads what a visitor sees, after the page's scripts have run.
  • Railway. Chosen after comparing seven hosts on the same six questions. The deciding one: does a two-minute audit fit inside the request time limit?
  • GitHub. All 16 test suites run automatically on every push, so a broken change shows up before anyone sees it.
04Key decisions

Three calls that shaped
the product

Each one started with something going wrong. Here's what I could have done, and why I went the way I did.


Decision 01

From scores to words

My first design showed numbers, and they weren't stable. The same page, run again with nothing changed, came back different.

What I chose, and why
Words, each backed by something the product can reproduce
Moving the maths into code didn't help, because what went into it kept moving. Research put a name on it: AI ratings drift towards the middle of any scale. So the AI-visibility card grades only structured data, with fixed checks against Google's own rules, and UX conversion gets a plain count, because I couldn't find an honest rule to grade it with.
Result: Same page, same structured-data verdict, every time. The UX card now reads "10 improvements found · 3 to fix first". Less satisfying than a grade, and true.

Decision 02

Only ask for what the evidence supports

The prompt asked the AI to judge whether AI crawlers were blocked, using information that isn't on the page. It also asked for an "expected uplift" on every finding, so the AI made one up.

"So why was the page audit built missing these crucial features?"

Me, pushing back on "one page is enough"
What I chose, and why
Only ask for what the evidence can support
One suggested fix told a plumber to advertise "established in 2014, over 1,200 jobs". Both made up, and both checkable against public records. A made-up trade registration can even be an offence, and the owner is exactly the person who'd paste it straight onto their site. So predictions are banned, suggested copy only uses facts from the page or a [placeholder], and credentials are never invented.
Result: Invented statistics went from 50% of findings to 0 of 62, checked across two runs. When the AI reaches for a credential now, it writes "[If Gas Safe registered: …]" and tells the owner to check it first.

Decision 03

Is this even the right page?

Auditing a national franchise's homepage, the tool wrote eight confident pages about a security-check screen. Seven checks passed, because each one asked whether the capture worked. None asked whether it was the right page.

What I chose, and why
Detect it, and refuse honestly before anything is spent
The AI had actually spotted the wall in its first finding, but findings were its only way to say so, and the verdict never read them. When I measured instead of guessing, the real cause was how the tool introduced itself to the site. The server's location had nothing to do with it.
Result: A refused page now costs $0 instead of about $0.20, and says why. Run again on the real site: 28 findings, every checkable fact true against the live page.
05Designing the report

One design,
on screen and on paper

An owner needs something they can keep, forward and print. I built it from the page they had already seen.


The decision

Where the PDF comes from

An audit lived only in the browser tab. A reload, or an old iPhone restarting Safari, and it was gone. An owner will also want to pass it on to whoever builds their site.

Rejected
Design a separate PDF report
My first version was a new serif document. It looked like every other AI report, and it was a second layout that would slowly drift away from the site.
What I chose, and why
Print the results page itself
The report loads the real results page, fills it with the page's own render functions and prints it. The print rules live in the same stylesheet as the site, so the Download button and my report script produce the same document. Change the site, and the next PDF changes with it.
Result: The site and the PDF can't disagree, because they are the same page. Any past audit re-renders as a PDF for $0, and results now survive a reload.

Details that carry the design
01
Old results say they're old
A report restored after a reload names the page and the time the audit ran, and its PDF carries the date the audit ran. Old findings that look fresh would mislead.
02
An honest card for a failed read
A blocked or failed audit once showed a confident "Needs Work" for a page nobody had read. It now says "We couldn't read this page" and gives the reason, as shown in 03.
03
Page breaks by measurement
On a 31-finding report, keeping every card whole left gaps at the foot of pages: 16 pages. Cards may now split only between labelled rows, so a label always stays with its text: 14 pages.
06How I worked with AI

The AI wrote the code,
I made the calls

Most of the code was written with Claude Code. My job was deciding what to build, and checking what came back.


Three habits
01
Rules in writing, added after each mistake
The AI works from a file of rules I keep. Ask before opening any file. Build nothing without a go-ahead. Anything that costs money gets its own question with the price, a rule added on 13 September after one "yes" to a two-part question paid for two audits.
02
Check the real output, and test the checker
The checker for invented numbers printed "0 ok" for nine days because it had never actually run. Since then, every check is first run against something built to fail it.
03
Catch the confident mistakes
A site blocking us was written down as a datacenter problem. The test had changed two things at once, and the real cause was how the tool introduced itself. I also caught a line about a test "failing the build" that overstated what the test did.

The decision

How to trust fast work

AI makes building fast. It makes being wrong fast too, and a wrong answer arrives sounding exactly as sure as a right one.

Rejected
Trust the tests passing
On 5 September every offline check passed. I ran one real audit and found two defects in a minute.
Rejected
Review every line by hand
Too slow to keep up, and reading code says less than seeing what the owner actually gets.
What I chose, and why
Check what the owner would see, and make checking free
A run with a deliberately invalid API key goes through the fetch, the browser and the page reading, then stops before anything is billed. A saved audit re-renders through the real results page for $0. Every test runs on every push, and again after every edit.
Result: 16 test suites guard the product. In one day of capture work, about 40 test runs cost nothing at all.
07Limits & reflection

Where it stands,
and what I'd keep

What the tool can't do yet, and three things the project settled for me.

Where it stands today
No real users yet
Every audit so far has been run by me. Sessions with 3 to 5 business owners come next.
Some findings still go wrong
In a close read of one 29-finding report, 4 carried something false or unsupported. An automatic check for each kind of error is the next build.
Narrow testing
6 sites across Shopify, Wix and WordPress, all in English. Other platforms and languages are untested.
One of three outputs is off
The improved-page generator stays switched off. Reflection 02 shows why.
01

Measure first, then let the AI explain

The biggest accuracy gain came from code, in a step that measures the page in a real browser before the AI sees anything: where the main button sits, what the page says about delivery, returns and reviews.

The first run with it turned a false "delivery information missing" critical into a correct "delivery information present". The model stopped guessing at things that could simply be measured.

If I built it again, I'd design what the tool measures first, and write the prompts second.
02

Design intent doesn't survive as HTML

The generator was meant to hand owners an improved version of their page. It kept the structure and lost the design.

A page's spacing, type and rhythm were chosen by someone, for reasons the markup never records. Asked to improve the page while keeping its look, the AI can only guess at those reasons. On Greenscents it rebuilt the product page in a plainer style and redrew the handwritten logo as plain text.

The original Greenscents product page for its organic washing-up liquid: a handwritten Greenscents logo, a large product photo with lemons, a row of thumbnail images, price, scent and size options, and a green Add to cart button.Original
The generator's version of the same page: the logo set in plain type, a single product photo with one thumbnail, the same price and options, and a black Add to Cart button.Generated

The product page, before and after the generator. Open full size

The original Greenscents logo: handwritten lettering with a leaf above the letters.Original
The generator's version of the logo: the word Greenscents in a plain, blurred sans-serif typeface.Generated

The brand's handwritten logo, and the generator's redraw.

A better model for this: the AI handles production, and the designer keeps authorship.
03

Shipping less protected the rest

The two audits promise precise findings about your page. The generator promised a better page, ready to use, and its output couldn't back that up. Released together, a redrawn logo would have undermined every accurate finding above it.

The same reasoning now keeps the full-site audit on the shelf. When the idea came up, my answer was that we still struggle to audit one page properly, so one page comes first.

The generator returns when its page beats the original.
Closing thought

What this project settled: the design work that matters most in an AI product happens before the model sees anything. Deciding what to measure, what to refuse, and what to leave out are design decisions. The prompts were the easy part.

08Ending

What shipped,
and what comes next

PageAudit is live and runs real audits. The next step is putting it in front of the owners it was built for.


Where it landed
~2 min
per audit, both audits run side by side (90 to 125 seconds)
$0.14–0.20
AI cost per audit, measured on every run
50% → 0%
findings with an invented statistic: half of 238, then 0 of 62
16
test suites that run on every change

What comes next
01
Sessions with owners
3 to 5 business owners run it on their own page while I watch, and a user-testing section joins this case study.
02
A check for every kind of error
Each type of mistake found so far gets a check that runs on every audit, so an error seen once is caught every time after.
03
Findings tied to evidence
Each finding quotes the page, and code confirms the quote is really there before the owner sees it. Designed; building it next.
Thanks for reading

I built PageAudit to answer one question for a business owner: what should I change on this page? Getting that answer right, and saying so when it can't, turned out to be the whole job.

© Copyright 2026

© Copyright 2026

© Copyright 2026