AI-executable UI review standard / 50 rules / 98 points
A review standard written for AI agents: every rule carries evidence requirements, a confidence model, and engineering tolerances. No taste debates. Deductions you can defend.
Eight scored dimensions split a fixed budget. The bar is the argument: visual stability, consistency and color carry more weight than information architecture, because that is where interfaces actually break. Hover a segment.
No vibes-based scoring. A finding is severity × confidence × rule weight. Weak evidence gets multiplied down to 0.15, not argued about. Pick a weight, pick a cell.
Step zero of every audit is classification. A dashboard is not judged like a landing page; a CTA rule makes no sense on documentation. Six archetypes gate which rules apply.
The full rulebook, with each rule's class and maximum deduction. Objective defects score hard; heuristics score with care; advisories never touch the number.
Most review checklists fail by enforcing folklore as law. This one names the folklore and strikes it out. These may surface as advisories in context, never as universal defects.
An agent runs the spec as a pipeline: map the page semantically before judging it, collect proof before deducting, sandbox interactions so the audit cannot break the page it is auditing.
Scores map to six grades. Everything left of the needle ships; everything right of it goes back. The needle is not negotiable, that is the point of having one.
A standard that cannot survive its own rules is an opinion. This page was audited with its own pipeline: semantic mapping, three breakpoints, DOM extraction, interaction checks. The receipt is on the right.
AUDIT PASSED · GRADE A