112 checks across 12 categories. Live checks run today and produce a real pass/fail/warn verdict. Pending checks are on the roadmap — each has a specific, named reason it isn't automated yet, not just "not done."
103 live
9 pending
12 categories
1. HTML Metadata 11/11 live
ID
Check
Priority
Status
Note
1.1
Title Tag
High
Live
1.2
Meta Description
High
Live
1.3
Meta Robots
High
Live
1.4
Canonical URL
High
Live
1.5
Character Encoding (Charset)
Medium
Live
1.6
Viewport Meta Tag
High
Live
1.7
HTML Language Attribute
High
Live
1.8
Open Graph Meta Tags
Medium
Live
1.9
Twitter Card Meta Tags
Low
Live
1.10
Favicon Declaration
High
Live
1.11
hreflang Attributes
High
Live
2. URL Structure 7/7 live
ID
Check
Priority
Status
Note
2.1
Descriptive & Human-Readable URLs
High
Live
2.2
URL Length
Medium
Live
2.3
Lowercase URLs
Medium
Live
2.4
Use Hyphens Instead of Underscores
Medium
Live
2.5
HTTPS
High
Live
2.6
URL Parameters
Medium
Live
2.7
Static vs Dynamic URLs
Medium
Live
3. Heading Structure 8/8 live
ID
Check
Priority
Status
Note
3.1
Heading Elements Overall
High
Live
3.2
H1 Tag
High
Live
3.3
Number of H1 Tags
Medium
Live
3.4
Heading Hierarchy
High
Live
3.5
Descriptive Headings
Medium
Live
3.6
Empty Heading Tags
Medium
Live
3.7
Skipped Heading Levels
Medium
Live
3.8
In-Page Anchor Links & Table of Contents
Medium
Live
4. Content Quality & Readability 11/15 live
ID
Check
Priority
Status
Note
4.1
Unique & Original Content
High
Live
4.2
Content Relevance (Search Intent)
High
Pending
Needs AI judgment — checklist itself says intent evaluation requires human review
4.3
Primary Topic & Keyword Usage
Medium
Live
4.4
Keyword Stuffing
High
Live
4.5
Content Freshness
Medium
Live
4.6
Content Depth & Completeness
High
Live
4.7
Readability
Medium
Live
4.8
Grammar & Spelling
Medium
Live
4.9
Duplicate Content
High
Live
4.10
Thin Content
High
Live
4.11
Helpful Content
High
Pending
Human review required — non-deterministic
4.12
E-E-A-T Signals
High
Live
4.13
Direct Answer Formatting (AEO)
High
Live
4.14
Intrusive Interstitials & Pop-ups
High
Pending
Required R&D
4.15
Above-the-Fold Content Visibility
High
Pending
Required R&D
5. Images 13/13 live
ID
Check
Priority
Status
Note
5.1
Image Alt Text
High
Live
5.2
Descriptive Image File Names
Medium
Live
5.3
Image Dimensions (Width & Height)
High
Live
5.4
Image File Formats
Medium
Live
5.5
Image Compression & File Size
High
Live
5.6
Responsive Images (srcset/sizes)
High
Live
5.7
Lazy Loading
High
Live
5.8
Decorative Images
Medium
Live
5.9
Image Captions
Low
Live
5.10
Structured Data for Images
Medium
Live
5.11
Image Sitemap
Medium
Live
5.12
Broken Images
High
Live
5.13
LCP Hero Image Optimization & Preloading
High
Live
6. Internal Linking 11/12 live
ID
Check
Priority
Status
Note
6.1
Internal Links Presence
High
Live
6.2
Descriptive Anchor Text
High
Live
6.3
Broken Internal Links
High
Live
6.4
Orphan Pages
High
Live
6.5
Crawl Depth
Medium
Live
6.6
Link Relevance
Medium
Pending
Human review required — non-deterministic
6.7
Excessive Internal Links
Medium
Live
6.8
Navigation Links
High
Live
6.9
Breadcrumb Navigation
Medium
Live
6.10
Footer Links
Low
Live
6.11
Redirecting Internal Links
Medium
Live
6.12
Canonical Target Links
Medium
Live
7. External Linking 12/13 live
ID
Check
Priority
Status
Note
7.1
External Links Presence
Medium
Live
7.2
Link to Authoritative Sources
Medium
Live
7.3
Broken External Links
High
Live
7.4
Descriptive External Anchor Text
Medium
Live
7.5
rel=nofollow for Untrusted Links
High
Live
7.6
rel=sponsored for Paid Links
High
Live
7.7
rel=ugc for User-Generated Links
Medium
Live
7.8
Excessive External Links
Medium
Live
7.9
External Links Opening in New Tab
Low
Live
7.10
HTTPS External Links
Medium
Live
7.11
Unsafe or Spammy External Links
High
Pending
Requires the Google Safe Browsing API — not yet integrated
7.12
Redirecting External Links
Low
Live
7.13
Security Attributes (rel=noopener noreferrer)
High
Live
8. Video & Audio Content 1/1 live
ID
Check
Priority
Status
Note
8.1
Video/Audio Transcripts & Captions
Medium
Live
9. Structured Data & Schema Infrastructure 9/11 live
ID
Check
Priority
Status
Note
9.1
JSON-LD Format Usage
High
Live
9.2
Content Parity (Visible Content Alignment)
High
Live
9.3
Required and Recommended Properties
High
Live
9.4
Use Specific Schema Types
Medium
Live
9.5
Avoid Irrelevant or Misleading Markup
High
Live
9.6
Keep Structured Data Up-to-Date
High
Pending
Not required — part of analytics
9.7
Don't Block Structured-Data Pages
High
Live
9.8
Entity Linking (@id) Consistency
Medium
Live
9.9
Validate Before and After Deployment (local syntax gate)
High
Live
9.10
Prioritize High-Impact Schema Types
High
Live
9.11
Monitor Rich Results & AI Overview Appearances
Medium
Pending
Not required — part of analytics
10. Mobile Infrastructure & Mobile-First Indexing 5/6 live
ID
Check
Priority
Status
Note
10.1
Mobile-First Indexing Parity
High
Live
10.2
Responsive Design & Layout (No Horizontal Scroll)
High
Live
10.3
Touch Target Sizing & Spacing (Min 48×48px)
High
Pending
Required R&D
10.4
Mobile Core Web Vitals
High
Live
10.5
Mobile Content Accessibility
High
Live
10.6
Mobile Crawlability
High
Live
11. Core Web Vitals & Performance Engineering 10/10 live
ID
Check
Priority
Status
Note
11.1
Largest Contentful Paint (LCP)
High
Live
11.2
Interaction to Next Paint (INP)
High
Live
11.3
Cumulative Layout Shift (CLS)
High
Live
11.4
Core Web Vitals Image Delivery
High
Live
11.5
JavaScript Optimization & Deferral
High
Live
11.6
CSS Optimization & Critical CSS
High
Live
11.7
Server Response Time (TTFB)
High
Live
11.8
Caching Infrastructure (Browser & Server)
High
Live
11.9
Minification (HTML, CSS, JavaScript)
Medium
Live
11.10
Resource Hints (Preload, Preconnect, Prefetch)
Medium
Live
12. Crawlability, Indexability & Architecture 5/5 live
ID
Check
Priority
Status
Note
12.1
robots.txt Configuration & Size Limits
High
Live
12.2
XML Sitemap Architecture & Validation
High
Live
12.3
Crawl Budget Optimization
High
Live
12.4
Pagination Handling
Medium
Live
12.5
Faceted Navigation & Parameter Control
High
Live
2. Data we collect
Everything below comes from the page itself, its directly-linked resources, and — only when explicitly enabled — a site-wide crawl or an external performance API. Nothing is collected beyond what a specific check needs.
Every audit, single page
Fetch data — final URL after redirects, HTTP status code, response headers, response time to first byte (TTFB), raw HTML
Parsed structure — title tag(s), every meta tag (name/property → content), the full heading outline (level, text, id), every image (src, alt, dimensions, loading attribute), every link (href, anchor text, rel, target, internal vs. external), every JSON-LD structured-data block, canonical URL, robots meta, HTML lang attribute, charset, viewport tag
Directly-linked resources — robots.txt, the XML sitemap, the favicon, the og:image, individual image files (for size/format checks), CSS and JS assets (for size, minification, and caching checks) — each fetched only if the relevant check needs it
Site-wide audits only (--site-wide)
Every page reachable from the entry URL, up to a configured page/depth limit, crawled via Crawl4AI
The full internal link graph — which pages link to which, and with what anchor text
A second crawl rendered with a Googlebot-Smartphone user agent at a mobile viewport, for mobile-parity checks
Performance checks only (when PAGESPEED_API_KEY is set)
Real-user Core Web Vitals field data (LCP, INP, CLS) from Google PageSpeed Insights / the Chrome UX Report
Not collected, ever: cookies, session state, form input, personal or user-identifiable data, or anything from behind a login. Nothing is retained after a run except the report file you asked for.
3. How results are judged (KPIs / thresholds)
Every number below is a config value, not something buried in code — thresholds can be tuned per client or industry without touching the check logic. This is the actual configuration the tool ships with.
Area
Threshold
Verdict
Title tag length
50–60 chars
Best-practice band; 10–60 is acceptable, outside that fails
Meta description length
150–160 chars
Best-practice band; 70–160 is acceptable, outside that fails
URL length
≤ 100 chars
Flagged past this length
URL path depth
≤ 6 segments
Flagged past this depth
Thin content
< 300 words
Failed as thin content
Shallow content depth
< 600 words
Warned as shallow coverage
Keyword stuffing
one term > 6% of words (min. 5 uses)
Failed as stuffing
Content freshness
> 730 days since update
Flagged for review, not auto-failed
Duplicate content
≥ 85% shingle similarity
Failed; 60–85% warned (site-wide only)
Crawl depth
≤ 3 clicks from homepage
Google's own site-structure guidance
Internal links per page
≤ 150
Flagged for manual review past this
Footer links
≤ 30
Flagged past this count
External links per page
≤ 100
Flagged as excessive past this
Image file size
≤ 300 KB
Flagged as oversized past this
Alt text length
≤ 125 chars
Flagged past this length
JS payload per page
≤ 500 KB
Flagged past this
CSS payload per page
≤ 100 KB
Flagged past this
Server response time (TTFB)
≤ 500ms good · ≤ 800ms target
Failed above 800ms
Static asset cache lifetime
≥ 1 month
Flagged if shorter
Resource hints per page
≤ 5
Flagged as excessive past this
Largest Contentful Paint (LCP)
≤ 2.5s good · > 4s poor
Google's published Core Web Vitals bands
Cumulative Layout Shift (CLS)
≤ 0.1 good · > 0.25 poor
Google's published Core Web Vitals bands
Interaction to Next Paint (INP)
≤ 200ms good
Google's published Core Web Vitals bands
robots.txt file size
≤ 500 KB
Google's own documented parsing limit
XML sitemap size
≤ 50,000 URLs · ≤ 50 MB
Google's own documented sitemap limits
Each individual result also carries a priority (High / Medium / Low, from the specification) which combines with its pass/fail/warn status into a final severity — Critical, High, Medium, or Low — used to sort recommendations by what to fix first.
4. Output formats
Every audit produces the same underlying result set — an overall score, per-category pass/fail/warn counts, and every individual check's evidence — rendered into one of three formats.
Console
Default — for a quick read in the terminal
Score, category breakdown, and the top issues ranked by severity, with the fix and reasoning for each.
JSON
For dashboards, pipelines, and re-audits
The full structured report — every check's status, evidence, and message. Nothing summarized away.
Markdown
For a shareable written report
The same content as console, formatted as a document — the format this page's data itself is drawn from.
Console output (excerpt)
======================================================================
ON-PAGE SEO AUDIT REPORT (Phase 1 — rule-based checks only)
======================================================================
URL: https://example.com/
Score: 75.5/100 (47 scored / 103 total checks)
-- 1. HTML Metadata (3/11 passed) --
[✗] 1.1 Title Tag FAIL
[✓] 1.2 Meta Description PASS
[!] 1.8 Open Graph Meta Tags WARN
...
======================================================================
RECOMMENDATIONS (16)
======================================================================
1. [Critical] 1.1 — Title Tag
Fix: Write one unique, accurate <title> per page, ideally 50-60 characters.
Why: The title is one of the strongest on-page SEO signals and directly
affects click-through rate from search results.
# On-Page SEO Audit — https://example.com/
**Score:** 75.5/100 (47 scored checks, 103 total)
**Run ID:** 44eb1b7d-fd50-4f4a-8c1f-a6bd7d4ed5d9
## Recommendations
### [Critical] 1.1 — Title Tag
- **Recommendation:** Write one unique, accurate <title> per page, ideally 50-60 characters.
- **Why it matters:** The title is one of the strongest on-page SEO signals and
directly affects click-through rate from search results.
## All Checks
### 1. HTML Metadata
- ❌ **1.1 Title Tag** — Title is 14 characters, shorter than the 50-60 best-practice band.
- ✅ **1.2 Meta Description** — Meta description is 152 characters, within the best-practice band.
Command line: audit.py <url> --format console|json|markdown [--out FILE]. Add --site-wide to crawl and audit an entire site instead of one page.