This project is a Python-based SEO crawler that mimics Googlebot (mobile-first by default) to audit websites for SEO, indexing, and technical best practices.
The tool works like a mini search engine:
-
Fetches pages with a Googlebot-like User-Agent (mobile or desktop).
-
Parses robots.txt, checks allowed/disallowed URLs, and detects sitemaps.
-
Audits key SEO signals:
-
Canonical tags, hreflang setup, meta robots, JSON-LD structured data
-
Titles, descriptions, headings (H1/H2)
-
Open Graph/Twitter metadata
-
Image alt tags, lazy loading, and dimensions (CLS risk)
-
Render-blocking scripts and excessive stylesheets
-
-
Detects common issues (e.g. missing viewport, duplicate hreflangs).
-
Generates human-readable console reports and structured JSON reports for deeper analysis.
-
Supports both single-page audits and site-wide BFS crawling with configurable depth.
This project shows how SEO audits can be automated in a developer-friendly way, combining web crawling with actionable insights for improving technical SEO and user experience.
Have an idea that needs to work?
Discuss the next practical step with De Wilde ICT Solutions.
Contact us