Python · standard library only
A config-driven pipeline that inventories every page across N websites, classifies each one into a shared taxonomy, and produces an N-way comparison — optionally framed as a gap report from a single site's point of view.
"How does our site compare to theirs, and what would it take to catch up?"
That question is usually answered by someone clicking around two websites for an afternoon and writing impressions. The output is unfalsifiable and impossible to price. This turns it into a count: here is every page they have, here is every page you have, here is the shape of the difference on one shared axis.
Name a subject site and the output is framed from its point of view — what it lacks relative to the references.
No subject. Every site is a peer and the output is a neutral matrix.
Every stage writes a JSON artifact and reads the previous one.
discover ──► urls_<host>.json sitemap, falling back to a bounded crawl │ enrich ──► pages_<host>.json tiered: sitemap → wp-json → <head> → full text │ classify ──► inventory_<host>.json universal buckets + vertical overlay │ compare ──► comparison.json + .md the N-way result
Because each stage reads the last one's output, any stage can be re-run, hand-corrected or debugged on its own. That matters more than it sounds: the network work is slow and rate-limited, and classification rules need repeated tuning. Separating them means tuning the rules costs nothing — you are not re-crawling anyone's website to fix a regex.
<head> is more expensive; pulling full text
is the last resort. Most pages never reach the expensive tier.
Every page on every site lands in exactly one bucket.
The comparison only works because every site maps onto the same axis. Per-vertical configs
relabel the buckets — offering becomes "Practice area" for a law firm —
and add subtypes underneath. They never replace the axis, because the moment two sites are
measured on different scales the comparison stops meaning anything.
Keeping unclassified as a real bucket is deliberate too. Pages that don't fit are
visible in the output rather than quietly forced into the nearest category, so a bad rule shows
up as a number instead of hiding as a wrong answer.
No dependency tree for a tool that runs against other people's websites.
The whole pipeline uses nothing outside Python's standard library. For a tool whose job is to fetch and parse arbitrary third-party pages, every dependency is both a supply-chain surface and a thing that breaks in eighteen months. It runs the same on any machine with Python on it.