thegian7 ~/catalog/…/site-inventory$

Python · standard library only

Site Inventory

A config-driven pipeline that inventories every page across N websites, classifies each one into a shared taxonomy, and produces an N-way comparison — optionally framed as a gap report from a single site's point of view.

The question it answers

"How does our site compare to theirs, and what would it take to catch up?"

That question is usually answered by someone clicking around two websites for an afternoon and writing impressions. The output is unfalsifiable and impossible to price. This turns it into a count: here is every page they have, here is every page you have, here is the shape of the difference on one shared axis.

Gap mode

Name a subject site and the output is framed from its point of view — what it lacks relative to the references.

Survey mode

No subject. Every site is a peer and the output is a neutral matrix.

Four stages, each one inspectable

Every stage writes a JSON artifact and reads the previous one.

 discover  ──►  urls_<host>.json        sitemap, falling back to a bounded crawl
    │
 enrich    ──►  pages_<host>.json       tiered: sitemap → wp-json → <head> → full text
    │
 classify  ──►  inventory_<host>.json   universal buckets + vertical overlay
    │
 compare   ──►  comparison.json + .md   the N-way result

Because each stage reads the last one's output, any stage can be re-run, hand-corrected or debugged on its own. That matters more than it sounds: the network work is slow and rate-limited, and classification rules need repeated tuning. Separating them means tuning the rules costs nothing — you are not re-crawling anyone's website to fix a regex.

The enrichment is tiered on purpose. A sitemap entry is cheap and often enough; a WordPress JSON endpoint is richer; parsing <head> is more expensive; pulling full text is the last resort. Most pages never reach the expensive tier.

One shared axis

Every page on every site lands in exactly one bucket.

home_utilitylegal offeringperson proofeditorial mediahub geo_landingunclassified

The comparison only works because every site maps onto the same axis. Per-vertical configs relabel the buckets — offering becomes "Practice area" for a law firm — and add subtypes underneath. They never replace the axis, because the moment two sites are measured on different scales the comparison stops meaning anything.

Keeping unclassified as a real bucket is deliberate too. Pages that don't fit are visible in the output rather than quietly forced into the nearest category, so a bad rule shows up as a number instead of hiding as a wrong answer.

Standard library only

No dependency tree for a tool that runs against other people's websites.

The whole pipeline uses nothing outside Python's standard library. For a tool whose job is to fetch and parse arbitrary third-party pages, every dependency is both a supply-chain surface and a thing that breaks in eighteen months. It runs the same on any machine with Python on it.

← the catalog·thegian7