TidyTools

Website-to-Markdown benchmark: Website Markdown Crawler vs Website Content Crawler

Method and aggregated results of a small benchmark run on 2026-09-30.

Sites

KeyStart URLKind
python-sphinxhttps://docs.python.org/3/tutorial/Sphinx docs (static HTML)
docusaurushttps://docusaurus.io/docs/Docusaurus docs (pre-rendered React)
nextjshttps://nextjs.org/docs/app/Next.js docs (JS-heavy)
intercom-helphttps://www.intercom.com/help/en/Help center (Intercom Articles)
cloudflare-bloghttps://blog.cloudflare.com/Blog with long posts and code

Settings: what a new user gets

Metrics

All computed by a small Python script (standard library only).

Results

Paragraph recall = sampled reference paragraphs found in the Markdown. "After" = our crawler after the fixes listed below; "blind" = our first run, before any change.

SiteWCC default: time / paragraphs found / $ per 1,000Ours after: time / paragraphs foundOurs blind: time / paragraph recall
python-sphinx56 s / 237 of 242 (98%) / $1.553 s / 239 of 242 (99%)4 s / 99%
docusaurus84 s / 258 of 273 (95%) / $1.5714 s / 144 of 148 (97%)4 s / 100%
nextjs200 s / 0 of 82 (0%) / $3.646 s / 82 of 82 (100%)4 s / 100%
intercom-help183 s / 106 of 309 (34%) / $3.6534 s / 143 of 146 (98%)91 s / 97%
cloudflare-blog159 s / 196 of 199 (98%) / $2.779 s / 179 of 179 (100%)4 s / 98%

Our crawler cost $1.00 per 1,000 pages on every site in these runs (all pages over HTTP; $1.50 since 2026-10-04). Each tool's recall is measured on its own 5 sample pages, so the reference counts differ when the two tools crawled different pages.

All runs (including WCC's cheapest mode on nextjs and a later 15-page clean-up round) with every metric:

Run names: <tool>-<mode>-<site>[-tag]. default = the settings a new user gets, cheap = the cheapest mode. No tag = the blind first run; -after = after the fixes; -q2before / -q2after = a later 15-page round before and after a second set of clean-ups.

Per site

What we changed between "blind" and "after"

Main-content extraction when a page has one clear <article> / <main>; code fences for bare <pre>; repair of tables with omitted closing tags; absolute links; removal of heading permalinks and one-line UI text ("Copy page", "Was this helpful?"); crawl order that puts in-content links before menus, sitemaps, translations, old versions and tag/author pages. These were tuned on the same five sites, so the "after" numbers are optimistic; the blind run is the fair comparison.

Where WCC is still better

Limits

One run per tool and site, 5 sample pages per site, 5 sites, one day. WCC's cost depends on your plan's compute-unit price and the memory you choose; we kept its default memory because the test is about defaults.

Reproduce

The method above is complete enough to repeat: run both Actors with the settings listed, on the same five start URLs, and compute the metrics as described. The raw crawl outputs contain text from the five sites, which belongs to their owners, so they are not published. Results will differ a little as the sites change. The per-run numbers are in the CSV above. Questions about the method: tidytools@yukai.uk.

← All tools