Crawlfast turns messy websites into clean, structured records — crawling, extraction schemas, and pipelines in one Docker-first workspace built for Advert.Cafe.
Everything you need to go from a URL to clean, queryable records — no glue code.
01
Crawl at scale
Playwright-powered crawling that renders JS, follows sitemaps, and respects duplicates — so you capture every page that matters, once.
02
Structured extraction
Define extraction schemas and turn raw HTML into typed records. No brittle scripts — just the fields you actually need.
03
Visual pipelines
Compose crawl → extract → transform → store as a visual pipeline. Edit nodes, branch logic, and re-run without redeploying.
04
Durable storage
Every crawl and result lands in S3-backed storage with full provenance, ready to push downstream to Advert.Cafe leads.
05
Fast iteration
Hot-reload services, instant test-on-website, and live metrics. Go from idea to working extractor in minutes.
06
Connected by design
Serper search/places, quality scoring, and webhooks out of the box — Crawlfast plugs straight into the Advert.Cafe suite.
01 /
Point
Drop in a URL or a search term. Crawlfast maps the site and queues the pages worth crawling.
02 /
Extract
Attach an extraction schema or pipeline. Render, parse, and classify into typed, deduplicated records.
03 /
Ship
Store results with full provenance and push them straight into Advert.Cafe leads and campaigns.
Interactive demo with sample data.
/cars/toyota-yaris
/cars/nissan-micra
/about
/el/cars/toyota-yaris
/el/cars/nissan-micra
All classified
Jump into the dashboard and run your first crawl, extractor, or pipeline in minutes.
Get Started