Web scraping vs buying datasets: the real cost comparison
Build vs buy for web data: scraping hours, proxy costs, maintenance burden versus one-time dataset prices. With concrete numbers from our own pipelines.
Sources compared
| Source | Access | Price | ISIN | Coverage | Update cadence |
|---|---|---|---|---|---|
| DIY scraping (Apify/Crawlee) | Build + run | $0.3-2/1k pages compute + dev hours | no | Whatever you build | You maintain it |
| Apify Store actors | Per-use | $1-5/1k results typical | no | Whatever actors exist | Maintained by dev |
| Dataset marketplaces (Bright Data, ScrapeHero, CoreSignal) | One-time or subscription | Enterprise: custom-quoted, often $100s+. Indie listings: $3-15 one-time | yes | Fixed snapshot | Frozen at purchase |
| Our datasets (vhsgreed) | One-time | $2-5 (launch pricing) | yes | Verified snapshots w/ methodology docs | Versioned (v4 = 4th revision) |
DIY scraping (Apify/Crawlee): Cheap at 1 run, expensive at maintenance forever.
Apify Store actors: Best for ongoing needs on popular targets.
Dataset marketplaces (Bright Data, ScrapeHero, CoreSignal): Cheapest for point-in-time analysis. Ask what QA was applied before you buy.
Our datasets (vhsgreed): Verification layer is the product, not just rows.
What our data adds
- Worked example from our own pipeline, 2026-09-24: 100 Swiss Zefix companies with full register details, 56 seconds runtime, $0.0124 platform compute. That is $0.12 per 1,000 companies; the compute is cheap, the maintenance is the cost.
- Maintenance evidence from a 30-scraper sweep the same day: 25 of 30 delivered data. The 5 failures were 2 upstream outages (crt.sh HTTP 502, Lithuanian KRS HTTP 500), 1 target silently dropping our search filter server-side, and 2 anti-bot blocks (Indeed, Google Jobs). Two of the five failed silently: exit code 0 with zero results.
- Targets change markup and behavior more often than buyers expect. Silent failures mean a scraper can look healthy and return nothing, so recurring scraping needs monitoring or a maintainer either way.
- Break-even: one-time analysis favors buying; recurring favors scraping or maintained actors
Get the data
Robotics Supply-Chain Intelligence 2026
$2
More comparisons
- Swedish stock market data sources: free vs paid compared
- Swedish company data: Bolagsverket vs allabolag vs SCB vs our CSV
- Robotics datasets: 12 sources compared (supply chain, stocks, standards)
- npm package security data: npm API vs Snyk vs OSV vs GitHub Advisory
- Hemnet data access options compared (listings, sold prices, ToS reality)