Turn any site's sitemap.xml, sitemap index, or robots.txt Sitemap directive into a clean URL dataset with loc, lastmod, changefreq and priority per row. Recursively expands sitemap indexes (bounded depth), follows redirects, and decompresses .xml.gz automatically.
Query the YC company directory by batch, industry, or status via the same Algolia index the website uses. Returns name, batch, status, industries, team size, website, one-liner, and YC profile URL as clean structured rows.
Pull Yahoo Finance quote snapshots and historical OHLCV candles for any list of tickers — equities, ETFs, indices, crypto, FX. We handle the retries, fingerprint rotation, and proxy routing so the data lands. One typed row per symbol, JSON or CSV export.
Search Yellow Pages business listings by keyword and US city — name, phone, address, categories, rating, review count, and website for every result. We rotate fingerprints and proxies so blocks don't stop your dataset. Local lead generation.
Detect brand-deal signals in a YouTube channel's recent Shorts — spoken sponsor mentions in the transcript, disclosure hashtags, @brand/domain mentions — and score each Short's sponsorship likelihood (0.0–1.0).
Search Zenodo's open research repository and export records — DOI, title, authors, resource type, publication date, license, access rights, file count and download stats — as clean JSON, CSV or Excel.
Scrape the top customer reviews from an Amazon product page by ASIN or URL — rating, title, body, verified-purchase flag, date, and helpful-vote count per review. Pay-Per-Event, no login, no API key.
Scrape licensed daycare and childcare facility registries across 5 US states (NY, CT, CO, DE, TX) via official state open-data APIs. Get facility name, license number, capacity, address, phone, and status — clean B2B leads, ready to export.
Scrape public Eventbrite events, category/search listings, and organizer profiles into structured rows: title, timing, venue, pricing, and ticket availability. Discover by keyword, location, or category, or supply direct event/listing/organizer URLs. No login required.
Scrape every job posting from any Recruitee-hosted employer career site via Recruitee's own public, unauthenticated JSON API, no login or browser required. One request per company returns titles, locations, salary, and full HTML descriptions.
Scrape SocialBlade creator/channel stats across YouTube, Twitch, Instagram, and TikTok — subscriber/follower counts, view counts, SocialBlade's letter grade, platform + category ranks, and 15-day daily growth history, normalized per handle.
Pull fresh building and construction permit records from any US city or county Socrata open-data portal. We normalize permit ID, type, status, address, valuation, and contact info into one typed row per permit — ready to feed your contractor CRM or lead pipeline.
Scrape UK property listings from Zoopla (for-sale or to-rent) — price, address, bedrooms, agent phone, floor plan, EPC, photos, full description. Export to JSON or CSV. We handle the blocks so your dataset stays clean.
Scrape used-car listings from AUTO.RIA (auto.ria.com), Ukraine's #1 car marketplace with ~300k live ads. Get price (USD/UAH), make, model, year, mileage, fuel, transmission, engine, body, color, VIN, location, seller type, photos, and full description — export to JSON or CSV.
Search the Open Library API (the Internet Archive's open book catalogue) and export structured book metadata — title, authors, ISBNs, subjects, publish year, cover URL, edition count, OpenLibrary ID — to JSON or CSV. We handle pagination and retries across 30M+ works.
Bulk Google Autocomplete (Suggest) scraper and API — run many seeds × languages × countries per run with optional A–Z expansion for long-tail keyword research. Export suggestions to JSON or CSV. Public Suggest endpoint, no auth, no key.
Search PubMed by query and export structured paper rows — title, authors, abstract, journal, DOI, PMID, MeSH terms, publication date — to JSON or CSV. A clean PubMed API wrapper that handles NCBI pagination, rate limits, and retries for research and ML pipelines.
Fetch Steam game prices across 60+ regions in one run via the Steam price API — compare USD-equivalent prices, spot regional discounts, and track arbitrage gaps — export to JSON or CSV. Public API, no login; we retry and pace so every region lands.
Scrape used-car listings from willhaben.at, Austria's #1 marketplace — price, make, model, year, mileage, fuel, transmission, power, body type, colour, seller, location, and photos. Export to JSON or CSV; optionally enrich each listing with its full description.
Pull crypto market data from CoinGecko — price, 24h change, market cap, volume, supply, ATH/ATL, exchanges — one clean typed row per coin, exported to JSON or CSV on a schedule. We handle the rate limits, retries, and proxy routing so the data lands. Built on the CoinGecko v3 API.
Scrape tech job postings from Dice.com — filter by keyword, location and radius, skill, or employment type. Get one normalized row per job with title, company, location, employment type, and optional skill tags, ready for CSV, JSON, or API export.
Search arXiv by query, category, or author and export structured paper metadata — title, authors, abstract, primary category, DOI, PDF URL, submitted and updated timestamps — to JSON or CSV. An arXiv API wrapper that handles pagination, retries, and rate-limit pacing for your pipeline.
Scrape used-car listings from Auto24.ee, Estonia's #1 car marketplace — make, model, year, price in EUR, mileage, fuel type, gearbox, engine power, body type, drivetrain, color, location, seller, and photos. Export to JSON or CSV; enrich each listing from its detail page.
Export posts from any public Bluesky custom or algorithm feed via the AT Protocol API — feed metadata and engagement counts — to JSON or CSV. No login needed; we page through and retry so the whole feed lands.