Detect the technologies behind any website: CMS, ecommerce platform, JavaScript frameworks, analytics, ad tech, CDN, hosting, email provider. 7,600+ fingerprints, versions and confidence scores. Clean JSON per website.
Scrape football matches from Flashscore across every league: fixtures, live scores, final results, halftime and second-half scores, extra time and penalty shootouts, goalscorers with assists, cards, substitutions and match statistics. No proxy, no login.
Check URLs in bulk or crawl a page's links: status codes, the full redirect chain hop by hop, redirect loops, https downgrades, soft 404s that answer 200 while being error pages, and certificate expiry. No proxy, no login.
Extract clean text and markdown from PDF, DOCX and HTML documents, with per-page text, document metadata and real line breaks. Detects scanned PDFs that have no text layer instead of returning an empty result. No proxy, no login.
Scrape baseball games from Flashscore: fixtures, live scores, final results, the full inning-by-inning line score and match statistics. MLB, NPB, KBO and every domestic league. No proxy, no login.
Scrape basketball games from Flashscore: fixtures, live scores, final results, quarter-by-quarter and overtime scores, team data and full match statistics. NBA, EuroLeague, NCAA and every domestic league. No proxy, no login.
Scrape ice hockey games from Flashscore: fixtures, live scores, final results, period-by-period scores, overtime and shootout outcomes, and match statistics. NHL, KHL, SHL and every domestic league. No proxy, no login.
Scrape ATP, WTA, ITF and Challenger tennis matches from Flashscore: fixtures, live scores, final results, set and tiebreak scores, player and tournament data, match statistics and full point-by-point sequences. No proxy, no login.
Extract business contact details from any website: email addresses, phone numbers in E.164 format, social profiles, postal addresses and the company name. Follows the site's own contact, about and imprint pages. No proxy, no login.
Give it a domain and it finds the sitemaps itself, from robots.txt, the homepage or the usual paths, follows sitemap index files, unpacks gzipped ones, and returns every URL with lastmod, priority, hreflang alternates and images. No proxy, no login.