Full archive history of any URL from the Internet Archive's official CDX API: every snapshot with date, archive link, HTTP status and content digest - or a per-URL summary with first/last capture, snapshots per year and unique content versions. Polite rate control built in. No credentials needed.
Bulk domain intelligence: registrar, age and expiry via official RDAP (structured WHOIS), SSL certificate expiry and trust, DNS records (A/AAAA/MX/NS/TXT), plus clear warnings like "expires in 30 days". Registrant contact details are never extracted - registry facts only.
Extract text from PDF URLs at scale: full text, per-page text, real structured tables (rows and columns as JSON, not text lines), document metadata and clean Markdown for LLM/RAG. The table extraction generic PDF text extractors don't have. No credentials needed.
Turn any RSS, Atom or RDF feed into clean, normalized JSON articles: ISO 8601 dates, plain-text and HTML content, authors, categories and podcast enclosures. Multiple feeds per run, broken feeds reported instead of crashing.
Extract every machine-readable block from web pages: JSON-LD (schema.org) with detected types, Open Graph, Twitter Card, standard meta tags, canonical URL, hreflang and favicons. One clean JSON item per URL. Fetches only the URLs you provide - no crawling. No credentials needed.
Detect the technology stack of any website: CMS, frameworks, JavaScript libraries, analytics, web servers and CDN - with versions where detectable and the evidence for every match. 7,500+ fingerprints (community Wappalyzer DB), one GET per URL, no crawling. No credentials needed.
Check thousands of URLs in one run: HTTP status, full redirect chain, final URL, latency and a clear ok/broken verdict with an error category (404, timeout, DNS, SSL, redirect loop). HEAD-first with GET fallback, retries on flaky errors. Checks only the URLs you provide - no crawling.