Input: PDF URLs, many per run. Output: one row per document — tables as rows and columns, text in reading order, Markdown for RAG, metadata, outline, form fields, key fields (invoice no., dates, totals, IBAN, VAT). No OCR: a scanned page is flagged and not charged.
Input: website URLs or bare domains, many per run. Output: one JSON row per site — CMS, ecommerce, JS framework, CDN, server, analytics, ad pixels, consent tool, security headers, pre-consent cookies — each detection carrying its evidence. Choose it when names alone are not enough.
Input: page URLs. Output: one row per URL with a link to the file — PNG, JPEG, WebP or PDF; full page, viewport, custom size or one CSS-selected element. Banners hidden, ads blocked, lazy images loaded. Device presets, retina, dark mode, locale. Failed URLs are not charged.
Input: domains or sitemap URLs. Output: one row per page URL with its lastmod, changefreq and priority, plus a report row per sitemap file read. Finds sitemaps via robots.txt, follows nested indexes, unpacks gzip. Filter by URL pattern or by date. Hard cap of 25,000 URLs per run.