Crawl a site and verify every link, internal and external: status codes, redirect targets, and exactly which page each broken link sits on. Dead-link report ready for CSV export. Pay per page crawled and link checked.
Parse RSS/Atom feeds into clean JSON items — and optionally fetch the FULL article text behind each link, not just the summary. Filter by age, cap per feed. Failures are free.
App ids or URLs in, user reviews out — from both the Apple App Store (public RSS) and Google Play. Rating, title, text, author, version and date per review. Store auto-detected. Charged per review returned.
URLs in, clean article JSON out: title, author, date, full text, word count, optional Markdown. No browser, no frills — fast extraction at the lowest price. Failures are free.
Company slug or careers URL in, open roles out — straight from each company's ATS (Greenhouse, Lever, Ashby) via official public JSON. Title, location, department, type, apply URL, posted date. Store auto-detected. Charged per role.
Search terms in, news articles out: headline, source, publish time, snippet, the resolved publisher URL, and the full article body extracted from the page. Uses Google News RSS (no key). Charged per article returned.
PDF URLs in, clean text out: full text, per-page text, title/author/page-count metadata. No browser, no OCR overhead — fast digital-PDF extraction. You only pay for PDFs that actually extract; failures are free.
Extract every URL from any sitemap: recursive sitemap-index expansion, gzipped sitemaps, robots.txt auto-discovery, include/exclude filters, lastmod output. Feed the URL list to any crawler or audit. Pay per sitemap file processed.
Watch pages or JSON APIs on a schedule and get precise diffs — added/removed paragraphs or changed JSON paths — plus optional webhook alerts. Pay only per check and per detected change.
URLs in, structured metadata out: title, description, canonical, Open Graph, Twitter Card, favicons, hreflang, robots and JSON-LD schema types. Clean object per page for link previews, SEO audits and social-card debugging.
URLs in, screenshots out: full-page or viewport PNG/JPEG, whole-page PDF, retina (2x) scaling, custom viewport, dark-mode, element capture and consent-banner hiding. Failed pages are reported and never charged.
Video URLs or IDs in, full transcripts out: plain text plus timestamped segments, language selection, auto-generated and manual captions. Videos without transcripts are reported and never charged.