Give it a list of websites, get back one clean contact profile per site: emails (incl. de-obfuscated), phone numbers, and social profiles (LinkedIn, X, Facebook, Instagram, YouTube, TikTok, GitHub) — each with the exact page it was found on. Precision-first: no junk emails, no false-positive phon...
Turn any list of RSS/Atom feeds — or just website homepages, feeds are auto-discovered — into a clean, deduplicated stream of news items, optionally filtered by keywords. Normalized output across RSS 2.0 and Atom: title, link, published date, author, categories, summary and full content as plain...
Export the complete product catalog of any Shopify store: titles, handles, vendors, tags, prices, variants with SKUs and availability, images and timestamps. Uses the store's own JSON endpoints (fast, no browser) and automatically falls back to product sitemaps to get past the 5,000-product pagin...
Crawl any website and get one clean Markdown document per page — ready for RAG pipelines, vector databases, LLM fine-tuning, or docs migration. Boilerplate (nav, footers, cookie banners) stripped, main content auto-detected, sitemap-seeded crawling, robots.txt respected. HARD page caps and flat p...
Scrape live job postings straight from company career boards (Greenhouse, Lever, Ashby, Workable, Recruitee) using their stable public JSON APIs — not fragile HTML. One unified schema across all five ATS providers: title, location, remote flag, department, compensation (where published), apply UR...