Normalizes the scattered public GitHub collections of leaked/published AI system prompts (ChatGPT, Claude, Cursor, Devin, v0, Perplexity, and more) into one deduplicated dataset with product, vendor, version, and leak-date fields. Passive: reads only already-public repositories.
Scrape classified ads from leboncoin.fr — title, price, location, images, attributes, and owner type. Pass any category or search URL.
Scrape auction property listings from leilaoimovel.com.br — Brazil's largest real estate auction portal. Extracts sale price, appraisal value, discount percentage, modality, closing date, bank, property type, address, and more. Filter by state, city, property type, and bank.
Scrapes employer reviews, ratings, and salary data from en-hyouban.com (エン カイシャの評判) — Japan's #2 employer-review platform. Extracts JSON-LD EmployerAggregateRating, per-dimension breakdowns, crowd-sourced salary, and employee reviews by URL or sitemap.
Scrape LinkedIn Ad Library ads by keyword, company ID, or country: advertiser name, ad creative, and the full EU DSA disclosure block — total and per-country impressions, audience targeting, and the legal entity paying for each ad.
Swiss business directory data: company names, addresses, phone numbers, emails, ratings, and opening hours for the Swiss SMB market. Filter by industry vertical. Structured JSON output ready for lead lists or CRM import.
Scrape crowd-sourced law school admissions data from LSD.Law — median LSAT and GPA, 25th/75th percentiles, acceptance rates, US News rankings, and application counts for all ABA-accredited law schools.
Scrape used heavy equipment listings from MachineryTrader.com across every major equipment category (excavators, skid steers, dozers, cranes, loaders, and more). Get pricing, full specs, dealer contact info, location, and photos for every for-sale listing.
Cross-ecosystem feed of confirmed malicious/typosquatted packages across npm, PyPI, crates.io, Go, Maven, NuGet, Packagist and RubyGems, sourced from the OSSF malicious-packages dataset (also feeds OSV.dev). Category taxonomy plus typosquat-target linkage for CI gating. Passive read only.
Scrapes the Mararun platform — the dominant Chinese marathon management SaaS. Returns event details for mararun-hosted Chinese marathons: name, date, city, registration windows, participant cap, organizer, and CAA/AIMS certification.
Scrape the complete Maskota.com.mx product catalog — Mexico's leading online pet retailer. Extracts product names, prices, brands, variants, tags, and category data for all pets across the Shopify-hosted store.
Search Meta Ad Library ads by keyword, advertiser page, country and ad type. Returns spend and impressions ranges, reach estimates, per-country reach, payer byline, state-media and AI-media disclosure flags, and creative-variant collation counts the standard scrapers omit.
Scrape the full MTGGoldfish Magic: The Gathering card price index. Extracts paper, online (MTGO), and foil prices for cards across all sets with weekly price-change data.
Extract live and completed government surplus auctions from Municibid — vehicles, equipment, and surplus gear from small-town and county agencies. Includes bid pricing, realized sold prices on ended lots, seller and location details, and vehicle specs like VIN and mileage.
Search Naver Place — Korea's dominant local directory — by region and category for business records: address, phone, hours, category, coordinates, visitor/blog review counts, booking/order flags, menu items, amenities, photos. Restaurants, salons, clinics, lodging, general businesses, nationwide.
Scrapes the National Dance Education Organization's college dance program directory, job board, and events. Returns institution names, locations, contact info, degree offerings, and faculty counts.
Scrape product listings from Newegg category and search pages. Extracts product title, brand, current price, was price, shipping, rating, review count, stock status, seller name (1P Newegg vs 3P marketplace), item number, and image URL.
Scrapes Noble Knight Games for out-of-print and collectible board games, RPGs, wargames, and miniatures. Extracts condition grades (New/Mint/NM/VG+/etc.), price per condition tier, publisher, category, stock, and condition price ladder. The go-to OOP/used tabletop catalog for collector valuation.
Look up French residents by name or phone number on PagesBlanches, the national residential directory. Returns full name, phone, and address (street, postal code, commune, department, region) per match. Covers all of France — search a name, a number, or narrow by commune, department, or postal code.
Scrape business listings from PagesJaunes.fr (French Yellow Pages) by keyword and location. Extracts business name, address, phone, email, website, opening hours, rating, and geolocation. Supports category-based search across any city, department, or region in France.
Scrape Italian business listings from PagineGialle by category and comune — name, address, phone, category, rating, opening hours, and partita IVA/email/website/servizi from each listing's detail page.
Scrape listings from PayPay Flea Market (Yahoo! Flea Market), Japan's second-largest C2C resale app. Extracts item id, title, price, condition, seller info, and images from embedded __NEXT_DATA__ JSON.
Convert PDF documents into structured JSON. Extracts text, tables, and fields from any PDF URL, and can optionally run an AI structuring pass that turns raw text into clean, organized JSON ready for automation or analysis. No API key needed.
Scrape the official Penguin Random House publisher catalog for book metadata: title, author, ISBN, imprint, format, publication date, price, description, praise blurbs, and more. Search-seeded crawl into detail pages returns primary-source data not available from consumer aggregators.