Convert any URL or raw HTML to clean markdown, plaintext, and reader-mode metadata. Choose Readability (articles), Turndown (verbatim), or Trafilatura (noisy pages). Outputs title, byline, published date, language, word count, and reading time. FREE — the one-call URL-to-markdown primitive.
Scrape AI/ML model metadata from the HuggingFace Hub. Extract model names, task types, download counts, likes, libraries, authors, tags, licenses, model sizes, and model card excerpts. Filter by task type, library, author, and search query.
Extract live salvage-vehicle auction listings from IAAI's public search. Search by keyword or filter by state. Returns stock number, partial VIN, year/make/model, damage type, odometer, title type, run-and-drive status, sale date, branch location, and more.
Scrapes Brazil's INPE satellite imagery catalog (CBERS-4, CBERS-4A, Amazonia-1) via the STAC API. Supports filtering by collection, bbox, date range, and cloud cover. Returns scene metadata including satellite, sensor, path/row, acquisition date, cloud cover, sun angles, thumbnail and download URLs.
Scrapes deforestation and fire data from INPE TerraBrasilis. Collects PRODES deforestation totals, DETER daily alerts, and fire records across Amazon, Cerrado, Pantanal, and other Brazilian biomes. Covers all states from 2015 to present. Essential for EUDR compliance and carbon offset research.
Scrape Instructables.com project metadata across all craft verticals including electronics, woodworking, sewing, cooking, and 3D printing. Returns title, author, category, difficulty, step count, materials summary, favorites, views, and comments counts per project.
Download and parse the IRS Tax Exempt Organization auto-revocation list. Returns all nonprofits that lost tax-exempt status for non-filing, including reinstatement records.
Search the iTunes/Apple Music catalog by keyword, artist, or album; look up tracks by Apple ID; and fetch Top Charts rankings for any country. Metadata and 30-second preview URLs only — no audio downloads.
Scrapes the public JailbreakBench leaderboard tracking attack-success-rate for jailbreak techniques (PAIR, GCG, AIM, and more) against open- and closed-source LLMs, with and without defenses (SmoothLLM, perplexity filter, etc). Snapshot each run to track technique-vs-model ASR movement over time.
Scrape listings from Jmty (ジモティー), Japan's dominant hyper-local classifieds and flea-market platform. Extract titles, prices, locations, categories, descriptions, images, and seller info by prefecture and category.
Job postings across Malaysia, Singapore, the Philippines, Indonesia, Hong Kong and Thailand — sourced from JobStreet and JobsDB (SEEK Asia). Get titles, companies, salaries, locations, classifications and full descriptions in one dataset.
Scrapes Indonesian public transit: KAI (PT Kereta Api Indonesia) station catalog (217 stations across Java, Sumatra, Sulawesi) with codes and city data, plus TransJakarta BRT corridor listing. For travel apps, logistics planners, and Indonesian transit integrations.
Extract Kavak's owned used-car inventory across Mexico, Brazil, Argentina and Chile: price, financing, mileage, spec, hub location and availability for every listed vehicle, normalised into one cross-country schema.
Scrapes Keep (gotokeep.com) — China's largest fitness platform with 200M+ users. Extracts yoga, meditation, qigong, HIIT and other fitness courses including title, difficulty level, participant count, equipment required, cover image, and category grouping.
Extract vehicle valuations from Kelley Blue Book. Get fair market price, MSRP, trade-in value, consumer reviews, specs, and trim details for any car by year, make, and model. Built for dealers, researchers, and automotive market analysis.
Scrape Kickstarter campaigns across the whole site — all 15 categories and 154 subcategories — with funding goal, amount pledged, percent funded, backer count, creator history, launch date and deadline. Pick any mix of categories, or crawl everything.
Scrapes the complete Kith product catalog from Shopify — prices, per-size availability, sale/markdown flags, and variant data for resale tracking and arbitrage.
Scrape classified listings from Kleinanzeigen.de (formerly eBay Kleinanzeigen). Extract listing details including title, price, description, location, seller info, images, and category-specific attributes.
Scrape Japanese intercity bus data from Kosokubus.com. Covers Willer Express, JR Bus (all 6 regional subsidiaries), Sakura Kotsu, Keio Bus, Odakyu City Bus, Meitetsu Bus, and 30+ operators. Returns seat class, amenities, overnight flag, fares in JPY, and availability.
Scrape Malaysian transit data: KTMB rail stations (ETS, Intercity, Komuter — ~150 stations with IDs and state groupings) and Easybook intercity bus routes (80+ routes with distance, duration, and city-pair metadata). Transit network graph for travel apps, MaaS platforms, and mapping services.
Extract attorney profiles, contact details, practice areas, and bios directly from law firm websites. Provide a list of law firm URLs and get structured attorney data including name, title, email, phone, education, bar admissions, and headshot.
Bulk-extract LeadIQ company profiles: industry, SIC/NAICS codes, employee bands, tech stack, key executives, continents of operation, and the observed corporate email-format pattern with confidence percentages. Covers roughly 1,045,000 companies with resumable crawls.
Normalizes the scattered public GitHub collections of leaked/published AI system prompts (ChatGPT, Claude, Cursor, Devin, v0, Perplexity, and more) into one deduplicated dataset with product, vendor, version, and leak-date fields. Passive: reads only already-public repositories.
Scrape classified ads from leboncoin.fr — title, price, location, images, attributes, and owner type. Pass any category or search URL.