Scrapes Brazil's INPE satellite imagery catalog (CBERS-4, CBERS-4A, Amazonia-1) via the STAC API. Supports filtering by collection, bbox, date range, and cloud cover. Returns scene metadata including satellite, sensor, path/row, acquisition date, cloud cover, sun angles, thumbnail and download URLs.
Scrapes deforestation and fire data from INPE TerraBrasilis. Collects PRODES deforestation totals, DETER daily alerts, and fire records across Amazon, Cerrado, Pantanal, and other Brazilian biomes. Covers all states from 2015 to present. Essential for EUDR compliance and carbon offset research.
Scrape Instructables.com project metadata across all craft verticals including electronics, woodworking, sewing, cooking, and 3D printing. Returns title, author, category, difficulty, step count, materials summary, favorites, views, and comments counts per project.
Download and parse the IRS Tax Exempt Organization auto-revocation list. Returns all nonprofits that lost tax-exempt status for non-filing, including reinstatement records.
Scrapes the public JailbreakBench leaderboard tracking attack-success-rate for jailbreak techniques (PAIR, GCG, AIM, and more) against open- and closed-source LLMs, with and without defenses (SmoothLLM, perplexity filter, etc). Snapshot each run to track technique-vs-model ASR movement over time.
Scrape listings from Jmty (ジモティー), Japan's dominant hyper-local classifieds and flea-market platform. Extract titles, prices, locations, categories, descriptions, images, and seller info by prefecture and category.
Job postings across Malaysia, Singapore, the Philippines, Indonesia, Hong Kong and Thailand — sourced from JobStreet and JobsDB (SEEK Asia). Get titles, companies, salaries, locations, classifications and full descriptions in one dataset.
Scrapes Indonesian public transit: KAI (PT Kereta Api Indonesia) station catalog (217 stations across Java, Sumatra, Sulawesi) with codes and city data, plus TransJakarta BRT corridor listing. For travel apps, logistics planners, and Indonesian transit integrations.
Extract Kavak's owned used-car inventory across Mexico, Brazil, Argentina and Chile: price, financing, mileage, spec, hub location and availability for every listed vehicle, normalised into one cross-country schema.
Scrapes Keep (gotokeep.com) — China's largest fitness platform with 200M+ users. Extracts yoga, meditation, qigong, HIIT and other fitness courses including title, difficulty level, participant count, equipment required, cover image, and category grouping.
Extract vehicle valuations from Kelley Blue Book. Get fair market price, MSRP, trade-in value, consumer reviews, specs, and trim details for any car by year, make, and model. Built for dealers, researchers, and automotive market analysis.
Scrape Kickstarter campaigns across the whole site — all 15 categories and 154 subcategories — with funding goal, amount pledged, percent funded, backer count, creator history, launch date and deadline. Pick any mix of categories, or crawl everything.
Scrape classified listings from Kleinanzeigen.de (formerly eBay Kleinanzeigen). Extract listing details including title, price, description, location, seller info, images, and category-specific attributes.
Scrape Japanese intercity bus data from Kosokubus.com. Covers Willer Express, JR Bus (all 6 regional subsidiaries), Sakura Kotsu, Keio Bus, Odakyu City Bus, Meitetsu Bus, and 30+ operators. Returns seat class, amenities, overnight flag, fares in JPY, and availability.
Scrape Malaysian transit data: KTMB rail stations (ETS, Intercity, Komuter — ~150 stations with IDs and state groupings) and Easybook intercity bus routes (80+ routes with distance, duration, and city-pair metadata). Transit network graph for travel apps, MaaS platforms, and mapping services.
Bulk-extract LeadIQ company profiles: industry, SIC/NAICS codes, employee bands, tech stack, key executives, continents of operation, and the observed corporate email-format pattern with confidence percentages. Covers roughly 1,045,000 companies with resumable crawls.
Normalizes the scattered public GitHub collections of leaked/published AI system prompts (ChatGPT, Claude, Cursor, Devin, v0, Perplexity, and more) into one deduplicated dataset with product, vendor, version, and leak-date fields. Passive: reads only already-public repositories.
Scrape classified ads from leboncoin.fr — title, price, location, images, attributes, and owner type. Pass any category or search URL.
Scrape auction property listings from leilaoimovel.com.br — Brazil's largest real estate auction portal. Extracts sale price, appraisal value, discount percentage, modality, closing date, bank, property type, address, and more. Filter by state, city, property type, and bank.
Scrape Lichess player profiles, full game histories (PGN + clocks + analysis), arena and Swiss tournament listings, and per-player rating history from Lichess's open REST API. No API key required. Supports bulk username input and all game variants.
Scrapes employer reviews, ratings, and salary data from en-hyouban.com (エン カイシャの評判) — Japan's #2 employer-review platform. Extracts JSON-LD EmployerAggregateRating, per-dimension breakdowns, crowd-sourced salary, and employee reviews by URL or sitemap.
Scrape LinkedIn Ad Library ads by keyword, company ID, or country: advertiser name, ad creative, and the full EU DSA disclosure block — total and per-country impressions, audience targeting, and the legal entity paying for each ad.
Scrape US federal lobbying disclosure filings from the Senate LDA database. Extract registrants, clients, lobbyists, activities, issue areas, and government entities. Filter by year, period, issue, registrant, or client.
Swiss business directory data: company names, addresses, phone numbers, emails, ratings, and opening hours for the Swiss SMB market. Filter by industry vertical. Structured JSON output ready for lead lists or CRM import.