Scrapes haken (派遣) and temp staffing job listings from hatarako.net — Japan's leading dispatch job aggregator. Extracts job title, staffing agency, hourly wage, occupation, location, work hours, contract type, start date, required skills, and job URL across all prefectures.
Scrape Indeed job listings by keyword and location. Extract job titles, companies, locations, salaries, job types, posting dates, and full descriptions. Perfect for market research, recruiting intelligence, and hiring trend analysis.
Scrapes the US insurance carrier directory from the NAIC Consumer Information Source — ~5,200 licensed carriers with NAIC code, carrier name, licensed states, lines of business, address, phone, and website. Covers PC, life, health, surplus-lines, and captive carriers.
Scrape Kenbiya (建美家 / kenbiya.com) — Japan's #2 investment-property portal
after Rakumachi. Extracts yield, monthly rent, occupancy, structure, and
broker data from ~20-40K active listings. Natural companion to
rakumachi-investment-scraper for full Japan investor-market coverage.
Scrape public room metadata from LINE OpenChat — Japan's largest public chat platform, also popular in Taiwan and Thailand. Extracts room name, description, member count, hashtags, region, and more from the public OpenChat directory. No LINE account required. Supports JP, TW, and TH markets.
Scrape live listings from magi (magi.camp) — Japan's leading C2C trading-card marketplace. Search by keyword (Pokémon, Yu-Gi-Oh!, One Piece, MTG, etc.) and collect listing prices, sold signals, favorite counts, badges, and image URLs.
Scrape property listings from Mexico's top portals — Inmuebles24 and Vivanuncios. Filter by operation (sale, rent), property type, and location. Returns price, area, bedrooms, and photos. Built for US investors, PropTech pipelines, and market analysts.
Bulk scraper for Mexico's DENUE business registry — 5.5M+ establishments with geo coordinates, SCIAN industry codes, employee ranges, contact info, and addresses. Filter by state, activity keyword, and entity type. Requires a free INEGI API token (email registration at inegi.org.mx).
Search and look up music metadata from MusicBrainz — the open music encyclopedia. Retrieve artists, release groups, recordings, labels, and works with full relationship graphs, genre tags, and cross-platform ID resolution (Spotify, Discogs, ISRC, Wikidata, and more). No API key required.
Scrape graded and raw trading card listings from MyCardPost — a fast-growing P2P card marketplace. Extracts title, price, grade, grader, sport, year, set, player, seller, and images from server-rendered detail pages.
Scrapes NewsletterHunt's cross-publication email archive. Extracts subject lines, sender, date, full email body HTML and plain text, plus the newsletter signup URL. Covers hundreds of publications and tens of thousands of archived issues — no login required.
Scrape the official NFL schedule from nfl.com. Returns every game for a season -- home/away teams, kickoff time, venue, broadcast network (CBS/FOX/NBC/ESPN/TNF), game status, primetime flags, and international game markers.
Extract European company records: registry identity, officers with appointment dates, multi-year published financials, merger and acquisition events, and register filings. Covers registers across Germany, Austria, Switzerland, UK, France, Benelux, and the Nordics.
Fetch active NWS weather alerts and multi-day forecasts for a list of coordinates (trailheads, campgrounds, peaks). Supply lat/lon pairs and get per-point active alerts plus 7-day forecast in one pass. Purpose-built for outdoor trip planning and trail-condition workflows. No API key required.
Scrapes Texas Railroad Commission (RRC) oil and gas well data. Returns API number, operator, lease, field, county, district, well type, on-schedule status, and drilled depth. Covers all 13 RRC districts and over 300k wells. Built for mineral rights research, energy analytics, and landman databases.
Scrapes the full Platzi course catalog (~1500+ courses across all categories). Returns structured data including title, teacher, rating, duration, level, syllabus topics, and certification status for every course. Covers Spanish and Portuguese tracks.
Normalizes published prompt-injection and jailbreak datasets from HuggingFace and GitHub research repos into one labeled corpus: technique, target model, defense bypassed, license, cross-source dedup. Defensive only — aggregates public data for guardrail/eval testing, never targets a live LLM.
Fetch crowd-sourced Linux and Steam Deck compatibility ratings from ProtonDB for any list of Steam game IDs. Returns compatibility tier, confidence level, score, and report count per game — joinable with the Steam catalog for full Deck-readiness analysis.
Extract government surplus auctions from PublicSurplus.com across every browse category - vehicles, heavy equipment, computers, furniture, industrial gear and more - from roughly 9,000 US and Canadian government sellers, with agency, location, bid, and sold-price data.
Scrape the QueryTracker literary agent directory — the #1 querying-author platform. Extracts agent name, agency, genres, query status, query method, social links, and last-updated date from every public agent profile. Ideal for query-CRM tools and literary data pipelines.
Scrape authenticated pre-owned luxury listings from RECLO, Japan's major luxury consignment marketplace. Returns product ID, title, brand, price in JPY, RECLO condition rank (N, S, A, B or C), product URL and primary image for every listing.
Scrapes the complete product catalog from REP Fitness and Titan Fitness — the two dominant direct-to-consumer strength-equipment retailers on Shopify. Returns normalized product records with titles, variants, pricing, availability, and images from both sources in a single dataset.
Scrape live job listings from Rozee.pk, Pakistan's dominant job board. Returns title, company, city, salary range, experience level, education requirement, functional area, industry, gender preference, skills, and full description for every posting in the unfiltered index.
Scrape rugs.com for catalog pricing, size variants, materials, construction, and availability across 60 000+ SKUs. One record per SKU with all size-specific price fields.