Normalizes published prompt-injection and jailbreak datasets from HuggingFace and GitHub research repos into one labeled corpus: technique, target model, defense bypassed, license, cross-source dedup. Defensive only — aggregates public data for guardrail/eval testing, never targets a live LLM.
Extract government surplus auctions from PublicSurplus.com across every browse category - vehicles, heavy equipment, computers, furniture, industrial gear and more - from roughly 9,000 US and Canadian government sellers, with agency, location, bid, and sold-price data.
Scrape authenticated pre-owned luxury listings from RECLO, Japan's major luxury consignment marketplace. Returns product ID, title, brand, price in JPY, RECLO condition rank (N, S, A, B or C), product URL and primary image for every listing.
Scrapes the complete product catalog from REP Fitness and Titan Fitness — the two dominant direct-to-consumer strength-equipment retailers on Shopify. Returns normalized product records with titles, variants, pricing, availability, and images from both sources in a single dataset.
Scrape live job listings from Rozee.pk, Pakistan's dominant job board. Returns title, company, city, salary range, experience level, education requirement, functional area, industry, gender preference, skills, and full description for every posting in the unfiltered index.
Scrape ranked startup lists from SeedTable — startups across 4,000+ city lists worldwide. Returns startup name, industry, location, profile URL, and logo for every ranked company, and filters to just the cities you care about.
Scrape curated highlight stories from public Snapchat profiles. Provide usernames and get direct media URLs, thumbnails, story titles, and creator metadata — one record per story snap.
Scrape job listings from StepStone across Germany, Austria, and Belgium. Get structured data on titles, companies, locations, salaries, work mode, contract type, and full job descriptions.
Scrapes highway and bridge construction bid-letting calendars from US state DOTs and federal SAM.gov. Covers upcoming bid openings, awarded contracts, engineer estimates, and DBE goals. Sources: Florida DOT and SAM.gov NAICS 237xxx solicitations.
Scrapes Wimbledon draws and match scores from wimbledon.com internal JSON feeds — the same endpoints the official site uses. Covers all five draws (MS/LS/MD/LD/XD) with per-set scores, seedings, and court assignments. Defaults to current year; no login required.
Scrapes Anytime Fitness club locations from their sitemap — covers 3,000+ US and international clubs. Extracts club name, address, contact details, coordinates, hours, and amenities from each location page.
Scrape property listings (apartments, houses, for-sale and rental) from ImmobilienScout24.de. Returns price, living space, rooms, address, realtor, features, and images per listing.
Scrape live Kick.com streamer leaderboards by category — viewer counts, stream titles, and channel profiles (bio, follower count, social links). Pick specific categories or let it auto-discover the top trending ones each run.
Look up French residents by name or phone number on PagesBlanches, the national residential directory. Returns full name, phone, and address (street, postal code, commune, department, region) per match. Covers all of France — search a name, a number, or narrow by commune, department, or postal code.
Scrape real estate listings and property data from Redfin — homes for sale, recently sold, and rentals. Returns price, beds, baths, sqft, address, coordinates, MLS status, year built, days on market, and sold prices. Search by city, ZIP, or Redfin URL. Covers MLS for-sale, sold, and rental listings.
Scrape attorney and law firm data from FindLaw Lawyer Directory to generate high-quality, targeted legal industry leads
Scrape attorney and law firm profiles from Martindale.com. Extract names, contact details, practice areas, firm info, and more. Filter by US state and legal practice area.
Scrape therapist and counselor profiles from PsychologyToday.com. Extract names, credentials, specialties, insurance accepted, fees, languages, addresses, phones, and bios. Filter by U.S. state, ZIP, or any category slug (specialty, insurance, language, treatment orientation).
Scrape 3D model listings from Cults3D, the leading paid 3D-model marketplace. Extracts title, creator, category, tags, price, license, like/download/collection counts, publish date, and thumbnail for each model. Supports filtering by one or more categories.
Extract company data from Dun & Bradstreet's public business directory. Search by name, industry, location, or NAICS/SIC code. Returns firmographic data: address, phone, website, industry classification, revenue range, employee count, founding year, company hierarchy, and key executives.
Scrape handelsregister.de for German companies: HR numbers, directors, shareholders (Gesellschafterliste), share capital, articles of association, branches. Pair with bundesanzeiger for full filings.
Scrape public Telegram channel messages without API credentials. Extracts text, media, reactions, views, and channel metadata from t.me preview pages.
Scrape the Swiss Federal Registry of Commerce (Zefix) for companies: UID, legal name, form, seat, purpose, address, auditors, branches, and full SHAB legal-events feed. Search by name, UID, or canton. KYC/KYB for banks, family offices, crypto, traders.
Scrape listings from OLX.pl — Poland's largest classifieds marketplace. Extract price, location, seller details, and category for any category or keyword search. Covers all OLX categories: goods, automotive, real estate, jobs, and services.