Scrape ORCID researcher registry. Modes: search profiles, researcher details by ORCID iD, works/publications, employment and education history. Extracts names, affiliations, DOIs, funding, peer reviews. Official Public API. For academic network analysis & research mapping.
Scrape arXiv research papers by keyword, category, or author. Extracts titles, abstracts, authors, citations, and metadata. Perfect for AI/ML research monitoring, literature reviews, and LLM training data collection.
Scrape full comment trees from Hacker News stories. Extracts threaded discussions with author, text, points, depth, timestamp. Fetch by story ID, URL, or top/new stories. Ideal for sentiment analysis, tech discourse research, and developer community insights. Uses official Firebase API.
Scrape Kaggle datasets marketplace. Modes: search by keyword/tag, dataset details (owner, license, file list, size, votes, downloads), trending, and user profiles. Extracts titles, descriptions, updated dates, usability scores. Ideal for ML dataset discovery and competitive landscape research.
Scrape Lobsters (lobste.rs) tech news stories and tags. Extract story titles, URLs, scores, comments, authors, and tag categories. Filter by tag, date, or popularity. Perfect for niche tech trend analysis and developer news aggregation.
Scrape npm the JavaScript package registry. Search packages, extract metadata, download statistics, dependencies, and version history. Track package popularity trends. Essential for JavaScript and Node.js ecosystem research and dependency analysis.
Scrape Bluesky posts, profiles, feeds and search results. Extract text, authors, engagement stats, media. No auth required. Social listening, trend monitoring, LLM training data.
Scrape Crates.io the Rust package registry. Extract crate metadata, versions, downloads, dependencies, and documentation links. Search by keyword or category. Ideal for Rust ecosystem analysis, dependency auditing, and package discovery.
Scrape GBIF species taxonomy & occurrence data. Extract scientific names, common names, kingdoms, habitats, and classifications from 2.4B+ biodiversity records. Perfect for ecological research, biodiversity monitoring, taxonomy databases, and citizen science apps.
Scrape Zenodo.org (CERN open research repository) for records, datasets, and software. Four modes: search with type/access filters, record details by DOI/ID, community browse, recent submissions. Extracts titles, authors, DOIs, files, stats. Uses official API. No auth, 60 req/min.
Scrape Medium articles by keyword, author, or tag. Extract titles, full text, claps, reading time, tags, author info, and publication metadata for content research, competitor analysis, and topic monitoring.
Extract transcripts and subtitles from YouTube videos. Get auto-generated or manual captions in any language. Bulk extraction from video URLs, channels, or playlists. Output as plain text, timestamped segments, or SRT. Perfect for content repurposing, SEO, and video analysis.
Scrape Product Hunt for product launches, upvotes, comments, and maker profiles. Discover trending startups, upcoming products, and daily/weekly featured launches. Track product performance and community engagement. Perfect for market research and startup scouting.
Scrape Crossref — largest DOI registry for academic literature. Modes: search works, DOI lookup, journal metadata, funder info, affiliation search. Extracts titles, authors, DOIs, ISSN, references, citations. Official REST API, no auth, 50 req/sec. For research & citation analysis.
Scrape GBIF (Global Biodiversity Information Facility) for species and occurrences. Modes: species search, occurrence records, dataset browse, country filters. Extracts scientific names, lat/long, dates, publishers, images, license. Official REST API, no auth.
Scrape computer science publications from DBLP. Search papers by keyword, get author profiles with publication lists, and retrieve venue/conference information. Access 6M+ publications from the largest CS bibliography.
Scrape Hacker News stories, comments, and user profiles. Extract trending tech news, top stories by score, new submissions, Ask HN, Show HN, and job posts. Filter by date, score, and comment count. Perfect for tech trend analysis, competitive intelligence, and content curation.
Scrape Open Library (Internet Archive) for books, authors, and editions. Modes: search by title/author/subject, book details by ISBN/OLID, author works, recent additions. Extracts titles, authors, ISBNs, covers, subjects, publish dates, editions. Uses official Search & Works API. No auth.
Scrape PyPI the Python Package Index. Extract package metadata, download statistics, version history, dependencies, and maintainer info. Track new releases and popularity trends. Perfect for Python ecosystem analysis and package research.
Scrape Stack Overflow questions, answers, tags, and user profiles. Search by keyword, tag, or date range. Extract vote counts, accepted answers, code snippets, and discussion threads. Ideal for developer knowledge mining and technical research.
Scrape Wikipedia across 300+ languages. Modes: full articles, summaries, search, random, recent changes, category browse. Extracts text, sections, references, images, links, infobox. Official MediaWiki API — stable, no auth. Great for research, knowledge graphs, content enrichment.
Scrape GitHub Trending repositories and developers. Extract repo names, descriptions, stars, forks, language, and daily/weekly/monthly trends. Discover rising open source projects and active contributors. Essential for tech scouting and OSS research.
Scrape Hugging Face models, datasets, and Spaces. Extracts metadata, downloads, likes, tags, and usage stats. Ideal for AI model discovery, competitive analysis, and tracking trending ML resources.
Scrape Semantic Scholar for academic papers, citations, abstracts, and author profiles. Search by topic, author, or venue. Extract citation graphs, reference lists, and research trends. Essential for literature reviews, academic research, and AI/ML paper discovery.