Point at any apt (Debian Packages) or yum/dnf (repomd/primary.xml) repository and extract every binary package's metadata — name, version, deps, checksums, size, license. Generic repo-index parser for SBOM, supply-chain audit, and mirror monitoring.
Point at ANY Matrix homeserver and extract its public room directory over the Client-Server API. One row per room: room ID, name, topic, member count, canonical alias, avatar, join rule, room type. Opaque next_batch pagination; optional token unlocks search/space filters.
Point at ANY MediaWiki wiki's Action API (Wikipedia, Fandom, wiki.gg, or any corporate/OSS wiki) and pull structured data: full-text search, enumerate pages, list category members, fetch page text + revisions, or read site info. Handles continue-token pagination. Pay per record.
Point at ANY Mobilizon instance and extract events + groups over its public GraphQL API. Browse or full-text-search the federated events/groups directory, or fetch one event/group by id. One clean row each: title, dates, category, location, organizer, tags, member counts.
Point at ANY OGC API - Coverages endpoint (the JSON/CoverageJSON successor to WCS) — pygeoapi, GNOSIS, ldproxy, GeoServer. List coverage collections, describe their schema (axes, params, CRS), or summarize CoverageJSON (axis shape + value count) as flat rows. Pay per record.
Point at ANY OGC API - Tiles endpoint (the JSON-native successor to WMTS) — ldproxy, GNOSIS, GeoServer, pygeoapi. Inventory vector/map tileset offerings, list tile matrix sets (CRS, zoom levels, axes), or discovery-scan a server's tiles capabilities as flat rows. Pay per record.
Point at ANY OGC API - EDR (Environmental Data Retrieval) service — Met Office, FMI, pygeoapi. List collections, then query environmental data by position, area, radius, trajectory, or cube. Flattens CoverageJSON / GeoJSON to rows (parameter, value, unit, lon/lat, time). Pay per record.
Point at ANY OHDSI WebAPI (the OMOP-CDM / ATLAS analytics REST API) and pull server info, data sources & daimons, saved cohort definitions, concept sets, and OMOP vocabulary concept search into clean flat rows. One generic runner for any observational-health host. Pay per record.
Point at ANY German OParl council-information system (Ratsinformationssystem) for flat rows — papers (Drucksachen), meetings, committees, people. One actor, every vendor (STERNBERG, more-rubin, CC e-gov): reads list URLs from the Body, follows links.next paging, skips dead endpoints. Pay per record.
Point at ANY OpenActive dataset site, data catalog, or RPDE feed and harvest UK sports & activity opportunity data — classes, sessions, facility slots and events from 100s of leisure operators (GLL, Everyone Active, Halo...). Follows the RPDE cursor, handles tombstones. Pay per item.
Scrape scholarly works (papers, preprints, books), authors, institutions & journals from the free public OpenAlex API. Search + filter by author, institution, concept, year, type & open-access; sort; auto-paginate. Clean columns + raw. Pay per record.
Point at ANY OpenDataSoft portal (Paris, RTE, data.opendatasoft.com + thousands more) and pull real ROW data via the Explore API v2.1. List the dataset catalog or extract records with ODSQL where/select/order_by/group_by + facets. Auto-paginates to flat rows + geo. Pay per record.
Point at ANY openEO back-end (the standard API for Earth-observation cloud back-ends — VITO/Terrascope, Copernicus Data Space, EODC). Extract collections, the process registry, capabilities, file formats, service types or UDF runtimes as flat rows + raw JSON. No auth. Pay per record.
Scrape the U.S. FDA's official openFDA API — drug & device adverse events, recalls/enforcement, 510(k) clearances, product labels, NDC, food & animal events. One actor over 18 official public datasets with the full search/count/sort grammar. Pay per record.
Point at any PeerTube instance and export its videos, channels and instance metadata from the public /api/v1 REST API. Federated, self-hosted YouTube alternative — one actor over thousands of instances.
Point at any Prometheus HTTP API v1 server (Prometheus, Thanos, Mimir, VictoriaMetrics, Grafana Agent) and export PromQL query results, range series, label values, scrape targets, rules, metric metadata and build info as clean rows. One actor over every Prometheus-compatible server.
Point at ANY W3C Reconciliation (OpenRefine) service and pull results: fetch the capability manifest, batch-match query strings to candidate entities with scores, autocomplete (suggest), and fetch property values (extend). Works with Wikidata, GND, Getty and any conforming endpoint.
Point at ANY RPKI validator's VRP export (rpki-client, Routinator, FORT, Cloudflare) and extract validated ROAs: origin ASN, prefix, maxLength, trust anchor, expiry. Optional ASN / prefix / trust-anchor filters. Auto-detects the array key. Seeded on Cloudflare & rpki-client.org.
Point at ANY RSS or Atom feed (news, blogs, podcasts, GitHub releases, YouTube) and get clean structured rows. Auto-detects RSS 2.0 / RSS 1.0 / Atom and normalizes every entry to one schema: title, link, content, author, categories, ISO-8601 dates, enclosures. Pay per item.
Scrape U.S. SEC EDGAR company filings + metadata from the official public JSON APIs. Query by ticker, CIK, or full-text keyword; filter by form type (10-K, 10-Q, 8-K…) and filing-date range. One clean record per filing with the index + document URLs. Pay per filing.
Give a ticker or CIK, get every reported XBRL financial fact for any US public company from the official SEC EDGAR company-facts API (data.sec.gov) — revenue, assets, EPS, cash, liabilities and every us-gaap + dei concept, flattened to one analysis-ready row per observation. No API key.
Point at ANY OGC SensorThings API v1.1 instance (FROST-Server, GOST, SensorUp and thousands of air-quality, smart-city and research deployments) and get flat sensor rows — observations, things with coordinates, datastreams with units — plus raw JSON. @iot.nextLink paging, OData filters. Pay per row.
Point at ANY Skosmos SKOS REST API and pull structured vocabulary data: list vocabularies, search concepts, resolve a concept URI to full detail (labels, broader/narrower/related, mappings), fetch top concepts and RDF types. Works with AGROVOC, ZBW/STW, Loterre and any conforming server.
Point at ANY Socrata open-data portal (NYC, Chicago, WA, hundreds of gov portals). Query rows via SoQL, list a domain's catalog, or pull a dataset's column schema. One runner for the whole SODA ecosystem.