Discoverability
Store classification
Categories
No public values available.
Store labels
No public values available.
Catalog ID mFh8ri33clPr5R4I8
Public Apify Actor
parseforge/dati-gov-it-italy-open-data-scraper
Scrape all 65,960 datasets from Italy's national open data portal dati.gov.it: full DCAT-AP_IT metadata, distributions, 399 publishers, licences and live link checks.
Active users / 30d
1
Total users
2
Runs / 30d
18
Total runs
31
Rating
5.0
1 reviews
Bookmarks
0
Usage history
Latest captured Active users / 30d
1
+1 since 28 Aug
19 of 30 UTC days captured
Latest captured Runs / 30d
18
+17 since 28 Aug
19 of 30 UTC days captured
1
Users / 7d
Discoverability
Categories
No public values available.
Store labels
No public values available.
Monetization
Pay per event
15 charge events configured
Actor Start
apify-actor-start
Charged when the Actor starts running. Number of events charged depends on Actor memory (one event per GB, minimum one event).
$0.054
Catalogue page scanned
search-page
One page of up to 200 datasets read from the dati.gov.it CKAN index. This is the fixed cost of paging through the catalogue, spread across every row that page yields.
$0.004
Dataset
1
Users / 30d
1
Users / 90d
dataset-row
One Italian open dataset with its full DCAT-AP_IT record: title, description, publishing organization, data holder with its IPA code, publisher and creator, EU themes and EuroVoc subthemes, tags, normalised licence, update frequency, national identifier, issue and modification dates, temporal coverage, geographical name with GeoNames link and bounding box, contact point, landing page, source catalogue it was harvested from, and every distribution summarised with its formats and total size.
$0.007
Distribution
resource-row
One downloadable file or service: its download URL, normalised and raw format, MIME type, byte size, checksum, licence and access rights, plus the identity, organization and data holder of the dataset it belongs to.
$0.003
Organization
organization-row
One of the 399 publishing bodies on the portal, with its IPA code, certified email, telephone, own open data website, Italian region, dataset count and the themes and keywords it publishes under.
$0.008
Theme
theme-row
One of the 13 EU data themes with its Italian and English titles, its EU theme code and how many datasets it carries.
$0.005
Tag
tag-row
One keyword from the catalogue vocabulary with the number of datasets carrying it, ranked, honouring whatever filters the run set.
$0.002
Licence
licence-row
One licence string as the portal stores it, with its normalised code, plain-English label, whether it is open, and how many datasets declare it. The portal records the same licence under several different spellings and this row exposes each one.
$0.004
Format
format-row
One distribution format spelling with its canonical name, whether it is machine readable, and its dataset count. The catalogue uses 92 spellings for about 34 real formats and the index is case sensitive, so the raw spellings matter.
$0.003
Data holder
holder-row
One titolare del dato, the body legally responsible for the data under Italian open data rules, with how many datasets it holds. There are 1,763 of them against 399 publishing organizations.
$0.002
Publisher
publisher-row
One DCAT publisher name with its dataset count. Usually a department inside a public body rather than the body itself.
$0.002
Source catalogue
source-catalog-row
One of the 328 regional, municipal and ministerial portals that dati.gov.it harvests, with how many datasets it contributes to the national catalogue.
$0.003
Distribution link checked
link-check
Optional. One distribution URL fetched from the publisher's own server to see whether it is actually alive: HTTP status, redirect target, content type, byte size and latency. dati.gov.it harvests metadata from 328 catalogues and never revalidates the links.
$0.004
Distribution bytes probed
file-probe
Optional. The first 64 KB of a distribution read with a range request, returning what the bytes really are, whether that matches the declared format, the CSV delimiter, the column headers and a sample row. Not charged when the file could not be read.
$0.012
Publisher profile attached
organization-profile
Optional. The publishing body's IPA code, contact email, telephone, website, Italian region and total dataset count added to a dataset row. Each organization is fetched once per run and reused across every dataset it published.
$0.008
Loading public Actor documentation
Reading the current public Store definition and reviews.
More from this builder
Tennis Abstract Scraper - Match History & Stats
781 active users · 13.6K runs / 30d
Reddit Scraper - Posts, Subreddit & Search Data API
145 active users · 3K runs / 30d
TennisExplorer Match Results Scraper
145 active users · 500 runs / 30d
TikTok Hashtag Scraper - View Counts & Top Posts
78 active users · 946 runs / 30d