Example datasets and public sharing¶
The public website shows what Qumasc produces without an account. Users’ datasets are private; the public pages (Explore, the data sources and template pages) show only example datasets published by the Qumasc team. This page is the runbook for preparing them and for the review of older public shares.
Rules¶
Who publishes. Administrators and data stewards (permission
datasets.examples) mark a dataset as an example withPUT /api/v1/admin/datasets/{dataset_id}/exampleand withdraw it withDELETEon the same address. A data steward marks own datasets, an administrator any dataset of a staff member. The owner must have a role ofsharing.public_roles: examples are datasets of the team, never of a regular user. Both actions are audited (datasets.example,datasets.example.remove; the values name no user).What qualifies. The same check as a public share: the licence of the dataset, every source that delivered its data and
sharing.public_enabledmust allow public sharing. Otherwise the answer is409 licence_forbids_sharingor409 sharing_restrictedwith the reason (for example data from Earth Data Hub, which may be shared between Qumasc users only, or GBIF records under CC BY-NC). When the settings later forbid public sharing of an example, the re-evaluation that follows the change (and the clean-up of the workers) withdraws it.What happens. The dataset is pinned (no expiry; the pin cannot be removed while it is an example), is listed by
GET /api/v1/public/examplesand is readable and previewable by anyone, also without an account, without becoming a public share. Visitors and other users see the publisher ofexamples.publisher(“Qumasc team”) instead of the owner; the owner’s identifier is replaced in every record, STAC item, provenance record and recipe shown to them. The owner and administrators see the dataset as before.Downloads of examples need sign-in (
401otherwise) unlessexamples.anonymous_downloadsis true. Signed-in users download them like shared datasets (counted in their download quota). Downloaded files of an example keep the permanent identifier of the account in the archive’s provenance files (never a name); single JSON and text assets are served with it replaced.Make your own like this. A dataset built from a template remembers the template parameters (
template_parameters, stored when a job is submitted with them and checked against the recipe).GET /api/v1/public/examples/{dataset_id}/parametersreturns them with the address of the pre-filled form; the script below stores them for every example it runs.
Setting |
Default |
Meaning |
|---|---|---|
|
|
Roles that may share publicly and whose datasets may become examples. |
|
|
Visitors without an account may download examples. |
|
|
Name shown instead of the owner of an example. |
|
|
|
|
|
Requests of the public catalogue per address or user. |
|
packaged texts |
Provider, coverage, resolution, update frequency, request limits and dataset identifiers of a source for the public catalogue. |
Preparing the examples¶
qumasc admin examples prepare <file> --user <staff user> (the same as
scripts/prepare_examples.py <file> --user <staff user>) reads a YAML list of examples, builds each
recipe from its template and parameters exactly as the simple mode builds it for that user, submits
the jobs as that user, waits for the workers of the service to finish them (--timeout, default
six hours) and marks every dataset made as an example. It prints one line per example: template,
status (example, refused, failed; planned with --dry-run), job, dataset and message.
Do not point it at a system by accident: it submits real jobs that contact the providers.
Each entry of the file has template, title, optionally description, parameters (the fields
of the form; missing ones get their defaults), avoid_connectors (sources of these connectors are
replaced by their fallback source of the recipe, e.g. Earth Data Hub by the CDS) and
source_parameters (parameters merged into a source of the recipe, e.g. a licence filter of GBIF
records). avoid_connectors and source_parameters are not part of the stored template
parameters: “make your own like this” offers the form as the simple mode builds it.
Staff account. Use an account of a role in
sharing.public_roles(administrator or data steward) with two-factor authentication. The examples belong to it and use its quotas.Provider keys of that account (Account > Provider keys,
PUT /api/v1/me/credentials/...): the CDS personal access token, with the licences of “ERA5-Land monthly averaged data” and “ERA5-Land hourly data” accepted on the CDS website; an S3 key pair of the Copernicus Data Space Ecosystem (NDVI); optionally a GBIF account (downloads with a DOI) and a BDL client identifier.Dry run:
qumasc admin examples prepare docs/admin/examples.yaml --user <name> --dry-runbuilds and validates every recipe and submits nothing.Run: the same without
--dry-run, on a host with the service’s database (QUMASC_DATABASE_URL) while its workers run. A refused or failed example is reported and the others go on; correct the entry and run a file with that entry alone.Combined example (second step).
regions-combinedreads earlier datasets of the same regions. After step 4, take the dataset identifiers of the two Lesser Poland examples of ERA5-Land (2023) and of the regional statistics, and run a second file:examples: - template: regions-combined title: Climate and regional statistics of Lesser Poland in one table description: >- Monthly ERA5-Land climate of 2023 joined with population, GDP and unemployment per NUTS 2024 region of Lesser Poland. An example dataset of the Qumasc team. parameters: area: {code_list: NUTS, version: "2024", levels: [2, 3], codes: [PL21]} period: {start: "2023-01", end: "2023-12"} climate: <dataset id of "Monthly temperature and precipitation 2023, Lesser Poland"> ndvi: <dataset id of "NDVI in summer 2023, Lesser Poland"> statistics: <dataset id of "Population, GDP and unemployment 2020 to 2023, Lesser Poland">
Check the website (
/explore) orGET /api/v1/public/examples: every card has a preview and a licence; withdraw an example withDELETE /api/v1/admin/datasets/{dataset_id}/example.
Proposed first set¶
The file docs/admin/examples.yaml holds the first step. All runs are small (one region of
Poland or one month for the country, one year of GBIF records of one species) and use only
sources whose data may be published: ERA5-Land from the CDS (CC BY 4.0), the NDVI of the Copernicus
Land Monitoring Service, Eurostat and GUS (CC BY 4.0), OpenStreetMap through Geofabrik (ODbL),
GBIF records under CC0 or CC BY, and the PRG boundaries. Earth Data Hub is never used.
Template |
Example |
Area and period |
Sources |
Note |
|---|---|---|---|---|
|
Mean July temperature 2024 per county |
Poland, TERYT 2025-01-01 powiaty, July 2024 |
CDS ERA5-Land monthly |
The map of the home page. |
|
Monthly temperature and precipitation 2023 |
Lesser Poland, NUTS 2024 PL21 (levels 2, 3), 2023 |
CDS ERA5-Land monthly |
Input of the combined example; carries the note on ERA5-Land precipitation. |
|
Heating degree days, winter 2024 |
PL21, January to March 2024 |
CDS ERA5-Land hourly ( |
The template reads Earth Data Hub first; the CDS fallback is used instead. |
|
NDVI in summer 2023 |
PL21, June to August 2023 |
CDSE, CLMS NDVI 300 m |
Needs the S3 key pair. |
|
Population, GDP and unemployment 2020 to 2023 |
PL21 |
Eurostat, GUS BDL |
Without a climate table. |
|
Roads, railways and amenities |
Opole Voivodeship, NUTS 2024 PL52 |
Geofabrik (OpenStreetMap) |
The smallest voivodeship extract. |
|
Roman snail records in Czechia 2023 |
Czechia, 2023 |
GBIF (CC0, CC BY only), CDS ERA5-Land monthly |
Not about Poland; the licence filter keeps the dataset publishable. |
|
Climate and regional statistics of Lesser Poland |
PL21, 2023 |
earlier examples |
Second step (identifiers known only after the first). |
|
none yet |
Reads hourly ERA5-Land from Earth Data Hub only: no publishable example until the template gets a CDS route. |
Templates that are Poland-only today. The forms of the region templates
(era5land-to-regions, era5land-daily-to-regions, degree-days-to-regions, ndvi-to-regions,
osm-to-regions, regional-statistics, regions-combined) accept Polish NUTS codes and TERYT
units only: the region boundaries of the simple mode come from the PRG of GUGiK (open licence),
while the GISCO boundaries of other countries carry the EuroGeographics non-commercial licence,
which would forbid public sharing; TERYT, the GUS Local Data Bank and the Geofabrik extracts of the
form are Polish as well. Full recipes in the advanced editor can use other countries. The point
template gbif-climate-points takes a country of simple_mode.countries (Poland, Czechia,
Slovakia, Germany, Lithuania) or any bounding box.