Skip to the content

Knowledge base

The methods behind the data: how Qumasc brings sources to the same regions and periods, and what its records mean.

  • Introduction

    What data fusion for high-performance data analytics means, why harmonised regional data are hard to make, and what Qumasc does and does not do.

  • Recipes and templates

    How a recipe describes a dataset, how templates and the simple and advanced modes build on it, and how the recipe hash makes a dataset reproducible.

  • Regions and code lists

    NUTS, LAU and TERYT, why every region set has a version, how correspondence tables convert between versions, and why the Polish NUTS regions are built from PRG boundaries.

  • From grid to region

    How Qumasc turns gridded data into regional values with area-weighted zonal statistics, and how the valid fraction, small regions and cells at borders are handled.

  • Time

    How Qumasc names periods, forms monthly and daily values, treats accumulated variables, computes degree days and represents fixed values without a period.

  • Data sources

    The families of data Qumasc reads, the connectors that read them, and the known characteristics of each that users should keep in mind.

  • Quality and flags

    What the quality record of a dataset contains, why every empty value carries a flag, what each flag means, and how notes on the data differ from warnings.

  • Licences

    How Qumasc derives the licence of a dataset from all its inputs, writes the attribution, decides on public sharing, and what changes when users read sources with their own provider keys.

  • Provenance and reproducibility

    How Qumasc records where every dataset comes from, with STAC, RO-Crate and PROV, how datasets are regenerated from their recipes, and how users are referenced in stored files.

  • Validation

    How the datasets of the first runs with real data compare with independent references, what the numbers mean, and why ERA5-Land precipitation lies above the gauge totals in Poland.

  • Output formats

    What a Qumasc dataset package contains, from long Parquet tables and GeoParquet regions to Zarr cubes, CSV, the quality record, the STAC item, the RO-Crate, the LICENSE file and the previews.

  • Publications and references

    Citations of the data sets Qumasc reads and of the reference data used in its validation, the method literature cited in its design notes and reports, and Qumasc's own publications.

  • Glossary

    The terms used on this website and in the application, with short definitions.