Products
collekt ships a source-adapter for each supported provider. An adapter knows how to fetch source-native files for a request and how to plan (dry-run) the same work. Adapters are registered on import (collekt.sources) and selected by a source’s kind.
Built-in adapters
| Kind | Produces | Access | Format |
|---|---|---|---|
cmems |
Gridded ocean fields (currents, SST, waves, biogeochemistry) | open | NetCDF |
ecmwf_open_data |
Gridded forecast wind | open | NetCDF (cropped from global GRIB2) |
era5 |
Gridded reanalysis wind | credentials (CDS) | NetCDF |
copernicus_dataspace |
Sentinel / Landsat product files | credentials | product (e.g. .zip) |
skytruth |
Oil-slick detections (features) | open | Parquet |
hozint |
Threat-intelligence reports (features) | credentials | Parquet |
gfw |
Fishing / encounter / port-visit / loitering / AIS-gap events (features) | credentials | Parquet |
eodyn |
eOdyn surface currents / drifters | open (preview archive) | NetCDF |
local |
Row-filtered subset of a local damast-managed tabular archive (features) |
none — local filesystem | Parquet |
Open sources need no secrets; credential-gated sources resolve their secrets from the environment or provider tools — see Credentials. The eOdyn adapter serves a historical preview archive (mode: archive) until the production API ships (mode: api). The local adapter reads an archive that already exists on disk rather than fetching from a remote provider — see Local data collections.
Bundled catalog
collekt ships a curated catalog of concrete datasets — the datasets AI4COPSEC uses, grouped into catalogs under collekt/conf/source/:
| Catalog | Sources |
|---|---|
cmems_global |
cmems_glorys_nrt, cmems_glorys_nrt_2d_hourly, cmems_glorys_nrt_3d_6h, cmems_glorys_nrt_total_currents, cmems_glorys_my, cmems_duacs_nrt, cmems_duacs_my, cmems_global_sst_nrt, cmems_global_sst_my, cmems_global_waves, cmems_global_chl_obs_nrt, cmems_global_chl_obs_my, cmems_global_bgc_model, cmems_global_wind_nrt |
cmems_eur |
cmems_duacs_eur_nrt, cmems_duacs_eur_my |
cmems_med |
cmems_med_currents_nrt, cmems_med_currents_nrt_15min, cmems_med_currents_nrt_2d_hourly, cmems_med_currents_nrt_3d_hourly, cmems_med_currents_my, cmems_med_currents_my_2d_hourly, cmems_med_waves, cmems_med_chl_obs, cmems_med_bgc_model |
cmems_ibi |
cmems_ibi_currents, cmems_ibi_currents_2d_hourly, cmems_ibi_currents_3d_hourly, cmems_ibi_waves, cmems_atl_chl_obs_nrt, cmems_atl_chl_obs_my, cmems_ibi_bgc_model |
cmems_nws |
cmems_nws_currents, cmems_nws_currents_2d_hourly, cmems_nws_currents_3d_hourly, cmems_nws_waves, cmems_nws_chl_obs, cmems_nws_bgc_model |
ecmwf |
ecmwf_open_data_forecast, era5_reanalysis |
eodyn |
eodyn_osmose_currents |
skytruth |
skytruth |
hozint |
hozint |
gfw |
gfw |
dataspace |
copernicus_s1_grd, copernicus_s2_l2a, copernicus_landsat |
Browse the catalog programmatically instead of reading the tables below:
import collekt
print(collekt.describe_datasets()) # human-readable, grouped by provider
collekt.available_datasets()["cmems_glorys_my"]["variables"]
# ['bottomT', 'mlotst', 'siconc', 'sithick', 'so', 'thetao', 'uo', 'usi', 'vo', 'vsi', 'zos']Both accept an optional conf_dir= to include a downstream catalog overlay (see Extending the catalog downstream).
CMEMS Dataset Snapshot
The table below mirrors the bundled src/collekt/conf/source/cmems_*.yaml catalog files as of 2026-09-21. Temporal coverage is the offline coverage declared in the repository on that date; rolling/NRT products intentionally keep an open end and can be refreshed manually with scripts/update_cmems_coverage.py. Resolution is inferred from the CMEMS dataset identifier when the catalog does not declare it as a separate field. Both this table and the catalog listing above are generated — run scripts/update_products_table.py rather than editing them.
| Source key | Dataset ID | Resolution | Native sampling | Spatial coverage | Temporal coverage | Variables | Depth | DOI / product |
|---|---|---|---|---|---|---|---|---|
cmems_glorys_nrt |
cmems_mod_glo_phy-cur_anfc_0.083deg_P1D-m |
0.083 deg | 24h |
-180.0/179.9 -80.0/90.0 | from 2022-06-01 (rolling/open on 2026-09-21) |
uo, vo |
yes | 10.48670/moi-00016 product |
cmems_glorys_nrt_2d_hourly |
cmems_mod_glo_phy_anfc_0.083deg_PT1H-m |
0.083 deg | 1h |
-180.0/179.9 -80.0/90.0 | from 2022-06-01 (rolling/open on 2026-09-21) |
so, thetao, uo, vo, zos |
no | 10.48670/moi-00016 product |
cmems_glorys_nrt_3d_6h |
cmems_mod_glo_phy-cur_anfc_0.083deg_PT6H-i |
0.083 deg | 6h |
-180.0/179.9 -80.0/90.0 | from 2022-06-01 (rolling/open on 2026-09-21) |
uo, vo |
yes | 10.48670/moi-00016 product |
cmems_glorys_nrt_total_currents |
cmems_mod_glo_phy_anfc_merged-uv_PT1H-i |
not declared | 1h |
-180.0/179.9 -80.0/90.0 | from 2020-11-01 (rolling/open on 2026-09-21) |
uo, utide, utotal, vo, vsdx, vsdy, vtide, vtotal |
no | 10.48670/moi-00016 product |
cmems_glorys_my |
cmems_mod_glo_phy_my_0.083deg_P1D-m |
0.083 deg | 24h |
-180.0/179.9 -80.0/90.0 | 1993-01-01 to 2026-06-23 |
bottomT, mlotst, siconc, sithick, so, thetao, uo, usi, vo, vsi, zos |
yes | 10.48670/moi-00021 product |
cmems_duacs_nrt |
cmems_obs-sl_glo_phy-ssh_nrt_allsat-l4-duacs-0.125deg_P1D |
0.125 deg | 24h |
-179.9/179.9 -89.9/89.9 | from 2024-07-01 (rolling/open on 2026-09-21) |
adt, err_sla, err_ugosa, err_vgosa, flag_ice, sla, ugos, ugosa, vgos, vgosa |
no | 10.48670/moi-00149 product |
cmems_duacs_my |
cmems_obs-sl_glo_phy-ssh_my_allsat-l4-duacs-0.125deg_P1D |
0.125 deg | 24h |
-179.9/179.9 -89.9/89.9 | 1993-01-01 to 2026-01-16 |
adt, err_sla, err_ugosa, err_vgosa, flag_ice, sla, tpa_correction, ugos, ugosa, vgos, vgosa |
no | 10.48670/moi-00148 product |
cmems_global_sst_nrt |
METOFFICE-GLO-SST-L4-NRT-OBS-SST-V2 |
not declared | 24h |
-180.0/180.0 -90.0/90.0 | from 2024-01-17 (rolling/open on 2026-09-21) |
analysed_sst, analysis_error, mask, sea_ice_fraction |
no | 10.48670/moi-00165 product |
cmems_global_sst_my |
METOFFICE-GLO-SST-L4-REP-OBS-SST |
not declared | 24h |
-180.0/180.0 -90.0/90.0 | 1981-10-01 to 2026-03-31 |
analysed_sst, analysis_error, mask, sea_ice_fraction |
no | 10.48670/moi-00168 product |
cmems_global_waves |
cmems_mod_glo_wav_anfc_0.083deg_PT3H-i |
0.083 deg | 3h |
-180.0/179.9 -80.0/90.0 | from 2022-11-01 (rolling/open on 2026-09-21) |
VCMX, VHM0, VHM0_SW1, VHM0_SW2, VHM0_WW, VMDR, VMDR_SW1, VMDR_SW2, VMDR_WW, VMXL, VPED, VSDX, VSDY, VTM01_SW1, VTM01_SW2, VTM01_WW, VTM02, VTM10, VTPK |
no | 10.48670/moi-00017 product |
cmems_global_chl_obs_nrt |
cmems_obs-oc_glo_bgc-plankton_nrt_l4-gapfree-multi-4km_P1D |
4 km | 24h |
-180.0/180.0 -90.0/90.0 | from 2026-09-05 (rolling/open on 2026-09-21) |
CHL, CHL_uncertainty, flags |
no | 10.48670/moi-00279 product |
cmems_global_chl_obs_my |
cmems_obs-oc_glo_bgc-plankton_my_l3-multi-4km_P1D |
4 km | 24h |
-180.0/180.0 -90.0/90.0 | 1997-09-04 to 2026-09-13 |
CHL, CHL_uncertainty, DIATO, DIATO_uncertainty, DINO, DINO_uncertainty, GREEN, GREEN_uncertainty, HAPTO, HAPTO_uncertainty, MICRO, MICRO_uncertainty, NANO, NANO_uncertainty, PICO, PICO_uncertainty, PROCHLO, PROCHLO_uncertainty, PROKAR, PROKAR_uncertainty, flags |
no | 10.48670/moi-00280 product |
cmems_global_bgc_model |
cmems_mod_glo_bgc-pft_anfc_0.25deg_P1D-m |
0.25 deg | 24h |
-180.0/179.8 -80.0/90.0 | from 2021-11-01 (rolling/open on 2026-09-21) |
chl, phyc |
yes | 10.48670/moi-00015 product |
cmems_global_wind_nrt |
cmems_obs-wind_glo_phy_nrt_l4_0.125deg_PT1H |
0.125 deg | 1h |
-179.9/179.9 -89.9/89.9 | from 2024-06-13 (rolling/open on 2026-09-21) |
air_density, eastward_stress, eastward_stress_bias, eastward_stress_sdd, eastward_wind, eastward_wind_bias, eastward_wind_sdd, northward_stress, northward_stress_bias, northward_stress_sdd, northward_wind, northward_wind_bias, northward_wind_sdd, number_of_observations, number_of_observations_divcurl, stress_curl, stress_curl_bias, stress_curl_dv, stress_divergence, stress_divergence_bias, stress_divergence_dv, wind_curl, wind_curl_bias, wind_curl_dv, wind_divergence, wind_divergence_bias, wind_divergence_dv |
no | 10.48670/moi-00305 product |
| Source key | Dataset ID | Resolution | Native sampling | Spatial coverage | Temporal coverage | Variables | Depth | DOI / product |
|---|---|---|---|---|---|---|---|---|
cmems_duacs_eur_nrt |
cmems_obs-sl_eur_phy-ssh_nrt_allsat-l4-duacs-0.0625deg_P1D |
0.0625 deg | 24h |
-30.0/42.0 20.0/66.0 | from 2024-07-01 (rolling/open on 2026-09-21) |
adt, err_sla, err_ugosa, err_vgosa, flag_ice, sla, ugos, ugosa, vgos, vgosa |
no | 10.48670/moi-00142 product |
cmems_duacs_eur_my |
cmems_obs-sl_eur_phy-ssh_my_allsat-l4-duacs-0.0625deg_P1D |
0.0625 deg | 24h |
-30.0/42.0 20.0/66.0 | 1993-01-01 to 2026-01-16 |
adt, err_sla, err_ugosa, err_vgosa, flag_ice, sla, tpa_correction, ugos, ugosa, vgos, vgosa |
no | 10.48670/moi-00141 product |
| Source key | Dataset ID | Resolution | Native sampling | Spatial coverage | Temporal coverage | Variables | Depth | DOI / product |
|---|---|---|---|---|---|---|---|---|
cmems_ibi_currents |
cmems_mod_ibi_phy_anfc_0.027deg-3D_P1D-m |
0.027 deg | 24h |
-19.1/5.1 26.2/56.1 | from 2022-11-23 (rolling/open on 2026-09-21) |
bottomT, mlotst, so, thetao, uo, vo, zos |
yes | 10.48670/moi-00027 product |
cmems_ibi_currents_2d_hourly |
cmems_mod_ibi_phy_anfc_0.027deg-2D_PT1H-m |
0.027 deg | 1h |
-19.1/5.1 26.2/56.1 | from 2022-11-23 (rolling/open on 2026-09-21) |
mlotst, thetao, ubar, uo, vbar, vo, zos |
no | 10.48670/moi-00027 product |
cmems_ibi_currents_3d_hourly |
cmems_mod_ibi_phy_anfc_0.027deg-3D_PT1H-m |
0.027 deg | 1h |
-19.1/5.1 26.2/56.1 | from 2024-11-10 (rolling/open on 2026-09-21) |
so, thetao, uo, vo |
yes | 10.48670/moi-00027 product |
cmems_ibi_waves |
cmems_mod_ibi_wav_anfc_0.027deg_PT1H-i |
0.027 deg | 1h |
-19.0/5.0 26.0/56.0 | from 2022-11-26 (rolling/open on 2026-09-21) |
VCMX, VHM0, VHM0_SW1, VHM0_SW2, VHM0_WW, VMDR, VMDR_SW1, VMDR_SW2, VMDR_WW, VMXL, VPED, VSDX, VSDY, VTM01_SW1, VTM01_SW2, VTM01_WW, VTM02, VTM10, VTPK |
no | 10.48670/moi-00025 product |
cmems_atl_chl_obs_nrt |
cmems_obs-oc_atl_bgc-plankton_nrt_l3-multi-1km_P1D |
1 km | 24h |
-46.0/13.0 20.0/66.0 | from 2026-09-04 (rolling/open on 2026-09-21) |
CHL, CHL_uncertainty, DIATO, DIATO_uncertainty, DINO, DINO_uncertainty, GREEN, GREEN_uncertainty, HAPTO, HAPTO_uncertainty, MICRO, MICRO_uncertainty, NANO, NANO_uncertainty, PICO, PICO_uncertainty, PROCHLO, PROCHLO_uncertainty, PROKAR, PROKAR_uncertainty, flags |
no | 10.48670/moi-00284 product |
cmems_atl_chl_obs_my |
cmems_obs-oc_atl_bgc-plankton_my_l3-multi-1km_P1D |
1 km | 24h |
-46.0/13.0 20.0/66.0 | 1997-09-04 to 2026-09-13 |
CHL, CHL_uncertainty, DIATO, DIATO_uncertainty, DINO, DINO_uncertainty, GREEN, GREEN_uncertainty, HAPTO, HAPTO_uncertainty, MICRO, MICRO_uncertainty, NANO, NANO_uncertainty, PICO, PICO_uncertainty, PROCHLO, PROCHLO_uncertainty, PROKAR, PROKAR_uncertainty, flags |
no | 10.48670/moi-00286 product |
cmems_ibi_bgc_model |
cmems_mod_ibi_bgc_anfc_0.027deg-3D_P1D-m |
0.027 deg | 24h |
-19.1/5.1 26.2/56.1 | from 2022-11-23 (rolling/open on 2026-09-21) |
chl, dissic, fe, nh4, no3, nppv, o2, ph, phyc, po4, si, spco2, zeu, zooc |
yes | 10.48670/moi-00026 product |
| Source key | Dataset ID | Resolution | Native sampling | Spatial coverage | Temporal coverage | Variables | Depth | DOI / product |
|---|---|---|---|---|---|---|---|---|
cmems_med_currents_nrt |
cmems_mod_med_phy-cur_anfc_4.2km_P1D-m |
4.2 km | 24h |
-17.3/36.3 30.2/46.0 | from 2024-08-27 (rolling/open on 2026-09-21) |
uo, vo |
yes | 10.48670/mds-00359 product |
cmems_med_currents_nrt_15min |
cmems_mod_med_phy-cur_anfc_4.2km_PT15M-i |
4.2 km | 15min |
-17.3/36.3 30.2/46.0 | from 2024-08-23 (rolling/open on 2026-09-21) |
uo, vo |
yes | 10.48670/mds-00359 product |
cmems_med_currents_nrt_2d_hourly |
cmems_mod_med_phy-cur_anfc_4.2km-2D_PT1H-m |
4.2 km | 1h |
-17.3/36.3 30.2/46.0 | from 2024-08-25 (rolling/open on 2026-09-21) |
uo, vo |
no | 10.48670/mds-00359 product |
cmems_med_currents_nrt_3d_hourly |
cmems_mod_med_phy-cur_anfc_4.2km-3D_PT1H-m |
4.2 km | 1h |
-17.3/36.3 30.2/46.0 | from 2025-10-06 (rolling/open on 2026-09-21) |
uo, vo |
yes | 10.48670/mds-00359 product |
cmems_med_currents_my |
cmems_mod_med_phy-cur_my_4.2km_P1D-m |
4.2 km | 24h |
-6.0/36.3 30.2/46.0 | 1987-01-01 to 2026-08-31 |
uo, vo |
yes | 10.48670/mds-00375 product |
cmems_med_currents_my_2d_hourly |
cmems_mod_med_phy-cur_my_4.2km_PT1H-m |
4.2 km | 1h |
-6.0/36.3 30.2/46.0 | 1987-01-01 to 2026-08-31 |
uo, vo |
no | 10.48670/mds-00375 product |
cmems_med_waves |
cmems_mod_med_wav_anfc_4.2km_PT1H-i |
4.2 km | 1h |
-18.1/36.3 30.2/46.0 | from 2021-11-30 (rolling/open on 2026-09-21) |
VCMX, VHM0, VHM0_SW1, VHM0_SW2, VHM0_WW, VMDR, VMDR_SW1, VMDR_SW2, VMDR_WW, VMXL, VPED, VSDX, VSDY, VTM01_SW1, VTM01_SW2, VTM01_WW, VTM02, VTM10, VTPK |
no | 10.48670/mds-00373 product |
cmems_med_chl_obs |
cmems_obs-oc_med_bgc-plankton_nrt_l3-multi-1km_P1D |
1 km | 24h |
-6.0/36.5 30.0/46.0 | from 2026-09-13 (rolling/open on 2026-09-21) |
CHL, CRYPTO, DIATO, DINO, GREEN, HAPTO, MICRO, NANO, PICO, PROKAR, QI_CHL, SENSORMASK, WTM |
no | 10.48670/moi-00297 product |
cmems_med_bgc_model |
cmems_mod_med_bgc-bio_anfc_4.2km_P1D-m |
4.2 km | 24h |
-5.5/36.3 30.2/46.0 | from 2024-07-23 (rolling/open on 2026-09-21) |
nppv, o2 |
yes | 10.48670/mds-00358 product |
| Source key | Dataset ID | Resolution | Native sampling | Spatial coverage | Temporal coverage | Variables | Depth | DOI / product |
|---|---|---|---|---|---|---|---|---|
cmems_nws_currents |
cmems_mod_nws_phy-cur_anfc_1.5km-3D_P1D-m |
1.5 km | 24h |
-16.0/13.0 46.0/62.7 | from 2024-08-02 (rolling/open on 2026-09-21) |
ubar, uo, vbar, vo, wo |
yes | 10.48670/moi-00054 product |
cmems_nws_currents_2d_hourly |
cmems_mod_nws_phy-cur_anfc_1.5km-2D_PT1H-i |
1.5 km | 1h |
-16.0/13.0 46.0/62.7 | from 2024-08-04 (rolling/open on 2026-09-21) |
uo, vo |
no | 10.48670/moi-00054 product |
cmems_nws_currents_3d_hourly |
cmems_mod_nws_phy-cur_anfc_1.5km-3D_PT1H-i |
1.5 km | 1h |
-16.0/13.0 46.0/62.7 | from 2024-09-17 (rolling/open on 2026-09-21) |
ubar, uo, vbar, vo, wo |
yes | 10.48670/moi-00054 product |
cmems_nws_waves |
cmems_mod_nws_wav_anfc_1.5km_PT1H-i |
1.5 km | 1h |
-16.0/13.0 46.0/62.7 | from 2024-08-06 (rolling/open on 2026-09-21) |
VCMX, VHM0, VHM0_SW1, VHM0_SW2, VHM0_WW, VMDR, VMDR_SW1, VMDR_SW2, VMDR_WW, VMXL, VPED, VSDX, VSDY, VTM01_SW1, VTM01_SW2, VTM01_WW, VTM02, VTM10, VTPK, forecast_period |
no | 10.48670/moi-00055 product |
cmems_nws_chl_obs |
cmems_obs_oc_nws_bgc_tur-spm-chl_nrt_l3-hr-mosaic_P1D-m |
not declared | 24h |
-12.0/13.0 48.0/62.0 | from 2020-01-01 (rolling/open on 2026-09-21) |
CHL, SPM, TUR |
no | 10.48670/moi-00118 product |
cmems_nws_bgc_model |
cmems_mod_nws_bgc-chl_anfc_7km-3D_P1D-m |
7 km | 24h |
-19.9/13.0 40.1/65.0 | from 2024-07-29 (rolling/open on 2026-09-21) |
chl |
yes | 10.48670/moi-00056 product |
Select concrete datasets with DatasetConfig, using provider-specific dataclasses whose fields match the YAML preset form:
config = collekt.DatasetConfig(
collekt.CMEMS("cmems_glorys_my", variables=["uo", "vo"], depth=[1.0, 1.1]),
collekt.CMEMS("cmems_duacs_my", variables=["ugos", "vgos"]),
collekt.SkyTruth(),
)The equivalent YAML is:
datasets:
- provider: cmems
key: cmems_glorys_my
variables: [uo, vo]
depth: [1.0, 1.1]
- provider: cmems
key: cmems_duacs_my
variables: [ugos, vgos]
- provider: skytruth
key: skytruthDatasetConfig resolves each key against the bundled catalog, validates selected variables against available_variables, and validates required provider parameters such as CMEMS depth ranges. Fetcher takes the output directory separately; it is runtime state, not dataset configuration.
The catalog is curated, not exhaustive — dataset IDs drift, so collekt doctor --online validates the CMEMS entries against the provider catalogue.
Each gridded source declares one temporal field: temporal_sampling, the native cadence of that dataset (15min, 1h, 3h, 6h, or 24h). Datasets are always fetched at that cadence — a request carries no sampling of its own, so thinning or aligning to a coarser common axis is a downstream Assembler step. When a provider offers multiple cadences or dimensionalities, model them as separate source entries (for example cmems_med_currents_nrt_15min, cmems_med_currents_nrt_2d_hourly, and cmems_med_currents_nrt) so each source maps to one provider dataset and one cadence, and the caller picks the cadence by picking the dataset.
Every model-currents product in the catalog now carries its subdaily variants alongside the daily one:
| Region | Daily | Subdaily |
|---|---|---|
| Global NRT | cmems_glorys_nrt |
cmems_glorys_nrt_2d_hourly, cmems_glorys_nrt_3d_6h, cmems_glorys_nrt_total_currents |
| Mediterranean NRT | cmems_med_currents_nrt |
cmems_med_currents_nrt_15min, cmems_med_currents_nrt_2d_hourly, cmems_med_currents_nrt_3d_hourly |
| Mediterranean reanalysis | cmems_med_currents_my |
cmems_med_currents_my_2d_hourly |
| IBI | cmems_ibi_currents |
cmems_ibi_currents_2d_hourly, cmems_ibi_currents_3d_hourly |
| North-West Shelf | cmems_nws_currents |
cmems_nws_currents_2d_hourly, cmems_nws_currents_3d_hourly |
The daily mean is the unsuffixed base entry, and every subdaily variant of it carries both its dimensionality and its cadence (_2d_hourly, _3d_6h). Dimensionality is in the key because it decides whether a caller must pass a depth: CMEMS ships its -2D datasets as surface fields, so cmems_med_currents_my needs a depth while cmems_med_currents_my_2d_hourly does not. A dataset holding a different quantity rather than another cadence of the same field is named for the quantity instead (cmems_glorys_nrt_total_currents). cmems_med_currents_nrt_15min predates this rule and keeps its released name.
Switching between cadences of the same product is then a one-key change: the sources share a path, resolve the same region on the same grid, and each output file is prefixed with its source name, so a daily and an hourly variant can also be collected side by side. Only the depth dimension differs, and the assembled dataset gains or loses its depth axis accordingly.
Regional Atlantic products use the CMEMS regional names: IBI (cmems_ibi) and European North-West Shelf (cmems_nws) are separate catalogs. Basin-level Atlantic ocean-colour observations keep cmems_atl in their source names.
Sources may also declare a coverage block:
coverage:
longitude: [-19.08, 5.08]
latitude: [26.17, 56.08]
temporal:
start: "2022-11-23"
end: null
kind: rollingThis is an offline planning hint. Static spatial bounds and stable temporal starts are shipped in the catalog; rolling/NRT products keep end: null because their latest date moves. When online catalogue access is available, CMEMS planning still prefers Copernicus Marine describe() and collekt doctor --online compares declared coverage with the provider catalogue.
Presets
Presets are plain YAML versions of DatasetConfig:
collekt fetch --bbox -6 20 35 45 --start 2023-06-15 \
--dataset-config westmed.yaml --output-dir data/collectionsThey name only desired datasets and provider parameters. The full source catalog remains package-maintained metadata under collekt/conf/source/.
Extending the catalog downstream
The bundled catalog is not exhaustive. A downstream project — for example a technological brick with its own ocean datasets — adds or overrides datasets with a configuration directory, passed as --conf-dir (CLI) or conf_dir= (Fetcher / DatasetConfig.resolve). It overlays the bundled configuration rather than replacing it: its default.yaml declares only additions, its source_catalogs are appended to the shipped ones, and its source/*.yaml files define new datasets or override bundled ones by key.
my_brick/conf/
default.yaml # source_catalogs: [ocean] (bundled catalogs stay available)
source/ocean.yaml # defines cmems / era5 / ... datasets by key
collekt fetch --bbox -6 20 35 45 --start 2023-06-15 \
--conf-dir my_brick/conf --dataset-config presets/currents.yaml \
--output-dir data/collections--conf-dir is the catalog (what datasets exist); --dataset-config is the selection (which of them to fetch, and with which variables and depth). A selected key must exist in the merged catalog, so a brick’s presets can reference both bundled and its own datasets. Adding a new provider (a new adapter kind) is a larger extension and is not configured this way.
Local data collections
The local adapter serves a tabular archive that already exists on disk as a collekt source, without fetching anything remotely. Each requested day’s rows are filtered to the request’s region and written out as Parquet; a day with no matching file, or no in-region rows, is skipped rather than treated as an error. Two ways to locate a day’s rows are supported, depending on whether the archive is already day-partitioned.
Because archive_root is a machine-local path, local sources are typically declared in a downstream conf_dir rather than the bundled catalog (see Extending the catalog downstream).
Day-partitioned archives (layout)
If the archive is already split into one file (or file set) per day — e.g. built with damast convert --save-as or AnnotatedDataFrame.export_partitioned — point layout at its partitioning spec; a day’s file(s) are then resolved from the filename alone, with nothing read until a match is found:
# my_brick/conf/source/ais.yaml
sources:
ais_archive:
kind: local
path: "local/ais"
dataset_id: "ais-daily-archive"
archive_root: "/data/ais/daily"
layout: "time:timestamp+daily:%Y-%m-%d"
region_columns: [lat, lon]
temporal_sampling: 24hpath— like any source, where collekt writes this source’s output, relative to the collection’s output directory (request_dir / path). It has nothing to do with the archive being read.archive_root— root of the pre-existing on-disk archive to read from. Unrelated topath:pathis collekt’s own output layout,archive_rootis the source’s input.layout— the archive’s partitioning spec, as produced bydamast.core.SaveAs/AnnotatedDataFrame.export_partitioned(atime:ortime+column:spec;localreads the time column out of it to scope each request day).region_columns—[lat_column, lon_column]used to filter rows to the request’s bounding box.
Rows are filtered by time as well as region in both modes, so layout only decides which files are opened — a partitioning coarser than the request (one file per month, say) still yields just that day’s rows for each day. If the archive’s files carry a compression suffix (AIS_2026-01-01.zst.parquet), they are matched despite expected_paths naming only AIS_2026-01-01.parquet. A region_columns or time_column entry the archive doesn’t actually have raises rather than skipping every day, which would otherwise be indistinguishable from an archive holding no data.
Unpartitioned archives (file_pattern + time_column)
If the archive is a flat set of files with no time-encoded naming, use file_pattern (a glob, resolved once per fetch rather than once per day) and time_column instead of layout; a day’s rows are then selected by filtering time_column, the same way region_columns filters by region:
sources:
ais_flat:
kind: local
path: "local/ais"
dataset_id: "ais-flat-archive"
archive_root: "/data/ais/flat"
file_pattern: "*.parquet"
time_column: timestamp
region_columns: [lat, lon]
temporal_sampling: 24hlayout and file_pattern are mutually exclusive; exactly one is required. Every matching file is scanned for every request day, so this mode leans on Parquet row-group statistics on time_column to skip irrelevant data — an unsorted archive, or one made of very few, very large files, is scanned in full for each day. Prefer layout when the archive can be (re-)partitioned (e.g. damast convert --save-as "time:timestamp+daily:%Y-%m-%d"); reach for file_pattern when it can’t.
That pushdown is Parquet-only: damast reads Parquet through polars.scan_parquet, but its NetCDF, CSV and HDF readers load each file in full, and file_pattern re-reads every matched file once per request day. For a non-Parquet archive of any size — especially NetCDF holding a dense, padded grid rather than a row per observation — convert it once (damast convert --save-as "time:<time_column>+daily:%Y-%m-%d") and point layout at the result.
Select either the same way:
config = collekt.DatasetConfig(
collekt.Local("ais_archive", variables=["mmsi", "timestamp", "lat", "lon"]),
)or as a preset:
datasets:
- provider: local
key: ais_archive
variables: [mmsi, timestamp, lat, lon]collekt fetch --bbox -6 20 35 45 --start 2026-01-01 \
--conf-dir my_brick/conf --dataset-config presets/ais.yaml \
--output-dir data/collectionsDependencies
collekt installs all provider clients and the gridded Assembler by default, so every source works out of the box after uv sync. Clients are imported lazily, so importing collekt stays fast and a source only pulls its client into memory when it actually runs. The ECMWF Open Data forecast subset, Skytruth, and eOdyn (in its current preview-archive mode) need no credentials; the other providers do — see Credentials.