credit.datasets.gen_2.gfs#

GFS and GDAS data loading for CREDIT Gen2.

This module provides GFSDataset, a PyTorch dataset for the public GFS/GDAS NetCDF files in Google Cloud Storage. A model run is represented by two files: an atmospheric file containing model-level fields and a surface file containing two-dimensional surface fields. GFSDataset discovers available runs, reads only the requested variables and levels, and returns CREDIT’s standard input/target sample dictionaries.

The dataset deliberately does not derive pressure, geopotential, or other diagnostics and does not regrid the native Gaussian grid. Those operations are handled by CREDIT postblocks or preblocks so the raw GFS state remains available to downstream processing.

Remote reads use anonymous obstore range requests, allowing xarray to read selected portions of the large NetCDF objects without downloading each file in full. Local mode reads files downloaded with credit.datasets.gen_2.gfs_download. Both modes use the same native layout:

{system}.YYYYMMDD/HH/atmos/{system}.tHHz.atmanl.nc
{system}.YYYYMMDD/HH/atmos/{system}.tHHz.sfcanl.nc

Here system is gdas by default and may be changed to gfs in the source configuration. Forecast files can be selected with forecast_hour.

Attributes#

Classes#

GFSDataset

Read GFS or GDAS atmospheric and surface NetCDF output.

Module Contents#

credit.datasets.gen_2.gfs.logger#
credit.datasets.gen_2.gfs.VALID_SYSTEMS#
credit.datasets.gen_2.gfs.VALID_LEVEL_TYPES#
class credit.datasets.gen_2.gfs.GFSDataset(data_config: dict[str, Any], return_target: bool = False, gfs_type: VALID_SYSTEMS | None = None)#

Bases: credit.datasets.gen_2.base_dataset.BaseDataset

Read GFS or GDAS atmospheric and surface NetCDF output.

GFSDataset reads the paired atmospheric and surface files published in the global-forecast-system Google Cloud bucket. Atmospheric files contain three-dimensional model-level fields and a small number of two-dimensional fields such as pressfc. Surface files contain the two-dimensional surface fields such as tmp2m and land. The dataset returns the same flat, slash-delimited tensor keys as the other Gen2 datasets; derivations, regridding, and vertical interpolation are left to later preblocks or postblocks.

Remote mode uses obstore for anonymous, range-based GCS reads. The remote NetCDF files are opened with xarray and h5netcdf because the obstore reader is a seekable file-like object. Local mode uses the netcdf4 xarray engine by default and expects files laid out like the public bucket. Availability is checked during initialization by default, so missing model runs are removed from datetimes rather than failing later during sampling.

Input settings:
dataset_type (str): Must be "gfs" when routed through

MultiSourceDataset.

system (str): Forecast system, either "gdas" or "gfs".

Defaults to "gdas". The aliases model and gfs_type are also accepted in source configuration.

mode (str): "remote" to read from GCS or "local" to read

downloaded files. Defaults to "remote" for this dataset.

base_path (str): Root directory for local files. Required when

mode is "local". The downloader creates the same {system}.YYYYMMDD/HH/atmos/ layout used by the bucket.

forecast_hour (int | None): Forecast lead hour. None selects the

analysis files atmanl.nc and sfcanl.nc. An integer selects files such as atmf003.nc and sfcf003.nc.

level_type (str): "model" selects one-based positions in the

pfull dimension. "pressure" selects the nearest values of the file’s pfull coordinate, in hPa. Defaults to "model".

level_coord (str): Atmospheric vertical dimension. Defaults to

"pfull".

levels (list[int | float] | None): Requested model-level indices or

pressure values, depending on level_type. None reads all atmospheric levels.

check_availability (bool): Check that the required atmospheric and

surface objects exist before adding a timestamp to datetimes. Defaults to True.

variables (dict): Field definitions grouped under prognostic,

dynamic_forcing, static, and diagnostic. Each field may contain vars_3D and/or vars_2D using native GFS names.

return_target (bool): Constructor argument controlling whether the

sample includes the next timestep under target.

Variables:
  • dataset_type (str) – The registered dataset type, "gfs".

  • system (str) – Active forecast system, "gdas" or "gfs".

  • mode (str) – Active storage mode, "remote" or "local".

  • base_path (str | None) – Expanded local storage root, if configured.

  • forecast_hour (int | None) – Active analysis or forecast lead hour.

  • level_type (str) – Active vertical-level interpretation.

  • level_coord (str) – Name of the atmospheric vertical dimension.

  • levels (list[int | float] | None) – Configured level selection.

  • datetimes (pandas.DatetimeIndex) – Available sampling timestamps after applying the configured clock and availability checks.

  • file_dict (dict) – Registered field types and their remote/local file source marker.

  • var_dict (dict) – Registered native GFS variables grouped by field type.

  • static_metadata (dict) – Calendar, grid, system, level, and forecast metadata exposed to MultiSourceDataset and downstream setup.

Example YAML configuration:

data:
  source:
    GDAS:
      dataset_type: "gfs"
      system: "gdas"
      mode: "remote"
      level_type: "model"
      levels: [1, 10, 30, 60, 90, 127]  # selected from the 127 available model levels (1-127)
      check_availability: true
      variables:
        prognostic:
          vars_3D: [tmp, ugrd, vgrd, spfh]
          vars_2D: [pressfc, tmp2m]
        dynamic_forcing: null
        static:
          vars_2D: [land, orog]
        diagnostic: null
  start_datetime: "2024-01-01T00:00:00"
  end_datetime: "2024-01-31T18:00:00"
  timestep: "6h"
  forecast_len: 1

For pressure-coordinate selection, change the source settings to level_type: "pressure" and provide values such as levels: [50, 100, 500, 850, 1000]. To read downloaded files, use mode: "local" and set base_path to the downloader’s output root.

Command-line usage:

# Download the configured files for local mode.
python -m credit.datasets.gen_2.gfs_download -c config/gfs.yml

# Instantiate GFSDataset directly from a YAML file.
python - <<'PY'
import yaml
from credit.datasets.gen_2.gfs import GFSDataset

with open("config/gfs.yml") as file:
    config = yaml.safe_load(file)
dataset = GFSDataset(config["data"], return_target=True)
print(len(dataset))
PY

# In normal training, MultiSourceDataset instantiates GFSDataset from
# the same data block.
credit_train_gen2 -c config/gfs.yml
system: VALID_SYSTEMS = ''#
mode#
forecast_hour: int | None#
base_path#
level_type: VALID_LEVEL_TYPES = ''#
levels: list[int | float] | None#
level_coord#
check_availability#
dataset_type = 'gfs'#
static_metadata#