credit.datasets.gen_2.gefs#

GEFS ensemble cube-sphere data loading for CREDIT Gen2.

This module provides GEFSDataset for the raw GEFS initialization files in the public gfs-ensemble-forecast-system Google Cloud bucket. Each selected ensemble member contains six atmospheric and six surface cube-sphere tiles. The dataset reads only configured variables, stacks members first, flattens the six tiles into tile_lat_lon, and leaves regridding and vertical interpolation to downstream CREDIT blocks.

The raw files are organized as:

gefs.YYYYMMDD/HH/atmos/init/{member}/gfs_ctrl.nc
gefs.YYYYMMDD/HH/atmos/init/{member}/gfs_data.tile{1..6}.nc
gefs.YYYYMMDD/HH/atmos/init/{member}/sfc_data.tile{1..6}.nc

Only initialization-time cube-sphere NetCDF files are supported. Forecast lead products in the bucket are different GRIB2 products and are intentionally outside this dataset’s scope.

Attributes#

Classes#

GEFSDataset

Read raw GEFS cube-sphere initialization data for selected members.

Module Contents#

credit.datasets.gen_2.gefs.logger#
class credit.datasets.gen_2.gefs.GEFSDataset(data_config: dict[str, Any], return_target: bool = False)#

Bases: credit.datasets.gen_2.base_dataset.BaseDataset

Read raw GEFS cube-sphere initialization data for selected members.

The selected members are stacked in the leading tensor dimension. Six cube-sphere tiles are flattened into one spatial dimension so a 3D tensor has shape (members, levels, 1, tile_lat_lon) and a 2D tensor has shape (members, 1, 1, tile_lat_lon). This leading member dimension is intentionally preserved through the Gen2 data pipeline and is compatible with the flattened-spatial handling in the Regridder preblock.

Raw staggered wind fields are requested with their native names u_s, v_s, u_w, and v_w. Users who want unstaggered winds should request the virtual variables u_a and v_a; these are computed from u_s and v_w respectively. There is no forecast_hour setting: this class reads only the cube-sphere initialization NetCDF files.

Input settings:
dataset_type (str): Must be "gefs" when routed through

MultiSourceDataset.

members (list[str] | None): Members to read, such as

["c00", "p01", "p02"]. Omitted configuration defaults to ["c00"]. An explicit empty list discovers and selects every member available in the first requested run.

mode (str): "remote" reads from Google Cloud Storage using

obstore. "local" reads the directory created by gefs_download.py. Defaults to "remote".

base_path (str): Root directory for local mode. Required when

mode is "local".

levels (list[int] | None): One-based model-level indices in the raw

lev dimension. The GEFS cube-sphere files contain 65 model levels. None selects all 65 levels. zh is converted from 66 interfaces to 65 model-level midpoints before this selection.

variables (dict): Field definitions grouped under prognostic,

dynamic_forcing, static, and diagnostic. Use native GEFS names in vars_3D and vars_2D.

return_target (bool): Constructor argument controlling whether the

sample includes the next initialization time under target.

Variables:
  • dataset_type (str) – The registered dataset type, "gefs".

  • members (list[str]) – Selected members in output stacking order.

  • mode (str) – Active storage mode, "remote" or "local".

  • base_path (str | None) – Expanded local storage root, if configured.

  • levels (list[int] | None) – Configured one-based model-level selection.

  • datetimes (pandas.DatetimeIndex) – Available initialization timestamps.

  • file_dict (dict) – Registered field types and their source markers.

  • var_dict (dict) – Registered native GEFS variables grouped by field.

  • static_metadata (dict) – Selected members, vertical-coordinate metadata, and the native unstructured cube-sphere grid.

Example YAML configuration:

data:
  source:
    GEFS:
      dataset_type: "gefs"
      mode: "remote"
      members: ["c00"]
      levels: [1, 65]
      variables:
        prognostic:
          vars_3D: [t]
          vars_2D: [ps]
        dynamic_forcing: null
        static: null
        diagnostic: null
  start_datetime: "2024-01-01T00:00:00"
  end_datetime: "2024-01-01T06:00:00"
  timestep: "6h"
  forecast_len: 1

An empty member list selects all members discovered in the first run:

members: []

Python usage:

import yaml
from credit.datasets.gen_2.gefs import GEFSDataset

with open("config/gefs.yml") as file:
    config = yaml.safe_load(file)
dataset = GEFSDataset(config["data"], return_target=True)
sample = dataset[(dataset.datetimes[0], 0)]
print(sample["input"]["GEFS/prognostic/3d/t"].shape)
members: list[str] = ['c00']#
mode#
engine#
base_path#
levels: list[int] | None#
dataset_type = 'gefs'#
static_metadata#