credit.datasets.gen_2.gfs#
GFS and GDAS data loading for CREDIT Gen2.
This module provides GFSDataset, a PyTorch dataset for the public GFS/GDAS
NetCDF files in Google Cloud Storage. A model run is represented by two files:
an atmospheric file containing model-level fields and a surface file
containing two-dimensional surface fields. GFSDataset discovers available
runs, reads only the requested variables and levels, and returns CREDIT’s
standard input/target sample dictionaries.
The dataset deliberately does not derive pressure, geopotential, or other diagnostics and does not regrid the native Gaussian grid. Those operations are handled by CREDIT postblocks or preblocks so the raw GFS state remains available to downstream processing.
Remote reads use anonymous obstore range requests, allowing xarray to read
selected portions of the large NetCDF objects without downloading each file in
full. Local mode reads files downloaded with
credit.datasets.gen_2.gfs_download. Both modes use the same native layout:
{system}.YYYYMMDD/HH/atmos/{system}.tHHz.atmanl.nc
{system}.YYYYMMDD/HH/atmos/{system}.tHHz.sfcanl.nc
Here system is gdas by default and may be changed to gfs in the
source configuration. Forecast files can be selected with forecast_hour.
Attributes#
Classes#
Read GFS or GDAS atmospheric and surface NetCDF output. |
Module Contents#
- credit.datasets.gen_2.gfs.logger#
- credit.datasets.gen_2.gfs.VALID_SYSTEMS#
- credit.datasets.gen_2.gfs.VALID_LEVEL_TYPES#
- class credit.datasets.gen_2.gfs.GFSDataset(data_config: dict[str, Any], return_target: bool = False, gfs_type: VALID_SYSTEMS | None = None)#
Bases:
credit.datasets.gen_2.base_dataset.BaseDatasetRead GFS or GDAS atmospheric and surface NetCDF output.
GFSDatasetreads the paired atmospheric and surface files published in theglobal-forecast-systemGoogle Cloud bucket. Atmospheric files contain three-dimensional model-level fields and a small number of two-dimensional fields such aspressfc. Surface files contain the two-dimensional surface fields such astmp2mandland. The dataset returns the same flat, slash-delimited tensor keys as the other Gen2 datasets; derivations, regridding, and vertical interpolation are left to later preblocks or postblocks.Remote mode uses
obstorefor anonymous, range-based GCS reads. The remote NetCDF files are opened with xarray and h5netcdf because the obstore reader is a seekable file-like object. Local mode uses thenetcdf4xarray engine by default and expects files laid out like the public bucket. Availability is checked during initialization by default, so missing model runs are removed fromdatetimesrather than failing later during sampling.- Input settings:
- dataset_type (str): Must be
"gfs"when routed through MultiSourceDataset.- system (str): Forecast system, either
"gdas"or"gfs". Defaults to
"gdas". The aliasesmodelandgfs_typeare also accepted in source configuration.- mode (str):
"remote"to read from GCS or"local"to read downloaded files. Defaults to
"remote"for this dataset.- base_path (str): Root directory for local files. Required when
modeis"local". The downloader creates the same{system}.YYYYMMDD/HH/atmos/layout used by the bucket.- forecast_hour (int | None): Forecast lead hour.
Noneselects the analysis files
atmanl.ncandsfcanl.nc. An integer selects files such asatmf003.ncandsfcf003.nc.- level_type (str):
"model"selects one-based positions in the pfulldimension."pressure"selects the nearest values of the file’spfullcoordinate, in hPa. Defaults to"model".- level_coord (str): Atmospheric vertical dimension. Defaults to
"pfull".- levels (list[int | float] | None): Requested model-level indices or
pressure values, depending on
level_type.Nonereads all atmospheric levels.- check_availability (bool): Check that the required atmospheric and
surface objects exist before adding a timestamp to
datetimes. Defaults toTrue.- variables (dict): Field definitions grouped under
prognostic, dynamic_forcing,static, anddiagnostic. Each field may containvars_3Dand/orvars_2Dusing native GFS names.- return_target (bool): Constructor argument controlling whether the
sample includes the next timestep under
target.
- dataset_type (str): Must be
- Variables:
dataset_type (str) – The registered dataset type,
"gfs".system (str) – Active forecast system,
"gdas"or"gfs".mode (str) – Active storage mode,
"remote"or"local".base_path (str | None) – Expanded local storage root, if configured.
forecast_hour (int | None) – Active analysis or forecast lead hour.
level_type (str) – Active vertical-level interpretation.
level_coord (str) – Name of the atmospheric vertical dimension.
levels (list[int | float] | None) – Configured level selection.
datetimes (pandas.DatetimeIndex) – Available sampling timestamps after applying the configured clock and availability checks.
file_dict (dict) – Registered field types and their remote/local file source marker.
var_dict (dict) – Registered native GFS variables grouped by field type.
static_metadata (dict) – Calendar, grid, system, level, and forecast metadata exposed to
MultiSourceDatasetand downstream setup.
Example YAML configuration:
data: source: GDAS: dataset_type: "gfs" system: "gdas" mode: "remote" level_type: "model" levels: [1, 10, 30, 60, 90, 127] # selected from the 127 available model levels (1-127) check_availability: true variables: prognostic: vars_3D: [tmp, ugrd, vgrd, spfh] vars_2D: [pressfc, tmp2m] dynamic_forcing: null static: vars_2D: [land, orog] diagnostic: null start_datetime: "2024-01-01T00:00:00" end_datetime: "2024-01-31T18:00:00" timestep: "6h" forecast_len: 1
For pressure-coordinate selection, change the source settings to
level_type: "pressure"and provide values such aslevels: [50, 100, 500, 850, 1000]. To read downloaded files, usemode: "local"and setbase_pathto the downloader’s output root.Command-line usage:
# Download the configured files for local mode. python -m credit.datasets.gen_2.gfs_download -c config/gfs.yml # Instantiate GFSDataset directly from a YAML file. python - <<'PY' import yaml from credit.datasets.gen_2.gfs import GFSDataset with open("config/gfs.yml") as file: config = yaml.safe_load(file) dataset = GFSDataset(config["data"], return_target=True) print(len(dataset)) PY # In normal training, MultiSourceDataset instantiates GFSDataset from # the same data block. credit_train_gen2 -c config/gfs.yml
- system: VALID_SYSTEMS = ''#
- mode#
- forecast_hour: int | None#
- base_path#
- level_type: VALID_LEVEL_TYPES = ''#
- levels: list[int | float] | None#
- level_coord#
- check_availability#
- dataset_type = 'gfs'#
- static_metadata#