credit.datasets.gen_2.gfs
=========================

.. py:module:: credit.datasets.gen_2.gfs

.. autoapi-nested-parse::

   gfs.py
   ------
   GFS and GDAS data loading for CREDIT Gen2.

   This module provides ``GFSDataset``, a PyTorch dataset for the public GFS/GDAS
   NetCDF files in Google Cloud Storage. A model run is represented by two files:
   an atmospheric file containing model-level fields and a surface file
   containing two-dimensional surface fields. ``GFSDataset`` discovers available
   runs, reads only the requested variables and levels, and returns CREDIT's
   standard input/target sample dictionaries.

   The dataset deliberately does not derive pressure, geopotential, or other
   diagnostics and does not regrid the native Gaussian grid. Those operations are
   handled by CREDIT postblocks or preblocks so the raw GFS state remains
   available to downstream processing.

   Remote reads use anonymous ``obstore`` range requests, allowing xarray to read
   selected portions of the large NetCDF objects without downloading each file in
   full. Local mode reads files downloaded with
   ``credit.datasets.gen_2.gfs_download``. Both modes use the same native layout::

       {system}.YYYYMMDD/HH/atmos/{system}.tHHz.atmanl.nc
       {system}.YYYYMMDD/HH/atmos/{system}.tHHz.sfcanl.nc

   Here ``system`` is ``gdas`` by default and may be changed to ``gfs`` in the
   source configuration. Forecast files can be selected with ``forecast_hour``.



Attributes
----------

.. autoapisummary::

   credit.datasets.gen_2.gfs.logger
   credit.datasets.gen_2.gfs.VALID_SYSTEMS
   credit.datasets.gen_2.gfs.VALID_LEVEL_TYPES


Classes
-------

.. autoapisummary::

   credit.datasets.gen_2.gfs.GFSDataset


Module Contents
---------------

.. py:data:: logger

.. py:data:: VALID_SYSTEMS

.. py:data:: VALID_LEVEL_TYPES

.. py:class:: GFSDataset(data_config: dict[str, Any], return_target: bool = False, gfs_type: VALID_SYSTEMS | None = None)

   Bases: :py:obj:`credit.datasets.gen_2.base_dataset.BaseDataset`


   Read GFS or GDAS atmospheric and surface NetCDF output.

   ``GFSDataset`` reads the paired atmospheric and surface files published in
   the ``global-forecast-system`` Google Cloud bucket. Atmospheric files
   contain three-dimensional model-level fields and a small number of
   two-dimensional fields such as ``pressfc``. Surface files contain the
   two-dimensional surface fields such as ``tmp2m`` and ``land``. The dataset
   returns the same flat, slash-delimited tensor keys as the other Gen2
   datasets; derivations, regridding, and vertical interpolation are left to
   later preblocks or postblocks.

   Remote mode uses ``obstore`` for anonymous, range-based GCS reads. The
   remote NetCDF files are opened with xarray and h5netcdf because the
   obstore reader is a seekable file-like object. Local mode uses the
   ``netcdf4`` xarray engine by default and expects files laid out like the
   public bucket. Availability is checked during initialization by default,
   so missing model runs are removed from ``datetimes`` rather than failing
   later during sampling.

   Input settings:
       dataset_type (str): Must be ``"gfs"`` when routed through
           ``MultiSourceDataset``.
       system (str): Forecast system, either ``"gdas"`` or ``"gfs"``.
           Defaults to ``"gdas"``. The aliases ``model`` and ``gfs_type``
           are also accepted in source configuration.
       mode (str): ``"remote"`` to read from GCS or ``"local"`` to read
           downloaded files. Defaults to ``"remote"`` for this dataset.
       base_path (str): Root directory for local files. Required when
           ``mode`` is ``"local"``. The downloader creates the same
           ``{system}.YYYYMMDD/HH/atmos/`` layout used by the bucket.
       forecast_hour (int | None): Forecast lead hour. ``None`` selects the
           analysis files ``atmanl.nc`` and ``sfcanl.nc``. An integer selects
           files such as ``atmf003.nc`` and ``sfcf003.nc``.
       level_type (str): ``"model"`` selects one-based positions in the
           ``pfull`` dimension. ``"pressure"`` selects the nearest values of
           the file's ``pfull`` coordinate, in hPa. Defaults to ``"model"``.
       level_coord (str): Atmospheric vertical dimension. Defaults to
           ``"pfull"``.
       levels (list[int | float] | None): Requested model-level indices or
           pressure values, depending on ``level_type``. ``None`` reads all
           atmospheric levels.
       check_availability (bool): Check that the required atmospheric and
           surface objects exist before adding a timestamp to ``datetimes``.
           Defaults to ``True``.
       variables (dict): Field definitions grouped under ``prognostic``,
           ``dynamic_forcing``, ``static``, and ``diagnostic``. Each field
           may contain ``vars_3D`` and/or ``vars_2D`` using native GFS names.
       return_target (bool): Constructor argument controlling whether the
           sample includes the next timestep under ``target``.

   :ivar dataset_type: The registered dataset type, ``"gfs"``.
   :vartype dataset_type: str
   :ivar system: Active forecast system, ``"gdas"`` or ``"gfs"``.
   :vartype system: str
   :ivar mode: Active storage mode, ``"remote"`` or ``"local"``.
   :vartype mode: str
   :ivar base_path: Expanded local storage root, if configured.
   :vartype base_path: str | None
   :ivar forecast_hour: Active analysis or forecast lead hour.
   :vartype forecast_hour: int | None
   :ivar level_type: Active vertical-level interpretation.
   :vartype level_type: str
   :ivar level_coord: Name of the atmospheric vertical dimension.
   :vartype level_coord: str
   :ivar levels: Configured level selection.
   :vartype levels: list[int | float] | None
   :ivar datetimes: Available sampling timestamps after
                    applying the configured clock and availability checks.
   :vartype datetimes: pandas.DatetimeIndex
   :ivar file_dict: Registered field types and their remote/local file
                    source marker.
   :vartype file_dict: dict
   :ivar var_dict: Registered native GFS variables grouped by field type.
   :vartype var_dict: dict
   :ivar static_metadata: Calendar, grid, system, level, and forecast
                          metadata exposed to ``MultiSourceDataset`` and downstream setup.

   :vartype static_metadata: dict

   Example YAML configuration::

       data:
         source:
           GDAS:
             dataset_type: "gfs"
             system: "gdas"
             mode: "remote"
             level_type: "model"
             levels: [1, 10, 30, 60, 90, 127]  # selected from the 127 available model levels (1-127)
             check_availability: true
             variables:
               prognostic:
                 vars_3D: [tmp, ugrd, vgrd, spfh]
                 vars_2D: [pressfc, tmp2m]
               dynamic_forcing: null
               static:
                 vars_2D: [land, orog]
               diagnostic: null
         start_datetime: "2024-01-01T00:00:00"
         end_datetime: "2024-01-31T18:00:00"
         timestep: "6h"
         forecast_len: 1

   For pressure-coordinate selection, change the source settings to
   ``level_type: "pressure"`` and provide values such as
   ``levels: [50, 100, 500, 850, 1000]``. To read downloaded files, use
   ``mode: "local"`` and set ``base_path`` to the downloader's output root.

   Command-line usage::

       # Download the configured files for local mode.
       python -m credit.datasets.gen_2.gfs_download -c config/gfs.yml

       # Instantiate GFSDataset directly from a YAML file.
       python - <<'PY'
       import yaml
       from credit.datasets.gen_2.gfs import GFSDataset

       with open("config/gfs.yml") as file:
           config = yaml.safe_load(file)
       dataset = GFSDataset(config["data"], return_target=True)
       print(len(dataset))
       PY

       # In normal training, MultiSourceDataset instantiates GFSDataset from
       # the same data block.
       credit_train_gen2 -c config/gfs.yml


   .. py:attribute:: system
      :type:  VALID_SYSTEMS
      :value: ''



   .. py:attribute:: mode


   .. py:attribute:: forecast_hour
      :type:  int | None


   .. py:attribute:: base_path


   .. py:attribute:: level_type
      :type:  VALID_LEVEL_TYPES
      :value: ''



   .. py:attribute:: levels
      :type:  list[int | float] | None


   .. py:attribute:: level_coord


   .. py:attribute:: check_availability


   .. py:attribute:: dataset_type
      :value: 'gfs'



   .. py:attribute:: static_metadata


