credit.datasets.gen_2.base_dataset
==================================

.. py:module:: credit.datasets.gen_2.base_dataset

.. autoapi-nested-parse::

   base_dataset.py
   -------------------------------------------------------
   AbstractBaseDataset and BaseDataset: A PyTorch Dataset class for:
   1. Type hinting and annotations throughout CREDIT
   2. Scaffolding the development of future datasets
   3. Provide a minimal implementation of a Dataset for testing
   4. Avoid redundant code across dataset classes

   BaseDataset handles configuration, sampling clocks, field registration, file
   mapping, history windows, metadata, and input/target sample assembly for a
   single data source. Subclasses customize file discovery and field extraction;
   MultiSourceDataset coordinates multiple BaseDataset instances.

   Temporal behavior is selected with the source-level ``temporal_mode`` option:

   * ``exact`` (default): require source timestamps to match the master clock.
   * ``persist``: use the latest source timestamp at or before each master-clock
     timestamp, which is useful for coarser-resolution data.
   * ``cyclic``: map each timestamp onto a representative ``cycle_year`` and
     resolve it within a repeating annual cycle, which is useful for climatology
     or seasonal forcing. February 29 is clamped to February 28 when the source
     calendar or cycle year has no leap day.

   Sources may use standard pandas calendars or supported CF calendars such as
   ``noleap``, ``all_leap``, and ``julian``. The sampling clock can use either a
   continuous ``start_datetime``/``end_datetime`` interval or non-contiguous
   ``date_ranges``; history and forecast margins are applied to each range.



Attributes
----------

.. autoapisummary::

   credit.datasets.gen_2.base_dataset.logger
   credit.datasets.gen_2.base_dataset.VALID_FIELD_TYPES


Classes
-------

.. autoapisummary::

   credit.datasets.gen_2.base_dataset.AbstractBaseDataset
   credit.datasets.gen_2.base_dataset.BaseDataset


Module Contents
---------------

.. py:data:: logger

.. py:data:: VALID_FIELD_TYPES

.. py:class:: AbstractBaseDataset(data_config: dict[str, Any], return_target: bool = False)

   Bases: :py:obj:`torch.utils.data.Dataset`\ [\ :py:obj:`Any`\ ]


   Abstract base dataset class based on PyTorch Dataset class for CREDIT.

   This class defines the expected methods and attributes for any dataset in CREDIT,
   but does not provide any implementation. The BaseDataset class inherits from this
   class and provides a minimal implementation. Any future dataset should inherit from
   either AbstractBaseDataset or BaseDataset depending on the level of functionality needed.

   For generality, the inheritance is from torch.utils.data.Dataset[Any], however
   there may be benefits to stricter typing than Any for consistency in the
   get item return, especially if torch supports dataset type accelerations in future
   releases.


   .. py:attribute:: curr_source_name
      :type:  str


   .. py:attribute:: dataset_type
      :type:  str


   .. py:attribute:: dt
      :type:  pandas.Timedelta


   .. py:attribute:: num_forecast_steps
      :type:  int


   .. py:attribute:: history_len
      :type:  int


   .. py:attribute:: start_datetime
      :type:  pandas.Timestamp


   .. py:attribute:: end_datetime
      :type:  pandas.Timestamp


   .. py:attribute:: calendar
      :type:  str


   .. py:attribute:: datetimes
      :type:  pandas.Index


   .. py:attribute:: return_target
      :type:  bool


   .. py:attribute:: mode
      :type:  str


   .. py:attribute:: file_dict
      :type:  dict[str, Any]


   .. py:attribute:: var_dict
      :type:  dict[str, Any]


   .. py:attribute:: static_metadata
      :type:  dict[str, Any]


   .. py:method:: __len__() -> int
      :abstractmethod:



   .. py:method:: __getitem__(args: tuple[pandas.Timestamp, int]) -> dict[str, Any]
      :abstractmethod:



   .. py:method:: init_register_all_fields() -> None
      :abstractmethod:



.. py:class:: BaseDataset(data_config: dict[str, Any], return_target: bool = False)

   Bases: :py:obj:`AbstractBaseDataset`


   PyTorch Dataset class for CREDIT that  enables:
   1. Type hinting and annotations throughout CREDIT
   2. Scaffolding the development of future datasets
   3. Provide a minimal implementation of a Dataset for testing

   Minimal YAML config for a dataset will have the following stucture::

       data:
         source:
           Example_Base:  # User-provided name (arbitrary key)
             # PARAMETERS FOR THIS DATASET TYPE
             # Ex: levels: [10, 20, 30]
             dataset_type: "base"  # Needs to match per type of dataset!
             variables:
               prognostic: null
               #  vars_3D: ['T', 'U', 'V', 'Q'] # Your 3D variables
               #  vars_2D: ['SP', 't2m'] # Your 2D variables
               # dynamic_forcing: null
               #  vars_3D: ...
               #  vars_2D: ...
               # static: null
               #  vars_3D: ...
               #  vars_2D: ...
               # diagnostic: null
               #  vars_3D: ...
               #  vars_2D: ...
             # OPTIONAL: Override the clock bounds for this dataset
             start_datetime: "2012-04-03T00:00Z"

           # <YourName2>_<DatasetType2>: # Multiple datasets (see multi_source)

         # These parameters set the overall clock of the sampler
         start_datetime: "2000-01-01T00:00:00Z" # The earliest datetime across datasets
         end_datetime: "2020-12-31T23:00:00Z" # The latest datetime across datasets
         timestep: "12h" # The smallest time interval for the clock
         forecast_len: 1 # The number of timesteps forward that need to be rolled out per sample


   .. py:attribute:: curr_source_name


   .. py:attribute:: curr_source_cfg


   .. py:attribute:: save_loc
      :type:  str | None


   .. py:attribute:: dt
      :type:  pandas.Timedelta


   .. py:attribute:: num_forecast_steps
      :type:  int


   .. py:attribute:: history_len
      :type:  int


   .. py:attribute:: calendar
      :type:  str
      :value: 'standard'



   .. py:attribute:: datetimes
      :type:  pandas.Index


   .. py:attribute:: return_target
      :type:  bool
      :value: False



   .. py:attribute:: mode
      :value: 'local'



   .. py:attribute:: temporal_mode
      :type:  str


   .. py:attribute:: cycle_year
      :type:  int | None


   .. py:attribute:: static_metadata
      :type:  dict[str, Any]


   .. py:attribute:: file_dict
      :type:  dict[str, Any]


   .. py:attribute:: var_dict
      :type:  dict[str, Any]


   .. py:method:: __len__() -> int

      For a CREDIT dataset, the length is the number of unique datetimes that can be sampled from.

      :returns: Dataset length
      :rtype: int



   .. py:method:: __getitem__(args: tuple[pandas.Timestamp, int]) -> dict[str, Any]

      Return a nested input/target sample dict.

      When ``temporal_mode == "persist"``, the requested timestamp *t* is
      snapped to the last native timestamp at-or-before *t* via
      ``pd.DatetimeIndex.asof()``.  The result is cached so that multiple
      fine-resolution master-clock ticks within the same native interval
      only trigger a single file read.

      :param args: ``(t, i)`` where *t* is the current timestamp (nanoseconds
                   or pd.Timestamp) and *i* is the within-sequence step index
                   produced by the sampler. When ``i == 0`` prognostic and static
                   fields are loaded in addition to dynamic forcing.

      :returns: Dict with keys ``"input"``, ``"metadata"``, and optionally
                ``"target"`` (when ``return_target=True``). Both ``"input"`` and
                ``"target"`` are dicts of per-variable tensors keyed by
                ``"{source}/{field_type}/{dim}/{varname}"``.



   .. py:method:: init_register_all_fields() -> None

      Initialize and register all fields for the dataset.

      :raises KeyError: If the config does not include any variables.



