credit.datasets.gen_2.base_dataset#
AbstractBaseDataset and BaseDataset: A PyTorch Dataset class for: 1. Type hinting and annotations throughout CREDIT 2. Scaffolding the development of future datasets 3. Provide a minimal implementation of a Dataset for testing 4. Avoid redundant code across dataset classes
BaseDataset handles configuration, sampling clocks, field registration, file mapping, history windows, metadata, and input/target sample assembly for a single data source. Subclasses customize file discovery and field extraction; MultiSourceDataset coordinates multiple BaseDataset instances.
Temporal behavior is selected with the source-level temporal_mode option:
exact(default): require source timestamps to match the master clock.persist: use the latest source timestamp at or before each master-clock timestamp, which is useful for coarser-resolution data.cyclic: map each timestamp onto a representativecycle_yearand resolve it within a repeating annual cycle, which is useful for climatology or seasonal forcing. February 29 is clamped to February 28 when the source calendar or cycle year has no leap day.
Sources may use standard pandas calendars or supported CF calendars such as
noleap, all_leap, and julian. The sampling clock can use either a
continuous start_datetime/end_datetime interval or non-contiguous
date_ranges; history and forecast margins are applied to each range.
Attributes#
Classes#
Abstract base dataset class based on PyTorch Dataset class for CREDIT. |
|
PyTorch Dataset class for CREDIT that enables: |
Module Contents#
- credit.datasets.gen_2.base_dataset.logger#
- credit.datasets.gen_2.base_dataset.VALID_FIELD_TYPES#
- class credit.datasets.gen_2.base_dataset.AbstractBaseDataset(data_config: dict[str, Any], return_target: bool = False)#
Bases:
torch.utils.data.Dataset[Any]Abstract base dataset class based on PyTorch Dataset class for CREDIT.
This class defines the expected methods and attributes for any dataset in CREDIT, but does not provide any implementation. The BaseDataset class inherits from this class and provides a minimal implementation. Any future dataset should inherit from either AbstractBaseDataset or BaseDataset depending on the level of functionality needed.
For generality, the inheritance is from torch.utils.data.Dataset[Any], however there may be benefits to stricter typing than Any for consistency in the get item return, especially if torch supports dataset type accelerations in future releases.
- curr_source_name: str#
- dataset_type: str#
- dt: pandas.Timedelta#
- num_forecast_steps: int#
- history_len: int#
- start_datetime: pandas.Timestamp#
- end_datetime: pandas.Timestamp#
- calendar: str#
- datetimes: pandas.Index#
- return_target: bool#
- mode: str#
- file_dict: dict[str, Any]#
- var_dict: dict[str, Any]#
- static_metadata: dict[str, Any]#
- abstractmethod __len__() int#
- abstractmethod __getitem__(args: tuple[pandas.Timestamp, int]) dict[str, Any]#
- abstractmethod init_register_all_fields() None#
- class credit.datasets.gen_2.base_dataset.BaseDataset(data_config: dict[str, Any], return_target: bool = False)#
Bases:
AbstractBaseDatasetPyTorch Dataset class for CREDIT that enables: 1. Type hinting and annotations throughout CREDIT 2. Scaffolding the development of future datasets 3. Provide a minimal implementation of a Dataset for testing
Minimal YAML config for a dataset will have the following stucture:
data: source: Example_Base: # User-provided name (arbitrary key) # PARAMETERS FOR THIS DATASET TYPE # Ex: levels: [10, 20, 30] dataset_type: "base" # Needs to match per type of dataset! variables: prognostic: null # vars_3D: ['T', 'U', 'V', 'Q'] # Your 3D variables # vars_2D: ['SP', 't2m'] # Your 2D variables # dynamic_forcing: null # vars_3D: ... # vars_2D: ... # static: null # vars_3D: ... # vars_2D: ... # diagnostic: null # vars_3D: ... # vars_2D: ... # OPTIONAL: Override the clock bounds for this dataset start_datetime: "2012-04-03T00:00Z" # <YourName2>_<DatasetType2>: # Multiple datasets (see multi_source) # These parameters set the overall clock of the sampler start_datetime: "2000-01-01T00:00:00Z" # The earliest datetime across datasets end_datetime: "2020-12-31T23:00:00Z" # The latest datetime across datasets timestep: "12h" # The smallest time interval for the clock forecast_len: 1 # The number of timesteps forward that need to be rolled out per sample
- curr_source_name#
- curr_source_cfg#
- save_loc: str | None#
- dt: pandas.Timedelta#
- num_forecast_steps: int#
- history_len: int#
- calendar: str = 'standard'#
- datetimes: pandas.Index#
- return_target: bool = False#
- mode = 'local'#
- temporal_mode: str#
- cycle_year: int | None#
- static_metadata: dict[str, Any]#
- file_dict: dict[str, Any]#
- var_dict: dict[str, Any]#
- __len__() int#
For a CREDIT dataset, the length is the number of unique datetimes that can be sampled from.
- Returns:
Dataset length
- Return type:
int
- __getitem__(args: tuple[pandas.Timestamp, int]) dict[str, Any]#
Return a nested input/target sample dict.
When
temporal_mode == "persist", the requested timestamp t is snapped to the last native timestamp at-or-before t viapd.DatetimeIndex.asof(). The result is cached so that multiple fine-resolution master-clock ticks within the same native interval only trigger a single file read.- Parameters:
args –
(t, i)where t is the current timestamp (nanoseconds or pd.Timestamp) and i is the within-sequence step index produced by the sampler. Wheni == 0prognostic and static fields are loaded in addition to dynamic forcing.- Returns:
Dict with keys
"input","metadata", and optionally"target"(whenreturn_target=True). Both"input"and"target"are dicts of per-variable tensors keyed by"{source}/{field_type}/{dim}/{varname}".
- init_register_all_fields() None#
Initialize and register all fields for the dataset.
- Raises:
KeyError – If the config does not include any variables.