credit.datasets.gen_2.tisr#

TISRDataset: PyTorch Dataset for Total Incident Solar Radiation (TISR) at the top of the atmosphere (TOA).

Sample structure returned by __getitem__:

{
    "input":    {<user_provided_name>: {"<user_provided_name>/dynamic_forcing/2d/tisr": tensor}},
    "target":   {<user_provided_name>: {}},  # empty since dynamic forcing is only input
    "metadata": {<user_provided_name>: {"input_datetime": int, "target_datetime": int}},
}

TISR only has a single variable and is 2D. Tensor shape (no batch dimension):

(1, 1, lat, lon)   — singleton level dim, consistent with CREDIT Gen2 2D convention

After DataLoader collation the batch dimension is prepended:

(batch, 1, 1, lat, lon)

Note that Total Incident Solar Radiation (TISR) and Total Solar Irradiance (TSI) are different physical quantities.

  • TSI is the total solar power per unit area measured on a plane perpendicular (at a 90 degree angle) to the sun’s rays. It is measured at TOA and at the mean Sun-Earth distance (1 AU), and it fluctuates slightly with the Sun’s 11-year solar cycle.

  • TISR is the actual amount of solar energy that hits a specific surface with any orientation. It can be measured at TOA or surface level, and it varies with time and location.

Attributes#

Classes#

TISRDataset

PyTorch Dataset for Total Incident Solar Radiation (TISR) at the top of the atmosphere (TOA).

Module Contents#

credit.datasets.gen_2.tisr.logger#
class credit.datasets.gen_2.tisr.TISRDataset(data_config: dict[str, Any], return_target: bool = False)#

Bases: credit.datasets.gen_2.base_dataset.BaseDataset

PyTorch Dataset for Total Incident Solar Radiation (TISR) at the top of the atmosphere (TOA).

Computations in this class are designed to mimic ERA5’s toa_incident_solar_radiation (tisr) variable (units: J/m2, see https://codes.ecmwf.int/grib/param-db/212) by interpolating ERA5-compatible Total Solar Irradiance (TSI) values to the requested timestamps, then integrating the product of the TSI, a solar scaling factor, and the cosine of the solar zenith angle over the specified period. Defaults to ERA5-compatible settings: an integration period of one hour with 360 integration bins.

While the default configuration targets ERA5 compatibility, both the integration period and bin count are configurable for other use cases. Input timestamps must fall within the TSI data range (1850-2299).

Note that the TISR dataset is typically used as a dynamic forcing/input variable rather than a target, so the return_target parameter is set to False by default. TISR dataset is not loading any data from local or remote files, but rather performing the computation on-the-fly (no need to specify loading mode like most other datasets). Because computation happens inside __getitem__, the dataset emits CPU tensors by default. When used with a multi-worker DataLoader (num_workers > 0), keep device="cpu" (the default) and let the training loop move each collated batch to the GPU; constructing CUDA tensors in worker subprocesses is unsupported by PyTorch. The device config key is provided mainly for single-process (num_workers=0) use.

Exactly one grid source must be configured: either latlon_grid_path (read from a NetCDF file) or both lat_spec and lon_spec (build a rectangular grid in-memory, no file read). Each spec is a [start, end, num_points] list with both endpoints inclusive.

See module docstring for full description of output format and file naming.

Example YAML configuration (grid read from file):

data:
    source:
        Example_TISR:  # User-provided name (arbitrary key)
            dataset_type: "tisr"
            variables:
                prognostic: null
                diagnostic: null
                dynamic_forcing:
                    var_2d: ['tisr']  # only accept 'tisr'
            num_integration_steps: 2160  # 360 steps per hour → 6h integration with 1h accumulation windows
            latlon_grid_path: "/glade/derecho/scratch/cbecker/test_CREDIT_data/era5_local_testing_data_onedeg_2021.nc"

    start_datetime: "2021-06-01"
    end_datetime: "2021-06-04"
    timestep: "6h"
    forecast_len: 1

Example YAML configuration (grid built in-memory from specs):

data:
    source:
        Example_TISR:
            dataset_type: "tisr"
            variables:
                prognostic: null
                diagnostic: null
                dynamic_forcing:
                    var_2d: ['tisr']
            num_integration_steps: 2160
            lat_spec: [90, -90, 721]      # [start, end, num_points], endpoints inclusive
            lon_spec: [0, 359.75, 1440]   # 0.25° grid; excludes the 360° wrap

    start_datetime: "2021-06-01"
    end_datetime: "2021-06-04"
    timestep: "6h"
    forecast_len: 1
dataset_type = 'tisr'#
static_metadata: dict[str, Any]#
num_integration_steps: int#
device: torch.device#
latlon_grid_path: str | None#