credit.datasets.gen_2.goes#

GOESDataset: PyTorch Dataset for GOES data with nested input/target structure.

Sample structure returned by __getitem__ (GOESDataset does not override this method; see BaseDataset._load_sample for the implementation). Note this is the per-source structure — when wrapped by MultiSourceDataset, an additional layer keyed by <user_provided_name> is added around each of “input”/”target”/”metadata”:

{
    "input":    {"<user_provided_name>/prognostic/2d/CMI_C04": tensor,
                 "<user_provided_name>/prognostic/2d/CMI_C07": tensor},
    "target":   {"<user_provided_name>/prognostic/2d/CMI_C04": tensor,
                 "<user_provided_name>/prognostic/2d/CMI_C07": tensor},  # only populated when return_target=True
    "metadata": {"input_datetime": int, "target_datetime": int},
}
All GOES variables are 2D. Tensor shape (no batch dimension):

(level, time, lat, lon) = (1, 1, lat, lon) — level is singleton since GOES has no vertical levels; time is singleton since each sample covers a single timestep. Consistent with CREDIT Gen2 convention, where 3D variables instead have shape (n_levels, time, lat, lon) (see e.g. credit/datasets/gen_2/era5.py).

After DataLoader collation the batch dimension is prepended:

(batch, level, time, lat, lon) = (batch, 1, 1, lat, lon)

Key features:
  • I/O: local (NetCDF) or remote (public, no-auth AWS S3) loading via mode.

  • Catalogs: a pre-built JSON catalog (file_catalog_path, from quality_check_goes.py) skips the directory/S3 scan and records per-timestamp availability – MISSING, QC_MASKED, SKIP, dropped from self.datetimes (see _filter_unavailable_timestamps); SKIP doubles as a hook for custom sampling on top of (never instead of) QC. Catalogs can be merged, and reused for a narrower time window, a smaller extent, or a subset of QC’d variables than they were built with (see _extent_covers, _variables_covers) – gated by five “invariant” config fields shared between the writer and its readers (see catalog_invariant_metadata).

  • Spatial: extent-based cropping (bbox or NW/SE corners) resolved against GOES’s curvilinear grid via nearest-neighbour search (see _build_spatial_slices); CONUS vs. full-disk products use different precomputed lat/lon grids under latlon2d_dir.

  • Temporal: only a single contiguous start_datetime-end_datetime window is supported (random subsampling, if needed, happens one level up via the trainer’s batches_per_epoch). Multi-step (forecast_len > 1) rollout validates that every target step is available, not just the input timestamp. GOES-16->19 (east) / GOES-17->18 (west) satellite transitions are detected and their ambiguous hour dropped automatically.

  • Data model: field types follow the CREDIT Gen2 convention – prognostic in input and target, dynamic_forcing in input every step, diagnostic in target only, static in input only (never target); rollout feeds back the model’s own prognostic predictions past step 0, with no disk read. Sample keys are "{source_name}/{field_type}/2d/{variable}"; tensors are float32.

Attributes#

Classes#

GOESDataset

PyTorch Dataset for GOES-R ABI Level-2 (L2) satellite imagery.

Module Contents#

credit.datasets.gen_2.goes.logger#
class credit.datasets.gen_2.goes.GOESDataset(data_config: dict[str, Any], return_target: bool = False)#

Bases: credit.datasets.gen_2.base_dataset.BaseDataset

PyTorch Dataset for GOES-R ABI Level-2 (L2) satellite imagery.

Field types follow CREDIT Gen2 conventions: prognostic variables appear in both input (at step 0) and target; dynamic_forcing appears in input at every step; diagnostic appears in target only; static appears in input at step 0 only, same timing as prognostic, but never in target. At step i > 0 the model’s own prognostic predictions are fed back — no disk read occurs for prognostic fields at those steps.

Supports loading directly from AWS S3 (remote mode) or from local NetCDF files (local mode). Spatial subsetting via extent is applied at load time on the curvilinear GOES grid.

See module docstring for full description of output format and file naming.

GOES imager projection background (for deriving the latlon2d_dir grids): https://www.star.nesdis.noaa.gov/atmospheric-composition-training/satellite_data_goes_imager_projection.php

Example YAML configuration (remote/S3 mode):

data:
    source:
        Example_GOES:  # user-provided name (arbitrary key)
            dataset_type: "goes"
            goes_position: "east"        # "east" (GOES-16/19) or "west" (GOES-17/18);
                                         # satellite transitions are handled automatically
            mode: "remote"               # streams directly from AWS S3 (public, no auth required)
            product: "ABI-L2-MCMIPC"    # CONUS; use "ABI-L2-MCMIPF" for full disk
            variables:
                prognostic:
                    vars_2D: ["CMI_C04", "CMI_C07", "CMI_C08", "CMI_C09", "CMI_C10", "CMI_C13"]
                diagnostic: null
                dynamic_forcing: null
            latlon2d_dir: "/glade/derecho/scratch/kevinyang/datasets/goes/"
            # Three extent forms (pick one):
            extent: {nw: [55, -130], se: [20, -60]}      # explicit NW/SE corners (more precise)
            # extent: [-130, -60, 20, 55]                # [lon_min, lon_max, lat_min, lat_max]
            # extent: null                                # no crop — load the full grid
            # this catalog file is generated by quality_check_goes.py
            file_catalog_path: "/path/to/goes_catalog_*.json"  # recommended for remote; avoids S3 listing
            # scan_tolerance: "3 minutes"  # optional; max gap between requested time and nearest file

    start_datetime: "2021-06-01"
    end_datetime: "2021-06-04"
    timestep: "6h"
    forecast_len: 1  # 1 = single-step training

For local mode, the same config applies with these differences: set mode: "local"; add a path key under variables.prognostic pointing to the local NetCDF directory to scan; file_catalog_path is optional rather than recommended, since a local directory scan is cheap.

Parameters:
  • data_config –

    Top-level experiment configuration dictionary. The relevant sub-keys are:

    • config["source"]["Example_GOES"]: user-provided source name.

      • dataset_type (str): has to be “goes” to trigger this dataset class.

      • goes_position (str): Satellite position. One of "east", "west". Defaults to "east".

      • mode (str): "local" or "remote" (S3). Defaults to "local".

      • product (str): ABI product string, e.g. "ABI-L2-MCMIPC".

      • extent (list or dict, optional): Spatial crop. Either [lon_min, lon_max, lat_min, lat_max] or {"nw": [lat, lon], "se": [lat, lon]}.

      • latlon2d_dir (str): Directory containing pre-computed lat/lon grid NetCDF files for GOES’s curvilinear (satellite projection) grid (see class docstring above for background and how to derive these).

      • file_catalog_path (str, optional): Path or glob pattern to a pre-built JSON file catalog. When matched, skips the directory scan entirely. A catalog covering a wide time range (e.g. a full year, built once via quality_check_goes.py) can be reused across many experiments — it only needs to be regenerated when mode, timestep, product, or goes_position change. For extent: an exact match always reuses the catalog; a smaller extent reuses it too, but only if inset from the catalog’s bounds by a safety margin (see _EXTENT_MARGIN_DEG) — too-close-to-boundary or larger requests are rejected, since the catalog’s QC never checked outside (or reliably near the edge of) its own extent (see _extent_covers). To train on a subset of a catalog’s range, do not trim the catalog itself; instead narrow config["start_datetime"] / config["end_datetime"] (below). _build_timestamps (see BaseDataset) derives the actual sample pool from those two values, and _load_file_catalog only looks up rows that fall inside them — the catalog’s full range does not have to match the training window. The dataset itself has no stride/random-subsample option — only a single contiguous start_datetime–end_datetime window. A form of random subsampling can still happen one level up, at the trainer: if trainer.batches_per_epoch is set smaller than a full epoch’s batch count, the sampler (which reshuffles every epoch via set_epoch) only draws that many batches, so each epoch trains on a random subset of the full window — but without direct control over which timestamps that is.

        Each catalog row’s file path may instead be one of three flag (sentinel) values, causing that timestamp to be dropped from self.datetimes (see _filter_unavailable_timestamps):

        • "MISSING" — no GOES file was found within scan_tolerance during the scan (written automatically).

        • "QC_MASKED" — the file failed QC, or failed to open at all (written automatically by quality_check_goes.py).

        • "SKIP" — manually added by hand-editing the catalog JSON (e.g. to exclude a timestamp for a reason outside the automated QC check). Not written by any script — if you regenerate a catalog from a fresh scan, any "SKIP" rows you’d added are lost, since the scan has no way to know about them. This can also be repurposed deliberately: an application with its own strategic sampling logic can assign "SKIP" to exactly the timestamps it wants left out, on top of (never instead of) running QC first — QC should always run before any such additional sampling.

      • scan_tolerance (str, optional): Maximum time difference between a requested timestamp and the nearest GOES file. Accepts any pandas.Timedelta-parseable string (e.g. "5 minutes"). Defaults to "3 minutes".

      • variables (dict): Mapping of field_type to variable spec.

    • config["timestep"] (str): Model timestep as a pandas.Timedelta-parseable string (e.g. "1h").

    • config["forecast_len"] (int): Number of autoregressive forecast steps.

    • config["start_datetime"] (str): Start of the data range.

    • config["end_datetime"] (str): End of the data range.

  • return_target – When True the sample also contains a "target" key populated with prognostic and diagnostic fields at t + dt. Defaults to False.

Variables:
  • datetimes (pd.DatetimeIndex) – Valid input times for which samples can be fetched.

  • file_dict (dict) – Maps each field type to a list of (period_start, period_end, file path) tuples built during initialization.

  • var_dict (dict) – Maps each field type to {"vars_3D": [], "vars_2D": [<variable names>]}. GOES fields are always 2D, so vars_3D is always empty.

  • y_slice (slice) – Row crop derived from extent (or slice(None) for the full grid).

  • x_slice (slice) – Column crop derived from extent (or slice(None) for the full grid).

Raises:
  • ValueError – If dataset_type is not "goes", if goes_position is not "east" or "west", if product does not end in "C" (CONUS) or "F" (full disk), or if mode is not "local" or "remote".

  • FileNotFoundError – If the lat/lon grid NetCDF cannot be found under latlon2d_dir.

  • See also init_register_all_fields and _register_field (both –

  • in BaseDataset), and this class's own _load_file_catalog, for –

  • errors raised during field registration and catalog loading. –

dataset_type = 'goes'#
goes_position: str#
product: str#
static_metadata: dict[str, Any]#
file_catalog_path: str#
tolerance#
extent#
datetimes#
latlon2d_dir: str#
catalog_invariant_metadata() → dict[str, Any]#

The five config fields relevant to catalog compatibility for this dataset.

Notably excludes variables (the requested channel lists), even though that’s also part of what makes a catalog compatible — see _validate_catalog_variables for that separate check, and why it’s separate: these five fields are all known immediately (read once from config, never change), so they can be checked as soon as a catalog is loaded. variables (self.var_dict) is instead built up incrementally, one field type at a time, and a catalog gets loaded in the middle of that process — checking variables here would risk comparing against a still-incomplete list. It’s checked separately, once that process finishes.

This dict is the single definition of “does a catalog match my config,” shared by one writer and two readers so they can’t drift out of sync:

  • Written into a new catalog’s JSON "metadata" by quality_check_goes.py when it builds one.

  • Read by this dataset’s own _load_file_catalog to decide whether an existing catalog is safe to load for training — does it cover what’s being requested?

  • Read by quality_check_goes.py’s pre-overwrite check to decide whether it’s about to clobber an existing catalog that’s already identical — is this the same catalog it’d be regenerating? This guards against silently losing hand-added "SKIP" rows.

mode, timestep, product, and goes_position must match a catalog exactly in all three contexts above. extent is looser only for the load check (see _load_file_catalog/_extent_covers): the catalog’s extent only needs to cover the requested extent (equal or a superset), since a smaller request is just a subdomain of what was already QC’d. The pre-overwrite check still compares extent for exact equality, since “is this the same catalog” is a different question than “does this catalog cover my request.”