credit.datasets.gen_2.goes#
GOESDataset: PyTorch Dataset for GOES data with nested input/target structure.
Sample structure returned by __getitem__ (GOESDataset does not override this method; see BaseDataset._load_sample for the implementation). Note this is the per-source structure — when wrapped by MultiSourceDataset, an additional layer keyed by <user_provided_name> is added around each of “input”/”target”/”metadata”:
{
"input": {"<user_provided_name>/prognostic/2d/CMI_C04": tensor,
"<user_provided_name>/prognostic/2d/CMI_C07": tensor},
"target": {"<user_provided_name>/prognostic/2d/CMI_C04": tensor,
"<user_provided_name>/prognostic/2d/CMI_C07": tensor}, # only populated when return_target=True
"metadata": {"input_datetime": int, "target_datetime": int},
}
- All GOES variables are 2D. Tensor shape (no batch dimension):
(level, time, lat, lon) = (1, 1, lat, lon) — level is singleton since GOES has no vertical levels; time is singleton since each sample covers a single timestep. Consistent with CREDIT Gen2 convention, where 3D variables instead have shape (n_levels, time, lat, lon) (see e.g.
credit/datasets/gen_2/era5.py).- After DataLoader collation the batch dimension is prepended:
(batch, level, time, lat, lon) = (batch, 1, 1, lat, lon)
- Key features:
I/O: local (NetCDF) or remote (public, no-auth AWS S3) loading via
mode.Catalogs: a pre-built JSON catalog (
file_catalog_path, fromquality_check_goes.py) skips the directory/S3 scan and records per-timestamp availability –MISSING,QC_MASKED,SKIP, dropped fromself.datetimes(see_filter_unavailable_timestamps);SKIPdoubles as a hook for custom sampling on top of (never instead of) QC. Catalogs can be merged, and reused for a narrower time window, a smallerextent, or a subset of QC’dvariablesthan they were built with (see_extent_covers,_variables_covers) – gated by five “invariant” config fields shared between the writer and its readers (seecatalog_invariant_metadata).Spatial:
extent-based cropping (bbox or NW/SE corners) resolved against GOES’s curvilinear grid via nearest-neighbour search (see_build_spatial_slices); CONUS vs. full-disk products use different precomputed lat/lon grids underlatlon2d_dir.Temporal: only a single contiguous
start_datetime-end_datetimewindow is supported (random subsampling, if needed, happens one level up via the trainer’sbatches_per_epoch). Multi-step (forecast_len> 1) rollout validates that every target step is available, not just the input timestamp. GOES-16->19 (east) / GOES-17->18 (west) satellite transitions are detected and their ambiguous hour dropped automatically.Data model: field types follow the CREDIT Gen2 convention –
prognosticin input and target,dynamic_forcingin input every step,diagnosticin target only,staticin input only (never target); rollout feeds back the model’s own prognostic predictions past step 0, with no disk read. Sample keys are"{source_name}/{field_type}/2d/{variable}"; tensors arefloat32.
Attributes#
Classes#
PyTorch Dataset for GOES-R ABI Level-2 (L2) satellite imagery. |
Module Contents#
- credit.datasets.gen_2.goes.logger#
- class credit.datasets.gen_2.goes.GOESDataset(data_config: dict[str, Any], return_target: bool = False)#
Bases:
credit.datasets.gen_2.base_dataset.BaseDatasetPyTorch Dataset for GOES-R ABI Level-2 (L2) satellite imagery.
Field types follow CREDIT Gen2 conventions:
prognosticvariables appear in both input (at step 0) and target;dynamic_forcingappears in input at every step;diagnosticappears in target only;staticappears in input at step 0 only, same timing asprognostic, but never in target. At stepi > 0the model’s own prognostic predictions are fed back — no disk read occurs for prognostic fields at those steps.Supports loading directly from AWS S3 (remote mode) or from local NetCDF files (local mode). Spatial subsetting via
extentis applied at load time on the curvilinear GOES grid.See module docstring for full description of output format and file naming.
GOES imager projection background (for deriving the
latlon2d_dirgrids): https://www.star.nesdis.noaa.gov/atmospheric-composition-training/satellite_data_goes_imager_projection.phpExample YAML configuration (remote/S3 mode):
data: source: Example_GOES: # user-provided name (arbitrary key) dataset_type: "goes" goes_position: "east" # "east" (GOES-16/19) or "west" (GOES-17/18); # satellite transitions are handled automatically mode: "remote" # streams directly from AWS S3 (public, no auth required) product: "ABI-L2-MCMIPC" # CONUS; use "ABI-L2-MCMIPF" for full disk variables: prognostic: vars_2D: ["CMI_C04", "CMI_C07", "CMI_C08", "CMI_C09", "CMI_C10", "CMI_C13"] diagnostic: null dynamic_forcing: null latlon2d_dir: "/glade/derecho/scratch/kevinyang/datasets/goes/" # Three extent forms (pick one): extent: {nw: [55, -130], se: [20, -60]} # explicit NW/SE corners (more precise) # extent: [-130, -60, 20, 55] # [lon_min, lon_max, lat_min, lat_max] # extent: null # no crop — load the full grid # this catalog file is generated by quality_check_goes.py file_catalog_path: "/path/to/goes_catalog_*.json" # recommended for remote; avoids S3 listing # scan_tolerance: "3 minutes" # optional; max gap between requested time and nearest file start_datetime: "2021-06-01" end_datetime: "2021-06-04" timestep: "6h" forecast_len: 1 # 1 = single-step training
For local mode, the same config applies with these differences: set
mode: "local"; add apathkey undervariables.prognosticpointing to the local NetCDF directory to scan;file_catalog_pathis optional rather than recommended, since a local directory scan is cheap.- Parameters:
data_config –
Top-level experiment configuration dictionary. The relevant sub-keys are:
config["source"]["Example_GOES"]: user-provided source name.dataset_type(str): has to be “goes” to trigger this dataset class.goes_position(str): Satellite position. One of"east","west". Defaults to"east".mode(str):"local"or"remote"(S3). Defaults to"local".product(str): ABI product string, e.g."ABI-L2-MCMIPC".extent(list or dict, optional): Spatial crop. Either[lon_min, lon_max, lat_min, lat_max]or{"nw": [lat, lon], "se": [lat, lon]}.latlon2d_dir(str): Directory containing pre-computed lat/lon grid NetCDF files for GOES’s curvilinear (satellite projection) grid (see class docstring above for background and how to derive these).file_catalog_path(str, optional): Path or glob pattern to a pre-built JSON file catalog. When matched, skips the directory scan entirely. A catalog covering a wide time range (e.g. a full year, built once viaquality_check_goes.py) can be reused across many experiments — it only needs to be regenerated whenmode,timestep,product, orgoes_positionchange. Forextent: an exact match always reuses the catalog; a smaller extent reuses it too, but only if inset from the catalog’s bounds by a safety margin (see_EXTENT_MARGIN_DEG) — too-close-to-boundary or larger requests are rejected, since the catalog’s QC never checked outside (or reliably near the edge of) its own extent (see_extent_covers). To train on a subset of a catalog’s range, do not trim the catalog itself; instead narrowconfig["start_datetime"]/config["end_datetime"](below)._build_timestamps(seeBaseDataset) derives the actual sample pool from those two values, and_load_file_catalogonly looks up rows that fall inside them — the catalog’s full range does not have to match the training window. The dataset itself has no stride/random-subsample option — only a single contiguousstart_datetime–end_datetimewindow. A form of random subsampling can still happen one level up, at the trainer: iftrainer.batches_per_epochis set smaller than a full epoch’s batch count, the sampler (which reshuffles every epoch viaset_epoch) only draws that many batches, so each epoch trains on a random subset of the full window — but without direct control over which timestamps that is.Each catalog row’s file path may instead be one of three flag (sentinel) values, causing that timestamp to be dropped from
self.datetimes(see_filter_unavailable_timestamps):"MISSING"— no GOES file was found withinscan_toleranceduring the scan (written automatically)."QC_MASKED"— the file failed QC, or failed to open at all (written automatically byquality_check_goes.py)."SKIP"— manually added by hand-editing the catalog JSON (e.g. to exclude a timestamp for a reason outside the automated QC check). Not written by any script — if you regenerate a catalog from a fresh scan, any"SKIP"rows you’d added are lost, since the scan has no way to know about them. This can also be repurposed deliberately: an application with its own strategic sampling logic can assign"SKIP"to exactly the timestamps it wants left out, on top of (never instead of) running QC first — QC should always run before any such additional sampling.
scan_tolerance(str, optional): Maximum time difference between a requested timestamp and the nearest GOES file. Accepts anypandas.Timedelta-parseable string (e.g."5 minutes"). Defaults to"3 minutes".variables(dict): Mapping of field_type to variable spec.
config["timestep"](str): Model timestep as apandas.Timedelta-parseable string (e.g."1h").config["forecast_len"](int): Number of autoregressive forecast steps.config["start_datetime"](str): Start of the data range.config["end_datetime"](str): End of the data range.
return_target – When
Truethe sample also contains a"target"key populated with prognostic and diagnostic fields att + dt. Defaults toFalse.
- Variables:
datetimes (pd.DatetimeIndex) – Valid input times for which samples can be fetched.
file_dict (dict) – Maps each field type to a list of
(period_start, period_end, file path)tuples built during initialization.var_dict (dict) – Maps each field type to
{"vars_3D": [], "vars_2D": [<variable names>]}. GOES fields are always 2D, sovars_3Dis always empty.y_slice (slice) – Row crop derived from
extent(orslice(None)for the full grid).x_slice (slice) – Column crop derived from
extent(orslice(None)for the full grid).
- Raises:
ValueError – If
dataset_typeis not"goes", ifgoes_positionis not"east"or"west", ifproductdoes not end in"C"(CONUS) or"F"(full disk), or ifmodeis not"local"or"remote".FileNotFoundError – If the lat/lon grid NetCDF cannot be found under
latlon2d_dir.See also init_register_all_fields and _register_field (both –
in BaseDataset), and this class's own _load_file_catalog, for –
errors raised during field registration and catalog loading. –
- dataset_type = 'goes'#
- goes_position: str#
- product: str#
- static_metadata: dict[str, Any]#
- file_catalog_path: str#
- tolerance#
- extent#
- datetimes#
- latlon2d_dir: str#
- catalog_invariant_metadata() dict[str, Any]#
The five config fields relevant to catalog compatibility for this dataset.
Notably excludes
variables(the requested channel lists), even though that’s also part of what makes a catalog compatible — see_validate_catalog_variablesfor that separate check, and why it’s separate: these five fields are all known immediately (read once from config, never change), so they can be checked as soon as a catalog is loaded.variables(self.var_dict) is instead built up incrementally, one field type at a time, and a catalog gets loaded in the middle of that process — checkingvariableshere would risk comparing against a still-incomplete list. It’s checked separately, once that process finishes.This dict is the single definition of “does a catalog match my config,” shared by one writer and two readers so they can’t drift out of sync:
Written into a new catalog’s JSON
"metadata"byquality_check_goes.pywhen it builds one.Read by this dataset’s own
_load_file_catalogto decide whether an existing catalog is safe to load for training — does it cover what’s being requested?Read by
quality_check_goes.py’s pre-overwrite check to decide whether it’s about to clobber an existing catalog that’s already identical — is this the same catalog it’d be regenerating? This guards against silently losing hand-added"SKIP"rows.
mode,timestep,product, andgoes_positionmust match a catalog exactly in all three contexts above.extentis looser only for the load check (see_load_file_catalog/_extent_covers): the catalog’s extent only needs to cover the requested extent (equal or a superset), since a smaller request is just a subdomain of what was already QC’d. The pre-overwrite check still comparesextentfor exact equality, since “is this the same catalog” is a different question than “does this catalog cover my request.”