credit.datasets.gen_2.goes
==========================

.. py:module:: credit.datasets.gen_2.goes

.. autoapi-nested-parse::

   goes.py
   -------------------------------------------------------
   GOESDataset: PyTorch Dataset for GOES data with nested input/target structure.

   Sample structure returned by __getitem__ (GOESDataset does not override this method;
   see BaseDataset._load_sample for the implementation). Note this is the per-source
   structure — when wrapped by MultiSourceDataset, an additional layer keyed by
   <user_provided_name> is added around each of "input"/"target"/"metadata"::

       {
           "input":    {"<user_provided_name>/prognostic/2d/CMI_C04": tensor,
                        "<user_provided_name>/prognostic/2d/CMI_C07": tensor},
           "target":   {"<user_provided_name>/prognostic/2d/CMI_C04": tensor,
                        "<user_provided_name>/prognostic/2d/CMI_C07": tensor},  # only populated when return_target=True
           "metadata": {"input_datetime": int, "target_datetime": int},
       }

   All GOES variables are 2D. Tensor shape (no batch dimension):
       (level, time, lat, lon) = (1, 1, lat, lon)
       — level is singleton since GOES has no vertical levels; time is singleton
       since each sample covers a single timestep. Consistent with CREDIT Gen2
       convention, where 3D variables instead have shape (n_levels, time, lat, lon)
       (see e.g. ``credit/datasets/gen_2/era5.py``).

   After DataLoader collation the batch dimension is prepended:
       (batch, level, time, lat, lon) = (batch, 1, 1, lat, lon)

   Key features:
       * **I/O**: local (NetCDF) or remote (public, no-auth AWS S3) loading via
         ``mode``.
       * **Catalogs**: a pre-built JSON catalog (``file_catalog_path``, from
         ``quality_check_goes.py``) skips the directory/S3 scan and records
         per-timestamp availability -- ``MISSING``, ``QC_MASKED``, ``SKIP``,
         dropped from ``self.datetimes`` (see ``_filter_unavailable_timestamps``);
         ``SKIP`` doubles as a hook for custom sampling on top of (never instead
         of) QC. Catalogs can be merged, and reused for a narrower time window, a
         smaller ``extent``, or a subset of QC'd ``variables`` than they were
         built with (see ``_extent_covers``, ``_variables_covers``) -- gated by
         five "invariant" config fields shared between the writer and its
         readers (see ``catalog_invariant_metadata``).
       * **Spatial**: ``extent``-based cropping (bbox or NW/SE corners) resolved
         against GOES's curvilinear grid via nearest-neighbour search (see
         ``_build_spatial_slices``); CONUS vs. full-disk products use different
         precomputed lat/lon grids under ``latlon2d_dir``.
       * **Temporal**: only a single contiguous ``start_datetime``-``end_datetime``
         window is supported (random subsampling, if needed, happens one level up
         via the trainer's ``batches_per_epoch``). Multi-step (``forecast_len`` >
         1) rollout validates that every target step is available, not just the
         input timestamp. GOES-16->19 (east) / GOES-17->18 (west) satellite
         transitions are detected and their ambiguous hour dropped automatically.
       * **Data model**: field types follow the CREDIT Gen2 convention --
         ``prognostic`` in input and target, ``dynamic_forcing`` in input every
         step, ``diagnostic`` in target only, ``static`` in input only (never
         target); rollout feeds back the model's own prognostic predictions past
         step 0, with no disk read. Sample keys are
         ``"{source_name}/{field_type}/2d/{variable}"``; tensors are ``float32``.



Attributes
----------

.. autoapisummary::

   credit.datasets.gen_2.goes.logger


Classes
-------

.. autoapisummary::

   credit.datasets.gen_2.goes.GOESDataset


Module Contents
---------------

.. py:data:: logger

.. py:class:: GOESDataset(data_config: dict[str, Any], return_target: bool = False)

   Bases: :py:obj:`credit.datasets.gen_2.base_dataset.BaseDataset`


   PyTorch Dataset for GOES-R ABI Level-2 (L2) satellite imagery.

   Field types follow CREDIT Gen2 conventions: ``prognostic`` variables appear in
   both input (at step 0) and target; ``dynamic_forcing`` appears in input
   at every step; ``diagnostic`` appears in target only; ``static`` appears in
   input at step 0 only, same timing as ``prognostic``, but never in target.
   At step ``i > 0`` the model's own prognostic predictions are fed back — no
   disk read occurs for prognostic fields at those steps.

   Supports loading directly from AWS S3 (remote mode) or from local
   NetCDF files (local mode). Spatial subsetting via ``extent``
   is applied at load time on the curvilinear GOES grid.

   See module docstring for full description of output format and file naming.

   GOES imager projection background (for deriving the ``latlon2d_dir`` grids):
   https://www.star.nesdis.noaa.gov/atmospheric-composition-training/satellite_data_goes_imager_projection.php

   Example YAML configuration (remote/S3 mode)::

       data:
           source:
               Example_GOES:  # user-provided name (arbitrary key)
                   dataset_type: "goes"
                   goes_position: "east"        # "east" (GOES-16/19) or "west" (GOES-17/18);
                                                # satellite transitions are handled automatically
                   mode: "remote"               # streams directly from AWS S3 (public, no auth required)
                   product: "ABI-L2-MCMIPC"    # CONUS; use "ABI-L2-MCMIPF" for full disk
                   variables:
                       prognostic:
                           vars_2D: ["CMI_C04", "CMI_C07", "CMI_C08", "CMI_C09", "CMI_C10", "CMI_C13"]
                       diagnostic: null
                       dynamic_forcing: null
                   latlon2d_dir: "/glade/derecho/scratch/kevinyang/datasets/goes/"
                   # Three extent forms (pick one):
                   extent: {nw: [55, -130], se: [20, -60]}      # explicit NW/SE corners (more precise)
                   # extent: [-130, -60, 20, 55]                # [lon_min, lon_max, lat_min, lat_max]
                   # extent: null                                # no crop — load the full grid
                   # this catalog file is generated by quality_check_goes.py
                   file_catalog_path: "/path/to/goes_catalog_*.json"  # recommended for remote; avoids S3 listing
                   # scan_tolerance: "3 minutes"  # optional; max gap between requested time and nearest file

           start_datetime: "2021-06-01"
           end_datetime: "2021-06-04"
           timestep: "6h"
           forecast_len: 1  # 1 = single-step training

   For local mode, the same config applies with these differences: set
   ``mode: "local"``; add a ``path`` key under ``variables.prognostic``
   pointing to the local NetCDF directory to scan; ``file_catalog_path`` is
   optional rather than recommended, since a local directory scan is cheap.

   :param data_config: Top-level experiment configuration dictionary. The relevant
                       sub-keys are:

                       - ``config["source"]["Example_GOES"]``: user-provided source name.

                         - ``dataset_type`` (str): has to be "goes" to trigger this dataset class.
                         - ``goes_position`` (str): Satellite position. One of ``"east"``, ``"west"``. Defaults to
                           ``"east"``.
                         - ``mode`` (str): ``"local"`` or ``"remote"`` (S3). Defaults to
                           ``"local"``.
                         - ``product`` (str): ABI product string, e.g.
                           ``"ABI-L2-MCMIPC"``.
                         - ``extent`` (list or dict, optional): Spatial crop. Either
                           ``[lon_min, lon_max, lat_min, lat_max]`` or
                           ``{"nw": [lat, lon], "se": [lat, lon]}``.
                         - ``latlon2d_dir`` (str): Directory containing pre-computed
                           lat/lon grid NetCDF files for GOES's curvilinear (satellite
                           projection) grid (see class docstring above for background
                           and how to derive these).
                         - ``file_catalog_path`` (str, optional): Path or glob pattern to
                           a pre-built JSON file catalog. When matched, skips the
                           directory scan entirely. A catalog covering a wide time range
                           (e.g. a full year, built once via ``quality_check_goes.py``)
                           can be reused across many experiments — it only needs to be
                           regenerated when ``mode``, ``timestep``, ``product``, or
                           ``goes_position`` change. For ``extent``: an exact match
                           always reuses the catalog; a *smaller* extent reuses it too,
                           but only if inset from the catalog's bounds by a safety
                           margin (see ``_EXTENT_MARGIN_DEG``) — too-close-to-boundary or
                           *larger* requests are rejected, since the catalog's QC never
                           checked outside (or reliably near the edge of) its own extent
                           (see ``_extent_covers``).
                           To train on a subset of a catalog's range, do **not** trim the
                           catalog itself; instead narrow ``config["start_datetime"]`` /
                           ``config["end_datetime"]`` (below). ``_build_timestamps``
                           (see ``BaseDataset``) derives the actual sample pool from
                           those two values, and ``_load_file_catalog`` only looks up
                           rows that fall inside them — the catalog's full range does not
                           have to match the training window. The dataset itself has no
                           stride/random-subsample option — only a single contiguous
                           ``start_datetime``–``end_datetime`` window. A form of random
                           subsampling can still happen one level up, at the trainer: if
                           ``trainer.batches_per_epoch`` is set smaller than a full
                           epoch's batch count, the sampler (which reshuffles every
                           epoch via ``set_epoch``) only draws that many batches, so
                           each epoch trains on a random subset of the full window —
                           but without direct control over which timestamps that is.

                           Each catalog row's file path may instead be one of three
                           flag (sentinel) values, causing that timestamp to be dropped from
                           ``self.datetimes`` (see ``_filter_unavailable_timestamps``):

                           * ``"MISSING"`` — no GOES file was found within
                             ``scan_tolerance`` during the scan (written automatically).
                           * ``"QC_MASKED"`` — the file failed QC, or failed to open at
                             all (written automatically by ``quality_check_goes.py``).
                           * ``"SKIP"`` — manually added by hand-editing the catalog
                             JSON (e.g. to exclude a timestamp for a reason outside the
                             automated QC check). Not written by any script — if you
                             regenerate a catalog from a fresh scan, any ``"SKIP"`` rows
                             you'd added are lost, since the scan has no way to know
                             about them. This can also be repurposed deliberately: an
                             application with its own strategic sampling logic can
                             assign ``"SKIP"`` to exactly the timestamps it wants left
                             out, on top of (never instead of) running QC first — QC
                             should always run before any such additional sampling.
                         - ``scan_tolerance`` (str, optional): Maximum time difference
                           between a requested timestamp and the nearest GOES file.
                           Accepts any ``pandas.Timedelta``-parseable string (e.g.
                           ``"5 minutes"``). Defaults to ``"3 minutes"``.
                         - ``variables`` (dict): Mapping of field_type to variable spec.

                       - ``config["timestep"]`` (str): Model timestep as a
                         ``pandas.Timedelta``-parseable string (e.g. ``"1h"``).
                       - ``config["forecast_len"]`` (int): Number of autoregressive
                         forecast steps.
                       - ``config["start_datetime"]`` (str): Start of the data range.
                       - ``config["end_datetime"]`` (str): End of the data range.
   :param return_target: When ``True`` the sample also contains a ``"target"``
                         key populated with prognostic and diagnostic fields at ``t + dt``.
                         Defaults to ``False``.

   :ivar datetimes: Valid input times for which samples can
                    be fetched.
   :vartype datetimes: pd.DatetimeIndex
   :ivar file_dict: Maps each field type to a list of
                    ``(period_start, period_end, file path)`` tuples built during
                    initialization.
   :vartype file_dict: dict
   :ivar var_dict: Maps each field type to
                   ``{"vars_3D": [], "vars_2D": [<variable names>]}``. GOES fields
                   are always 2D, so ``vars_3D`` is always empty.
   :vartype var_dict: dict
   :ivar y_slice: Row crop derived from ``extent`` (or ``slice(None)``
                  for the full grid).
   :vartype y_slice: slice
   :ivar x_slice: Column crop derived from ``extent`` (or
                  ``slice(None)`` for the full grid).

   :vartype x_slice: slice

   :raises ValueError: If ``dataset_type`` is not ``"goes"``, if ``goes_position``
       is not ``"east"`` or ``"west"``, if ``product`` does not end in
       ``"C"`` (CONUS) or ``"F"`` (full disk), or if ``mode`` is not
       ``"local"`` or ``"remote"``.
   :raises FileNotFoundError: If the lat/lon grid NetCDF cannot be found under
       ``latlon2d_dir``.
   :raises See also ``init_register_all_fields`` and ``_register_field`` (both:
   :raises in ``BaseDataset``), and this class's own ``_load_file_catalog``, for:
   :raises errors raised during field registration and catalog loading.:


   .. py:attribute:: dataset_type
      :value: 'goes'



   .. py:attribute:: goes_position
      :type:  str


   .. py:attribute:: product
      :type:  str


   .. py:attribute:: static_metadata
      :type:  dict[str, Any]


   .. py:attribute:: file_catalog_path
      :type:  str


   .. py:attribute:: tolerance


   .. py:attribute:: extent


   .. py:attribute:: datetimes


   .. py:attribute:: latlon2d_dir
      :type:  str


   .. py:method:: catalog_invariant_metadata() -> dict[str, Any]

      The five config fields relevant to catalog compatibility for this dataset.

      Notably excludes ``variables`` (the requested channel lists), even
      though that's also part of what makes a catalog compatible — see
      ``_validate_catalog_variables`` for that separate check, and why it's
      separate: these five fields are all known immediately (read once from
      config, never change), so they can be checked as soon as a catalog is
      loaded. ``variables`` (``self.var_dict``) is instead built up
      incrementally, one field type at a time, and a catalog gets loaded
      *in the middle of* that process — checking ``variables`` here would
      risk comparing against a still-incomplete list. It's checked
      separately, once that process finishes.

      This dict is the single definition of "does a catalog match my config,"
      shared by one writer and two readers so they can't drift out of sync:

      * Written into a new catalog's JSON ``"metadata"`` by
        ``quality_check_goes.py`` when it builds one.
      * Read by this dataset's own ``_load_file_catalog`` to decide whether
        an *existing* catalog is safe to load for training — does it cover
        what's being requested?
      * Read by ``quality_check_goes.py``'s pre-overwrite check to decide
        whether it's about to clobber an existing catalog that's already
        identical — is this the *same* catalog it'd be regenerating? This
        guards against silently losing hand-added ``"SKIP"`` rows.

      ``mode``, ``timestep``, ``product``, and ``goes_position`` must match
      a catalog exactly in all three contexts above. ``extent`` is looser
      only for the load check (see ``_load_file_catalog``/``_extent_covers``):
      the catalog's extent only needs to cover the requested extent (equal
      or a superset), since a smaller request is just a subdomain of what
      was already QC'd. The pre-overwrite check still compares ``extent``
      for exact equality, since "is this the same catalog" is a different
      question than "does this catalog cover my request."



