credit.datasets.gen_2.grid_utils#

Everything about horizontal grid geometry for gen2: coordinate-pair detection, rectilinear-vs-curvilinear classification, and GridSchema (the real-coordinate contract for output, mirroring ChannelSchema in channel_utils.py).

find_coord_pair / infer_grid_type#

Small, dependency-free helpers for locating a lon/lat coordinate pair in an xr.Dataset (by name) and classifying it as rectilinear (1D) or curvilinear (2D). Used by the per-dataset grid detection in local.py/era5.py, and by credit.grid.scrip_from_netcdf (SCRIP-format grid generation for ESMF regridding) — this is the shared home for both rather than duplicating the logic in each.

GridSchema#

ForecastWriter previously fabricated output lat/lon from model.image_height/image_width via a global [-90, 90] x [0, 360) linspace — wrong for regional domains and for curvilinear sources (e.g. HRRR), which need a real 2D lat/lon field, not a fabricated 1D one. There is no fabricated fallback anywhere in this pipeline anymore: if a real grid can’t be found or resolved, callers raise rather than guess.

Two distinct grid concepts:

  • Native input grid — each dataset class (LocalDataset, era5.py, goes.py, hrrr.py) exposes the real lat/lon it read from its own files via self.static_metadata["grid"] (a debugging aid, inspectable directly on the live dataset — not necessarily what ends up in the output file). The same dataset class also best-effort persists it to {save_loc}/{source}_grid_schema.nc the moment it’s known, via write_source_grid_schema_if_missing — including from inside a DataLoader worker subprocess, which has filesystem access even though it can’t propagate Python object state back to the main training process.

  • Resolved output grid — what ForecastWriter actually writes. The model produces one flat tensor at one fixed (H, W), so there is exactly one output grid per run: the (single, in practice) source’s native grid, overridden by an active Regridder preblock’s real destination grid when regridding is in play (the common case where regridding is not used is the default: the native grid passes straight through unchanged).

Lifecycle#

Training/rollout setup resolves the schema once — via GridSchema.resolve using the live dataset + preblocks (with a disk fallback to each source’s {source}_grid_schema.nc when the live process never saw the data itself) — and, only when an active Regridder actually changes the grid, saves the result to {save_loc}/output_grid_schema.nc (skipped when no regridder is active: the single source’s own file already is the effective output grid, so a second identical copy would be pure duplication). Later runs (or a re-run without training) load whichever file is present instead of re-resolving, via GridSchema.load_or_resolve.

Scope: rectilinear and curvilinear only (no unstructured). Projection/CRS metadata (e.g. HRRR’s Lambert Conformal Conic) is deferred to a follow-on — plain lat/lon coordinates are CF-valid on their own.

Attributes#

Classes#

GridSchema

The resolved horizontal output grid: shared across every variable/source

Functions#

find_coord_pair(ds)

Find a lon/lat coordinate pair in an xr.Dataset.

infer_grid_type(lat, lon)

Classify a lat/lon coordinate pair as "rectilinear" or "curvilinear".

write_source_grid_schema_if_missing(→ None)

Best-effort persist one source's native grid to {save_loc}/{source}_grid_schema.nc.

Module Contents#

credit.datasets.gen_2.grid_utils.logger#
credit.datasets.gen_2.grid_utils.find_coord_pair(ds)#

Find a lon/lat coordinate pair in an xr.Dataset.

Searches _COORD_CANDIDATES in order across both ds.coords and ds.data_vars.

Return type:

(lon_array, lat_array, lon_name, lat_name)

Raises:

ValueError if no recognised pair is found. –

credit.datasets.gen_2.grid_utils.infer_grid_type(lat, lon)#

Classify a lat/lon coordinate pair as “rectilinear” or “curvilinear”.

1D lat and lon -> rectilinear. 2D lat and lon -> curvilinear.

Returns:

str

Return type:

“rectilinear” or “curvilinear”

Raises:

ValueError if lat/lon are not both 1D or both 2D. –

credit.datasets.gen_2.grid_utils.SOURCE_GRID_SCHEMA_FILENAME = '{source}_grid_schema.nc'#
credit.datasets.gen_2.grid_utils.OUTPUT_GRID_SCHEMA_FILENAME = 'output_grid_schema.nc'#
credit.datasets.gen_2.grid_utils.GridType#
credit.datasets.gen_2.grid_utils.write_source_grid_schema_if_missing(source_name: str, grid: dict[str, Any] | None, save_loc: str | None) → None#

Best-effort persist one source’s native grid to {save_loc}/{source}_grid_schema.nc.

Called from each dataset class right after static_metadata["grid"] is first populated — including from inside a DataLoader worker subprocess under num_workers > 0, since workers have their own filesystem access even though they can’t propagate Python object state back to the main process. Guarded by file existence, so redundant calls (e.g. multiple workers independently resolving the same grid) are cheap after the first successful write. Failures (no save_loc, read-only filesystem, …) are logged and swallowed — this must never break the data-loading path it’s piggybacked onto.

A source can report a native grid_type this module doesn’t support as a resolved output grid (e.g. "unstructured" — see the module docstring’s “Scope” note): that’s expected, not a failure, so it’s logged at info level and skipped here rather than attempting (and swallowing the inevitable failure of) a GridSchema construction. Such a source still needs an active Regridder preblock for GridSchema.resolve to produce a resolvable output grid (see its docstring).

class credit.datasets.gen_2.grid_utils.GridSchema(grid_type: GridType, lat: numpy.ndarray, lon: numpy.ndarray, origin: Literal['native', 'regridded'] = 'native')#

The resolved horizontal output grid: shared across every variable/source in one output file (the model produces one flat tensor at one fixed shape).

Parameters:
  • grid_type – "rectilinear" or "curvilinear".

  • lat – 1D (rectilinear) or 2D (y, x) (curvilinear) latitude array.

  • lon – 1D (rectilinear) or 2D (y, x) (curvilinear) longitude array.

  • origin – "native" (a source’s own grid, unchanged) or "regridded" (an active Regridder preblock’s destination grid). Set by .resolve(); defaults to "native" for direct construction. Used by callers (e.g. the trainer) to decide whether this schema needs its own output_grid_schema.nc — a "native" schema is already fully covered by the source’s own {source}_grid_schema.nc.

grid_type: GridType#
lat#
lon#
origin = 'native'#
classmethod resolve(dataset: Any, ic_preblocks=None, step_preblocks=None, save_loc: str | None = None) → GridSchema#

Resolve the effective output grid from a live dataset + preblocks.

Starts from the dataset’s native grid (static_metadata["grid"], falling back to each source’s persisted {source}_grid_schema.nc in save_loc when the live process doesn’t have it); if an active Regridder preblock is found, its real destination grid is used instead. When no regridder is active (the common case), the native grid passes through unchanged and the returned schema’s origin is "native".

Raises:

ValueError – if no native grid is available, if grids disagree (see _native_grid / _find_regridder), or if the native grid_type (e.g. "unstructured") isn’t directly resolvable and no active Regridder preblock provides a rectilinear/curvilinear destination.

save(path: str) → None#

Write the schema as NetCDF (atomically: temp file + rename).

The staging file is per-process. Every rank — and every DataLoader worker under num_workers > 0 — persists the grid it resolved (see write_source_grid_schema_if_missing), so several processes routinely write the same path at once. A single shared <path>.tmp made them collide inside HDF5, surfacing as [Errno 13] Permission denied for every writer but the first. With one temp file each, the writes are independent and the renames are atomic; the content is identical, so last-writer-wins is harmless.

classmethod load(path: str) → GridSchema#
classmethod load_or_resolve(dataset: Any, ic_preblocks=None, step_preblocks=None, save_loc: str | None = None) → GridSchema | None#

Load output_grid_schema.nc from save_loc (only ever written when regridding was used), else resolve live from dataset (native grid, with a disk fallback to each source’s own {source}_grid_schema.nc, overridden by an active Regridder).

Returns None (with a warning) when neither is possible — callers should treat that as fatal; there is no fabricated fallback.