credit.datasets.gen_2.hrrr#

HRRRDataset: PyTorch Dataset for HRRR GRIB2 data.

Supports three HRRR products (VALID_PRODUCTS):

  • "wrfprs" — pressure-level output (default, ~200 MB/file)

  • "wrfnat" — native/hybrid-sigma level output (~200 MB/file, ~65 levels)

  • "wrfsubh" — 15-minute sub-hourly surface output (surface vars only)

Tensor keys follow the pattern {user_provided_name}/{hrrr_product}/{field_type}/{dim}/{varname} where hrrr_product is product-specific:

  • "wrfprs" → {user_provided_name}/wrfprs/{field_type}/{dim}/{varname}

  • "wrfnat" → {user_provided_name}/wrfnat/{field_type}/{dim}/{varname}

  • "wrfsubh" → {user_provided_name}/wrfsubh/{field_type}/2d/{varname}

dim is "3d" for multi-level variables and "2d" for surface variables.

Tensor shapes (before DataLoader batching):

3D variables: (n_levels, 1, y, x) 2D variables: (1, 1, y, x)

The y / x spatial dimensions correspond to HRRR’s native Lambert Conformal Conic grid; if extent is specified they reflect the cropped sub-domain rather than the full CONUS grid (~1059 x 1799).

Two S3 path layouts are handled automatically:

v1/v2 (before 2018-07-12):

s3://noaa-hrrr-bdp-pds/hrrr.{YYYYMMDD}/hrrr.t{HH}z.{product}f{FF:02d}.grib2

v3/v4 (2018-07-12 onward):

s3://noaa-hrrr-bdp-pds/hrrr.{YYYYMMDD}/conus/hrrr.t{HH}z.{product}f{FF:02d}.grib2

GRIB2 reading#

Both local and remote modes use the same .idx + byte-range pipeline:

Remote mode:

  1. Fetch the sidecar .idx inventory (~100 KB) to get exact byte offsets for every GRIB message.

  2. Issue one Obstore get_ranges request to pull all the relevant data fields (based on the byte ranges from the idx file).

Local mode:

  1. Reads the .idx sidecar from disk,

  2. Uses file.seek() + file.read() — identical byte-range approach, no full-file scan.

The .idx sidecar must be present alongside the grib2; download it with hrrr_download.py.

For a typical training sample (5 vars x 6 levels ≈ 30 messages) remote mode transfers ~3 MB instead of ~200 MB (~60-100x reduction).

Variable lookup is driven by VAR_REGISTRY. Extend it at import time to add variables without subclassing:

from credit.datasets.gen_2.hrrr import VAR_REGISTRY
VAR_REGISTRY["MYVAR"] = {
    "shortName": "myvar", "typeOfLevel": "isobaricInhPa",
    "idx_name": "MYVAR", "idx_level": None,
}

Example YAML (wrfprs, local mode):

data:
  source:
    Example_HRRR:  # User-provided name (arbitrary key)
      dataset_type: "HRRR"
      # product: "wrfprs" # Optional for PRS product. Default is "wrfprs".
      mode: "local"
      base_path: "/data/hrrr"
      forecast_hour: 0
      levels: [250, 500, 700, 850, 925, 1000]
      variables:
        prognostic:
          vars_3D: [T, U, V, Q, GH]
          vars_2D: [t2m]
      extent: [-130, -60, 20, 55]

  start_datetime: "2021-06-01"
  end_datetime:   "2021-06-05"
  timestep:       "1h"
  forecast_len:   0

Example YAML (wrfnat, remote mode):

data:
  source:
    Example_HRRR_NAT:  # User-provided name (arbitrary key)
      dataset_type: "HRRR"
      product: "wrfnat" # Options: "wrfprs" (default), "wrfnat", "wrfsubh"
      mode: "remote"
      forecast_hour: 0
      levels: [10, 20, 30, 40, 50]   # hybrid level indices 1-65
      variables:
        prognostic:
          vars_3D: [T, U, V, Q]

  start_datetime: "2022-01-01"
  end_datetime:   "2022-01-31"
  timestep:       "1h"
  forecast_len:   0

Example YAML (wrfsubh, remote mode — 15-min output):

data:
  source:
    Example_HRRR_SUBH:  # User-provided name (arbitrary key)
      dataset_type: "HRRR"
      product: "wrfsubh" # Options: "wrfprs" (default), "wrfnat", "wrfsubh"
      mode: "remote"
      variables:
        prognostic:
          vars_2D: [t2m, sp, refc]

  start_datetime: "2022-01-01 00:15"
  end_datetime:   "2022-01-31 00:00"
  timestep:       "15min"
  forecast_len:   0

Attributes#

Classes#

HRRRDataset

CREDIT Dataset for HRRR GRIB2 data (wrfprs / wrfnat / wrfsubh).

Module Contents#

credit.datasets.gen_2.hrrr.logger#
credit.datasets.gen_2.hrrr.VAR_REGISTRY: dict[str, dict[str, str | None]]#
credit.datasets.gen_2.hrrr.VALID_PRODUCTS#
class credit.datasets.gen_2.hrrr.HRRRDataset(data_config: dict[str, Any], return_target: bool = False)#

Bases: credit.datasets.gen_2.base_dataset.BaseDataset

CREDIT Dataset for HRRR GRIB2 data (wrfprs / wrfnat / wrfsubh).

Implements the same field-type semantics as BaseDataset:

  • prognostic — input at step 0 and target (autoregressive rollout)

  • dynamic_forcing — input at every step; never a target

  • diagnostic — target only

  • static — input at step 0; never a target, applies to all steps

Both modes use pygrib for GRIB2 decoding. Remote mode fetches the .idx sidecar and issues parallel HTTP Range requests — no full file download required.

See module docstring for full output format, tensor shapes, and YAML configuration examples.

Variables:
  • dataset_type – Tensor key - “HRRR”

  • product – Active HRRR product ("HRRR_PRS" / "wrfprs", "HRRR_NAT" / "wrfnat", or "HRRR_SUBH" / "wrfsubh") with default value "HRRR_PRS".

  • datetimes – DatetimeIndex of valid initialization timestamps.

  • static_metadata – Dataset-level metadata for MultiSourceDataset.

dataset_type#
product: VALID_PRODUCTS#
mode: str#
base_path: str | None#
forecast_hour: int#
extent: list[float] | None#
global_levels: list[int] | None#
num_decompress_workers: int#
static_metadata: dict[str, Any]#