credit.datasets.gen_2.hrrr#
HRRRDataset: PyTorch Dataset for HRRR GRIB2 data.
Supports three HRRR products (VALID_PRODUCTS):
"wrfprs"— pressure-level output (default, ~200 MB/file)"wrfnat"— native/hybrid-sigma level output (~200 MB/file, ~65 levels)"wrfsubh"— 15-minute sub-hourly surface output (surface vars only)
Tensor keys follow the pattern {user_provided_name}/{hrrr_product}/{field_type}/{dim}/{varname}
where hrrr_product is product-specific:
"wrfprs"→{user_provided_name}/wrfprs/{field_type}/{dim}/{varname}"wrfnat"→{user_provided_name}/wrfnat/{field_type}/{dim}/{varname}"wrfsubh"→{user_provided_name}/wrfsubh/{field_type}/2d/{varname}
dim is "3d" for multi-level variables and "2d" for surface variables.
- Tensor shapes (before DataLoader batching):
3D variables:
(n_levels, 1, y, x)2D variables:(1, 1, y, x)
The y / x spatial dimensions correspond to HRRR’s native Lambert
Conformal Conic grid; if extent is specified they reflect the cropped
sub-domain rather than the full CONUS grid (~1059 x 1799).
Two S3 path layouts are handled automatically:
- v1/v2 (before 2018-07-12):
s3://noaa-hrrr-bdp-pds/hrrr.{YYYYMMDD}/hrrr.t{HH}z.{product}f{FF:02d}.grib2
- v3/v4 (2018-07-12 onward):
s3://noaa-hrrr-bdp-pds/hrrr.{YYYYMMDD}/conus/hrrr.t{HH}z.{product}f{FF:02d}.grib2
GRIB2 reading#
Both local and remote modes use the same .idx + byte-range pipeline:
Remote mode:
Fetch the sidecar
.idxinventory (~100 KB) to get exact byte offsets for every GRIB message.Issue one Obstore get_ranges request to pull all the relevant data fields (based on the byte ranges from the idx file).
Local mode:
Reads the
.idxsidecar from disk,Uses
file.seek()+file.read()— identical byte-range approach, no full-file scan.
The .idx sidecar must be present alongside the grib2;
download it with hrrr_download.py.
For a typical training sample (5 vars x 6 levels ≈ 30 messages) remote mode transfers ~3 MB instead of ~200 MB (~60-100x reduction).
Variable lookup is driven by VAR_REGISTRY. Extend it at import
time to add variables without subclassing:
from credit.datasets.gen_2.hrrr import VAR_REGISTRY
VAR_REGISTRY["MYVAR"] = {
"shortName": "myvar", "typeOfLevel": "isobaricInhPa",
"idx_name": "MYVAR", "idx_level": None,
}
Example YAML (wrfprs, local mode):
data:
source:
Example_HRRR: # User-provided name (arbitrary key)
dataset_type: "HRRR"
# product: "wrfprs" # Optional for PRS product. Default is "wrfprs".
mode: "local"
base_path: "/data/hrrr"
forecast_hour: 0
levels: [250, 500, 700, 850, 925, 1000]
variables:
prognostic:
vars_3D: [T, U, V, Q, GH]
vars_2D: [t2m]
extent: [-130, -60, 20, 55]
start_datetime: "2021-06-01"
end_datetime: "2021-06-05"
timestep: "1h"
forecast_len: 0
Example YAML (wrfnat, remote mode):
data:
source:
Example_HRRR_NAT: # User-provided name (arbitrary key)
dataset_type: "HRRR"
product: "wrfnat" # Options: "wrfprs" (default), "wrfnat", "wrfsubh"
mode: "remote"
forecast_hour: 0
levels: [10, 20, 30, 40, 50] # hybrid level indices 1-65
variables:
prognostic:
vars_3D: [T, U, V, Q]
start_datetime: "2022-01-01"
end_datetime: "2022-01-31"
timestep: "1h"
forecast_len: 0
Example YAML (wrfsubh, remote mode — 15-min output):
data:
source:
Example_HRRR_SUBH: # User-provided name (arbitrary key)
dataset_type: "HRRR"
product: "wrfsubh" # Options: "wrfprs" (default), "wrfnat", "wrfsubh"
mode: "remote"
variables:
prognostic:
vars_2D: [t2m, sp, refc]
start_datetime: "2022-01-01 00:15"
end_datetime: "2022-01-31 00:00"
timestep: "15min"
forecast_len: 0
Attributes#
Classes#
CREDIT Dataset for HRRR GRIB2 data (wrfprs / wrfnat / wrfsubh). |
Module Contents#
- credit.datasets.gen_2.hrrr.logger#
- credit.datasets.gen_2.hrrr.VAR_REGISTRY: dict[str, dict[str, str | None]]#
- credit.datasets.gen_2.hrrr.VALID_PRODUCTS#
- class credit.datasets.gen_2.hrrr.HRRRDataset(data_config: dict[str, Any], return_target: bool = False)#
Bases:
credit.datasets.gen_2.base_dataset.BaseDatasetCREDIT Dataset for HRRR GRIB2 data (wrfprs / wrfnat / wrfsubh).
Implements the same field-type semantics as BaseDataset:
prognostic— input at step 0 and target (autoregressive rollout)dynamic_forcing— input at every step; never a targetdiagnostic— target onlystatic— input at step 0; never a target, applies to all steps
Both modes use
pygribfor GRIB2 decoding. Remote mode fetches the.idxsidecar and issues parallel HTTP Range requests — no full file download required.See module docstring for full output format, tensor shapes, and YAML configuration examples.
- Variables:
dataset_type – Tensor key - “HRRR”
product – Active HRRR product (
"HRRR_PRS" / "wrfprs","HRRR_NAT" / "wrfnat", or"HRRR_SUBH" / "wrfsubh") with default value"HRRR_PRS".datetimes – DatetimeIndex of valid initialization timestamps.
static_metadata – Dataset-level metadata for MultiSourceDataset.
- dataset_type#
- product: VALID_PRODUCTS#
- mode: str#
- base_path: str | None#
- forecast_hour: int#
- extent: list[float] | None#
- global_levels: list[int] | None#
- num_decompress_workers: int#
- static_metadata: dict[str, Any]#