credit.datasets.gen_2.hrrr_download

credit.datasets.gen_2.hrrr_download#

Standalone utility for downloading HRRR prs GRIB2 data from AWS S3 to local disk.

Downloads are embarrassingly parallel: each timestamp is an independent task dispatched to a ThreadPoolExecutor. Both the grib2 file and its .idx sidecar are downloaded so that HRRRDataset in local mode can use byte-range reads rather than scanning the full file.

Downloaded files follow the native HRRR directory layout used by HRRRDataset in local mode, so they are immediately usable without any renaming:

v3/v4 (2018-07-12+): {base_path}/hrrr.{YYYYMMDD}/conus/hrrr.t{HH}z.{product}f{FF:02d}.grib2
v1/v2 (before):      {base_path}/hrrr.{YYYYMMDD}/hrrr.t{HH}z.{product}f{FF:02d}.grib2

After downloading, switch mode to "local" in the config.

Usage as script:

python -m credit.datasets.gen_2.hrrr_download -c config/my_conf.yaml --num-workers 8

Or programmatically:

from credit.datasets.gen_2.hrrr_download import download_hrrr
download_hrrr(config['data'], num_workers=8, overwrite=False)

Config section used (data.source):

data:
  source:
    Example_HRRR:
      dataset_type: "hrrr"
      product: "wrfprs" # Options: "wrfprs", "wrfnat", "wrfsubh"
      mode: "local"          # mode to use after download
      base_path: "$SCRATCH/data/hrrr"
      forecast_hour: 0
  start_datetime: "2022-01-01"
  end_datetime:   "2022-01-31"
  timestep:       "1h"
  forecast_len:   0

Attributes#

Functions#

download_hrrr(→ None)

Download HRRR grib2 + .idx files from AWS S3 to local disk using obstore.

Module Contents#

credit.datasets.gen_2.hrrr_download.logger#
credit.datasets.gen_2.hrrr_download.download_hrrr(data_config: dict[str, Any], num_workers: int = 4, overwrite: bool = False) → None#

Download HRRR grib2 + .idx files from AWS S3 to local disk using obstore.

Each timestamp is downloaded in parallel using a ThreadPoolExecutor. Both the grib2 file and its .idx sidecar are fetched so that HRRRDataset in mode: "local" can use fast byte-range reads.

Parameters:
  • data_config (dict[str, Any]) – Top-level data config dict (same object passed to HRRRDataset).

  • num_workers (int, optional) – Number of parallel download workers. Each worker opens its own connection. Default 4.

  • overwrite (bool, optional) – Re-download files that already exist on disk. Default False (skip existing files).

Raises:
  • ImportError – If obstore is not installed.

  • KeyError – If the config is missing required fields.

  • ValueError – If product is not a recognised HRRR product.

credit.datasets.gen_2.hrrr_download.parser#