credit.models.wxformer.crossformer

Contents

credit.models.wxformer.crossformer#

Attributes#

Classes#

CubeEmbedding

UpBlock

Base class for all neural network modules.

UpBlockPS

Base class for all neural network modules.

CrossEmbedConvBranch

nn.Sequential(ZeroPad2d, Conv2d) for one CrossEmbedLayer kernel branch.

CrossEmbedLayer

Base class for all neural network modules.

DynamicPositionBias

Base class for all neural network modules.

LayerNorm

Base class for all neural network modules.

FeedForward

Base class for all neural network modules.

Attention

Attention module for the CrossFormer model.

Transformer

Base class for all neural network modules.

CrossFormer

Base class for all neural network modules.

Functions#

cast_tuple(val[, length])

apply_spectral_norm(model)

icnr_init_(weight, scale[, init])

ICNR init for a sub-pixel conv feeding nn.PixelShuffle (Aitken et al. 2017).

crossembed_pad_total(→ int)

Total zero-padding a CrossEmbedLayer conv branch applies per spatial dim.

crossembed_out_size(→ int)

Spatial output size of one CrossEmbedLayer conv branch.

migrate_legacy_state_dict(→ dict)

Migrate a legacy checkpoint in place for model, or raise if it can't be.

Module Contents#

credit.models.wxformer.crossformer.logger#
credit.models.wxformer.crossformer.cast_tuple(val, length=1)#
credit.models.wxformer.crossformer.apply_spectral_norm(model)#
class credit.models.wxformer.crossformer.CubeEmbedding(img_size, patch_size, in_chans, embed_dim, norm_layer=nn.LayerNorm)#

Bases: torch.nn.Module

Parameters:
  • img_size – T, Lat, Lon

  • patch_size – T, Lat, Lon

img_size#
patches_resolution#
embed_dim#
proj#
forward(x: torch.Tensor)#
credit.models.wxformer.crossformer.icnr_init_(weight, scale, init=nn.init.kaiming_normal_)#

ICNR init for a sub-pixel conv feeding nn.PixelShuffle (Aitken et al. 2017).

Initializes the conv weight so that, immediately after PixelShuffle(scale), the output equals a nearest-neighbor upsample of a single initialized sub-kernel. All scale**2 sub-pixel channels start identical, which removes the checkerboard grid pattern present at initialization with default init.

Parameters:
  • weight – conv weight of shape (out_ch * scale**2, in_ch, kh, kw).

  • scale – PixelShuffle upscale factor.

  • init – in-place initializer applied to the sub-kernel.

class credit.models.wxformer.crossformer.UpBlock(in_chans, out_chans, num_groups, num_residuals=2, attention_type=None, reduction=32, spatial_kernel=7, fsdp2_shard=True)#

Bases: torch.nn.Module

Base class for all neural network modules.

Your models should also subclass this class.

Modules can also contain other Modules, allowing them to be nested in a tree structure. You can assign the submodules as regular attributes:

import torch.nn as nn
import torch.nn.functional as F


class Model(nn.Module):
    def __init__(self) -> None:
        super().__init__()
        self.conv1 = nn.Conv2d(1, 20, 5)
        self.conv2 = nn.Conv2d(20, 20, 5)

    def forward(self, x):
        x = F.relu(self.conv1(x))
        return F.relu(self.conv2(x))

Submodules assigned in this way will be registered, and will also have their parameters converted when you call to(), etc.

Note

As per the example above, an __init__() call to the parent class must be made before assignment on the child.

Variables:

training (bool) – Boolean represents whether this module is in training or evaluation mode.

conv#
output_channels#
b#
attention = None#
forward(x)#
class credit.models.wxformer.crossformer.UpBlockPS(in_ch, out_ch, num_groups, scale=2, num_residuals=2, fsdp2_shard=True)#

Bases: torch.nn.Module

Base class for all neural network modules.

Your models should also subclass this class.

Modules can also contain other Modules, allowing them to be nested in a tree structure. You can assign the submodules as regular attributes:

import torch.nn as nn
import torch.nn.functional as F


class Model(nn.Module):
    def __init__(self) -> None:
        super().__init__()
        self.conv1 = nn.Conv2d(1, 20, 5)
        self.conv2 = nn.Conv2d(20, 20, 5)

    def forward(self, x):
        x = F.relu(self.conv1(x))
        return F.relu(self.conv2(x))

Submodules assigned in this way will be registered, and will also have their parameters converted when you call to(), etc.

Note

As per the example above, an __init__() call to the parent class must be made before assignment on the child.

Variables:

training (bool) – Boolean represents whether this module is in training or evaluation mode.

conv#
ps#
sharp#
b#
forward(x)#
credit.models.wxformer.crossformer.crossembed_pad_total(kernel: int, stride: int) → int#

Total zero-padding a CrossEmbedLayer conv branch applies per spatial dim.

credit.models.wxformer.crossformer.crossembed_out_size(size: int, kernel: int, stride: int) → int#

Spatial output size of one CrossEmbedLayer conv branch.

Standard conv arithmetic with the layer’s asymmetric zero-padding; with pad_total = kernel - stride this is kernel-independent (floor(size/stride)), which is why the cat() across kernel sizes works. Shared with the credit begin wizard’s grid-spec search so the two cannot drift apart.

class credit.models.wxformer.crossformer.CrossEmbedConvBranch(*args: torch.nn.modules.module.Module)#
class credit.models.wxformer.crossformer.CrossEmbedConvBranch(arg: collections.OrderedDict[str, torch.nn.modules.module.Module])

Bases: torch.nn.Sequential

nn.Sequential(ZeroPad2d, Conv2d) for one CrossEmbedLayer kernel branch.

A distinctly-named nn.Sequential subclass – behaves identically to a plain nn.Sequential (same forward, same auto-indexed “0”/”1” children, so state_dict keys are unaffected) – purely so domain-parallel conversion (credit/domain_parallel/convert.py) can recognize this exact pattern via isinstance and replace it as a unit instead of matching only the inner Conv2d. See DomainParallelCrossEmbedBranch for why that distinction matters: the inner Conv2d alone doesn’t carry enough information to redo this branch’s padding correctly under domain parallelism.

class credit.models.wxformer.crossformer.CrossEmbedLayer(dim_in, dim_out, kernel_sizes, stride=2)#

Bases: torch.nn.Module

Base class for all neural network modules.

Your models should also subclass this class.

Modules can also contain other Modules, allowing them to be nested in a tree structure. You can assign the submodules as regular attributes:

import torch.nn as nn
import torch.nn.functional as F


class Model(nn.Module):
    def __init__(self) -> None:
        super().__init__()
        self.conv1 = nn.Conv2d(1, 20, 5)
        self.conv2 = nn.Conv2d(20, 20, 5)

    def forward(self, x):
        x = F.relu(self.conv1(x))
        return F.relu(self.conv2(x))

Submodules assigned in this way will be registered, and will also have their parameters converted when you call to(), etc.

Note

As per the example above, an __init__() call to the parent class must be made before assignment on the child.

Variables:

training (bool) – Boolean represents whether this module is in training or evaluation mode.

convs#
forward(x)#
credit.models.wxformer.crossformer.migrate_legacy_state_dict(model: torch.nn.Module, state_dict: dict) → dict#

Migrate a legacy checkpoint in place for model, or raise if it can’t be.

The same migrations the load_state_dict pre-hooks perform, applied up front. Needed for the FSDP2 path: set_model_state_dict reconciles the checkpoint against the model’s current parameter names before calling load_state_dict, so the hooks fire too late to help there.

Walks the model to find real CrossEmbedLayer instances rather than pattern- matching key names, so sibling architectures that define their own unwrapped convs (camulator, crossformer_downscaling) are untouched.

Parameters:
  • model – the instantiated model the checkpoint is destined for. Pass the unwrapped module (getattr(model, "module", model)) so the names line up with the checkpoint’s.

  • state_dict – checkpoint state dict; mutated in place.

Returns:

the same state_dict, for convenience.

Return type:

dict

Raises:

RuntimeError – if the checkpoint holds the removed ConvTranspose2d decoder.

class credit.models.wxformer.crossformer.DynamicPositionBias(dim)#

Bases: torch.nn.Module

Base class for all neural network modules.

Your models should also subclass this class.

Modules can also contain other Modules, allowing them to be nested in a tree structure. You can assign the submodules as regular attributes:

import torch.nn as nn
import torch.nn.functional as F


class Model(nn.Module):
    def __init__(self) -> None:
        super().__init__()
        self.conv1 = nn.Conv2d(1, 20, 5)
        self.conv2 = nn.Conv2d(20, 20, 5)

    def forward(self, x):
        x = F.relu(self.conv1(x))
        return F.relu(self.conv2(x))

Submodules assigned in this way will be registered, and will also have their parameters converted when you call to(), etc.

Note

As per the example above, an __init__() call to the parent class must be made before assignment on the child.

Variables:

training (bool) – Boolean represents whether this module is in training or evaluation mode.

layers#
forward(x)#
class credit.models.wxformer.crossformer.LayerNorm(dim, eps=1e-05)#

Bases: torch.nn.Module

Base class for all neural network modules.

Your models should also subclass this class.

Modules can also contain other Modules, allowing them to be nested in a tree structure. You can assign the submodules as regular attributes:

import torch.nn as nn
import torch.nn.functional as F


class Model(nn.Module):
    def __init__(self) -> None:
        super().__init__()
        self.conv1 = nn.Conv2d(1, 20, 5)
        self.conv2 = nn.Conv2d(20, 20, 5)

    def forward(self, x):
        x = F.relu(self.conv1(x))
        return F.relu(self.conv2(x))

Submodules assigned in this way will be registered, and will also have their parameters converted when you call to(), etc.

Note

As per the example above, an __init__() call to the parent class must be made before assignment on the child.

Variables:

training (bool) – Boolean represents whether this module is in training or evaluation mode.

eps = 1e-05#
g#
b#
forward(x)#
class credit.models.wxformer.crossformer.FeedForward(dim, mult=4, dropout=0.0, tp_col='layers.1', tp_row='layers.4')#

Bases: torch.nn.Module

Base class for all neural network modules.

Your models should also subclass this class.

Modules can also contain other Modules, allowing them to be nested in a tree structure. You can assign the submodules as regular attributes:

import torch.nn as nn
import torch.nn.functional as F


class Model(nn.Module):
    def __init__(self) -> None:
        super().__init__()
        self.conv1 = nn.Conv2d(1, 20, 5)
        self.conv2 = nn.Conv2d(20, 20, 5)

    def forward(self, x):
        x = F.relu(self.conv1(x))
        return F.relu(self.conv2(x))

Submodules assigned in this way will be registered, and will also have their parameters converted when you call to(), etc.

Note

As per the example above, an __init__() call to the parent class must be made before assignment on the child.

Variables:

training (bool) – Boolean represents whether this module is in training or evaluation mode.

layers#
forward(x)#
class credit.models.wxformer.crossformer.Attention(dim, attn_type, window_size, dim_head=32, dropout=0.0, tp_col='to_qkv', tp_row='to_out')#

Bases: torch.nn.Module

Attention module for the CrossFormer model.

Tensor parallelism opt-in: to_qkv is column-parallel (output sharded), to_out is row-parallel (input sharded, all_reduce).

This module performs either short-range or long-range attention on the input tensor. It uses a dynamic positional bias to incorporate relative positional information.

Parameters:
  • dim (int) – Input dimension.

  • attn_type (str) – Type of attention, either “short” or “long”.

  • window_size (int) – Size of the attention window.

  • dim_head (int, optional) – Dimension of each attention head. Defaults to 32.

  • dropout (float, optional) – Dropout rate. Defaults to 0.0.

heads#
scale = 0.1767766952966369#
attn_type#
window_size#
norm#
dropout#
to_qkv#
to_out#
dpb#
forward(x)#

Forward pass of the Attention module.

Parameters:

x (torch.Tensor) – Input tensor of shape (batch, dim, height, width).

Returns:

Output tensor of the same shape as input.

Return type:

torch.Tensor

class credit.models.wxformer.crossformer.Transformer(dim, *, local_window_size, global_window_size, depth=4, dim_head=32, attn_dropout=0.0, ff_dropout=0.0, fsdp2_shard=True)#

Bases: torch.nn.Module

Base class for all neural network modules.

Your models should also subclass this class.

Modules can also contain other Modules, allowing them to be nested in a tree structure. You can assign the submodules as regular attributes:

import torch.nn as nn
import torch.nn.functional as F


class Model(nn.Module):
    def __init__(self) -> None:
        super().__init__()
        self.conv1 = nn.Conv2d(1, 20, 5)
        self.conv2 = nn.Conv2d(20, 20, 5)

    def forward(self, x):
        x = F.relu(self.conv1(x))
        return F.relu(self.conv2(x))

Submodules assigned in this way will be registered, and will also have their parameters converted when you call to(), etc.

Note

As per the example above, an __init__() call to the parent class must be made before assignment on the child.

Variables:

training (bool) – Boolean represents whether this module is in training or evaluation mode.

layers#
forward(x)#
class credit.models.wxformer.crossformer.CrossFormer(image_height: int = 640, patch_height: int = 1, image_width: int = 1280, patch_width: int = 1, frames: int = 2, output_frames: int = 1, channels: int = 4, surface_channels: int = 7, input_only_channels: int = 3, output_only_channels: int = 0, levels: int = 15, dim: tuple = (64, 128, 256, 512), depth: tuple = (2, 2, 8, 2), dim_head: int = 32, global_window_size: tuple = (5, 5, 2, 1), local_window_size: int = 10, cross_embed_kernel_sizes: tuple = ((4, 8, 16, 32), (2, 4), (2, 4), (2, 4)), cross_embed_strides: tuple = (4, 2, 2, 2), attn_dropout: float = 0.0, ff_dropout: float = 0.0, use_spectral_norm: bool = True, attention_type: str = None, interp: bool = True, upsample_with_ps: bool = True, padding_conf: dict = None, post_conf: dict = None, **kwargs)#

Bases: credit.models.base_model.BaseModel

Base class for all neural network modules.

Your models should also subclass this class.

Modules can also contain other Modules, allowing them to be nested in a tree structure. You can assign the submodules as regular attributes:

import torch.nn as nn
import torch.nn.functional as F


class Model(nn.Module):
    def __init__(self) -> None:
        super().__init__()
        self.conv1 = nn.Conv2d(1, 20, 5)
        self.conv2 = nn.Conv2d(20, 20, 5)

    def forward(self, x):
        x = F.relu(self.conv1(x))
        return F.relu(self.conv2(x))

Submodules assigned in this way will be registered, and will also have their parameters converted when you call to(), etc.

Note

As per the example above, an __init__() call to the parent class must be made before assignment on the child.

Variables:

training (bool) – Boolean represents whether this module is in training or evaluation mode.

image_height = 640#
image_width = 1280#
patch_height = 1#
patch_width = 1#
upsample_with_ps = True#
frames = 2#
output_frames = 1#
channels = 4#
surface_channels = 7#
levels = 15#
use_spectral_norm = True#
use_interp = True#
use_padding#
use_post_block#
input_only_channels = 3#
base_input_channels = 70#
input_channels = 140#
base_output_channels = 67#
output_channels = 67#
layers#
cube_embedding#
up_block1#
up_block2#
up_block3#
up_block4#
forward(x)#
rk4(x)#
credit.models.wxformer.crossformer.image_height = 180#