GB-LSR: Local Spectral Decoding with a Learned Global Bandwidth
Trained weights for GB-LSR, from GB-LSR: Local Spectral Decoding with a Learned Global Bandwidth for Arbitrary-Scale Super-Resolution. GB-LSR partitions the image domain into a fixed grid of non-overlapping square patches. Each patch carries coefficients for a truncated Fourier basis, predicted by a single linear projection from shared convolutional-encoder features, and one trainable scalar bandwidth is shared across every patch and every image. As in earlier local spectral decoders, decoding at a continuous coordinate is a fixed-size basis contraction whose cost is set by the spectral cutoff; GB-LSR learns the bandwidth of that basis instead of fixing it.
Code: https://github.com/KempnerInstitute/gblsr.
Models
One checkpoint per configuration, from training seed 0. The paper's tables report means over three training seeds; the PSNR of the released checkpoints differs from those means by at most 0.021 dB.
Arbitrary-scale super-resolution
RDN encoder, trained on DIV2K at scales from x1 to x4. PSNR-Y (dB) and LPIPS at x4. Learned s: the bandwidth after training.
| Folder | Model | Params (M) | Set5 | Set14 | B100 | Urban100 | DIV2Kval | Mean LPIPS (lower is better) | Learned s |
|---|---|---|---|---|---|---|---|---|---|
gblsr-scalar-asr |
GB-LSR-Scalar-ASR (base) | 22.02 | 32.19 | 28.72 | 27.66 | 26.43 | 30.67 | 0.2761 | 0.881 |
gblsr-scalar-asr-noLE |
GB-LSR-Scalar-ASR-noLE | 22.02 | 32.19 | 28.71 | 27.67 | 26.45 | 30.68 | 0.2747 | 0.794 |
gblsr-scalar-asr-nf96-noLE |
GB-LSR-Scalar-ASR-nf96+noLE | 24.93 | 32.23 | 28.74 | 27.68 | 26.47 | 30.70 | 0.2743 | 0.795 |
gblsr-scalar-asr-nf48-noLE |
GB-LSR-Scalar-ASR-nf48+noLE | 20.61 | 32.18 | 28.72 | 27.67 | 26.40 | 30.66 | not scored | 0.784 |
noLE: trained and evaluated without the local ensemble. nf96, nf48: RDN encoder with 96 or 48 base feature channels instead of 64. Scored as in the paper: PIL bicubic low-resolution inputs; PSNR-Y on the BT.601 luminance with a border as wide as the scale removed; Set5, Set14, B100, and Urban100 ground truth cropped at the bottom and right to a multiple of 12 pixels; DIV2Kval is the DIV2K validation split, scored on luminance and so not comparable with published DIV2K numbers; LPIPS-AlexNet on the RGB output.
Latency at x4 (ms per image) and peak allocated GPU memory, from the paper's timing session on one NVIDIA H200 at batch size 1: median of 50 timed passes, averaged over the images and the three seeds' models. Speed: the base model's latency divided by the row's, geometric mean over the three test sets.
| Folder | Set14 | B100 | Urban100 | Peak memory, Urban100 (MiB) | Speed vs base |
|---|---|---|---|---|---|
gblsr-scalar-asr |
34.8 | 25.6 | 109.4 | 899 | 1.00x |
gblsr-scalar-asr-noLE |
17.7 | 14.6 | 52.5 | 900 | 1.93x |
gblsr-scalar-asr-nf96-noLE |
20.1 | 16.0 | 55.7 | 1,109 | 1.76x |
gblsr-scalar-asr-nf48-noLE |
18.8 | 14.4 | 52.9 | 730 | 1.90x |
Design-study models (native reconstruction)
GB-LSR-Scalar at patch size 32 and 16, spectral cutoff 16, feature width 128, trained on DTD and DIV2K. PSNR-RGB (dB) on images standardized to 256x256 as in the paper: a center crop, after upsampling any image with a side under 256. Latent: encoder feature values per image. Learned s is the implementation's value, as in the paper's design study (Appendix A.1); the bandwidth of the paper's Section 3.1 is this value times P/(P-1) for patch size P. Latency and memory as above, per 256x256 image. The design study compares these models with other GB-LSR variants only.
| Folder | Patch size | Params | Latent | Kodak | Set14 | Urban100 | Learned s | Latency (ms) | Peak memory (MiB) |
|---|---|---|---|---|---|---|---|---|---|
gblsr-scalar-p32 |
32 | 989,955 | 8,192 | 26.32 | 25.34 | 22.55 | 0.758 | 1.367 | 117.8 |
gblsr-scalar-p16 |
16 | 842,115 | 32,768 | 32.44 | 31.01 | 28.09 | 0.658 | 1.358 | 117.3 |
Results in brief
Against the authors' released LIIF, LTE, and SRNO checkpoints on the same RDN encoder, scored under one protocol and timed in one session per scale, the base model runs 1.25x faster than LIIF-RDN at x4 and as fast as SRNO-RDN, and trails the three by 0.07 to 0.79 dB PSNR-Y in distribution. Without the local ensemble, PSNR-Y stays within seed variation and the speedup over LIIF-RDN rises to 2.41x at x4 and 3.00x at x8, at the cost of value jumps at cell boundaries of 0.22 gray levels (of 255) on average at x4. The GB-LSR models scored with LPIPS (base, noLE, nf96+noLE) have the highest (worst) mean LPIPS at x4 of the methods compared. The paper gives the full comparison and its scope.
Usage
Install the package at the commit these files were checked with, and the Hub client:
pip install "git+https://github.com/KempnerInstitute/gblsr@6a3cd604c2ba9e868daf01727551d70a76932e1c" huggingface_hub
Arbitrary-scale super-resolution:
import json, torch
from huggingface_hub import snapshot_download
from safetensors.torch import load_file
from gblsr import GBLSRScalarASR
root = snapshot_download("KempnerInstituteAI/gblsr", revision="arxiv-v2", allow_patterns="gblsr-scalar-asr/*")
cfg = json.load(open(f"{root}/gblsr-scalar-asr/config.json"))
model = GBLSRScalarASR(encoder_cfg=cfg["encoder_cfg"], decoder_cfg=cfg["decoder_cfg"]).eval()
model.load_state_dict(load_file(f"{root}/gblsr-scalar-asr/model.safetensors"))
lr = torch.rand(1, 3, 64, 64) # RGB in [0, 1]
with torch.no_grad(): # any output size; tile_q chunks the decoder's queries
hr = model.predict_full(lr, H_q=256, W_q=256, tile_q=30000).clamp(0, 1)
Design-study model:
import json, torch
from huggingface_hub import snapshot_download
from safetensors.torch import load_file
from gblsr import BasisConfig, EncoderConfig, ModelConfig, build_model
root = snapshot_download("KempnerInstituteAI/gblsr", revision="arxiv-v2", allow_patterns="gblsr-scalar-p32/*")
b = json.load(open(f"{root}/gblsr-scalar-p32/config.json"))["build"]
model = build_model(
ModelConfig(arm=b["arm"], image_size=b["image_size"], patch_size=b["patch_size"],
basis=BasisConfig(patch_size=b["patch_size"], p_max=b["p_max"], s_e_range=tuple(b["s_e_range"])),
encoder=EncoderConfig(d_feat=b["d_feat"])),
bandwidth_mode=b["bandwidth_mode"], adapt_order=b["adapt_order"]).eval()
model.load_state_dict(load_file(f"{root}/gblsr-scalar-p32/model.safetensors"))
with torch.no_grad():
recon = model(torch.rand(1, 3, 256, 256))["recon"].clamp(0, 1) # 256x256 RGB
examples/infer.py runs either kind of model on an image file. SHA256SUMS gives the sha256 of this
card, the example, and every model file.
Training
Super-resolution: 1,000,000 steps on DIV2K with an L1 loss on 48x48 low-resolution patches; each example is a square crop of side round(48k) pixels with k drawn uniformly from [1, 4], downsampled to 48x48 with PIL bicubic, and the loss is taken on 2,304 of its pixels; batch size 16, random horizontal flips and 90 degree rotations, Adam at a learning rate of 1e-4 halved at 200,000, 400,000, 600,000, and 800,000 steps. The bandwidth starts at s = 1.0.
Design study: 1,000,000 steps on 256x256 crops from a mixture of DTD and DIV2K, batch size 8, mean squared error loss, AdamW (beta = (0.9, 0.95), no weight decay) at a constant learning rate of 2e-4 with gradient-norm clipping at 1.0. The bandwidth passes through a log-space sigmoid onto [0.25, 2.0] and starts at s = 0.707.
License and attribution
Released under CC BY-NC 4.0. The super-resolution models are trained on
DIV2K, released "for academic research purpose only";
the design-study models are trained on DTD, released
for research purposes, and DIV2K. Use of the weights must respect those terms. The gblsr code is
BSD-3-Clause.
Versions
- arXiv v2 (this version, tag
arxiv-v2): the 2,000-step native modelgblsr-scalaris replaced by the 1,000,000-step design-study modelsgblsr-scalar-p32andgblsr-scalar-p16; the four super-resolution weight files are unchanged. - arXiv v1: tag
arxiv-v1.
Citation
@article{shad2026gblsr,
title = {GB-LSR: Local Spectral Decoding with a Learned Global Bandwidth for
Arbitrary-Scale Super-Resolution},
author = {Shad, Max and Khoshnevis, Naeem},
journal = {arXiv preprint arXiv:2606.19617},
year = {2026}
}