ProPolar: Progressive Polar Decomposition
for Implicit Neural Representations

Pureum Kim, Younggeon Ryu, Dongyoon Lee,
Hae Beom Lee†, Kyong Hwan Jin†
Korea University, IPA Lab
†Corresponding authors
NeurIPS 2026
Overview of ProPolar with rank and learning-rate schedules, low-rank polar orthogonalization, vanilla NeRF tables, reconstructions for image fitting, SDF, NeRF, and fine-tuning, and effective-rank and frequency-band plots.
Click to zoom

ProPolar at a glance. Rank and learning-rate schedules with low-rank polar orthogonalization, vanilla NeRF results on LLFF and Synthetic scenes, reconstructions for image fitting, SDF, NeRF, and fine-tuning, and effective-rank and frequency-band analyses.

Abstract

Implicit neural representations (INRs) capture global low-frequency structure during the early stage of training and then refine localized high-frequency details, afterwards. However, standard optimizers are agnostic to the coarse-to-fine learning behavior of INRs. Such optimizers accumulate a gradient matrix in momentum buffer and retain components misaligned with the dominant low-frequency directions in the early stage. We propose a progressively growing rank scheduler and a rank-aware learning-rate scheduler for enhancing a subspace-based momentum optimization, where a momentum matrix is updated within a low-rank subspace. The rank schedule expands the subspace during training, and the scheduler holds the peak learning rate through rank growth so that the added subspace contributes to fine-detail recovery. Experiments on image fitting, single-image super-resolution (SISR), and neural radiance field optimization show that the proposed method improves convergence and peak reconstruction quality over competitive optimizers. The largest gains appear in the high-frequency refinement stage, mirroring the coarse-to-fine progression that the rank schedule targets. A significant study on hyperparameter optimization (HPO) on NeRF further confirms that these gains persist under matched search budgets.

TL;DR ProPolar orthogonalizes each momentum matrix within a low-rank subspace, grows the subspace rank by a cosine schedule, and holds the peak learning rate until the rank saturates. Early updates concentrate on dominant directions tied to low-frequency structure, and later updates admit directions for fine detail.

Why ProPolar? Fixed Rank in Coarse-to-Fine Training

INR training follows a coarse-to-fine progression. Global low-frequency structure forms first, and high-frequency detail follows. Capacity demand grows with training, so a low-rank update subspace suits the early stage and a high-rank subspace suits the later stage.

Standard optimizers keep one update representation for the whole run.

They are agnostic to the training stage. A full-rank momentum matrix accumulates every gradient component, including components misaligned with the dominant low-frequency directions in the early stage. A fixed rank assigns capacity to less informative directions early and constrains the update subspace during detail refinement.

ProPolar aligns the update subspace with the training stage. The subspace rank starts from a data-dependent value and grows by a cosine schedule. A rank-aware learning-rate schedule holds the peak step size through rank growth so that the added directions contribute to fine-detail recovery.

How ProPolar Works

ProPolar overview. The rank schedule grows over training, the rank-aware WSD schedule holds the peak learning rate, and the momentum matrix is orthogonalized in a low-rank subspace.
Overview. The rank schedule \(r_t\) expands the momentum subspace from dominant low-rank directions to higher-rank directions. The orthonormal basis \(Q_t\in\mathbb{R}^{m\times r_t}\) defines the active subspace for projected polar orthogonalization. The rank-aware WSD schedule \(\gamma_t\) holds the peak step size during rank growth, and decay begins after the rank saturates.

ProPolar updates every hidden-layer weight matrix \(W\in\mathbb{R}^{m\times n}\) with the five steps below and applies Adam to non-matrix parameters.

Step 1. Accumulate Nesterov momentum.

For each hidden-layer matrix, ProPolar keeps a momentum matrix \(M_t\) with coefficient \(\beta\).

\[ M_t = \beta M_{t-1} + \nabla_W \mathcal{L}_t\big(W_{t-1} - \beta M_{t-1}\big) \]

Step 2. Set the rank for this iteration.

During the first \(T_0\) iterations, the rank \(r_t\) is the smallest \(k\) whose spectral energy \(E_t(k)\) reaches the threshold \(\rho\), and \(r^{\mathrm{init}} = \lceil \operatorname{median}(r_0,\dots,r_{T_0-1}) \rceil\). The rank then grows from \(r^{\mathrm{init}}\) to \(r^{\max}\) over \(T_g\) iterations by a cosine schedule and stays at \(r^{\max}\) afterward.

\[ E_t(k) = \frac{\sum_{i=1}^{k}\sigma_{t,i}^2}{\sum_{i=1}^{d}\sigma_{t,i}^2}, \qquad r_t = \min\big\{r^{\max},\ \min\{k : E_t(k)\ge\rho\}\big\}, \qquad 0\le t\lt T_0 \] \[ r_t = \Big\lceil r^{\mathrm{init}} + \big(r^{\max}-r^{\mathrm{init}}\big)\,\frac{1-\cos(\pi s_t)}{2} \Big\rceil, \qquad s_t = \min\Big\{1,\ \frac{t-T_0}{T_g-1}\Big\}, \qquad t\ge T_0 \]

Step 3. Build a low-rank basis.

A Gaussian sketch \(\Omega_t\) and a thin QR decomposition give an orthonormal basis \(Q_t\) for the dominant column space of \(M_t\). By randomized SVD analysis, \(Q_tQ_t^{\top}M_t\) approximates the best rank-\(r_t\) truncation of \(M_t\).

\[ \Omega_t \sim \mathcal{N}(0,1)^{n\times r_t}, \qquad Q_t = \mathrm{qf}\big(M_t\Omega_t\big) \in \mathbb{R}^{m\times r_t} \]

Step 4. Orthogonalize in the reduced space.

Five Newton–Schulz iterations approximate the polar factor \(\mathcal{P}\) of the \(r_t\times n\) reduced matrix, and \(Q_t\) maps the result back to the parameter space. Full-space orthogonalization amplifies every singular direction equally, including tail directions misaligned with the dominant update subspace. The reduced polar factor restricts the update to the span of \(Q_t\).

\[ \hat M_t = Q_t\,\mathcal{P}\big(Q_t^{\top} M_t\big), \qquad W_{t+1} = (1-\eta_t\lambda)\,W_t - \eta_t\,\hat M_t \]

Step 5. Hold the peak learning rate until rank saturation.

The rank-aware warmup-stable-decay schedule, raWSD, sets \(\eta_t = \gamma_t\,\eta_{\max}\). Decay starts after all tracked layers reach \(r^{\max}\), so directions added during rank growth receive full-size steps.

\[ T_d = \min\big\{T-1,\ \min\{t\ge T_0 : r_t = r^{\max}\} + \Delta_{\mathrm{sat}}\big\} \] \[ \gamma_t = \begin{cases} \dfrac{t+1}{T_0}, & 0\le t\lt T_0 \\[6pt] 1, & T_0\le t\lt T_d \\[6pt] \gamma_{\min} + (1-\gamma_{\min})\,\dfrac{T-1-t}{\max\{1,\ T-1-T_d\}}, & T_d\le t\lt T \end{cases} \]

Explore the two schedules.

Move the sliders to see how the decay start follows rank saturation. Before \(T_0\) the rank follows the energy threshold, and the plot draws it at \(r^{\mathrm{init}}\).

Rank saturates at iteration –. Learning-rate decay starts at iteration –.

Illustrative settings with T = 5,000, rmax = 256, and γmin = 0.

Spectral Insight

Two measurements on image overfitting motivate the rank schedule, and two ablations test its design.

Effective rank grows late in training.

The normalized entropy effective rank of the momentum matrix \(M_t\) and the update \(\Delta W_t = W_{t+1} - W_t\) stays low for most of training and rises in later iterations. A low value means that a few leading singular values account for most of the spectrum.

A low-rank approximation suffices in the early stage, and more singular directions become non-negligible later. This pattern matches a rank that grows over training.

Normalized entropy effective rank of the momentum matrix and the update over 5,000 iterations for three layers. Both stay low and rise near the end.
Normalized entropy effective rank of \(M_t\) in the top panel and \(\Delta W_t\) in the bottom panel for three hidden layers.

Head directions carry fine detail.

Let \(\bar M_t = U_tV_t^{\top}\) denote the full polar factor of \(M_t\) and \(\bar M_{t,k} = U_{t,k}V_{t,k}^{\top}\) its rank-\(k\) head. With \(k = 128\), head-only updates reduce low-frequency residual energy by orders of magnitude and continue to reduce mid- and high-frequency residuals. Tail-only updates stay near the initial level in mid and high bands. Update components for fine-detail recovery concentrate in the head.

Relative residual energy in low, mid, and high frequency bands for head-only and tail-only updates. Head-only curves fall steeply in all bands.
Relative residual energy in low, mid, and high frequency bands over 5,000 iterations of image overfitting on Kodak, for head-only and tail-only updates.

Increasing rank gives the best reconstruction.

Four rank schedules share the image overfitting setting. Increasing rank reaches 27.12 dB, followed by fixed rank at 26.58 dB, full rank at 26.47 dB, and decreasing rank at 26.06 dB.

Ground truth, full image
Ground truth, region in the red box
Ground truthReference
Increasing rank, full image
Increasing rank, region in the red box
Increasing rank27.12 dB
Fixed rank, full image
Fixed rank, region in the red box
Fixed rank26.58 dB
Full rank, full image
Full rank, region in the red box
Full rank26.47 dB
Decreasing rank, full image
Decreasing rank, region in the red box
Decreasing rank26.06 dB

The second row magnifies the region in the red box.

raWSD pairs with rank growth.

With raWSD, ProPolar improves over its non-WSD counterpart in all eight NeRF scenes. Muon with WSD drops in six of eight scenes. Holding the peak step size during rank growth matches the learning-rate schedule to the expanding subspace.

PSNR per NeRF scene for ProPolar with and without WSD. WSD raises PSNR in every scene.
ProPolar with and without WSD.
PSNR per NeRF scene for Muon with and without WSD. WSD lowers PSNR in six of eight scenes.
Muon with and without WSD.

Results

ProPolar changes the optimizer and keeps the architecture and training configuration of each baseline. Light-blue cells mark ProPolar. Bold marks the best value and underline marks the second best within each comparison.

Image fitting on Kodak24

Mean over 24 images.

ModelOptimizerPSNR ↑SSIM ↑LPIPS ↓
ReLU MLPAdam22.550.54040.5835
ProPolar28.700.78550.3417
SIRENAdam33.670.91040.0754
ProPolar35.050.93200.0751

Super-resolution ×4 on DIV2K

PSNR ↑ across INR architectures.

ModelAdamMuonProPolar
ReLU MLP19.6522.0123.47
SIREN24.3224.7324.73
WIRE23.9724.6224.68
FINER25.2325.1925.45

Neural radiance fields

MethodFernFlowerFortressHornsLeavesOrchidsRoomT-RexAvg.
Adam26.3827.4731.7427.7921.9221.1031.9626.9026.91
MoFaSGD26.2427.3031.4227.3221.5620.9832.0126.7126.69
Muon26.8428.1232.2128.6022.1621.4233.1727.7927.53
ProPolar27.2228.6832.8229.0222.5421.8033.7228.4428.03
Adam0.83220.85660.90360.87020.77370.71350.95150.90000.8501
MoFaSGD0.82000.84350.89460.84890.74540.69340.94810.88840.8352
Muon0.85070.87770.91840.89340.78800.73180.96040.91750.8672
ProPolar0.85840.88800.92700.89840.80140.75080.96290.92800.8768

Scene-level results at 100K iterations with the original hyperparameter configuration of vanilla NeRF.

Training cost

MethodGPU memory, GBTime per iteration, msTotal time, sOptimizer step, ms
Adam2.9796.63966.355.23
Muon2.97107.221072.1619.75
Shampoo, F = 104.57166.861668.6288.25
Shampoo, F = 1004.57106.541065.4428.00
ProPolar2.97112.701126.9625.45

Lego scene with 1.19M trainable parameters, timed over 10K iterations after 50 warm-up iterations. Like Muon, ProPolar adds optimizer time over Adam through the polar step. GPU memory matches Adam.

The paper reports vanilla NeRF at 200K iterations and the TPE search spaces in the appendix.

Qualitative Comparison

Drag the divider to compare each baseline with ProPolar on the same scene. Zoom magnifies the region that the paper marks with a red box.

Kodak image fitted by a ReLU MLP with Adam
Kodak image fitted by a ReLU MLP with ProPolar
Image fitting, ReLU MLPKodak. The black diagonal band shows per-pixel absolute error.
Kodak image fitted by SIREN with Adam
Kodak image fitted by SIREN with ProPolar
Image fitting, SIRENKodak. The black diagonal band shows per-pixel absolute error.
Armadillo surface from a neural SDF with Adam
Armadillo surface from a neural SDF with ProPolar
SDF, ArmadilloReLU MLP on the Stanford 3D Scanning Repository.
Dragon surface from a neural SDF with Adam
Dragon surface from a neural SDF with ProPolar
SDF, DragonReLU MLP on the Stanford 3D Scanning Repository.
Novel view of the LLFF T-Rex scene with Muon
Novel view of the LLFF T-Rex scene with ProPolar
NeRF, T-RexReal-world LLFF scene.
Novel view of the Blender Ship scene with Muon
Novel view of the Blender Ship scene with ProPolar
NeRF, ShipSynthetic Blender scene.
Lego scene after NeRF fine-tuning with Muon
Lego scene after NeRF fine-tuning with ProPolar
Fine-tuning, LegoRefinement stage after coarse reconstruction.

BibTeX

@inproceedings{kim2026propolar,
  title     = {{ProPolar}: Progressive Polar Decomposition for Implicit Neural Representations},
  author    = {Pureum Kim and Younggeon Ryu and Dongyoon Lee and Hae Beom Lee and Kyong Hwan Jin},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}