- Published on
Fourier Spectra of the Loss Landscape: Spectral Navigation to Minima?
Fourier Spectra of the Loss Landscape: Spectral Navigation to Minima?
Draft idea note — literature survey completed 2026-09-06, no experiments yet. Sibling to fourier-spectral-shapes: that note applies spectral thinking to shapes; this one applies it to the loss surface itself.
1. The core idea
Take a trained-or-training neural network and freeze everything except a small set of weights (say two). Sweep those two weights over a bounded box and evaluate the loss on the fixed dataset. You get a 2D surface. Intuitively it looks extremely rough — "noisy" — with occasional places where the loss plummets. Gradient descent is very good at finding those plunges. But a rough deterministic surface is still a well-behaved mathematical object: any bounded, continuous function on a box has a Fourier series. So the pitch:
The loss surface over a bounded weight range is a signal with a Fourier spectrum. If we could learn that spectrum cheaply, we could find where many modes interfere constructively to make small — i.e., jump to deep minima without walking there.
Three sub-ideas embedded in this:
- The surface has a spectrum. Restrict (or the full over a box) and expand it as . Low frequencies = large-scale basin structure; high frequencies = the "noise" that makes the surface look rough.
- The spectrum is usable for optimization. The value of at any point is the coherent sum of all modes. Minima should sit where dominant modes align in phase to produce a deep trough — "constructive interference" of low-loss contributions. Read the phases, predict the trough, skip the walking.
- The spectrum can be learned cheaply. One full training run yields one trajectory, not a surface. Dense sampling of the surface would cost dataset-passes × grid-points. Question: can the spectral configuration be inferred from a single pass / few probes instead?
The survey below asks, for each piece: does this exist in the literature, and under what assumptions does it actually work?
2. The questions, made precise
- Q1 (theory): Is there theoretical work that uses the global Fourier/spectral structure of the loss surface (not Hessian eigenvalues) to characterize where minima are, or to find them?
- Q2 (interference): Is there work that uses mode phase / "constructive interference" reasoning over weight space to locate low-loss regions?
- Q3 (sampling): Is there work on estimating the spectral configuration of a loss surface from very few evaluations — ideally from a single GD trajectory?
- Q4 (design): Is the spectrum "configurable"? i.e., can architecture, loss, or noise schedule shape the loss surface's frequency content to make optimization easier?
Headline result of the survey: no paper was found that does exactly (Q1+Q2+Q3 together) — full-spectrum weight-space navigation from a cheaply learned Fourier model. But every component exists somewhere, several 2025-2026 papers are strikingly close in spirit, and the literature is clear about why the naive version fails and which restricted version could work.
3. Survey, by cluster
3.1 What a 2-weight slice actually looks like (landscape geometry)
The premise "sampling loss vs 2 weights gives a noisy plane with plunges" is half right — and the half that is wrong matters.
- Li et al., Visualizing the Loss Landscape of Neural Nets (NeurIPS 2018) — the canonical method: 1D/2D slices through weight space along random directions, but with filter normalization (per-layer scaling) to remove scale-invariance artifacts. Findings: at large scale along training-relevant directions, ResNet-style landscapes look like smooth bowls with benign structure; naive un-normalized plots are misleading, and sharpness read off raw 1D interpolations is unreliable. Key methodological warning for this idea: which two weights / which directions you slice along completely determines what the "surface" looks like.
- Chunyuan Li et al., Measuring the Intrinsic Dimension of Objective Landscapes (ICLR 2018) — you only need ~ random directions to reach near-full training accuracy; objective landscapes have low intrinsic dimension. Consequence: most single-weight axes are nearly flat, and a random 2-weight slice of a big network is typically dominated by a near-constant plane + fine-scale texture — the dramatic structure lives in a few special (training-aligned, filter-normalized) directions.
- Choromanska et al., The Loss Surfaces of Multilayer Networks (AISTATS 2015) — spin-glass-style theory: number of minima grows exponentially with dimension, but low-loss minima are rare; typical higher-lying minima proliferate. So "places where the loss plummets" are not exotic — they are exactly what over-parameterized GD latches onto.
- Dauphin et al., Identifying and attacking the saddle point problem (NeurIPS 2014) — the dominant obstructions in high-D loss surfaces are saddles, not local minima; early training is mostly saddle-escape dynamics, which is why noise/curvature tricks matter.
- No-bad-basin theory: Kawaguchi (Deep Learning without Poor Local Minima, 2016) for deep linear nets; Sohl-Dickstein et al. (Eliminating all bad Local Minima from Loss Landscapes without even adding an Extra Unit, 2019) — any landscape can be provably cleared of bad minima under a ridge-like regularization; Liu et al. (Loss landscapes and optimization in over-parameterized non-linear systems and neural networks, 2020). Message: for over-parameterized models the interesting question is not "do bad minima exist" but "how does the optimizer flow to the good region" — a question about large-scale shape, which is precisely the low-frequency content.
3.2 The large-scale shape: basins, wedges, mode connectivity
If you smooth away the fine texture, what remains is a remarkably connected structure:
- Garipov et al., Loss Surfaces, Mode Connectivity, and Fast Ensembling (NeurIPS 2018) and Draxler et al., Essentially No Barriers in Neural Network Energy Landscape (2018) — distinct SGD minima are connected by low-loss paths (mode connectivity); for common architectures the connecting paths are nearly flat in train and test loss.
- Fort & Scherlis, The Large-Scale Structure of Neural Network Loss Landscapes (NeurIPS 2019) — phenomenological "wedge" model: the landscape is a set of high-dimensional wedges forming one large interconnected structure toward which optimization is drawn; hyperparameters change the path more than the destination region.
- Frankle et al., Linear Mode Connectivity and the Lottery Ticket Hypothesis (ICML 2020) — networks become stable to SGD noise early: after that, optimization outcome is a linearly connected region, not a point.
- Baldassi et al. (2016) / Entropy-SGD / SAM — see 3.5; flat minima are entropically favored and occupy huge connected volumes.
Reading for this idea: the deep, robust minima are not isolated Dirac wells — they are broad, connected, low-loss regions. That is exactly what a low-frequency-dominated spectrum looks like. The high-frequency "noise" is decoration on top of a smooth skeleton.
3.3 Spectral bias & the F-Principle — the one place where "spectrum + optimization dynamics" is rigorously developed
The most substantial body of work connecting Fourier analysis to neural optimization is about learning dynamics in function space, and it directly supports your intuition that "GD is good at finding the big plunges":
- Rahaman et al., On the Spectral Bias of Neural Networks (ICML 2019) — DNNs fit target functions low-frequency-first; spectral bias explains why Fourier features / position encodings are needed for high-frequency targets.
- Xu et al., Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks (2019) + Luo et al., Theory of the Frequency Principle for General Deep Neural Networks (2019) — during training, error energy concentrates in progressively higher frequencies; convergence per frequency is governed by the eigenvalues of a frequency-domain operator. The 2022 survey Xu, Overview frequency principle/spectral bias in deep learning collects the whole line (including the F-Principle under general loss functions, Xu 2018).
- Gousia et al./2025-adjacent, Gradient Descent as a Shrinkage Operator for Spectral Bias (2025) — GD can be reinterpreted as a shrinkage operator on the network Jacobian's singular values: learning rate × iterations implicitly select a bandwidth (how many frequency components stay alive). This is the cleanest statement that "GD itself is a low-pass-filtered spectral process."
- Important caveat: this entire literature analyzes the spectrum of the input-output function or the loss in frequency space over the data, not the spectrum of the loss as a function of weights. The weight-space analogue (spectrum of along optimization-relevant directions) has essentially no theory — an open gap this idea would occupy.
3.4 The other "spectral" literature: Hessian eigenvalue spectra (local, not Fourier)
When ML papers say "spectrum of the loss landscape," ~95% of the time they mean the eigenvalue spectrum of the Hessian at a point — local curvature, not global Fourier structure. It is worth knowing this literature so the two are not conflated:
- Sagun et al. (2016, 2017) — Hessian of over-parameterized nets is low rank + bulk; eigenvalues concentrate near zero with a few outliers.
- Ghorbani et al., An Investigation into NN Optimization via Hessian Eigenvalue Density (ICML 2019) — density evolves during training; outliers appear late.
- Papyan, The Full Spectrum of Deepnet Hessians at Scale (2019) — two-component structure: bulk + spikes; blockwise structure across layers.
- Xie et al., On the Power-Law Hessian Spectrums in Deep Learning (2022); Martin & Mahoney, Implicit Self-Regularization (RMT evidence, 2018-19) — well-trained Hessians show power-law / heavy-tailed spectra: no natural frequency cutoff, self-similar structure. Consistent with fractal roughness (3.5).
- Wei et al., How noise affects the Hessian spectrum in overparameterized neural networks (2019); Singh, How the Hessian-Spectrum of Neural Networks Depends on Data (2026); Gabdullin et al., Effects of Hessian eigenvalue spectral density type on generalization analysis (2025) — spectral shape is controlled by noise and data; flatness measures only make sense relative to spectral density type.
- Dinh et al., Sharp Minima Can Generalize For Deep Nets (ICML 2017) — the classic caution: raw flatness is not causally tied to generalization (reparameterization changes curvature without changing the function).
- Calvo, Spectral Asymptotics of Neural Network Loss Landscapes: An Exact Decomposition of the Curvature Exponent (2026) — despite the title, this is Hessian-land: a Weyl-law-style analysis of how Hessian eigenvalues scale with gradient singular values across layer types (α ≈ 2 conv, ≈ 1 attention). Excellent recent theory — but again local curvature spectra, not Fourier spectra of the global surface.
Relevance: (a) when searching, "spectral" mostly retrieves this cluster; (b) power-law, scale-free structure at the local level is the microscopic origin of the "noisy" look you see at fine scales.
3.5 The surface is multiscale / fractal — and data noise creates fake small-scale geometry
Recent work directly characterizes the roughness you're imagining:
- Multifractal loss landscapes (Nature Communications, 2025) — loss landscapes are multifractal: clustered degenerate minima, multiscale structure, edge-of-stability dynamics; optimizers navigate to smooth solution spaces housing flatter minima because of, not despite, this structure.
- Rolling Ball Optimizer (2025) — explicitly motivated by the same observation you made: landscapes are "highly complex and textured, even fractal-like... noise in the training data propagates forward and gives rise to unrepresentative small-scale geometry." Their fix: optimize a scale-space smoothed version of the loss (simulate a rigid ball of radius rolling on the surface; provably a smoothing of ; radius = granularity dial). Key sentence: the large-scale geometry of the loss landscape is less data-specific than its fine-grained structure, and easier to optimize.
- Landscaper (2026) — arbitrary-dimensional loss landscape analysis via Hessian-based subspace construction + topological data analysis; introduces SMAD (Saddle-Minimum Average Distance) as a landscape smoothness metric and exposes basin hierarchy.
- Taxonomizing local vs global structure (2021) — systematic study of thousands of models: local smoothness metrics correlate with generalization, but local and global structure are largely independent axes.
Reading: the spectrum of a real loss surface is scale-free (power-law-ish), not sparse, not band-limited. There is structure at every scale, and the fine scales are partly data-noise artifacts.
3.6 Smoothing / coarse-to-fine = the theorem-backed "low-pass" program
If the low-frequency envelope is the useful part, the literature has already built optimizers that exploit it — under the name of smoothing, continuation, or graduated optimization:
- Hazan, Levy & Shalev-Shwartz, On Graduated Optimization for Stochastic Non-Convex Problems (ICML 2016) — the theory anchor: for functions whose coarse-scale (smoothed) versions are "nice," a coarse-to-fine schedule provably converges to a global optimum in gradient steps. This is the formal statement of "low-pass the landscape, then descend."
- Rolling Ball Optimizer (2025) — modern implementation of exactly that principle (3.5), beating SGD/SAM/Entropy-SGD on small benchmarks.
- Entropy-SGD (Chaudhari et al., 2019) — optimizes a smoothed (local-entropy) version of the loss; noise term biases toward wide valleys. SAM (Foret et al., 2021) — flatness via worst-case perturbation. SWA (Izmailov et al., 2018) — weight averaging lands in wide, connected optima. All three are implicit low-pass filters on basin structure.
- Stochastic global search theory: SGLD (Welling & Teh, ICML 2011) and annealing (Hajek 1988: logarithmic cooling ⇒ convergence to global min in probability) — noise as the search mechanism; recent sharp result Azizian & Lelarge, The global convergence time of SGD in non-convex landscapes (2025) — via large deviations theory, time-to-global-minimum is dominated by the highest barrier × noise statistics (Eyring–Kramers-type exponential scaling). So in full generality, noise-driven global search is exponentially slow unless barriers are low or noise is shaped — which is precisely why the connected low-barrier structure of 3.2 matters, and why graduated smoothing can beat raw annealing.
Reading: "use only the coarse (low-frequency) structure to find the basin, then descend locally" is a proven strategy. The full-spectrum version of your idea is unnecessary for this; the low-pass version is what all these methods approximate.
3.7 Fourier methods literally applied to NN loss surfaces (2026) + the quantum analogue
These are the closest literal hits to your idea:
- FourierPathFinder (arXiv 2602.21276, 2026) — "Neural network optimization strategies and the topography of the loss landscape." A general-purpose algorithm that finds low-height paths between two points on a loss/energy landscape by expanding the connecting path in a Fourier basis, with a regularization penalty controlling path smoothness. Used to compare SGD vs quasi-Newton solutions: SGD minima are separated by lower barriers and sit in "smoother, more interconnected regions." So Fourier series are being used on NN loss landscapes — but of paths between known endpoints, for topography analysis, not to locate unknown minima.
- Dynamical loss functions (arXiv 2410.10690, 2024) — per-class loss contributions are made to oscillate periodically over training, which "globally alters the loss landscape without affecting the global minima" and improves learning. Demonstrates you can sculpt landscape structure from outside (relevant to Q4).
- Quantum machine learning — the strongest existing analogue of the whole program. Schuld et al., The effect of data encoding on the expressive power of variational QML models (2021): variational quantum models are provably truncated Fourier series — the accessible frequency spectrum is fixed by the data-encoding circuit, and the trainable parameters only set the coefficients. The spectrum is analytically known a priori, and there is a whole literature on designing/expanding it (data re-uploading) and on trainability failures caused by spectral structure — e.g. "Fourier locking" in data re-uploading classifiers, fixed via spectral homotopy (2026). This is the only field where "loss-as-Fourier-series + use the spectrum to fix optimization" is a working research program — because the circuits are engineered so the spectrum is known and controllable. Classical nets have no such handle: their effective spectrum is emergent, data-dependent, and only characterized statistically (3.3, 3.5).
- Fourier-parameterized networks: Fourier Multi-Component and Multi-Layer Neural Networks (2025) report that nets whose parameterization is Fourier-based have substantially more favorable optimization landscapes, especially for high-frequency targets — empirical support for Q4 (parameterization shapes the effective landscape).
3.8 Learning a spectrum from few samples: the machinery for Q3
Your Q3 is a sampling/recovery problem: how many evaluations of are needed to learn its spectrum well enough to trust its minima? The mathematics here is classical and unforgiving, plus one modern miracle:
- The obstacle — aliasing / no band limit. A function's Fourier coefficients can only be estimated down to the scale set by your sampling density (Nyquist–Shannon). Real loss surfaces are not band-limited: with ReLU activations is piecewise-polynomial (algebraic coefficient decay, infinitely many modes); with smooth activations it is analytic in for fixed data (exponential decay in principle), but minibatch noise + the fractal small-scale structure (3.5) put a noise floor on high frequencies. Anything above the sampling cutoff aliases into the low frequencies and corrupts exactly the modes you care about. This is the core technical reason "sample a coarse grid, FFT it, trust the troughs" fails.
- The miracle — sparsity. If the spectrum were -sparse, compressed sensing / sparse FFT recovers it from samples: FFAST (Hassanieh et al. 2012) and high-dimensional sparse Fourier algorithms (Choi et al. 2016; multiscale/noisy variant 2019). Same idea powers sparse polynomial chaos expansions (Blatman & Sudret 2011), the standard UQ method for fitting spectral surrogates of expensive simulators from few runs. But real loss spectra are power-law, not sparse (3.5) — CS does not directly apply; it applies to projected low-dimensional slices if those turn out to be dominated by a few modes (an empirical question worth testing — see §6).
- The practical workhorse — Bayesian optimization with spectral kernels. Wilson & Adams, Gaussian Process Kernels for Pattern Discovery and Extrapolation (ICML 2013): spectral-mixture kernels literally learn the power spectrum of the objective from samples and use it to predict where minima are; BO is the standard answer to "optimize a function you can only query rarely." It works up to ~tens of dimensions — not weights. Snoek et al. (2012) made BO practical for ML hyperparameters.
- Dimensionality reduction first: active subspaces (Constantine 2014) find the few spectrally dominant input directions of a function from gradient samples — the natural way to pick which "two weights" (directions, not raw weights) carry the structure; complements intrinsic-dimension results (3.1).
4. Direct answers to Q1-Q4
- Q1 (Fourier-of-the-landscape theory): No paper found that performs global optimization of NN training from the weight-space loss surface's Fourier spectrum. The rigorous spectral theory of NN optimization is either (a) function-space spectral bias / F-Principle (3.3), or (b) local Hessian eigenvalue spectra (3.4). Fourier analysis has only just entered loss-landscape work directly (FourierPathFinder, 2026 — paths, not minima; 3.7). Note also: even given the exact spectrum, finding the global minimum of a generic multivariate trigonometric polynomial is not fundamentally easier than the original problem — the Fourier basis is a reparameterization, not a relaxation; it only pays off when the spectrum is compressible (few dominant modes) or the dimension is low, or in special regimes (e.g., near-NTK where the landscape is nearly quadratic).
- Q2 (phase/interference navigation): No direct work. Nearest neighbors: mode connectivity (deep minima are connected low-loss regions, i.e., phase-coherent at large scale — 3.2), graduated optimization and RBO (descend the coarse envelope — 3.6), annealing/SGLD (noise-driven basin hopping — 3.6), and the QML Fourier-locking line where spectral structure explicitly causes/repairs optimization failure (3.7).
- Q3 (cheap spectral learning): No work estimating an NN loss surface's spectrum from a single run. A GD trajectory is a 1D curve — it constrains the landscape locally along a path (curvature/roughness of what you traversed), not globally; learning global modes requires off-trajectory queries, which is the expensive grid you wanted to avoid. The adjacent machinery exists: spectral-mixture GP/BO (few-query spectral learning, low-dim), sparse FFT/CS and sparse PCE (few-query sparse spectral learning), active subspaces (which directions matter). All assume low dimension and/or spectral sparsity/smoothness — both assumptions need empirical testing on real loss slices before the idea has legs.
- Q4 (configurable spectrum): Partial yes, no unifying theory. Parameterization changes the effective landscape (Fourier-component nets train far more easily — 3.7); time-varying loss sculpting demonstrably reshapes the landscape without moving global minima (dynamical loss functions — 3.7); smoothing radius (RBO), noise schedule (SGLD/annealing), and flatness bias (SAM/Entropy-SGD) all act on the coarse structure. What does not exist is a framework that says "here is the loss surface's spectrum; here is how a design choice moves its energy."
5. Where the idea breaks (honest critique)
- No band limit + noise floor → aliasing. You can never sample densely enough (dataset-pass cost) to resolve the frequencies that dominate the "noisy" look, and the unresolved tail aliases into and corrupts the low-frequency basin signal. Any workable version must first smooth (RBO/graduated style), i.e., deliberately discard the high frequencies — at which point you have reinvented graduated optimization.
- Spectra are power-law, not sparse. Scale-free roughness (multifractal landscape, power-law Hessians) means no few-modes representation exists at the global scale; compressed-sensing-style sample savings don't materialize without added structure.
- Curse of dimension in the coefficient space. Full weight space has ~+ dims; even the coefficient description is astronomically large. Only projected slices (2-weight planes, active subspaces, filter-normalized directions) are tractable — and on those, structure concentrates in special directions (intrinsic dimension), so "any two weights" is usually the wrong two.
- Phase is meaningless beyond the resolved band. "Constructive interference of modes" is only computable for modes you can actually estimate; for the rest, phase is garbage/aliased.
- Even with the exact spectrum, global optimization of the trigonometric surrogate is hard in high dimension; the spectrum only pays if it is compressible and you search the surrogate in low dimension.
- One run ≠ global knowledge. A single GD trajectory samples a 1D manifold of weight space. Spectral inference from it gives you the spectrum along the path (useful for detecting regime changes — cf. grokking/eigenvalue early-warning literature) but cannot reveal basins it never approached. You must spend queries off-trajectory — the budget question is how few, and the honest answer from BO theory is "exponential in intrinsic dimension without strong priors" (no-free-lunch / information-based complexity).
- A random 2-weight slice is mostly flat + fine noise (intrinsic dimension, filter-normalization findings) — the dramatic "plunges" you'd find in a naive slice are partly minibatch-eval noise and zoomed-in fractal texture, not navigable structure. The interesting surface only appears along training-aligned, scale-normalized directions.
6. What survives — concrete research directions
The idea is not dead; it reduces to a restricted, empirical, low-frequency program:
- Probe the spectrum, empirically (cheapest first experiment). Take a real small-scale training run (MNIST/CIFAR subset, full-batch). Choose a filter-normalized direction pair (Li et al. recipe) or the top-2 active-subspace directions. Grid-sample at ~ points, FFT, and look: Is the 2D spectrum power-law? Does the phase structure of the top modes place the basin correctly? This directly tests Q1's empirical premise for ~1 GPU-hour.
- Trajectory spectral probes. During one SGD run, every steps evaluate loss at along a few directions → estimate the local power spectrum along the path. Does its low-frequency part predict the next basin (the weight-space analogue of the F-Principle)? Ties into 3.3 — a genuinely open direction.
- Low-pass basin finding = graduated optimization with a spectral dial. Compare RBO / Gaussian-smoothed descent / explicit spectral truncation of a probed slice: does spectral truncation beat spatial smoothing at the same budget?
- Sparse spectral surrogate in an active subspace. Project to ~2-6 dominant directions; fit a sparse-Fourier / sparse-PCE surrogate (Blatman-Sudret machinery) from a few hundred full-batch probes; grid-search the surrogate; validate by GD-from-candidate vs multistart. This is the most faithful testable version of "learn the spectrum, find the interference trough."
- Configurable-spectrum testbed. Train Fourier-parameterized (FMMNN-style) networks and measure whether their loss spectra are genuinely more low-pass/compressible than standard nets — testing Q4 with a real dial.
7. Open questions
- Is the low-frequency envelope of a real loss surface band-limited in any practical sense (energy below frequency for above some data-dependent cutoff)? RBO's success suggests yes at the scale of optimization steps; nobody has measured it directly.
- Does the F-Principle have a weight-space dual: is there a "frequency" ordering of weight-space directions along which SGD energy moves during training? (Active-subspace/Hessian alignment results are suggestive.)
- How many off-trajectory probes are needed to locate a basin's attractor (not its exact minimum) to within GD's catchment? That is the real sample-complexity question, and it is much softer than "find the global min."
- Do flat/connected minima (3.2) correspond to phase-coherent low-frequency structure — i.e., is "flatness" literally "low high-frequency energy" in the weight-space Fourier sense?
- Can the dynamical-loss-functions trick (3.7) be re-derived as deliberate spectral sculpting of the loss surface rather than an empirical curiosity?
8. Survey method & negative-result statement
- When: 2026-09-06. Engines: arXiv API (full-metadata queries over title/abstract, ~30 query families including:
"loss landscape" AND "fourier","loss landscape" AND "spectral","loss landscape" AND "spectrum","loss landscape" AND "global minimum","frequency principle","spectral bias" AND "gradient descent","loss landscape" AND "smoothing","loss landscape" AND "trajectory","global optimization" AND "fourier"(math.OC),"sparse fourier" AND "high-dimensional", plus targeted web searches for the newest items) and web search (Google-indexed, incl. arXiv HTML listings). - Explicit negative result: exhaustive search located no paper that (a) computes the Fourier spectrum of a neural network's weight-space loss surface over a bounded box and (b) uses that spectrum — phases or otherwise — to locate or jump to low-loss regions, and (c) learns that spectrum from a single run or few evaluations. Nearest literal neighbors, each missing at least one of (a)-(c): FourierPathFinder (2026, Fourier series of paths, endpoints given, no global-min location); Rolling Ball Optimizer (2025, spatial smoothing, no Fourier model); graduated optimization (2016, smoothing theory, no spectral estimation); spectral-mixture GP / BO (spectral learning but low-dim black-box regime, not NN weight space); QML Fourier-series models (exact spectra by construction — classical nets lack this); spectral-bias / F-Principle line (spectra of learned functions, not of loss surfaces). If you believe a counterexample exists, the query families above are the ones to re-run.
9. Sources
Landscape geometry & visualization
- Li et al., Visualizing the Loss Landscape of Neural Nets, NeurIPS 2018 — https://arxiv.org/abs/1712.09913
- Li et al., Measuring the Intrinsic Dimension of Objective Landscapes, ICLR 2018 — https://arxiv.org/abs/1804.08838
- Choromanska et al., The Loss Surfaces of Multilayer Networks, AISTATS 2015 — https://arxiv.org/abs/1412.0233
- Dauphin et al., Identifying and attacking the saddle point problem, NeurIPS 2014 — https://arxiv.org/abs/1406.2572
- Kawaguchi, Deep Learning without Poor Local Minima, NeurIPS 2016 — https://arxiv.org/abs/1605.07110
- Sohl-Dickstein et al., Eliminating all bad Local Minima from Loss Landscapes without even adding an Extra Unit, 2019 — https://arxiv.org/abs/1901.03909
- Liu et al., Loss landscapes and optimization in over-parameterized non-linear systems and neural networks, 2020 — https://arxiv.org/abs/2003.00307
Large-scale structure & mode connectivity
- Garipov et al., Loss Surfaces, Mode Connectivity, and Fast Ensembling, NeurIPS 2018 — https://arxiv.org/abs/1802.10026
- Draxler et al., Essentially No Barriers in Neural Network Energy Landscape, 2018 — https://arxiv.org/abs/1803.00885
- Fort & Scherlis, Large Scale Structure of Neural Network Loss Landscapes, NeurIPS 2019 — https://arxiv.org/abs/1906.04724
- Frankle et al., Linear Mode Connectivity and the Lottery Ticket Hypothesis, ICML 2020 — https://arxiv.org/abs/1912.05671
- Xie et al., Evaluating Loss Landscapes from a Topology Perspective, 2024 — https://arxiv.org/abs/2411.09807
Spectral bias / frequency principle (function-space spectra of training)
- Rahaman et al., On the Spectral Bias of Neural Networks, ICML 2019 — https://arxiv.org/abs/1806.08734
- Xu et al., Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks, 2019 — https://arxiv.org/abs/1901.06523
- Luo et al., Theory of the Frequency Principle for General Deep Neural Networks, 2019 — https://arxiv.org/abs/1906.09235
- Xu, Overview frequency principle/spectral bias in deep learning, 2022 — https://arxiv.org/abs/2201.07395
- Xu et al., Frequency Principle in Deep Learning with General Loss Functions, 2018 — https://arxiv.org/abs/1811.10146
- Xu et al., Training behavior of deep neural network in frequency domain, 2018 — https://arxiv.org/abs/1807.01251
- Gradient Descent as a Shrinkage Operator for Spectral Bias, 2025 — https://arxiv.org/abs/2504.18207
Hessian / curvature spectra (the other "spectral")
- Sagun et al., Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond, 2016 — https://arxiv.org/abs/1611.07476 ; Empirical Analysis of the Hessian of Over-Parametrized NNs, 2017 — https://arxiv.org/abs/1706.04454
- Ghorbani et al., An Investigation into NN Optimization via Hessian Eigenvalue Density, ICML 2019 — https://arxiv.org/abs/1901.10159
- Papyan, The Full Spectrum of Deepnet Hessians at Scale, 2019 — https://arxiv.org/abs/1811.07062
- Xie et al., On the Power-Law Hessian Spectrums in Deep Learning, 2022 — https://arxiv.org/abs/2201.13011
- Martin & Mahoney, Implicit Self-Regularization in Deep Neural Networks: Evidence from RMT, 2018 — https://arxiv.org/abs/1810.01075
- Wei et al., How noise affects the Hessian spectrum in overparameterized neural networks, 2019 — https://arxiv.org/abs/1910.00195
- Singh et al., How the Hessian-Spectrum of Neural Networks Depends on Data, 2026 — https://arxiv.org/abs/2607.13631
- Gabdullin et al., Effects of Hessian eigenvalue spectral density type on generalization analysis, 2025 — https://arxiv.org/abs/2504.17618
- Dinh et al., Sharp Minima Can Generalize For Deep Nets, ICML 2017 — https://arxiv.org/abs/1703.04933
- Calvo et al., Spectral Asymptotics of NN Loss Landscapes: An Exact Decomposition of the Curvature Exponent, 2026 — https://arxiv.org/abs/2606.02596
Roughness, multiscale structure, smoothing
- Optimization on multifractal loss landscapes explains... deep learning, Nature Communications 2025 — https://www.nature.com/articles/s41467-025-58532-9
- Rolling Ball Optimizer: Learning by ironing out loss landscape wrinkles, 2025 — https://arxiv.org/abs/2505.19527
- Landscaper: Loss Landscapes Through Multi-Dimensional Topological Analysis, 2026 — https://arxiv.org/abs/2602.07135
- Taxonomizing local versus global structure in neural network loss landscapes, 2021 — https://arxiv.org/abs/2107.11228
- Asymptotic Smoothing of the Lipschitz Loss Landscape in Overparameterized One-Hidden-Layer ReLU Networks, 2026 — https://arxiv.org/abs/2602.17596
Coarse-to-fine / smoothing / noise-based global search
- Hazan, Levy, Shalev-Shwartz, On Graduated Optimization for Stochastic Non-Convex Problems, ICML 2016 — https://arxiv.org/abs/1503.03712
- Chaudhari et al., Entropy-SGD: Biasing Gradient Descent Into Wide Valleys, ICLR 2019 — https://arxiv.org/abs/1611.01838
- Foret et al., Sharpness-Aware Minimization, ICLR 2021 — https://arxiv.org/abs/2010.01412
- Izmailov et al., Averaging Weights Leads to Wider Optima and Better Generalization (SWA), UAI 2018 — https://arxiv.org/abs/1803.05407
- Keskar et al., On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima, ICLR 2017 — https://arxiv.org/abs/1609.04836
- Welling & Teh, Bayesian Learning via Stochastic Gradient Langevin Dynamics, ICML 2011 — http://www.cs.toronto.edu/~welling/publications/papers/bayesian_via_sgd.pdf
- Hajek, Cooling schedules for optimal annealing, Math. of OR 1988 — https://pubsonline.informs.org/doi/10.1287/moor.13.2.311
- Azizian & Lelarge, The global convergence time of SGD in non-convex landscapes: Sharp estimates via large deviations, 2025 — https://arxiv.org/abs/2503.16398
Fourier on loss surfaces / loss sculpting / quantum analogue
- Yu et al., Neural network optimization strategies and the topography of the loss landscape (FourierPathFinder), 2026 — https://arxiv.org/abs/2602.21276
- Lavin Pallero et al., Dynamical loss functions shape landscape topography and improve learning, 2024 — https://arxiv.org/abs/2410.10690
- Schuld, Sweke, Meyer, The effect of data encoding on the expressive power of variational QML models, Phys. Rev. A 2021 — https://arxiv.org/abs/2008.08605
- Overcoming Fourier Locking in Quantum Data Re-uploading Classifiers via Spectral Homotopy, 2026 — https://arxiv.org/abs/2607.11013
- Fourier Multi-Component and Multi-Layer Neural Networks: Unlocking High-Frequency Potential, 2025 — https://arxiv.org/abs/2502.18959
Learning spectra / functions from few samples (Q3 machinery)
- Hassanieh et al., Nearly Optimal Sparse Fourier Transform (FFAST), STOC 2012 — https://arxiv.org/abs/1201.2501
- Choi et al., High-Dimensional Sparse Fourier Algorithms, 2016 — https://arxiv.org/abs/1606.07407 ; multiscale noisy variant — https://arxiv.org/abs/1907.03692
- Blatman & Sudret, Adaptive sparse polynomial chaos expansion based on least angle regression, J. Comput. Phys. 2011 — https://doi.org/10.1016/j.jcp.2010.12.009
- Wilson & Adams, Gaussian Process Kernels for Pattern Discovery and Extrapolation (spectral mixture kernels), ICML 2013 — https://arxiv.org/abs/1302.4245
- Snoek et al., Practical Bayesian Optimization of Machine Learning Algorithms, NeurIPS 2012 — https://arxiv.org/abs/1206.2944
- Constantine, Computing active subspaces with Monte Carlo, 2014 — https://arxiv.org/abs/1408.0545
- Nyquist–Shannon sampling theorem (aliasing/band-limit background) — https://en.wikipedia.org/wiki/Nyquist%E2%80%93Shannon_sampling_theorem
- Candès, Romberg, Tao, Robust uncertainty principles: exact signal reconstruction from highly incomplete frequency information, IEEE Trans. Inf. Theory 2006 (compressed sensing) — https://ieeexplore.ieee.org/document/1580791