Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

306191121 · Jun 202019922001200920182026
48 results for Minimum norm

Minimum-norm solutions generalize well in over-parametrized neural networks.

problem Generalization error in over-parametrized neural networks.
method Analyzing three models: random feature model, two-layer neural network, and residual network.
result Generalization error for minimum-norm solutions is comparable to Monte Carlo rate, up to logarithmic terms.

The study analyzes robustness of estimators in linear models with adversarial errors.

problem Analyzing robustness of estimators in linear models with adversarial errors.
method Develops a general theory for minimum norm interpolating estimators and RERM in linear models without conditions on errors.
result Quantitative bound for the prediction error relating it to Rademacher complexity, norm of minimum norm interpolator of errors, and subdifferential size.

Study shows unusual non-monotonic risk behavior in minimum-norm interpolants for various data scaling.

problem Understanding the risk behavior of minimum-norm interpolants in RKHS for different data scaling.
method Analysis of spectral properties of the random kernel matrix restricted to eigen-spaces of the population covariance operator.
result Minimum-norm interpolants in RKHS exhibit multiple descent in risk for d=nαd = n^α with α(0,1)α\in(0,1).

The paper studies the minimum ℓ₁-norm interpolator's risk behavior in over-parameterized settings.

problem Understanding the risk behavior of minimum ℓ₁-norm interpolators in high-dimensional settings.
method Exact characterization of the risk behavior through a system of two non-linear equations.
result Observation of a multi-descent phenomenon in the generalization risk of the minimum ℓ₁-norm interpolator.

Batching stabilizes risk in high-dimensional linear regression models.

problem Stability and risk behavior in high-dimensional overparameterized linear regression.
method Minimum-norm overparameterized linear regression model with batch-partitioning.
result Optimal batch size is inversely proportional to noise level and overparametrization ratio, leading to stable risk behavior.

Interpolation with Laplace kernel fails in low dimensions but succeeds in high dimensions.

problem Consistency of interpolation with Laplace kernels in low-dimensional settings.
method Minimum-norm interpolation in Reproducing Kernel Hilbert Space (RKHS) with Laplace kernel.
result Consistency of interpolation is a high-dimensional phenomenon.

The paper explores why a specific type of predictor works well in noisy data.

problem Understanding why a specific type of predictor (minimum-norm interpolator) works well in noisy data.
method The paper uses uniform convergence and zero-error predictors in a norm ball to explain the success of the minimum-norm interpolator.
result The minimum-norm interpolator is consistent, and this can be explained by uniform convergence of zero-error predictors in a norm ball.

We introduce a norm on the space of test configurations, which we call the minimum norm. We conjecture that uniform K-stability with respect to this norm is equivalent to the existence of a constant scalar curvature Kähler metric. This notion of uniform K-stability is analogous to coercivity of the Mabuchi functional. …

2014-12-01abs ↗pdf ↗

Proposes a filtering method for cluster analysis using 0\ell_0-norm regularization.

problem Improving cluster analysis by filtering data.
method Minimizes a least squares function with a weighted 0\ell_0-norm penalty, approximated by smooth non-convex functions.
result The proposed method can enhance existing clustering techniques.

Weight normalization and reparametrized gradient descent adaptively regularize weights and converge to minimum l2 norm solutions.

problem Adapting to non-convex weight normalization for convergence to minimum l2 norm solutions.
method Weight normalization and reparametrized projected gradient descent (rPGD) for overparametrized least-squares regression.
result rPGD converges close to the minimum l2 norm solution, even for far-from-zero initializations.

Inflating the minimum norm interpolator improves linear regression generalization error.

problem Highly anisotropic covariances and diverging d/nd/n in linear regression.
method Inflating the minimum 2\ell_2 norm interpolator by a constant greater than one.
result Inflating the minimum norm interpolator improves generalization error.

Adaptive methods sometimes outperform SGD in over-parameterized DNNs.

problem Generalization in over-parameterized DNNs with adaptive methods.
method Comparing SGD and adaptive methods on over-parameterized linear regression.
result Minimum weight norm is not always the best indicator of good generalization.

New algorithm finds minimum weight norm solutions in deep neural networks.

problem Training over-parameterized deep neural networks efficiently and improving generalization.
method Minnorm training method that minimizes the sum of weights' norms while fitting training data.
result Faster convergence to minimum-norm solutions and better generalization performance.

Study shows interpolating predictor's risk is optimal in low-dimensional factor regression models.

problem Understanding the risk of interpolating predictors in high-dimensional factor regression models.
method Detailed finite-sample analysis of minimum-norm interpolating predictor's risk in factor regression models.
result The risk of the minimum-norm interpolating predictor approaches optimal benchmarks in low-dimensional factor regression models.

The paper analyzes the risk of a least squares estimator under a spike covariance model.

problem Risk analysis of the least squares estimator under a spike covariance model.
method Assumes spike covariance matrices, studies risk as d/nightarrowd/n ightarrow \infty.
result Risk of the minimum norm least squares estimator vanishes compared to the null estimator.

Deep linear networks can closely approximate interpolants without improving risk.

problem Understanding the risk bounds of deep linear networks compared to minimum 2\ell_2-norm solutions.
method Bounding excess risk of interpolating deep linear networks trained using gradient flow.
result Deep linear networks can closely approximate or match minimum 2\ell_2-norm solutions in terms of risk.

Estimates generalization error for two-layer ReLU NNs through minimum norm solutions.

problem Estimating generalization error for two-layer ReLU NNs trained by mean squared error.
method Uses minimum norm solutions and Neural Tangent Kernel (NTK) regime to derive generalization error bounds.
result Derives an a priori generalization error bound for two-layer ReLU NNs without requiring exponentially large number of neurons.

Paper provides a performance guarantee for spectral clustering.

problem Finding the global solution to the minimum ratio cut problem.
method Two-step spectral clustering method with a rounding step, analyzed using two-to-infinity norm perturbation bounds.
result Spectral clustering is guaranteed to output the global solution under certain conditions.

Riemannian cubics are critical points for the L2L^2 norm of acceleration of curves in Riemannian manifolds MM. In the present paper the LL^\infty norm replaces the L2L^2 norm, and a less direct argument is used to derive necessary conditions analogous to those for Riemannian cubics. The necessary conditions are exami…

2011-04-13abs ↗pdf ↗

Study how optimization methods' choices affect the solutions they find.

problem Characterize the solutions found by optimization methods under different potentials and norms.
method Examined mirror descent, natural gradient descent, and steepest descent for underdetermined linear regression and separable linear classification.
result The specific global minimum reached by an algorithm can be characterized by the potential or norm of the optimization geometry, independent of hyperparameters.

The paper explores how over-parameterized linear regression models generalize without violating learning theory principles.

problem Understanding how over-parameterized linear regression models generalize without violating learning theory principles.
method The paper uses the predictive normalized maximum likelihood (pNML) learner to investigate the minimum norm solution of over-parameterized linear regression models.
result The model generalizes well when the test sample lies in a subspace spanned by eigenvectors associated with large eigenvalues of the training data.

Task shift from classification to regression is possible in overparameterized linear models with limited additional data.

problem Transferability of latent knowledge from classification to regression in overparameterized linear models.
method Investigation of task shift in overparameterized linear regression, zero-shot and few-shot cases, with a focus on minimum-norm interpolation.
result Minimum-norm interpolators can transfer latent knowledge from classification to regression with limited additional data.

The study analyzes how covariance estimation errors affect the global minimum-variance portfolio under heavy-tailed distributions.

problem The impact of covariance estimation errors on the global minimum-variance portfolio under heavy-tailed distributions.
method Characterization of covariance-estimation error's effect on GMVP suboptimality, derivation of regret identity and bound, application to heavy-tailed returns.
result The decision geometry of GMVP regret is invariant to a (p-1)-dimensional projection of the error matrix, with invariance to the covariance-scale direction as an exact special case.

Optimal ridge penalty can be negative or zero in high-dimensional data.

problem Overfitting in high-dimensional underdetermined linear regression.
method Simulations and real-life data analysis with minimum-norm estimator.
result Optimal ridge penalty can be negative, contradicting conventional wisdom.

Develops exact convex optimization formulations for neural networks.

problem Training two-layer neural networks with rectified linear units.
method Uses semi-infinite duality and minimum norm regularization to develop exact convex optimization formulations.
result Shows equivalence of ReLU networks trained with weight decay to block 1\ell_1 penalized convex models.

The paper characterizes functions of shallow ReLU NN denoisers under minimal norm constraints.

problem Understanding the theoretical success of neural network denoisers.
method Characterization of functions realized by shallow ReLU NN denoisers under minimal norm constraints.
result The functions realized by shallow ReLU NN denoisers are contractive toward clean data points and generalize better than the empirical MMSE estimator at low noise levels.

Exact expressions for double descent and implicit regularization in over-parameterized models.

problem Understanding the generalization error of over-parameterized models like deep neural networks.
method Surrogate random design to replace standard i.i.d. design, leading to exact expressions for mean squared error and implicit regularization.
result Exact non-asymptotic expressions for double descent and implicit regularization in over-parameterized models.

This study explains gradient flow dynamics in neural networks for small initialisation.

problem Understanding the training dynamics of neural networks for small initialisation.
method Analysis of gradient flow dynamics for one-hidden layer ReLU networks with orthogonal inputs.
result Gradient flow converges to zero loss and characterizes implicit bias towards minimum variation norm.

SAM optimizes deep networks by oscillating between sides of the minimum.

problem Improving performance of deep networks.
method Gradient-based optimization method that oscillates between sides of the minimum.
result SAM effectively performs gradient descent on the spectral norm of the Hessian, encouraging drift towards wider minima.

Solves a triangulation problem by showing minimum tetrahedra equals minimum integral 3-chain.

problem Finding the minimum number of tetrahedra to extend a triangulation of a 2-sphere to a 3-ball.
method Relates the minimum number of tetrahedra to the minimum integral 3-chain norm, proving them equal and showing how to achieve the minimum.
result The minimum number of tetrahedra needed to extend a triangulation of a 2-sphere to a 3-ball equals the minimum integral 3-chain norm.

To a compact Riemann surface of genus g can be assigned a principally polarized abelian variety (PPAV) of dimension g, the Jacobian of the Riemann surface. The Schottky problem is to discern the Jacobians among the PPAVs. Buser and Sarnak showed, that the square of the first successive minimum, the squared norm of the …

2010-08-12abs ↗pdf ↗

The paper studies co-Hamiltonian diffeomorphisms on compact cosymplectic manifolds.

problem Fix-point theory and co-Hamiltonian diffeomorphisms on compact cosymplectic manifolds.
method Fix-point theory, Arnold's conjecture, co-Hofer norms, topologies, approximations lemmas.
result Minimum number of fix points for co-Hamiltonian diffeomorphisms is at least 1.

The paper analyzes boosting and minimum-1\ell_1-norm classifiers in high dimensions.

problem Understanding the generalization error and optimal Bayes error in boosting.
method High-dimensional asymptotic theory, Gaussian comparison techniques, uniform deviation argument.
result Precise characterizations of boosting test error and optimal Bayes error.

New approach finds minimum width for deep, narrow MLPs.

problem Finding the minimum width for deep, narrow MLPs to approximate continuous functions.
method Proposes a framework to simplify finding minimum width into determining a geometrical function w(dx,dy)w(d_x, d_y) based on input and output dimensions.
result Proves that w(dx,dy)w(d_x, d_y) equals the optimal minimum width for deep, narrow MLPs to achieve universality.

The study examines how shallow neural nets converge to training samples or manifold points during diffusion.

problem Understanding when and how shallow neural nets converge to training samples or manifold points during diffusion.
method Analysis of shallow ReLU neural network denoisers trained with minimal 2\ell^2 norm, comparing score flow and diffusion flow.
result Probability flow converges to training points, sums of training points, or manifold points, depending on the diffusion time scheduler.