Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

255176101 · May 202619922001200920182026
48 results for identity parameterization

Identity parameterization simplifies deep learning models and improves performance.

problem Designing deep neural networks with stable and expressive architectures.
method Theoretical analysis and empirical validation of residual networks with identity parameterization.
result Residual networks with identity parameterization have no spurious local optima and universal expressivity.

We briefly review the current situation with various relations between knot/braid polynomials (Chern-Simons correlation functions), ordinary and extended, considered as functions of the representation and of the knot topology. These include linear skein relations, quadratic Plucker relations, as well as "differential" …

2012-08-10abs ↗pdf ↗

Stochastic descent methods achieve optimal performance in certain conditions.

problem Optimizing deep learning models for general loss functions and nonlinear models.
method Revisits minimax properties of stochastic gradient descent and mirror descent.
result Stochastic mirror descent (and gradient descent) is minimax optimal under certain conditions.

Over-parameterized models reduce Out-of-Distribution (OOD) generalization loss.

problem Understanding how over-parameterized models handle non-trivial distributional shifts.
method Investigating random feature models and examining non-trivial natural distributional shifts.
result Increasing model parameterization reduces OOD loss.

The least absolute shrinkage and selection operator (lasso) and ridge regression produce usually different estimates although input, loss function and parameterization of the penalty are identical. In this paper we look for ridge and lasso models with identical solution set. It turns out, that the lasso model with shri…

2014-01-10abs ↗pdf ↗

Improved covariance matrix estimation for portfolio optimization with guaranteed PSD and controlled conditioning.

problem Guaranteeing positive semidefinite ness and controlling spectral conditioning in IQ estimators.
method Introducing squeezing identity and atomic-IQ parameterization to construct structured channel matrices with PSD guarantees and analytic eigen floor for conditioning control.
result Atomic-IQ improves Sharpe ratios and delivers a more stable risk profile compared to standard estimators.

Ensembles of neural networks improve training dynamics and performance.

problem Improving neural network performance through model size increase.
method Defining collegial ensembles (CE) as multiple independent models trained as a single model, and using theoretical results on NTK to optimize architecture search.
result CE dynamics simplify and scale favorably, resembling wide models, and can be efficiently implemented using group convolutions and block diagonal layers.

Gradient descent with identity initialization learns positive definite linear transformations efficiently.

problem Learning positive definite linear transformations using gradient descent.
method Gradient descent with identity initialization, analyzing the population quadratic loss.
result Gradient descent with identity initialization efficiently learns positive definite linear transformations.

This paper was first written in 1990, but was never published. In it, the author presents a novel approach to the study of constant curvature spacetimes in 2+1 dimensions. A parameterization of flat 2+1-dimensional domains of dependence is given in terms of measured geodesic laminations. There is also an interesting re…

2007-06-11abs ↗pdf ↗

FedProx tackles heterogeneity in federated learning networks.

problem Significant variability in systems characteristics and non-identically distributed data in federated networks.
method FedProx is a framework that generalizes and re-parametrizes FedAvg, introducing modifications to handle both systems and statistical heterogeneity.
result FedProx demonstrates significantly more stable and accurate convergence behavior than FedAvg, improving test accuracy by 22% on average in highly heterogeneous settings.

Improved standard parameterization yields well-defined neural tangent kernel.

problem Extrapolation of standard parameterization to infinite width is problematic.
method Proposed an improved extrapolation of the standard parameterization.
result Improved standard parameterization yields similar accuracy to NTK parameterization but with better correspondence to finite width networks.

GANs improve stochastic parameterization of the Lorenz '96 model.

problem Improving stochastic parameterizations for sub-grid processes.
method Developed a GAN-based stochastic parameterization for the Lorenz '96 model.
result GAN configurations outperform a bespoke parameterization in skillful forecasts and climate simulations.

Study on likelihood functions, associative equations, and Frobenius manifolds.

problem Maximum likelihood estimation and associativity equations in statistical models.
method Analyzes the cone of concentration matrices, log-likelihood function, and Frobenius manifolds.
result Maximum likelihood degree is indexed by components of Frobenius residuals.

Parallel algorithm for conformal parameterization of 3D surfaces.

problem Computational difficulties with high-resolution 3D surface meshes.
method Partitioning surfaces into subdomains, parallel local parameterization, partial welding for boundary integration, solving Laplace equation.
result Significant improvement in computational time and accuracy compared to existing methods.

Bayesian models' singular fluctuation is shown to be akin to specific heat, influencing model complexity and generalization.

problem Understanding the thermodynamic interpretation of singular fluctuation in Bayesian models.
method Showed singular fluctuation as the curvature of Bayesian free energy and variance of log-likelihood observable under a Gibbs posterior.
result Singular fluctuation is the statistical analogue of specific heat, controlling model complexity and generalization.

We study R-covered foliations of 3-manifolds from the point of view of their transverse geometry. For an R-covered foliation in an atoroidal 3-manifold M, we show that M-tilde can be partially compactified by a canonical cylinder S^1_univ x R on which pi_1(M) acts by elements of Homeo(S^1) x Homeo(R), where the S^1 fac…

1999-03-29abs ↗pdf ↗

The current paper discusses some new results about conformal polynomic surface parameterizations. A new theorem is proved: Given a conformal polynomic surface parameterization of any degree it must be harmonic on each component. As a first geometrical application, every surface that admits a conformal polynomic paramet…

2012-05-25abs ↗pdf ↗

The paper presents efficient methods for identifying causal graphs with latent variables.

problem Recovering causal graphs with latent variables while minimizing intervention costs.
method Two intervention cost models (linear and identity) are considered. Algorithms are provided for both models.
result Upper bounds on the number of interventions needed for recovery, and approximation factors for the linear cost model.

The paper analyzes how over-parameterization affects GD convergence in matrix sensing problems.

problem Matrix sensing problem with over-parameterized gradient descent.
method Analyzes symmetric and asymmetric parameterizations, provides lower bounds and convergence rates.
result Over-parameterization slows down GD convergence, but asymmetric parameterization can speed up convergence.

Novel framework for policy optimization with general parameterization and linear convergence.

problem Lack of theoretical guarantees for policy optimization with general parameterization schemes.
method Mirror descent approach for policy optimization with general parameterization.
result First result of linear convergence for policy-gradient-based method with general parameterization.

Improved loss scaling for stochastic momentum algorithms in high dimensions.

problem Improving loss scaling for stochastic momentum algorithms in high dimensions.
method Dimension-adapted Nesterov acceleration (DANA) scales momentum hyperparameters based on model size and data complexity.
result DANA improves loss scaling exponents across various data and target complexities.

We introduce a new parameterization method for deep learning layers using spectral tensor train decomposition.

problem Efficiency and stability in deep learning models with weight matrix compression.
method Spectral Tensor Train Parameterization (STTP) of weight matrices.
result Improved compression and training stability in neural networks.

Develops a method for conformal parameterization of point clouds without fixed boundaries.

problem Desirable distortion in fixed-boundary parameterizations of point clouds.
method Free-boundary conformal parameterization method involving approximation of point cloud Laplacian and boundary treatment.
result High-quality point cloud meshing achieved through the proposed method.

Paper introduces NICc for fast cluster-based validation of prediction models.

problem Validation of prediction models on clustered data.
method Derived NICc to approximate leave-one-cluster-out deviance for standard regression models.
result NICc provides more accurate model size and variable selection, especially with strong clustering.

Improves performance in various machine learning tasks by reparameterizing subset sampling.

problem Stochastic optimization involving subset sampling is not reparameterizable.
method Continuous relaxation of subset sampling to provide reparameterization gradients.
result Improves performance in instance-wise feature selection, deep stochastic k-nearest neighbors, and parametric t-SNE.

Gradient descent achieves good generalization for over-parameterized deep ReLU networks.

problem Understanding good generalization in over-parameterized deep neural networks.
method Algorithm-dependent generalization error bound for deep ReLU networks using gradient descent.
result Gradient descent with proper initialization can achieve arbitrarily small generalization error for over-parameterized DNNs.

Least squares regression shows unexpected double descent in under-parameterized models.

problem Understanding the generalization of under-parameterized models in regression.
method Analyzing the spectrum and eigenvectors of the sample covariance matrix.
result Least squares regression can exhibit a peak in generalization in the under-parameterized regime, contrary to previous explanations.

The paper proposes methods for volumetric parameterization of 3D solid manifolds.

problem Complex structure of solid manifolds makes conventional approaches ineffective.
method Incorporates models to preserve geometric structure, achieve density equalization, and balance distortions.
result Various 3D manifold parameterizations with different properties can be achieved.

Fast algorithm for spherical parameterization of surfaces with adaptive remeshing.

problem Efficiently parameterizing genus-0 closed surfaces with user-defined quasiconformal distortion.
method Proposes a fast algorithm for spherical quasiconformal parameterization.
result Effective for adaptive surface remeshing in computer graphics and animations.

Study shows how over-parameterized classifiers can still perform well on noisy data.

problem Understanding how maximum margin classifiers perform in over-parameterized settings with noisy data.
method Analyzes maximum margin classifiers on sub-Gaussian mixtures, providing risk bounds.
result Characterizes conditions for 'benign overfitting' in linear classification problems.

This is a survey of the theory of complex projective (CP^1) structures on compact surfaces. After some preliminary discussion and definitions, we concentrate on three main topics: (1) Using the Schwarzian derivative to parameterize the moduli space (2) Thurston's parameterization of the moduli space using grafting (3) …

2009-02-11abs ↗pdf ↗

Gradient descent recovers low-rank matrices from corrupted measurements with double over-parameterization.

problem Robust recovery of low-rank matrices from grossly corrupted measurements.
method Gradient descent with discrepant learning rates for double over-parameterized models.
result Gradient descent with discrepant learning rates provably recovers the underlying matrix without prior knowledge on rank or sparsity.

This paper discovers new identities linking geodesic and orthogeodesic lengths on hyperbolic surfaces.

problem Understanding relationships between geodesic and orthogeodesic lengths on hyperbolic surfaces.
method Investigates a broad family of identities involving lengths of all closed geodesics and orthogeodesics.
result Introduces new identities that include lengths of all closed geodesics, contrasting with previous identities.

CAVI speeds up Bayesian MIDAS regression by 107x-1,772x with similar accuracy.

problem Efficiently estimating Bayesian MIDAS regression models with many predictors.
method Coordinate Ascent Variational Inference (CAVI) for linear MIDAS regression.
result CAVI produces posterior means nearly identical to Gibbs sampling with significant speedup.

Gradient descent slows significantly in over-parameterized single neuron learning.

problem Learning a single neuron with over-parameterization and square loss.
method Analysis of gradient descent dynamics, proving convergence rates and lower bounds.
result Over-parameterization can exponentially slow down the convergence rate of gradient descent.