Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

57114170227 · Jun 202019922001200920182026
48 results for nondifferentiable regularizers

Representer theorem extended to Banach spaces for nondifferentiable regularizers.

problem Conditions for a representer theorem in Banach spaces.
method Generalized representer theorem to Banach spaces for nondifferentiable regularizers.
result Necessary and sufficient conditions for a representer theorem in Banach spaces.

Study nondifferentiable metrics in general relativity, resolving causality issues and limits evolution scenarios.

problem Causality issues and evolution scenarios in black hole interiors with closed timelike geodesics.
method Method of equivalence on Courant algebroids to derive new differential invariants.
result Resolved causality issues and limited evolution scenarios for gravitational collapse.

We introduce a recursive adaptive group lasso algorithm for real-time penalized least squares prediction that produces a time sequence of optimal sparse predictor coefficient vectors. At each time index the proposed algorithm computes an exact update of the optimal 1,\ell_{1,\infty}-penalized recursive least squares (R…

2011-01-29abs ↗pdf ↗

This manuscript develops the theory of agglomerative clustering with Bregman divergences. Geometric smoothing techniques are developed to deal with degenerate clusters. To allow for cluster models based on exponential families with overcomplete representations, Bregman divergences are developed for nondifferentiable co…

2012-06-27abs ↗pdf ↗

Study efficient derivative computation for nondifferentiable maps in machine learning.

problem Efficiently compute derivatives of fixed-point of nondifferentiable contractions.
method Iterative Differentiation (ITD), Approximate Implicit Differentiation (AID), and New Stochastic Implicit Differentiation (NSID).
result Established convergence rates for ITD, AID, and NSID, matching or improving smooth setting rates.

In this paper, we consider an 0\ell_{0}-norm penalized formulation of the generalized eigenvalue problem (GEP), aimed at extracting the leading sparse generalized eigenvector of a matrix pair. The formulation involves maximization of a discontinuous nonconcave objective function over a nonconvex constraint set, and is…

2014-08-28abs ↗pdf ↗

Algorithm checks local optimality and escapes saddles in ReLU networks.

problem Checking local optimality and escaping saddles in ReLU networks with nondifferentiable points.
method Polyhedral geometry to reduce complexity, exploiting convex and nonconvex QPs.
result Algorithm efficiently solves local optimality and saddle point issues in ReLU networks.

The paper optimizes bridge-type estimators for sparse models using pathwise methods.

problem Sparse parametric models with adaptive coefficients and multiple penalties.
method Pathwise optimization with accelerated proximal gradient descent and blockwise alternating optimization.
result Efficient computation of the full solution path for adaptive bridge estimators.

AEN-SAEs address feature starvation in sparse autoencoders by stabilizing the geometric alignment of sparse coding.

problem Feature starvation in sparse autoencoders, leading to unstable and misaligned representations.
method Adaptive Elastic Net SAEs (AEN-SAEs) combine 2\ell_2 and 1\ell_1 terms to stabilize the sparse coding map and control feature interactions.
result AEN-SAEs mitigate feature starvation without heuristic resampling, maintaining competitive reconstruction abilities.

Large-scale generalized linear array models (GLAMs) can be challenging to fit. Computation and storage of its tensor product design matrix can be impossible due to time and memory constraints, and previously considered design matrix free algorithms do not scale well with the dimension of the parameter vector. A new des…

2015-10-12abs ↗pdf ↗

We present a distributed proximal-gradient method for optimizing the average of convex functions, each of which is the private local objective of an agent in a network with time-varying topology. The local objectives have distinct differentiable components, but they share a common nondifferentiable component, which has…

2012-10-08abs ↗pdf ↗

A model verifies during classification to reduce memorization, matching baseline accuracy with fewer parameters.

problem Classification systems require memorizing all classes, leading to increased memory usage and poor sample efficiency.
method Iterative nondifferentiable queries for verification during classification, balancing recognition and verification.
result The model can match baseline accuracy while using fewer parameters, but requires careful balance between recognition and verification.

This paper analyzes Stein variational gradient descent for Bayesian inference.

problem Sampling or approximating high-dimensional probability distributions.
method Iterated steepest descent steps with a reproducing kernel Hilbert space norm.
result Performance gains of certain nondifferentiable kernels with adjusted tails.

GD-trained shallow ReLU nets learn Lipschitz functions with noise.

problem Learning Lipschitz functions with additive noise in overparameterized neural networks.
method Gradient Descent (GD) with early stopping, focusing on the Neural Tangent Kernel (NTK).
result Early-stopped GD achieves minimax optimal rates for learning Lipschitz functions.

Paper proposes a new method for SP with covariates using PADR and ERM.

problem Stochastic programming with covariate information.
method Empirical risk minimization (ERM) with nonconvex piecewise affine decision rules (PADR).
result The method provides theoretical consistency and computational tractability for nonconvex SP problems.

Piecewise linear activations create many spurious local minima in neural networks.

problem Understanding the loss surface of neural networks with piecewise linear activations.
method Proved the existence of infinite spurious local minima and partitioned the loss surface into smooth cells.
result Piecewise linear activations create many spurious local minima that are invariant under a continuous path.

A new approach for efficient batch multiobjective optimization using Thompson sampling.

problem Inefficient batch multiobjective optimization due to expensive oracles and hard inner optimization.
method Proposes a Thompson sampling approach (qextttPOTSq exttt{POTS}) that chooses Pareto optimal candidates sequentially.
result Empirically superior performance compared to classical evolutionary approaches and MOBO.

A novel algorithm converges for solving a specific matrix decomposition problem.

problem Nonlinear matrix decomposition with ReLU function for sparse data.
method Introduced a reparametrization of the Latent-RMD model and developed eBCD for convergence proof.
result eBCD converges and outperforms state-of-the-art methods on various data sets.

ParMAC optimizes nested functions in distributed systems, improving parallelism and speed.

problem Optimizing complex, nested machine learning models with large datasets and distributed computing.
method ParMAC introduces auxiliary coordinates for parallel training of submodels and coordinates, reducing communication overhead.
result ParMAC achieves high parallelism and low communication overhead, facilitating fast training of binary autoencoders.

AIR-Net adapts low-rank regularization dynamically for better image completion.

problem Fixed low-rank regularization limits adaptability to different images.
method AIR-Net uses adaptive and implicit regularization parameterized by a dynamic Laplacian matrix.
result AIR-Net enhances implicit regularization and outperforms fixed methods in non-uniform missing data scenarios.

The article explores toric spaces of regular polyhedra, highlighting rational and non-rational cases.

problem Exploring toric spaces associated with regular convex polyhedra.
method Symplectic and complex toric spaces associated with five regular convex polyhedra.
result The regular dodecahedron and icosahedron cannot be treated via standard toric geometry.

Regularized deep networks improve generalization and robustness.

problem Improving generalization and robustness of deep neural networks.
method Input gradient regularization combined with Lipschitz and adversarial robustness.
result Regularized models show improved adversarial robustness and generalization.

Gradient descent implicitly regularizes neural networks by penalizing large loss gradients.

problem How to optimize deep neural networks without explicit regularization.
method Backward error analysis to calculate implicit gradient regularization and demonstrate its effectiveness empirically.
result Implicit gradient regularization biases gradient descent toward flat minima, improving model robustness and test errors.

Choquet regularization improves exploration in RL.

problem Improving exploration in reinforcement learning.
method Introducing Choquet regularizers to measure and manage exploration, reformulating RL problems and deriving explicit solutions.
result Explicit optimal distributions and Choquet regularizers for various exploratory samplers.

Gradient-coherent strong regularization improves deep neural networks' generalization.

problem Deep neural networks overfit with strong L1/L2 regularization.
method Imposes regularization only when gradients are coherent, using stochastic gradient descent.
result Significantly improves accuracy and compression (up to 9.9x).

The paper explores optimal regularizers for data sources, linking them to star bodies.

problem Understanding optimal regularizers for data sources.
method Investigates optimal regularizers for data distributions using star bodies and dual Brunn-Minkowski theory.
result Identifies optimal regularizers and assesses amenability to convex regularization.

A triangulation of a connected closed surface is called weakly regular if the action of its automorphism group on its vertices is transitive. A triangulation of a connected closed surface is called degree-regular if each of its vertices have the same degree. Clearly, a weakly regular triangulation is degree-regular. In…

2004-03-25abs ↗pdf ↗

The paper proves existence and multiplicity of affine connections on regular manifolds.

problem Existence and multiplicity of affine connections on regular manifolds.
method Regularity theory and properties of the structural presheaf.
result The space of regular affine connections is an affine space of the space of regular End(TM)\operatorname{End}(TM)-valued 1-forms.

Study uses property elicitation to understand how fairness regularizers affect optimal decisions.

problem Understanding how fairness regularizers change the optimal decision in predictive algorithms.
method Property elicitation to analyze the relationship between loss, regularization, and optimal decision.
result Necessary and sufficient condition for when a property changes with the addition of a regularizer.

Fiedler regularization uses spectral graph theory to improve neural network performance.

problem Improving neural network performance by penalizing weights based on connectivity.
method Uses the Fiedler value of the neural network's graph as a regularization tool, providing theoretical and computational methods.
result Demonstrates Fiedler regularization's effectiveness in improving neural network performance.

New input gradient regularization improves adversarial robustness efficiently.

problem Improving adversarial robustness in machine learning models.
method Derive robustness bounds, implement scaleable input gradient regularization, avoid double backpropagation.
result Input gradient regularization is competitive with adversarial training and avoids gradient obfuscation.

Dropout is a simple but effective technique for learning in neural networks and other settings. A sound theoretical understanding of dropout is needed to determine when dropout should be applied and how to use it most effectively. In this paper we continue the exploration of dropout as a regularizer pioneered by Wager,…

2014-12-15abs ↗pdf ↗

We establish continuous maximal regularity results for parabolic differential operators acting on sections of tensor bundles on Riemannian manifolds. As an application, we show that solutions to the Yamabe flow instantaneously regularize and become real analytic in space and time. The regularity result is obtained by i…

2013-09-09abs ↗pdf ↗

Improved optimal regularity for harmonic almost complex structures.

problem Establishing optimal regularity for harmonic almost complex structures.
method Quantitative stratification method and rectifiability of singular strata.
result Optimal regularity theory for energy minimizing harmonic almost complex structures.