Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

98197295393 · Jun 202019922001200920172026
48 results for Transformed ℓ1 Regularization

The p-Laplacian Transformer improves transformer models by assigning higher attention weights to tokens in close proximity.

problem The self-attention mechanism in transformers does not effectively distinguish attention weights between tokens in close and non-close proximity.
method Proposes a novel class of transformers, p-Laplacian Transformers, that use pp-Laplacian regularization to assign higher attention weights to tokens in close proximity.
result Empirically demonstrates that p-Laplacian Transformers outperform baseline transformers on various benchmark datasets.

Transformer learns context and regularization for ICL in inverse problems.

problem Learning context and effective regularization for transformer-based in-context learning (ICL) in inverse problems.
method Introduced a linear transformer to learn inverse mapping from contextual examples to weight vectors, addressing rank-deficient problems.
result Transformer implicitly learns a prior distribution and effective regularization strategy, outperforming traditional methods.

A new method matches measures across different spaces using cost-regularized optimal transport.

problem Matching measures in different spaces without aligned data.
method Cost-regularized optimal transport formulation to match measures across two Euclidean spaces.
result Demonstrated applicability to single-cell spatial transcriptomics/multiomics matching tasks.

The paper investigates polynomial alternatives to softmax in transformer models.

problem The effectiveness of softmax attention in transformers is questioned.
method The authors explore polynomial activations as alternatives to softmax, focusing on their ability to regularize the attention matrix.
result Certain polynomials can serve as effective substitutes for softmax in transformer applications, achieving strong performance.

Sparse transformer architecture improves accuracy and speed in generative modeling and inverse problems.

problem Improving accuracy and speed in generative modeling and inverse problems.
method Proposes a sparse transformer architecture using regularized Wasserstein proximal operator with L1L_1 prior.
result Sparse transformer achieves higher accuracy and faster convergence than classical methods.

Transformations of macroeconomic data affect machine learning forecasts, especially with regularization and nonlinearity.

problem The impact of data transformations on machine learning forecasts in macroeconomic contexts.
method Review and propose new data transformations, empirically evaluate their effects, and compare traditional and moving average rotations.
result Traditional factors should almost always be included as predictors, and moving average rotations can provide important gains.

Injectivity of geodesic X-ray transform on low-regularity manifolds.

problem Injectivity of geodesic X-ray transform on manifolds with low regularity.
method Calculus of differential and curvature operators on non-smooth structures.
result Injectivity of geodesic X-ray transform on simple Riemannian manifolds with C1,1C^{1,1}-regularity.

Geometric structures on quaternionic unit ball for slice regular Möbius transformations.

problem No new problem introduced.
method Introducing Hermitian, Riemannian, and Kähler-like structures on quaternionic unit ball using regular Möbius transformations.
result Geometric structures are natural generalizations of complex setup and solve problems not achieved by other geometries.

Researchers improve transformer networks' optimization and understanding.

problem Improving the understanding and optimization of transformer networks.
method Introducing a convex alternative to the self-attention mechanism and reformulating the training problem as a convex optimization problem.
result Revealed an implicit regularization mechanism that promotes sparsity across tokens.

Paper proves injectivity of non-abelian X-ray transform on certain spaces.

problem Injectivity of non-abelian X-ray transform on asymptotically hyperbolic spaces.
method Gauge equivalence for unitary connections and skew-Hermitian Higgs fields.
result Injectivity result for non-abelian X-ray transform over skew-Hermitian Higgs fields.

Mixup improves model accuracy and calibration through data transformation and random perturbation.

problem Improving model accuracy and calibration in machine learning.
method Interprets Mixup as empirical risk minimization with data transformation and random perturbation.
result Mixup induces multiple known regularization schemes that prevent overfitting and overconfident predictions.

In order to understand the linearization problem around a leaf of a singular foliation, we extend the familiar holonomy map from the case of regular foliations to the case of singular foliations. To this aim we introduce the notion of holonomy transformation. Unlike the regular case, holonomy transformations can not be…

2012-05-27abs ↗pdf ↗

Study anisotropic obstacle problem for minimal surfaces using Cahn-Hoffman transform.

problem Anisotropic obstacle problem for minimal surfaces.
method Cahn-Hoffman transform to convert to isotropic problem with generalized Robin boundary condition.
result Optimal regularity of the solution and C1,1C^{1,1} regularity of the free boundary.

Geometrically transforms nonconservative dynamics to linearize Kepler and Manev systems.

problem Regularizing and linearizing nonconservative central force dynamics.
method Projective transformation and conformal scaling in configuration and phase spaces.
result Full linearization of Kepler and Manev dynamics in any finite dimension.

Study proves solenoidal injectivity for tensor fields on curved manifolds with low regularity.

problem Injectivity for tensor fields on negatively curved manifolds with low regularity metrics.
method Pestov energy estimates for transport equation on non-smooth unit sphere bundle, keeping track of regularity, and using functions with more vertical than horizontal regularity.
result Proves solenoidal injectivity for tensor fields on simple Riemannian manifolds with C1,1C^{1,1} metrics and non-positive sectional curvature.

Transformers with linear space and time complexity for accurate attention estimation.

problem Efficiently estimating attention in large-scale tasks without relying on priors.
method Performers use Fast Attention Via positive Orthogonal Random features (FAVOR+) for linear approximation of softmax attention.
result Performers achieve competitive results on various tasks, demonstrating the effectiveness of their attention-learning approach.

Study the Lax equation in infinite-dimensional Lie algebras and Lie groups.

problem Investigate the Lax equation in infinite-dimensional Lie algebras and Lie groups.
method Derived integral expansions and generalized Baker-Campbell-Hausdorff formula for Lie groups.
result Explicit representation of product integral in terms of exponential map.

Strongly polynomial algorithm for approximate Forster transforms and halfspace learning.

problem Computing approximate Forster transforms and halfspace learning.
method Strongly polynomial time algorithm for approximate Forster transforms and halfspace learning.
result First strongly polynomial time algorithm for distribution-free PAC learning of halfspaces.

Regularizes GAMs to improve interpretability by reducing concurvity.

problem Susceptibility of GAMs to concurvity reduces interpretability.
method Proposes a regularizer to penalize pairwise correlations of non-linearly transformed features.
result Improves interpretability and reduces concurvity without sacrificing prediction quality.

Let MM be a complete non-compact Riemannian manifold. In this paper, we derive sufficient conditions on metric perturbation for stability of LpL^p-boundedness of the Riesz transform, p(2,)p\in (2,\infty). We also provide counter-examples regarding in-stability for LpL^p-boundedness of Riesz transform.

2018-08-06abs ↗pdf ↗

Transformers model contextual relations using probabilistic measures, revealing their expressive power.

problem Lack of clear understanding of Transformer's ability to model contextual relations.
method Introduced a measure-theoretic framework connecting softmax attention and entropy-regularized optimal transport.
result Transformer architectures can approximate arbitrary contextual relations, and the choice of normalization affects how these relations are represented.

We show injectivity of the X-ray transform and the dd-plane Radon transform for distributions on the nn-torus, lowering the regularity assumption in the recent work by Abouelaz and Rouvière. We also show solenoidal injectivity of the X-ray transform on the nn-torus for tensor fields of any order, allowing the tensor…

2014-02-25abs ↗pdf ↗

This paper analyzes convergence of large-scale Transformers with weight decay.

problem Understanding optimization guarantees in large-scale Transformer training.
method Construct mean-field limit, show gradient flow convergence to PDE, demonstrate global minimum consistency.
result Gradient flow reaches global minimum in large-scale Transformers with small weight decay.

Given a slice regular function f:ΩHHf:Ω\subset\mathbb{H}\to \mathbb{H}, with ΩRΩ\cap\mathbb{R}\neq \emptyset, it is possible to lift it to a surface in the twistor space CP3\mathbb{CP}^{3} of S4H{}\mathbb{S}^4\simeq \mathbb{H}\cup \{\infty\} (see~\cite{gensalsto}). In this paper we show that the same result is true if one rem…

2016-05-27abs ↗pdf ↗

This work improves Fourier pricing for multi-asset options using RQMC with domain transformation.

problem Efficiently pricing multi-asset options in high dimensions with Fourier methods.
method Randomized quasi-Monte Carlo (RQMC) with domain transformation to handle singularities.
result RQMC with domain transformation provides accurate and scalable Fourier pricing for multi-asset options.

In this paper, we study a fast approximation method for {\it large-scale high-dimensional} sparse least-squares regression problem by exploiting the Johnson-Lindenstrauss (JL) transforms, which embed a set of high-dimensional vectors into a low-dimensional space. In particular, we propose to apply the JL transforms to …

2015-07-18abs ↗pdf ↗

Following Burstall and Hertrich-Jeromin we study the Ribaucour transformation of Legendre submanifolds in Lie sphere geometry. We give an explicit parametrization of the resulted Legendre submanifold F^\hat{F} of a Ribaucour transformation, via a single real function ττ which represents the regular Ribaucour sphere co…

2013-03-08abs ↗pdf ↗

GTMs model complex multivariate data with varying conditional independencies.

problem Modeling multivariate data with intricate marginals and complex dependency structures.
method Semiparametric approach using penalized splines and lasso regularization.
result GTMs accurately learn complex dependencies and identify conditional independencies.

The version of Marsden-Ratiu reduction theorem for Nambu-Poisson manifolds by a regular distribution has been studied by Ibaˊn~\acute{\text{a}}\tilde{\text{n}}ez et al. In this paper we show that the reduction is always ensured unless the distribution is zero. Next we extend the more general Falceto-Zambon Poisson reduct…

2017-02-06abs ↗pdf ↗

Transformers preserve support and can approximate any continuous map.

problem Understanding the mathematical properties of transformers.
method Characterizing maps between measures that can be represented as transformers and proving their properties.
result Transformers preserve support and have uniformly continuous Fréchet derivatives.

This study analyzes how one-layer transformers learn regular language recognition tasks.

problem Understanding how one-layer transformers solve regular language recognition tasks like even pairs and parity check.
method Theoretical analysis of training dynamics and gradient descent for a one-layer transformer.
result A one-layer transformer can solve even pairs directly but needs CoT for parity check. Training phases show rapid growth in attention layer followed by logarithmic growth in linear layer.

Classifies regularity for Lagrangian mean curvature type equations.

problem Classifying regularity for Lagrangian mean curvature type equations.
method Generalized constant rank theorem for Legendre transform, constructed convex solutions, and showed regularity conditions.
result Optimal regularity conditions for Lagrangian mean curvature type equations.

Extends optimal regularity and compactness to vector bundles over non-Riemannian manifolds.

problem Optimal regularity and compactness for connections on vector bundles.
method Derive RT-equations, establish existence theory, handle curvature up to L1L^1.
result Optimal regularity and compactness extended to vector bundles over non-Riemannian manifolds.

Dual optimization connects ERM-fDR to normalization function.

problem Empirical risk minimization with f-divergence regularization.
method Dual formulation, Legendre-Fenchel transform, implicit function theorem, nonlinear ODE.
result Computational method to calculate normalization function efficiently.

This work studies clustering in transformer models, proving exponential convergence to a single token state.

problem Understanding the long-term behavior of tokens in transformer models.
method Investigates mean-field transformer models under specific conditions to prove exponential convergence to a single state.
result Transformer models synchronize exponentially fast to a single token state with explicit rates.

We present authors' new theory of the RT-equations, nonlinear elliptic partial differential equations which determine the coordinate transformations which smooth connections ΓΓ to optimal regularity, one derivative smoother than the Riemann curvature tensor Riem(Γ){\rm Riem}(Γ). As one application we extend Uhlenbeck compa…

2018-12-14abs ↗pdf ↗

Smooth manifold structure on Möbius transformations of quaternionic ball identified.

problem Identifying the manifold structure of Möbius transformations of quaternionic unit ball.
method Realizing M(B)\mathcal{M}(\mathbb{B}) as a quotient of Sp(1,1)\mathrm{Sp}(1,1) and using Lie group properties.
result The manifold M(B)\mathcal{M}(\mathbb{B}) is diffeomorphic to R4imesS3\mathbb{R}^4 imes S^3.