Regularity results for geodesic X-ray transform on nonsmooth manifolds
problem Geodesic X-ray transform on nonsmooth simple manifolds
method Symbol smoothing arguments and pseudodifferential operators with low regularity symbols
result Improved injectivity results for Lp functions The p-Laplacian Transformer improves transformer models by assigning higher attention weights to tokens in close proximity.
problem The self-attention mechanism in transformers does not effectively distinguish attention weights between tokens in close and non-close proximity.
method Proposes a novel class of transformers, p-Laplacian Transformers, that use p-Laplacian regularization to assign higher attention weights to tokens in close proximity. result Empirically demonstrates that p-Laplacian Transformers outperform baseline transformers on various benchmark datasets.
Novel regularization for Vision Transformers improves model generalization and sparsity.
problem Improving generalization and sparsity in Vision Transformers.
method Likelihood-guided variational Ising-based regularization.
result Improved generalization and sparsity in Vision Transformers.
We study ray transforms on spherically symmetric manifolds with a piecewise C1,1 metric. Assuming the Herglotz condition, the X-ray transform is injective on the space of L2 functions on such manifolds. We also prove injectivity results for broken ray transforms (with and without periodicity) on such manifolds …
Transformer learns context and regularization for ICL in inverse problems.
problem Learning context and effective regularization for transformer-based in-context learning (ICL) in inverse problems.
method Introduced a linear transformer to learn inverse mapping from contextual examples to weight vectors, addressing rank-deficient problems.
result Transformer implicitly learns a prior distribution and effective regularization strategy, outperforming traditional methods.
A new method matches measures across different spaces using cost-regularized optimal transport.
problem Matching measures in different spaces without aligned data.
method Cost-regularized optimal transport formulation to match measures across two Euclidean spaces.
result Demonstrated applicability to single-cell spatial transcriptomics/multiomics matching tasks.
The paper investigates polynomial alternatives to softmax in transformer models.
problem The effectiveness of softmax attention in transformers is questioned.
method The authors explore polynomial activations as alternatives to softmax, focusing on their ability to regularize the attention matrix.
result Certain polynomials can serve as effective substitutes for softmax in transformer applications, achieving strong performance.
Sparse transformer architecture improves accuracy and speed in generative modeling and inverse problems.
problem Improving accuracy and speed in generative modeling and inverse problems.
method Proposes a sparse transformer architecture using regularized Wasserstein proximal operator with L1 prior. result Sparse transformer achieves higher accuracy and faster convergence than classical methods.
Transforms solutions of Davey-Stewartson II equation geometrically.
problem Solving the Davey-Stewartson II equation.
method Moutard transform and spinor representation of surfaces.
result Constructs examples of solutions with smooth initial data losing regularity.
This work provides theoretical and empirical evidence that invariance-inducing regularizers can increase predictive accuracy for worst-case spatial transformations (spatial robustness). Evaluated on these adversarially transformed examples, we demonstrate that adding regularization on top of standard or adversarial tra…
Transformations of macroeconomic data affect machine learning forecasts, especially with regularization and nonlinearity.
problem The impact of data transformations on machine learning forecasts in macroeconomic contexts.
method Review and propose new data transformations, empirically evaluate their effects, and compare traditional and moving average rotations.
result Traditional factors should almost always be included as predictors, and moving average rotations can provide important gains.
On simple geodesic disks of constant curvature, we derive new functional relations for the geodesic X-ray transform, involving a certain class of elliptic differential operators whose ellipticity degenerates normally at the boundary. We then use these relations to derive sharp mapping properties for the X-ray transform…
Injectivity of geodesic X-ray transform on low-regularity manifolds.
problem Injectivity of geodesic X-ray transform on manifolds with low regularity.
method Calculus of differential and curvature operators on non-smooth structures.
result Injectivity of geodesic X-ray transform on simple Riemannian manifolds with C1,1-regularity. Geometric structures on quaternionic unit ball for slice regular Möbius transformations.
problem No new problem introduced.
method Introducing Hermitian, Riemannian, and Kähler-like structures on quaternionic unit ball using regular Möbius transformations.
result Geometric structures are natural generalizations of complex setup and solve problems not achieved by other geometries.
Researchers improve transformer networks' optimization and understanding.
problem Improving the understanding and optimization of transformer networks.
method Introducing a convex alternative to the self-attention mechanism and reformulating the training problem as a convex optimization problem.
result Revealed an implicit regularization mechanism that promotes sparsity across tokens.
Paper proves injectivity of non-abelian X-ray transform on certain spaces.
problem Injectivity of non-abelian X-ray transform on asymptotically hyperbolic spaces.
method Gauge equivalence for unitary connections and skew-Hermitian Higgs fields.
result Injectivity result for non-abelian X-ray transform over skew-Hermitian Higgs fields.
New sampling method using regularized Wasserstein proximal for Gibbs distributions.
problem Sampling from Gibbs distributions with numerical stability and efficiency.
method Preconditioned regularized Wasserstein proximal operator.
result Discrete-time convergence analysis and explicit bias characterization.
Mixup improves model accuracy and calibration through data transformation and random perturbation.
problem Improving model accuracy and calibration in machine learning.
method Interprets Mixup as empirical risk minimization with data transformation and random perturbation.
result Mixup induces multiple known regularization schemes that prevent overfitting and overconfident predictions.
In order to understand the linearization problem around a leaf of a singular foliation, we extend the familiar holonomy map from the case of regular foliations to the case of singular foliations. To this aim we introduce the notion of holonomy transformation. Unlike the regular case, holonomy transformations can not be…
Injective X-ray transform on Heisenberg group for regular functions.
problem Injectivity of X-ray transform on sub-Riemannian manifolds.
method Group Fourier Transform and analysis of taming metrics.
result Sufficiently regular functions on Heisenberg group are determined by their line integrals.
Study anisotropic obstacle problem for minimal surfaces using Cahn-Hoffman transform.
problem Anisotropic obstacle problem for minimal surfaces.
method Cahn-Hoffman transform to convert to isotropic problem with generalized Robin boundary condition.
result Optimal regularity of the solution and C1,1 regularity of the free boundary. Geometrically transforms nonconservative dynamics to linearize Kepler and Manev systems.
problem Regularizing and linearizing nonconservative central force dynamics.
method Projective transformation and conformal scaling in configuration and phase spaces.
result Full linearization of Kepler and Manev dynamics in any finite dimension.
Study proves solenoidal injectivity for tensor fields on curved manifolds with low regularity.
problem Injectivity for tensor fields on negatively curved manifolds with low regularity metrics.
method Pestov energy estimates for transport equation on non-smooth unit sphere bundle, keeping track of regularity, and using functions with more vertical than horizontal regularity.
result Proves solenoidal injectivity for tensor fields on simple Riemannian manifolds with C1,1 metrics and non-positive sectional curvature. Transformers with linear space and time complexity for accurate attention estimation.
problem Efficiently estimating attention in large-scale tasks without relying on priors.
method Performers use Fast Attention Via positive Orthogonal Random features (FAVOR+) for linear approximation of softmax attention.
result Performers achieve competitive results on various tasks, demonstrating the effectiveness of their attention-learning approach.
Regularizes attention scores in vision transformers using bootstrapping.
problem Noisy and diffused attention maps in ViT limit interpretability.
method Statistical learning techniques, bootstrapping of attention scores.
result Improves shrinkage and sparsity of attention scores.
The use of convex regularizers allows for easy optimization, though they often produce biased estimation and inferior prediction performance. Recently, nonconvex regularizers have attracted a lot of attention and outperformed convex ones. However, the resultant optimization problem is much harder. In this paper, for a …
Study the Lax equation in infinite-dimensional Lie algebras and Lie groups.
problem Investigate the Lax equation in infinite-dimensional Lie algebras and Lie groups.
method Derived integral expansions and generalized Baker-Campbell-Hausdorff formula for Lie groups.
result Explicit representation of product integral in terms of exponential map.
Strongly polynomial algorithm for approximate Forster transforms and halfspace learning.
problem Computing approximate Forster transforms and halfspace learning.
method Strongly polynomial time algorithm for approximate Forster transforms and halfspace learning.
result First strongly polynomial time algorithm for distribution-free PAC learning of halfspaces.
Regularizes GAMs to improve interpretability by reducing concurvity.
problem Susceptibility of GAMs to concurvity reduces interpretability.
method Proposes a regularizer to penalize pairwise correlations of non-linearly transformed features.
result Improves interpretability and reduces concurvity without sacrificing prediction quality.
Let M be a complete non-compact Riemannian manifold. In this paper, we derive sufficient conditions on metric perturbation for stability of Lp-boundedness of the Riesz transform, p∈(2,∞). We also provide counter-examples regarding in-stability for Lp-boundedness of Riesz transform.
Transformers model contextual relations using probabilistic measures, revealing their expressive power.
problem Lack of clear understanding of Transformer's ability to model contextual relations.
method Introduced a measure-theoretic framework connecting softmax attention and entropy-regularized optimal transport.
result Transformer architectures can approximate arbitrary contextual relations, and the choice of normalization affects how these relations are represented.
We show injectivity of the X-ray transform and the d-plane Radon transform for distributions on the n-torus, lowering the regularity assumption in the recent work by Abouelaz and Rouvière. We also show solenoidal injectivity of the X-ray transform on the n-torus for tensor fields of any order, allowing the tensor…
This paper analyzes convergence of large-scale Transformers with weight decay.
problem Understanding optimization guarantees in large-scale Transformer training.
method Construct mean-field limit, show gradient flow convergence to PDE, demonstrate global minimum consistency.
result Gradient flow reaches global minimum in large-scale Transformers with small weight decay.
Given a slice regular function f:Ω⊂H→H, with Ω∩R=∅, it is possible to lift it to a surface in the twistor space CP3 of S4≃H∪{∞} (see~\cite{gensalsto}). In this paper we show that the same result is true if one rem…
This work improves Fourier pricing for multi-asset options using RQMC with domain transformation.
problem Efficiently pricing multi-asset options in high dimensions with Fourier methods.
method Randomized quasi-Monte Carlo (RQMC) with domain transformation to handle singularities.
result RQMC with domain transformation provides accurate and scalable Fourier pricing for multi-asset options.
In this paper, we study a fast approximation method for {\it large-scale high-dimensional} sparse least-squares regression problem by exploiting the Johnson-Lindenstrauss (JL) transforms, which embed a set of high-dimensional vectors into a low-dimensional space. In particular, we propose to apply the JL transforms to …
Following Burstall and Hertrich-Jeromin we study the Ribaucour transformation of Legendre submanifolds in Lie sphere geometry. We give an explicit parametrization of the resulted Legendre submanifold F^ of a Ribaucour transformation, via a single real function τ which represents the regular Ribaucour sphere co…
GTMs model complex multivariate data with varying conditional independencies.
problem Modeling multivariate data with intricate marginals and complex dependency structures.
method Semiparametric approach using penalized splines and lasso regularization.
result GTMs accurately learn complex dependencies and identify conditional independencies.
The version of Marsden-Ratiu reduction theorem for Nambu-Poisson manifolds by a regular distribution has been studied by Ibaˊn~ez et al. In this paper we show that the reduction is always ensured unless the distribution is zero. Next we extend the more general Falceto-Zambon Poisson reduct…
Transformers preserve support and can approximate any continuous map.
problem Understanding the mathematical properties of transformers.
method Characterizing maps between measures that can be represented as transformers and proving their properties.
result Transformers preserve support and have uniformly continuous Fréchet derivatives.
This study analyzes how one-layer transformers learn regular language recognition tasks.
problem Understanding how one-layer transformers solve regular language recognition tasks like even pairs and parity check.
method Theoretical analysis of training dynamics and gradient descent for a one-layer transformer.
result A one-layer transformer can solve even pairs directly but needs CoT for parity check. Training phases show rapid growth in attention layer followed by logarithmic growth in linear layer.
Classifies regularity for Lagrangian mean curvature type equations.
problem Classifying regularity for Lagrangian mean curvature type equations.
method Generalized constant rank theorem for Legendre transform, constructed convex solutions, and showed regularity conditions.
result Optimal regularity conditions for Lagrangian mean curvature type equations.
Extends optimal regularity and compactness to vector bundles over non-Riemannian manifolds.
problem Optimal regularity and compactness for connections on vector bundles.
method Derive RT-equations, establish existence theory, handle curvature up to L1. result Optimal regularity and compactness extended to vector bundles over non-Riemannian manifolds.
In this paper we combine two important extensions of ordinary least squares regression: regularization and optimal scaling. Optimal scaling (sometimes also called optimal scoring) has originally been developed for categorical data, and the process finds quantifications for the categories that are optimal for the regres…
Dual optimization connects ERM-fDR to normalization function.
problem Empirical risk minimization with f-divergence regularization.
method Dual formulation, Legendre-Fenchel transform, implicit function theorem, nonlinear ODE.
result Computational method to calculate normalization function efficiently.
This work studies clustering in transformer models, proving exponential convergence to a single token state.
problem Understanding the long-term behavior of tokens in transformer models.
method Investigates mean-field transformer models under specific conditions to prove exponential convergence to a single state.
result Transformer models synchronize exponentially fast to a single token state with explicit rates.
We present authors' new theory of the RT-equations, nonlinear elliptic partial differential equations which determine the coordinate transformations which smooth connections Γ to optimal regularity, one derivative smoother than the Riemann curvature tensor Riem(Γ). As one application we extend Uhlenbeck compa…
Smooth manifold structure on Möbius transformations of quaternionic ball identified.
problem Identifying the manifold structure of Möbius transformations of quaternionic unit ball.
method Realizing M(B) as a quotient of Sp(1,1) and using Lie group properties. result The manifold M(B) is diffeomorphic to R4imesS3.