We analyze Darboux transformations in very general settings for multidimensional linear partial differential operators. We consider all known types of Darboux transformations, and present a new type. We obtain a full classification of all operators that admit Wronskian type Darboux transformations of first order and a …
Transformation Equivariant Representations (TERs) aim to capture the intrinsic visual structures that equivary to various transformations by expanding the notion of {\em translation} equivariance underlying the success of Convolutional Neural Networks (CNNs). For this purpose, we present both deterministic AutoEncoding…
Transformed quadrics from 2D to higher dimensions.
problem Generalizing quadric transformations to higher dimensions.
method Bianchi's Hazzidakis transformation method.
result Generalization to higher dimensional quadrics.
This paper explains how model invariance improves generalization using data transformations.
problem Understanding why model invariance leads to better generalization performance.
method Introducing sample cover induced by transformations and refining generalization bounds.
result The sample covering number can be used to evaluate and select suitable data transformations.
The paper offers generalization bounds for Transformers that ignore sequence length.
problem Developing generalization bounds for Transformers that are independent of sequence length.
method Covering number approach to upper bound Rademacher complexity of bounded linear transformations.
result Theoretical bounds for Transformer generalization are independent of sequence length.
The notion of a generalized harmonic inverse mean curvature surface in the Euclidean four-space is introduced. A backward Bäcklund transform of a generalized harmonic inverse mean curvature surface is defined. A Darboux transform of a generalized harmonic inverse mean curvature surface is constructed by a backward Bäck…
We define a transformation on harmonic maps from a Riemann surface into the 2-sphere which depends on a complex parameter, the so-called mu-Darboux transformation. In the case when the harmonic map N is the Gauss map of a constant mean curvature surface f and the parameter is real, the mu-Darboux transformation of -N i…
Study shows Transformers can generalize to varying task lengths.
problem Understanding when and how Transformers can generalize to different input lengths.
method Proposed a unifying framework and introduced the RASP-Generalization Conjecture.
result Transformers tend to length generalize on tasks if solvable by short RASP programs.
Transformers trained on random classification tasks generalize well and can overfit without error.
problem Understanding how transformers generalize and overfit in-context.
method Analysis of implicit regularization during gradient descent training.
result Transformers can overfit without error and still generalize well.
Data augmentation (DA) is fundamental against overfitting in large convolutional neural networks, especially with a limited training dataset. In images, DA is usually based on heuristic transformations, like geometric or color transformations. Instead of using predefined transformations, our work learns data augmentati…
Transformers can simulate MLE for Bayesian network sequences.
problem Understanding transformers' capabilities in Bayesian network sequence generation.
method In-context maximum likelihood estimation (MLE) for autoregressive sequence generation.
result A simple transformer model can estimate Bayesian network probabilities and generate new samples.
Using Cartan's Method of Equivalence, we prove an upper bound for the generality of generic rank-1 Bäcklund transformations relating two hyperbolic Monge-Ampère systems. In cases when the Bäcklund transformation admits a symmetry group whose orbits have codimension 1, 2, or 3, we obtain classification results and new e…
The spherical Radon transform on the unit sphere can be regarded as a member of the analytic family of suitably normalized generalized cosine transforms. We derive new formulas for these transforms and apply them to study classes of intersections bodies in convex geometry.
We prove Transformers can learn diverse Gröbner bases.
problem Training Transformers for Gröbner basis computation.
method Prove generality of dataset generation algorithm; propose extended algorithm.
result Datasets are sufficiently general for diverse Gröbner bases learning.
New bounds show transformers need longer training for length generalization.
problem Understanding when transformers can generalize to longer inputs.
method Analyzing different settings of transformers, providing quantitative bounds.
result Transformers need training data longer than previously thought for length generalization.
Legendre transformations link related integrable hierarchies.
problem Understanding relationships between integrable hierarchies.
method Legendre-type transformations of generalized Frobenius manifolds.
result Linear reciprocal transformations link related hierarchies.
Unified framework for learning function representations using INRs and Transformers.
problem Scalability and efficiency limitations in existing generative models.
method Integrates INRs and Transformer-based hypernetworks into latent variable models.
result Improved scalability, expressiveness, and generalization over existing models.
StrokeCoder uses Transformers to generate images from single examples.
problem Creating diverse images from a single example.
method Transformer Neural Network learns from a single path-based example to generate a set of images.
result The model can generate a large set of deviated images that still represent the original image's style and concept.
Paper proposes linear transformers for efficient in-context learning without context length limitations.
problem Quadratic complexity of softmax transformers limits data processing speed.
method Investigates linear transformers under domain generalization, showing they learn mappings from context distributions to response functions.
result Linear transformers achieve in-context learning with a linear complexity in context length, offering a dimension-independent convergence rate.
In trying to generalize Bianchi's Bäcklund transformation of quadrics to Bäcklund transformations of isometric deformations of other (classes of) surfaces, we investigate basic features of the isometric deformation of surfaces via the Bäcklund transformation with isometric correspondence of leaves of a general nature (…
New invertible transformations improve flow-based generative models.
problem Improving flow-based generative models for better performance.
method Proposed new invertible transformations and coupling layers.
result New coupling layers achieve better results in IDF.
Study shows statistical biases can mislead transformer models, impairing their generalization.
problem Statistical biases in transformers affect their ability to generalize.
method Evaluated transformer models on synthetic algorithmic tasks with varying statistical biases.
result Statistical biases lead to overestimation of transformer models' generalization capabilities.
We give a construction of a Poisson transform mapping density valued differential forms on generalized flag manifolds to differential forms on the corresponding Riemannian symmetric spaces, which can be described entirely in terms of finite dimensional representations of reductive Lie groups. Moreover, we will explicit…
New transforms improve signal classification and data analysis.
problem Improving signal classification and data analysis.
method Algebraic generative models and transport transforms.
result Classes of signals are transformed into convex sets, simplifying classification.
New analysis shows PE in Transformers increases generalization gap and vulnerability.
problem Understanding the impact of PE on Transformer generalization and robustness.
method Generalization analysis and adversarial Rademacher bounds for a single-layer Transformer with trainable PE.
result PE systematically enlarges the generalization gap and makes models more vulnerable to attacks.
Sparse transformer architecture improves accuracy and speed in generative modeling and inverse problems.
problem Improving accuracy and speed in generative modeling and inverse problems.
method Proposes a sparse transformer architecture using regularized Wasserstein proximal operator with L1 prior. result Sparse transformer achieves higher accuracy and faster convergence than classical methods.
Transformers can learn optimal regression mixtures efficiently.
problem Limited adoption of tailored regression methods due to their model-specific nature.
method Constructed a generative process for a mixture of linear regressions and used transformers to learn optimal predictors.
result Transformers achieve low mean-squared error and make predictions close to the optimal procedure.
This paper precisely estimates transformer derivatives for explicit learning guarantees.
problem Computing fully-explicit generalization bounds for transformers with precise higher-order derivative estimates.
method Analyzes and estimates all higher-order derivatives of transformers with multiple attention heads and layer normalization.
result Obtains explicit pathwise generalization bounds for transformers learning from non-i.i.d. samples.
Extends cohomology theory for infinite volume transformation groups.
problem Generalizing cohomology for infinite volume transformation groups.
method Introduces norm-controlled cohomology as a generalization of bounded cohomology.
result Establishes norm-controlled cohomology for infinite volume transformation groups.
Transformers fine-tuned on synthetic data boost tabular data classification performance.
problem Improving tabular data classification accuracy.
method Fine-tuning ICL-transformers on synthetic datasets with complex decision boundaries.
result Fine-tuned ICL-transformers outperform regular neural networks on real-world datasets.
Let (M,g) be an analytic, compact, Riemannian manifold with boundary, of dimension n >= 2. We study a class of generalized Radon transforms, integrating over a family of hypersurfaces embedded in M, satisfying the Bolker condition [23]. Using analytic microlocal analysis, we prove a microlocal regularity theorem for ge…
This work characterizes benign overfitting in Vision Transformers.
problem Understanding generalization of Vision Transformers when trained to overfit.
method Gradient descent on a data distribution model, focusing on self-attention layer and softmax.
result Established a condition to distinguish between small and large test errors based on signal-to-noise ratio.
New linear flows using exponential of linear transformations improve generative models.
problem Improving generative models in machine learning.
method Developed convolution exponentials and generalized Sylvester Flows using the exponential of linear transformations.
result Convolution exponentials and Convolutional Sylvester Flows outperform other models in log-likelihood.
This paper investigates how transformers can learn to generalize to unseen examples in context.
problem Understanding how transformers can generalize to unseen examples in a prompt.
method Gradient descent analysis of one-layer multi-head transformers for in-context learning.
result The training loss for a one-layer multi-head transformer converges linearly to a global minimum, effectively learning ridge regression over basis functions.
The classical Liouville Theorem on conformal transformations determines local conformal transformations on the Euclidean space of dimension ≥3. Its natural adaptation to the general framework of Riemannian structures is the 2-rigidity of conformal transformations, that is such a transformation is fully determined…
Transformers excel at sparse token selection, surpassing FCNs in both worst and average cases.
problem Sparse token selection task
method One-layer transformer trained with gradient descent
result Transformers learn sparse token selection and exhibit strong out-of-distribution length generalization
The reparameterization trick has become one of the most useful tools in the field of variational inference. However, the reparameterization trick is based on the standardization transformation which restricts the scope of application of this method to distributions that have tractable inverse cumulative distribution fu…
We interpret the setting for a Radon transform as a submanifold of the space of generalized functions, and compute its extrinsic curvature: it is the Hessian composed with the Radon transform.
We show that the quantum field theoretical formulation of the τ-function theory has a geometrical interpretation within the classical transformation theory of conjugate nets. In particular, we prove that i) the partial charge transformations preserving the neutral sector are Laplace transformations, ii) the basic ver…
Develops support theorem for analytic transforms in tomography.
problem Analytic wave front set resolution for integral transforms.
method Microlocal analysis, double fibration framework, wave packet transforms.
result Uniqueness and support theorems for analytic transforms.
Diffusion Transformer captures spatial-temporal dependencies in sequential data.
problem Capturing rich spatial and temporal dependencies in sequential data.
method Established theoretical guarantees for diffusion transformers learning Gaussian process data.
result Spatial-temporal dependencies are captured within attention layers of diffusion transformers.
Study on SignGD optimization of two-layer transformer on noisy data.
problem Understanding how SignGD optimizes transformers and its generalization.
method Analysis of a two-layer transformer with SignGD on a linearly separable noisy dataset.
result SignGD converges fast but has poor generalization on noisy data.
Study excess risk in statistical inference with transformations.
problem Excess risk in estimating random variables from feature vectors and transformations.
method Characterize lossless transformations, develop test statistics, and information-theoretic bounds.
result Strongly consistent partitioning test statistic for lossless transformations.
The paper provides convergence guarantees for ODE-based generative models using transformers.
problem Theoretical guarantees for ODE-based generative models.
method A pre-trained autoencoder maps inputs to a latent space, and a transformer predicts the velocity field.
result The distribution of samples generated via estimated ODE flow converges to the target distribution in Wasserstein-2 distance.
The general framework of Legendre transformation is extended to the case of symplectic groupoids, using an appropriate generalization of the notion of generating function (of a Lagrangian submanifold).
Study proves uniqueness for ray transform on surfaces with obstacles.
problem Uniqueness of functions and 1-forms on surfaces with reflecting obstacles.
method Broken ray transform on twisted geodesics with nonpositive curvature and reflecting boundary.
result Proves uniqueness result for sums of functions and 1-forms.
GT-PCA improves PCA for image and time series data.
problem Lack of robustness to transformations in PCA.
method GT-PCA is a neural network that estimates components invariant to specific transformations.
result GT-PCA outperforms alternative methods in synthetic and real data experiments.
We introduce a nonlocal transformation to generate exact solutions of the constant astigmatism equation zyy+(1/z)xx+2=0. The transformation is related to the special case of the famous Bäcklund transformation of the sine-Gordon equation with the Bäcklund parameter λ=±1. It is also a nonlocal symmetry…