Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

2775538301,106 · Jun 202019922001200920172026
48 results for Generative Transformation

This paper explains how model invariance improves generalization using data transformations.

problem Understanding why model invariance leads to better generalization performance.
method Introducing sample cover induced by transformations and refining generalization bounds.
result The sample covering number can be used to evaluate and select suitable data transformations.

The paper offers generalization bounds for Transformers that ignore sequence length.

problem Developing generalization bounds for Transformers that are independent of sequence length.
method Covering number approach to upper bound Rademacher complexity of bounded linear transformations.
result Theoretical bounds for Transformer generalization are independent of sequence length.

The notion of a generalized harmonic inverse mean curvature surface in the Euclidean four-space is introduced. A backward Bäcklund transform of a generalized harmonic inverse mean curvature surface is defined. A Darboux transform of a generalized harmonic inverse mean curvature surface is constructed by a backward Bäck…

2012-11-20abs ↗pdf ↗

Study shows Transformers can generalize to varying task lengths.

problem Understanding when and how Transformers can generalize to different input lengths.
method Proposed a unifying framework and introduced the RASP-Generalization Conjecture.
result Transformers tend to length generalize on tasks if solvable by short RASP programs.

Data augmentation (DA) is fundamental against overfitting in large convolutional neural networks, especially with a limited training dataset. In images, DA is usually based on heuristic transformations, like geometric or color transformations. Instead of using predefined transformations, our work learns data augmentati…

2019-09-21abs ↗pdf ↗

Transformers can simulate MLE for Bayesian network sequences.

problem Understanding transformers' capabilities in Bayesian network sequence generation.
method In-context maximum likelihood estimation (MLE) for autoregressive sequence generation.
result A simple transformer model can estimate Bayesian network probabilities and generate new samples.

Using Cartan's Method of Equivalence, we prove an upper bound for the generality of generic rank-1 Bäcklund transformations relating two hyperbolic Monge-Ampère systems. In cases when the Bäcklund transformation admits a symmetry group whose orbits have codimension 1, 2, or 3, we obtain classification results and new e…

2019-02-12abs ↗pdf ↗

The spherical Radon transform on the unit sphere can be regarded as a member of the analytic family of suitably normalized generalized cosine transforms. We derive new formulas for these transforms and apply them to study classes of intersections bodies in convex geometry.

2006-02-24abs ↗pdf ↗

New bounds show transformers need longer training for length generalization.

problem Understanding when transformers can generalize to longer inputs.
method Analyzing different settings of transformers, providing quantitative bounds.
result Transformers need training data longer than previously thought for length generalization.

Legendre transformations link related integrable hierarchies.

problem Understanding relationships between integrable hierarchies.
method Legendre-type transformations of generalized Frobenius manifolds.
result Linear reciprocal transformations link related hierarchies.

StrokeCoder uses Transformers to generate images from single examples.

problem Creating diverse images from a single example.
method Transformer Neural Network learns from a single path-based example to generate a set of images.
result The model can generate a large set of deviated images that still represent the original image's style and concept.

Paper proposes linear transformers for efficient in-context learning without context length limitations.

problem Quadratic complexity of softmax transformers limits data processing speed.
method Investigates linear transformers under domain generalization, showing they learn mappings from context distributions to response functions.
result Linear transformers achieve in-context learning with a linear complexity in context length, offering a dimension-independent convergence rate.

Study shows statistical biases can mislead transformer models, impairing their generalization.

problem Statistical biases in transformers affect their ability to generalize.
method Evaluated transformer models on synthetic algorithmic tasks with varying statistical biases.
result Statistical biases lead to overestimation of transformer models' generalization capabilities.

We give a construction of a Poisson transform mapping density valued differential forms on generalized flag manifolds to differential forms on the corresponding Riemannian symmetric spaces, which can be described entirely in terms of finite dimensional representations of reductive Lie groups. Moreover, we will explicit…

2016-04-01abs ↗pdf ↗

New analysis shows PE in Transformers increases generalization gap and vulnerability.

problem Understanding the impact of PE on Transformer generalization and robustness.
method Generalization analysis and adversarial Rademacher bounds for a single-layer Transformer with trainable PE.
result PE systematically enlarges the generalization gap and makes models more vulnerable to attacks.

Sparse transformer architecture improves accuracy and speed in generative modeling and inverse problems.

problem Improving accuracy and speed in generative modeling and inverse problems.
method Proposes a sparse transformer architecture using regularized Wasserstein proximal operator with L1L_1 prior.
result Sparse transformer achieves higher accuracy and faster convergence than classical methods.

Transformers can learn optimal regression mixtures efficiently.

problem Limited adoption of tailored regression methods due to their model-specific nature.
method Constructed a generative process for a mixture of linear regressions and used transformers to learn optimal predictors.
result Transformers achieve low mean-squared error and make predictions close to the optimal procedure.

This paper precisely estimates transformer derivatives for explicit learning guarantees.

problem Computing fully-explicit generalization bounds for transformers with precise higher-order derivative estimates.
method Analyzes and estimates all higher-order derivatives of transformers with multiple attention heads and layer normalization.
result Obtains explicit pathwise generalization bounds for transformers learning from non-i.i.d. samples.

Transformers fine-tuned on synthetic data boost tabular data classification performance.

problem Improving tabular data classification accuracy.
method Fine-tuning ICL-transformers on synthetic datasets with complex decision boundaries.
result Fine-tuned ICL-transformers outperform regular neural networks on real-world datasets.

This work characterizes benign overfitting in Vision Transformers.

problem Understanding generalization of Vision Transformers when trained to overfit.
method Gradient descent on a data distribution model, focusing on self-attention layer and softmax.
result Established a condition to distinguish between small and large test errors based on signal-to-noise ratio.

New linear flows using exponential of linear transformations improve generative models.

problem Improving generative models in machine learning.
method Developed convolution exponentials and generalized Sylvester Flows using the exponential of linear transformations.
result Convolution exponentials and Convolutional Sylvester Flows outperform other models in log-likelihood.

This paper investigates how transformers can learn to generalize to unseen examples in context.

problem Understanding how transformers can generalize to unseen examples in a prompt.
method Gradient descent analysis of one-layer multi-head transformers for in-context learning.
result The training loss for a one-layer multi-head transformer converges linearly to a global minimum, effectively learning ridge regression over basis functions.

The classical Liouville Theorem on conformal transformations determines local conformal transformations on the Euclidean space of dimension 3\geq 3. Its natural adaptation to the general framework of Riemannian structures is the 2-rigidity of conformal transformations, that is such a transformation is fully determined…

2014-11-20abs ↗pdf ↗

The reparameterization trick has become one of the most useful tools in the field of variational inference. However, the reparameterization trick is based on the standardization transformation which restricts the scope of application of this method to distributions that have tractable inverse cumulative distribution fu…

2019-11-06abs ↗pdf ↗

We interpret the setting for a Radon transform as a submanifold of the space of generalized functions, and compute its extrinsic curvature: it is the Hessian composed with the Radon transform.

1992-10-01abs ↗pdf ↗

Diffusion Transformer captures spatial-temporal dependencies in sequential data.

problem Capturing rich spatial and temporal dependencies in sequential data.
method Established theoretical guarantees for diffusion transformers learning Gaussian process data.
result Spatial-temporal dependencies are captured within attention layers of diffusion transformers.

Study excess risk in statistical inference with transformations.

problem Excess risk in estimating random variables from feature vectors and transformations.
method Characterize lossless transformations, develop test statistics, and information-theoretic bounds.
result Strongly consistent partitioning test statistic for lossless transformations.

The paper provides convergence guarantees for ODE-based generative models using transformers.

problem Theoretical guarantees for ODE-based generative models.
method A pre-trained autoencoder maps inputs to a latent space, and a transformer predicts the velocity field.
result The distribution of samples generated via estimated ODE flow converges to the target distribution in Wasserstein-2 distance.

The general framework of Legendre transformation is extended to the case of symplectic groupoids, using an appropriate generalization of the notion of generating function (of a Lagrangian submanifold).

1996-12-04abs ↗pdf ↗

GT-PCA improves PCA for image and time series data.

problem Lack of robustness to transformations in PCA.
method GT-PCA is a neural network that estimates components invariant to specific transformations.
result GT-PCA outperforms alternative methods in synthetic and real data experiments.

We introduce a nonlocal transformation to generate exact solutions of the constant astigmatism equation zyy+(1/z)xx+2=0z_{yy} + (1/z)_{xx} + 2 = 0. The transformation is related to the special case of the famous Bäcklund transformation of the sine-Gordon equation with the Bäcklund parameter λ=±1λ= \pm1. It is also a nonlocal symmetry…

2011-11-08abs ↗pdf ↗