Proposes a Complex Transformer for complex-valued sequence modeling.
problem Lack of deep learning models for complex-valued data.
method Develops a Complex Transformer using transformer backbone with specialized attention and encoder-decoder networks.
result Achieves state-of-the-art performance on complex-valued datasets.
Transformers prefer simpler explanations in hierarchical tasks.
problem Navigating tasks with varying complexity levels.
method Well-controlled testbeds based on Markov chains and linear regression.
result Transformers favor the least complex sufficient explanation when presented with simpler data.
One-layer transformers can't solve induction heads task efficiently.
problem Solving the induction heads task efficiently with one-layer transformers.
method Communication complexity argument showing exponential size requirement.
result No one-layer transformer can solve the induction heads task efficiently.
Transformers become faster by linearizing self-attention.
problem Quadratic complexity of transformers makes them slow for long sequences.
method Expressed self-attention as a linear dot-product and used matrix product associativity to reduce complexity.
result Linear transformers are up to 4000x faster on long sequences.
Various complexes of differential operators are constructed on complex projective space via the Penrose transform, which also computes their cohomology.
Deep learning model estimates uncertainty in complex regression tasks.
problem Uncertainty quantification in probabilistic regression predictions.
method Combines statistical and deep learning transformation models using gradient descent.
result State-of-the-art performance on small datasets and complex image data.
Transformers show strengths and weaknesses in complexity analysis.
problem Understanding the strengths and limitations of attention layers in transformers.
method Analysis of representation power through complexity parameters and task-specific constructions.
result Transformers can solve sparse averaging tasks with logarithmic complexity, but triple detection tasks require linear complexity.
Transformer architecture struggles with complex tasks due to limitations in function composition.
problem Transformer architecture's limitations in handling complex tasks.
method Used Communication Complexity to prove limitations in composing functions.
result Transformer layer is incapable of handling large domain functions, even when domains are small.
CT improves neural network performance on cell complex data.
problem Improving predictive performance of neural networks on complex data.
method Introducing the Cellular Transformer (CT) that generalizes graph-based transformers to cell complexes.
result CT achieves state-of-the-art performance on cell complex datasets without complex enhancements.
Linformer reduces transformer complexity to linear, improving efficiency.
problem High cost of training and deploying large transformer models for long sequences.
method Approximates self-attention with low-rank matrix, proposing Linformer with O(n) complexity. result Linformer performs similarly to standard transformers but is more memory- and time-efficient.
New transformations preserve link isotopy, showing complexity differences.
problem Link isotopy preservation with complexity differences.
method Introducing multiflypes of rectangular diagrams of links.
result Two diagrams of same complexity not related by simpler moves.
Transformers fine-tuned on synthetic data boost tabular data classification performance.
problem Improving tabular data classification accuracy.
method Fine-tuning ICL-transformers on synthetic datasets with complex decision boundaries.
result Fine-tuned ICL-transformers outperform regular neural networks on real-world datasets.
New bounds adaptively control spectral complexity of trained Transformers.
problem Understanding why Transformers generalize well in machine learning.
method Spectrum-adaptive post hoc generalization bounds for multi-layer Transformers.
result Bounds adaptively trade off spectral complexity against dimension and depth factors.
We develop some theory of double fibration transforms where the cycle space is a smooth manifold and apply it to complex projective space.
The paper finds transformation formulas for quaternionic complex structures.
problem Quaternionic projective invariance of k-Cauchy-Fueter complex. method Explicit transformation formulae under mSL(n+1,H). result Quaternionic projectively invariant operator and defining density.
The paper explores a B-field transform of complex structures on complex tori.
problem Deforming complex structures on complex tori using B-field transformations.
method Constructing holomorphic line bundles with integrable connections and interpreting them as deformations of complex tori by flat gerbes.
result Homological mirror symmetry between deformed and original complex tori.
Transformers capture combinatorial tasks with bounded error and logarithmic sample dependence.
problem Capturing complex combinatorial tasks with bounded error and sample efficiency.
method Formal definition of algorithmic capture, empirical analysis of infinite-width transformers, upper bounds on computational complexity.
result Transformers exhibit an inductive bias favoring simpler algorithmic procedures over higher complexity ones.
GTMs model complex multivariate data with varying conditional independencies.
problem Modeling multivariate data with intricate marginals and complex dependency structures.
method Semiparametric approach using penalized splines and lasso regularization.
result GTMs accurately learn complex dependencies and identify conditional independencies.
This paper constructs Poisson transforms and analyzes their properties on complex hyperbolic spaces.
problem Understanding discrete series representations of SU(n+1,1) using differential forms.
method Constructing Poisson transforms and analyzing their boundary asymptotics and intertwining properties with the Rumin complex.
result The constructed transforms realize the direct sum of all discrete series representations of SU(n+1,1).
We define a transformation on harmonic maps from a Riemann surface into the 2-sphere which depends on a complex parameter, the so-called mu-Darboux transformation. In the case when the harmonic map N is the Gauss map of a constant mean curvature surface f and the parameter is real, the mu-Darboux transformation of -N i…
Simple models are preferred over complex models, but over-simplistic models could lead to erroneous interpretations. The classical approach is to start with a simple model, whose shortcomings are assessed in residual-based model diagnostics. Eventually, one increases the complexity of this initial overly simple model a…
The paper offers generalization bounds for Transformers that ignore sequence length.
problem Developing generalization bounds for Transformers that are independent of sequence length.
method Covering number approach to upper bound Rademacher complexity of bounded linear transformations.
result Theoretical bounds for Transformer generalization are independent of sequence length.
Paper introduces a new method for Transformers with linear complexity.
problem No efficient relative positional encoding for linear Transformer models.
method Stochastic Positional Encoding (SPE) that replaces classical RPE.
result SPE behaves like RPE and performs well on benchmarks.
We introduce complex generalizations of the classical Legendre transform, operating on Kähler metrics on a compact complex manifold. These Legendre transforms give explicit local isometric symmetries for the Mabuchi metric on the space of Kähler metrics around any real analytic Kähler metric, answering a question origi…
Let G be a semisimple Lie group with finite center, K⊂G a maximal compact subgroup, and P⊂G a parabolic subgroup. Following ideas of P.Y.\ Gaillard, one may use G-invariant differential forms on G/K×G/P to construct G-equivariant Poisson transforms mapping differential forms on G/P to …
We consider complex projective space with its Fubini-Study metric and the X-ray transform defined by integration over its geodesics. We identify the kernel of this transform acting on symmetric tensor fields.
A new algorithm reduces time complexity for binary time series classification.
problem High time complexity of ensemble shapelet transform limits its application.
method Introduces short isometric shapelet transform with two strategies: fixed shapelet length and single linear classifier.
result Demonstrates superior performance and reduced time complexity.
We consider the generalized Segal-Bargmann transform, defined in terms of the heat operator, for a noncompact symmetric space of the complex type. For radial functions, we show that the Segal-Bargmann transform is a unitary map onto a certain L^2 space of meromorphic functions. For general functions, we give an inversi…
We present a version of the Penrose transform which relates compactly supported cohomology on a complex or CR manifold Z to kernels and cokernels of differential operators on a parameter space X of compact complex submanifolds of Z.
Looped transformers outperform standard transformers in complex reasoning tasks due to a specific loss landscape geometry.
problem Understanding why looped transformers outperform standard transformers in complex reasoning tasks.
method Explained through loss landscape geometry, distinguishing between U-shaped and V-shaped valleys, and proposing SHIFT training strategy.
result Looped transformers' recursive architecture induces a River-V-Valley landscape, leading to better loss convergence and complex pattern learning.
Transformer predicts Ethereum prices using cross-currency correlation and sentiment analysis.
problem Predicting Ethereum cryptocurrency prices with limited data.
method Transformer-based neural network with cross-currency correlation and sentiment analysis.
result Transformer model outperforms other models on some parameters.
It is shown that the heat operator in the Hall coherent state transform for a compact Lie group K is related with a Hermitian connection associated to a natural one-parameter family of complex structures on T∗K. The unitary parallel transport of this connection establishes the equivalence of (geometric) quantizati…
Transformers can outperform feedforward and recurrent networks due to dynamic sparsity.
problem Understanding when and why Transformers outperform other neural network architectures.
method Analyzing a sequence-to-sequence data generating model with dynamic sparsity, proving sample complexity differences between feedforward, recurrent, and Transformers.
result Transformers can learn dynamic sparsity models with lower sample complexity than feedforward and recurrent networks.
Develops first robustness verification for complex Transformers.
problem Certify prediction behavior of Transformers with complex self-attention layers.
method Resolves challenges of cross-nonlinearity and cross-position dependency in Transformers.
result Certified robustness bounds are significantly tighter than Interval Bound Propagation.
The paper constructs a complex for the Dirac operator in 4 dimensions.
problem Constructing a complex for the Dirac operator in 4 dimensions.
method Using the Penrose transform, the paper constructs a relative BGG complex and its direct image.
result An explicit construction of a complex starting with the Dirac operator in any number of variables.
Resurgence of Joyce structures gauged to a standard form using gauge transformations.
problem Resurgence of Joyce structures in complex hyperkähler geometry.
method Gauge transformations and Borel transforms to show resurgent behavior.
result Established the resurgent behavior of infinitesimal gauge transformations for Joyce structures.
We show that for any complete connected Kähler manifold the index of the group of complex affine transformations in the group of c-projective transformations is at most two unless the Kähler manifold is isometric to complex projective space equipped with a positive constant multiple of the Fubini-Study metric. This est…
ALT improves TSC by capturing complex patterns in time series data.
problem Challenges in traditional TSC methods with time series complexity and variability.
method ALT incorporates variable-length shifted time windows to enhance LLT for better feature representation.
result ALT achieves state-of-the-art performance with few hyperparameters.
Let G be a linear connected complex reductive Lie group. The purpose of this paper is to give explicit symplectic isomorphisms from twisted cotangent bundles of the complex generalized flag varieties, whose transition functions are given by affine transformations instead of linear transformations, onto the complex co…
Paper transforms a complex equation into simpler forms for analysis.
problem Analyzing a fourth-order dispersive flow equation on Kähler manifolds.
method Developed the generalized Hasimoto transformation to simplify the equation.
result Explicit expressions derived for three examples of compact Kähler manifolds.
The paper classifies complex Dirac structures on flag manifolds.
problem Classifying invariant complex Dirac structures on flag manifolds.
method Described using roots of the Lie algebra and classified under B-transformations. result All invariant complex Dirac structures with constant real index on a maximal flag manifold are described.
Sumformer simplifies Transformers to handle long sequences efficiently.
problem Quadratic complexity of Transformers limits their use with long sequences.
method Introducing Sumformer, a simple architecture that universally approximates equivariant sequence-to-sequence functions.
result Sumformer achieves the first universal approximation results for Linformer and Performer.
Efficiently accelerates attention calculation for Transformers with relative positional encoding.
problem Quadratic complexity of attention in long sequences.
method Kernelized attention with Fast Fourier Transform (FFT) for RPE.
result Achieves O(n log n) time complexity, mitigates training instability, and outperforms other models.
We classify the simplest rational elements in a twisted loop group, and prove that dressing actions of them on proper indefinite affine spheres give the classical Tzitzéica transformation and its dual. We also give the group point of view of the Permutability Theorem, construct complex Tzitzéica transformations, and di…
Study of groups acting on complex projective varieties.
problem Classifying groups of birational transformations on complex projective varieties.
method Free, properly discontinuous, cocompact action on open sets of complex projective varieties.
result Classification in dimension two.
Simplified calculus for semimartingales makes complex transformations easier.
problem Complex transformations of semimartingales.
method Unified treatment of transformations for real and complex semimartingales.
result Unified calculus for semimartingales simplifies various transformations.
Novel deep learning model for multivariate time series prediction.
problem Challenges in multivariate time series prediction with correlations and complex temporal patterns.
method Temporal Tensor Transformation Network (TTNT) that transforms multivariate time series into tensors for improved feature extraction.
result TTNT outperforms state-of-the-art methods in window-based predictions across various tasks.
Enhances stock movement prediction using Higher Order Transformers for multimodal time-series data.
problem Predicting stock movements in financial markets with complex dynamics.
method Introduced Higher Order Transformers, extending self-attention and transformer architecture to capture complex market dynamics. Employed low-rank tensor decomposition and kernel attention to manage computational complexity. Integrated technical and fundamental analysis from historical prices and tweets.
result Demonstrated effectiveness of the method on the Stocknet dataset, improving stock movement prediction.