This paper investigates efficient Transformers and finds they scale with problem size.
problem Finding suitable replacements for standard Transformers in large-scale tasks.
method Modeling efficient Transformers (Sparse and Linear) as Dynamic Programming problems and analyzing their reasoning capabilities.
result Efficient Transformers scale with problem size, but can be more efficient for certain DP problems.
Transformers improve Finnish language modeling, achieving lower perplexity scores.
problem Improving language modeling for Finnish using deep learning models.
method Used BERT and Transformer-XL models in a sub-word setting, compared to LSTM.
result Transformer-XL outperforms LSTM, achieving a 27% better perplexity score.
Transformation Equivariant Representations (TERs) aim to capture the intrinsic visual structures that equivary to various transformations by expanding the notion of {\em translation} equivariance underlying the success of Convolutional Neural Networks (CNNs). For this purpose, we present both deterministic AutoEncoding…
Transformers interpret as probabilistic mixtures, offering new insights.
problem Understanding Transformers from a probabilistic perspective.
method Modeling Transformers as mixtures of Gaussian models.
result Transformers can be seen as maximum posterior probability estimators.
This work studies clustering in transformer models, proving exponential convergence to a single token state.
problem Understanding the long-term behavior of tokens in transformer models.
method Investigates mean-field transformer models under specific conditions to prove exponential convergence to a single state.
result Transformer models synchronize exponentially fast to a single token state with explicit rates.
CHOOSE enhances shallow Transformers for wireless symbol detection.
problem Improving wireless symbol detection with shallow Transformers.
method Introducing autoregressive latent reasoning steps within hidden space.
result Lightweight Transformers achieve comparable performance to deep models.
This paper explains how model invariance improves generalization using data transformations.
problem Understanding why model invariance leads to better generalization performance.
method Introducing sample cover induced by transformations and refining generalization bounds.
result The sample covering number can be used to evaluate and select suitable data transformations.
Data is said to follow the transform (or analysis) sparsity model if it becomes sparse when acted on by a linear operator called a sparsifying transform. Several algorithms have been designed to learn such a transform directly from data, and data-adaptive sparsifying transforms have demonstrated excellent performance i…
B-cos transformers explain Vision Transformers' decisions.
problem Lack of holistic explanations for transformer outputs.
method Formulate each component as dynamic linear, allowing a single linear transform for summarization.
result Bcos-ViTs are highly interpretable and competitive on ImageNet.
XR-Transformer accelerates XMC by recursively fine-tuning on multi-resolution objectives.
problem Efficiently classifying texts with large label sets.
method Recursive multi-resolution fine-tuning of transformers.
result XR-Transformer achieves 20x faster training time and 54% Precision@1 on Amazon-3M.
Study compares LSTM and Transformer models in financial time series prediction.
problem Comparing LSTM and Transformer models for financial time series prediction.
method Various LSTM-based and Transformer-based models compared on financial tasks; DLSTM and new Transformer architecture designed.
result Transformer-based models show limited advantage in absolute price sequence prediction, while LSTM-based models perform better on difference sequences.
Bayesian model for discrete data with conditional transformations.
problem Handling discrete ordinal and count data with excess zeros.
method Bayesian framework with conditional transformation functions and modular MCMC algorithm.
result Flexible modeling of linear and nonlinear covariate effects for ordinal and count data.
Transformers improve time series modeling by capturing long-range dependencies.
problem Capturing long-range dependencies in time series data.
method Summarized and reviewed adaptations of Transformers for time series analysis.
result Transformers enhance time series forecasting, anomaly detection, and classification.
Transformer learns long-term dependencies from real-world data.
problem Sample inefficiency in deep reinforcement learning.
method Transformer architecture applied to autoregressive real-world episodes.
result Transformer-based world model generates meaningful experience.
Transformer-MGK replaces redundant heads with Gaussian key mixtures, improving efficiency and performance.
problem Redundant attention heads in transformers degrade performance and efficiency.
method Transformer-MGK replaces redundant heads with a mixture of Gaussian keys.
result Transformer-MGK accelerates training and inference, reduces parameters and FLOPs, and achieves comparable or better accuracy.
Fusion of transformer networks using optimal transport for improved performance.
problem Improving performance of transformer-based models through fusion.
method Exploiting optimal transport for soft alignment of transformer components.
result Consistently outperforms vanilla fusion and individual parent models.
This paper proposes a new RV prediction model using neural distributional transformation and co-training.
problem Predicting skewed and fat-tailed realized volatility (RV) is challenging.
method The paper uses a neural distributional transformation and co-training to predict RV. It jointly trains the transformation and prediction model using a maximum-likelihood objective function.
result The proposed method significantly outperforms other methods on a dataset of 100 stocks.
The paper models musical motif transformations in Beethoven's works.
problem Understanding how motifs transform in symbolic music.
method Developed a probabilistic framework using Conditional Random Fields.
result Identified patterns of motif transformations and their co-occurrences.
A new model improves CT image quality from low-dose scans.
problem Improving CT image quality from low-dose scans.
method Multi-layer Residual Sparsifying Transform (MRST) learning model for low-dose CT reconstruction.
result The MRST model outperforms conventional methods in maintaining subtle details.
Transformation models are a very important tool for applied statisticians and econometricians. In many applications, the dependent variable is transformed so that homogeneity or normal distribution of the error holds. In this paper, we analyze transformation models in a high-dimensional setting, where the set of potent…
This study evaluates the robustness of transformation-based ensemble defense against evasion attacks.
problem Understanding the reasons behind the robustness improvement in transformation-based ensemble defense.
method Designing two adaptive attacks to evaluate transformation-based ensemble defense, conducting experiments to analyze robustness.
result The robustness improvement is mainly from irreversible transformations rather than the ensemble of models.
Deep learning model estimates uncertainty in complex regression tasks.
problem Uncertainty quantification in probabilistic regression predictions.
method Combines statistical and deep learning transformation models using gradient descent.
result State-of-the-art performance on small datasets and complex image data.
Paper uses Time Series Transformer for bank stability prediction.
problem Predicting bank stability using complex financial data.
method Time Series Transformer model with self-attention mechanism.
result Time Series Transformer model outperforms other models in MSE and MAE.
Transformers can learn optimal regression mixtures efficiently.
problem Limited adoption of tailored regression methods due to their model-specific nature.
method Constructed a generative process for a mixture of linear regressions and used transformers to learn optimal predictors.
result Transformers achieve low mean-squared error and make predictions close to the optimal procedure.
New copula models capture volatility and directionality in financial time series.
problem Modeling financial return series with volatility and serial correlation.
method Stationary d-vine copula processes with v-transforms for stochastic volatility and directionality.
result Models can rival and sometimes outperform GARCH family models.
Transformer architecture improved with credibility mechanism for better model performance.
problem Improving predictive models in tabular data.
method Introducing a credibility mechanism to the Transformer architecture.
result Credibility Transformer leads to superior predictive models compared to state-of-the-art models.
Paper proposes SERT model for US stock pricing, outperforming standard models during market shocks.
problem Capturing patterns of temporal sparsity in asset pricing during market fluctuations.
method Introduces SERT model based on pre-trained Transformer, compares with standard models in three periods.
result SERT model achieves highest out-of-sample R2 (11.94\% and 11.47\%) during extreme market fluctuations. Transformers solve Gaussian Mixture Models without supervision.
problem Solving Gaussian Mixture Models (GMMs) unsupervised.
method Proposes TGMM, a transformer-based framework for GMM tasks.
result Transformers can effectively solve GMM tasks, improving upon classical methods.
Sparse transformer architecture improves accuracy and speed in generative modeling and inverse problems.
problem Improving accuracy and speed in generative modeling and inverse problems.
method Proposes a sparse transformer architecture using regularized Wasserstein proximal operator with L1 prior. result Sparse transformer achieves higher accuracy and faster convergence than classical methods.
Learned data models based on sparsity are widely used in signal processing and imaging applications. A variety of methods for learning synthesis dictionaries, sparsifying transforms, etc., have been proposed in recent years, often imposing useful structures or properties on the models. In this work, we focus on sparsif…
Effective theory for Transformer initialization improves model performance.
problem Improving performance of Transformers at initialization.
method Effective-theory analysis of signal propagation in wide and deep Transformers.
result Particular width scalings of initialization and training hyperparameters.
Unified framework for learning function representations using INRs and Transformers.
problem Scalability and efficiency limitations in existing generative models.
method Integrates INRs and Transformer-based hypernetworks into latent variable models.
result Improved scalability, expressiveness, and generalization over existing models.
Recent works have highlighted the strength of the Transformer architecture on sequence tasks while, at the same time, neural architecture search (NAS) has begun to outperform human-designed models. Our goal is to apply NAS to search for a better alternative to the Transformer. We first construct a large search space in…
Paper proposes Multi-Transformer for more accurate stock volatility forecasts.
problem Accurate equity risk models needed for effective risk management.
method Introduces Multi-Transformer neural network architecture, adapted from Transformer models.
result Empirical results show Multi-Transformer leads to more accurate risk measures.
New transforms improve signal classification and data analysis.
problem Improving signal classification and data analysis.
method Algebraic generative models and transport transforms.
result Classes of signals are transformed into convex sets, simplifying classification.
Transformers interpreted as probabilistic Laplacian Eigenmaps steps.
problem Improving transformer performance through probabilistic interpretation.
method Probabilistic Laplacian Eigenmaps model derivation and graph diffusion step.
result Subtracting identity from attention matrix improves transformer performance.
Tracr compiles programs into transformer models for interpretability.
problem Uncertainty in understanding transformer model outputs due to unknown learned programs.
method Tracr compiles human-readable programs into known structure transformer models.
result Known structure of Tracr-compiled models serves as ground-truth for interpretability.
Study reveals Transformer's expressive power and mechanisms.
problem Understanding the approximation properties of Transformer for sequence modeling.
method Systematic study of Transformer's components and their combined effects, establishing approximation rates.
result Reveals roles of critical parameters in Transformer, such as number of layers and attention heads.
Enhances non-life insurance pricing models using transformer models.
problem Improving predictive power of non-life insurance pricing models.
method Enhances actuarial non-life models with transformer models for tabular data.
result Transformer models outperform benchmark models in claim frequency prediction.
New invertible transformations improve flow-based generative models.
problem Improving flow-based generative models for better performance.
method Proposed new invertible transformations and coupling layers.
result New coupling layers achieve better results in IDF.
We describe a method for training accurate Transformer machine-translation models to run inference using 8-bit integer (INT8) hardware matrix multipliers, as opposed to the more costly single-precision floating-point (FP32) hardware. Unlike previous work, which converted only 85 Transformer matrix multiplications to IN…
Standard Transformers approximate Hölder functions and achieve optimal nonparametric regression rate.
problem Approximating Hölder functions and achieving optimal nonparametric regression rate with Transformers.
method Using the size tuple and dimension vector metrics, the paper characterizes Transformer structures and derives upper bounds for their Lipschitz constant and memorization capacity.
result Standard Transformers achieve the minimax optimal rate in nonparametric regression for Hölder target functions.
Paper proves efficiency of MARL with transformers, addressing agent complexity.
problem Theoretical understanding of MARL with many agents and limited relational reasoning.
method Set Transformer for relational reasoning, model-free and model-based MARL algorithms.
result Provable efficiency of MARL algorithms, suboptimality gaps independent of number of agents.
We propose a novel multilinear dynamical system (MLDS) in a transform domain, named L-MLDS, to model tensor time series. With transformations applied to a tensor data, the latent multidimensional correlations among the frontal slices are built, and thus resulting in the computational independence in the tra…
A framework uses free probability to analyze Transformer models.
problem Understanding the dynamics and complexity of Transformer-based language models.
method Formal operator-theoretic analysis using free probability theory.
result Entropy-based generalization bounds derived under freeness assumptions.
Mixtures of Gaussians, factor analyzers (probabilistic PCA) and hidden Markov models are staples of static and dynamic data modeling and image and video modeling in particular. We show how topographic transformations in the input, such as translation and shearing in images, can be accounted for in these models by inclu…
A great number of deep learning based models have been recently proposed for automatic music composition. Among these models, the Transformer stands out as a prominent approach for generating expressive classical piano performance with a coherent structure of up to one minute. The model is powerful in that it learns ab…
Study shows statistical biases can mislead transformer models, impairing their generalization.
problem Statistical biases in transformers affect their ability to generalize.
method Evaluated transformer models on synthetic algorithmic tasks with varying statistical biases.
result Statistical biases lead to overestimation of transformer models' generalization capabilities.