The paper introduces models to learn generalized transformation equivariant representations.
problem Capturing intrinsic visual structures equivariant to various transformations.
method Deterministic and probabilistic AutoEncoding Transformations (AET and AVT) models trained to learn visual representations from generic groups of transformations.
result Generalized TERs (GTERs) that are equivariant to transformations in a more general fashion.
This paper investigates efficient Transformers and finds they scale with problem size.
problem Finding suitable replacements for standard Transformers in large-scale tasks.
method Modeling efficient Transformers (Sparse and Linear) as Dynamic Programming problems and analyzing their reasoning capabilities.
result Efficient Transformers scale with problem size, but can be more efficient for certain DP problems.
Transformers improve Finnish language modeling, achieving lower perplexity scores.
problem Improving language modeling for Finnish using deep learning models.
method Used BERT and Transformer-XL models in a sub-word setting, compared to LSTM.
result Transformer-XL outperforms LSTM, achieving a 27% better perplexity score.
New filter bank sparsifying transforms outperform patch-based methods for image denoising.
problem Improving image denoising performance using data-adaptive sparsifying transforms.
method Proposes a new transform learning framework using undecimated perfect reconstruction filter banks, allowing independent filter length choice.
result Filter bank sparsifying transforms outperform existing patch-based methods for image denoising.
The paper analyzes transformation models in high-dimensional settings.
problem Analyzing transformation models in high-dimensional data.
method Proposed an estimator for transformation parameter and showed asymptotic normality.
result The proposed estimator works well in small samples and tests the log-wage transformation.
Transformers interpret as probabilistic mixtures, offering new insights.
problem Understanding Transformers from a probabilistic perspective.
method Modeling Transformers as mixtures of Gaussian models.
result Transformers can be seen as maximum posterior probability estimators.
ETs improve model robustness to transformations in images.
problem Improving model robustness to predefined transformations.
method Equivariant Transformers (ETs) incorporating functions equivariant to continuous transformation groups.
result ETs achieve up to 15% relative improvement in error rate on image classification tasks.
New multi-layer transform models improve image denoising.
problem Improving image denoising techniques.
method Developed multi-layer transform learning algorithms.
result Multi-layer models outperform single-layer schemes in image denoising.
Proposes a Complex Transformer for complex-valued sequence modeling.
problem Lack of deep learning models for complex-valued data.
method Develops a Complex Transformer using transformer backbone with specialized attention and encoder-decoder networks.
result Achieves state-of-the-art performance on complex-valued datasets.
This work studies clustering in transformer models, proving exponential convergence to a single token state.
problem Understanding the long-term behavior of tokens in transformer models.
method Investigates mean-field transformer models under specific conditions to prove exponential convergence to a single state.
result Transformer models synchronize exponentially fast to a single token state with explicit rates.
CHOOSE enhances shallow Transformers for wireless symbol detection.
problem Improving wireless symbol detection with shallow Transformers.
method Introducing autoregressive latent reasoning steps within hidden space.
result Lightweight Transformers achieve comparable performance to deep models.
This paper explains how model invariance improves generalization using data transformations.
problem Understanding why model invariance leads to better generalization performance.
method Introducing sample cover induced by transformations and refining generalization bounds.
result The sample covering number can be used to evaluate and select suitable data transformations.
B-cos transformers explain Vision Transformers' decisions.
problem Lack of holistic explanations for transformer outputs.
method Formulate each component as dynamic linear, allowing a single linear transform for summarization.
result Bcos-ViTs are highly interpretable and competitive on ImageNet.
XR-Transformer accelerates XMC by recursively fine-tuning on multi-resolution objectives.
problem Efficiently classifying texts with large label sets.
method Recursive multi-resolution fine-tuning of transformers.
result XR-Transformer achieves 20x faster training time and 54% Precision@1 on Amazon-3M.
Study compares LSTM and Transformer models in financial time series prediction.
problem Comparing LSTM and Transformer models for financial time series prediction.
method Various LSTM-based and Transformer-based models compared on financial tasks; DLSTM and new Transformer architecture designed.
result Transformer-based models show limited advantage in absolute price sequence prediction, while LSTM-based models perform better on difference sequences.
Evolved Transformer improves on Transformer architecture for language tasks.
problem Improving Transformer architecture for sequence tasks.
method Evolutionary architecture search with warm starting and dynamic resource allocation.
result Evolved Transformer achieves state-of-the-art BLEU scores and reduces parameter count.
Bayesian model for discrete data with conditional transformations.
problem Handling discrete ordinal and count data with excess zeros.
method Bayesian framework with conditional transformation functions and modular MCMC algorithm.
result Flexible modeling of linear and nonlinear covariate effects for ordinal and count data.
Transformers improve time series modeling by capturing long-range dependencies.
problem Capturing long-range dependencies in time series data.
method Summarized and reviewed adaptations of Transformers for time series analysis.
result Transformers enhance time series forecasting, anomaly detection, and classification.
Transformer learns long-term dependencies from real-world data.
problem Sample inefficiency in deep reinforcement learning.
method Transformer architecture applied to autoregressive real-world episodes.
result Transformer-based world model generates meaningful experience.
Alternative approach to model selection using transformation analysis.
problem Over-simplistic models lead to erroneous interpretations.
method Step-wise complexity reduction to identify simpler, better-interpretable models.
result Transformation models improve model fit and interpretability.
Transformer-MGK replaces redundant heads with Gaussian key mixtures, improving efficiency and performance.
problem Redundant attention heads in transformers degrade performance and efficiency.
method Transformer-MGK replaces redundant heads with a mixture of Gaussian keys.
result Transformer-MGK accelerates training and inference, reduces parameters and FLOPs, and achieves comparable or better accuracy.
Paper explores EEG-based speech recognition using transformers, showing faster training and better performance for smaller vocabularies.
problem Continuous speech recognition using EEG features.
method Transformer-based ASR model compared to RNN-based models.
result Transformer models perform better for smaller vocabularies but RNN models outperform them for larger vocabularies.
Improved Transformer language models using dynamic evaluation.
problem Language model perplexity and accuracy improvements.
method Combining Transformers with dynamic evaluation techniques.
result Significant improvement in language model performance (e.g., 0.99 to 0.94 bits/char).
Fusion of transformer networks using optimal transport for improved performance.
problem Improving performance of transformer-based models through fusion.
method Exploiting optimal transport for soft alignment of transformer components.
result Consistently outperforms vanilla fusion and individual parent models.
Transformer improves pop piano composition by incorporating beat-based structure.
problem Generating expressive pop piano compositions with coherent rhythmic structure.
method Improved data representation for Transformers, incorporating beat-bar-phrase structure.
result Composes pop piano music with better rhythmic structure than existing models.
This paper proposes a new RV prediction model using neural distributional transformation and co-training.
problem Predicting skewed and fat-tailed realized volatility (RV) is challenging.
method The paper uses a neural distributional transformation and co-training to predict RV. It jointly trains the transformation and prediction model using a maximum-likelihood objective function.
result The proposed method significantly outperforms other methods on a dataset of 100 stocks.
Paper trains models to resist string transformations.
problem Vulnerability of NLP models to adversarial string transformations.
method Combines search and abstraction techniques for robust training.
result Trained models resist combinations of user-defined transformations.
Ouroboros accelerates training of large Transformer models.
problem Training large Transformer-based language models is slow and resource-intensive.
method Proposes a novel model-parallel algorithm to speed up training.
result Demonstrates significantly faster training speed compared to data parallelism.
The paper models musical motif transformations in Beethoven's works.
problem Understanding how motifs transform in symbolic music.
method Developed a probabilistic framework using Conditional Random Fields.
result Identified patterns of motif transformations and their co-occurrences.
A new model improves CT image quality from low-dose scans.
problem Improving CT image quality from low-dose scans.
method Multi-layer Residual Sparsifying Transform (MRST) learning model for low-dose CT reconstruction.
result The MRST model outperforms conventional methods in maintaining subtle details.
This study evaluates the robustness of transformation-based ensemble defense against evasion attacks.
problem Understanding the reasons behind the robustness improvement in transformation-based ensemble defense.
method Designing two adaptive attacks to evaluate transformation-based ensemble defense, conducting experiments to analyze robustness.
result The robustness improvement is mainly from irreversible transformations rather than the ensemble of models.
Deep learning model estimates uncertainty in complex regression tasks.
problem Uncertainty quantification in probabilistic regression predictions.
method Combines statistical and deep learning transformation models using gradient descent.
result State-of-the-art performance on small datasets and complex image data.
Transformers can learn optimal regression mixtures efficiently.
problem Limited adoption of tailored regression methods due to their model-specific nature.
method Constructed a generative process for a mixture of linear regressions and used transformers to learn optimal predictors.
result Transformers achieve low mean-squared error and make predictions close to the optimal procedure.
Paper uses Time Series Transformer for bank stability prediction.
problem Predicting bank stability using complex financial data.
method Time Series Transformer model with self-attention mechanism.
result Time Series Transformer model outperforms other models in MSE and MAE.
New copula models capture volatility and directionality in financial time series.
problem Modeling financial return series with volatility and serial correlation.
method Stationary d-vine copula processes with v-transforms for stochastic volatility and directionality.
result Models can rival and sometimes outperform GARCH family models.
Neural network models transform physical systems into latent Gaussian distributions.
problem Simplifying and solving classical Hamiltonian systems.
method Symplectic neural networks for canonical transformations.
result Captures nonlinear collective modes in latent space.
Paper proposes SERT model for US stock pricing, outperforming standard models during market shocks.
problem Capturing patterns of temporal sparsity in asset pricing during market fluctuations.
method Introduces SERT model based on pre-trained Transformer, compares with standard models in three periods.
result SERT model achieves highest out-of-sample R2 (11.94\% and 11.47\%) during extreme market fluctuations. Transformers solve Gaussian Mixture Models without supervision.
problem Solving Gaussian Mixture Models (GMMs) unsupervised.
method Proposes TGMM, a transformer-based framework for GMM tasks.
result Transformers can effectively solve GMM tasks, improving upon classical methods.
Transformer architecture improved with credibility mechanism for better model performance.
problem Improving predictive models in tabular data.
method Introducing a credibility mechanism to the Transformer architecture.
result Credibility Transformer leads to superior predictive models compared to state-of-the-art models.
Improved machine translation with INT8 hardware using a novel training method.
problem Training accurate machine translation models with limited hardware precision.
method Convert all Transformer matrix multiplications to 8-bit integer (INT8) without sacrificing accuracy.
result INT8 Transformer models achieve BLEU scores 99.3% to 100% relative to FP32 models.
Sparse transformer architecture improves accuracy and speed in generative modeling and inverse problems.
problem Improving accuracy and speed in generative modeling and inverse problems.
method Proposes a sparse transformer architecture using regularized Wasserstein proximal operator with L1 prior. result Sparse transformer achieves higher accuracy and faster convergence than classical methods.
Unified framework for learning function representations using INRs and Transformers.
problem Scalability and efficiency limitations in existing generative models.
method Integrates INRs and Transformer-based hypernetworks into latent variable models.
result Improved scalability, expressiveness, and generalization over existing models.
Transformers can approximate any sequence-to-sequence function, surprising given their complexity.
problem Understanding the expressive power of Transformer models for sequence-to-sequence functions.
method Established that Transformers are universal approximators of continuous permutation equivariant sequence-to-sequence functions with compact support, and extended this to arbitrary functions using positional encodings.
result Transformers are universal approximators of arbitrary continuous sequence-to-sequence functions on a compact domain.
Effective theory for Transformer initialization improves model performance.
problem Improving performance of Transformers at initialization.
method Effective-theory analysis of signal propagation in wide and deep Transformers.
result Particular width scalings of initialization and training hyperparameters.
Paper proposes Multi-Transformer for more accurate stock volatility forecasts.
problem Accurate equity risk models needed for effective risk management.
method Introduces Multi-Transformer neural network architecture, adapted from Transformer models.
result Empirical results show Multi-Transformer leads to more accurate risk measures.
New transforms improve signal classification and data analysis.
problem Improving signal classification and data analysis.
method Algebraic generative models and transport transforms.
result Classes of signals are transformed into convex sets, simplifying classification.
Transformer improves sequence generation with insertion and deletion phases.
problem Sequence generation challenges in machine translation.
method Insertion-Deletion Transformer with iterative insertion and deletion phases.
result Significant BLEU score improvement over insertion-only models.
Transformer models outperform recurrent ones in modeling hierarchical data.
problem Modeling hierarchical structure in data.
method Introducing Multiresolution Transformer Networks leveraging self-attention.
result Multiresolution Transformer Networks significantly outperform state-of-the-art models on query suggestion datasets.