New unsupervised learning technique learns independent kernels for better machine learning tasks.
problem Improving unsupervised representation learning for machine learning tasks.
method Stacking convolutional transforms using alternating proximal minimization scheme.
result DCTL outperforms shallow version CTL on benchmark datasets.
In many state-of-the-art compression systems, signal transformation is an integral part of the encoding and decoding process, where transforms provide compact representations for the signals of interest. This paper introduces a class of transforms called graph-based transforms (GBTs) for video compression, and proposes…
Deep neural networks for natural language processing tasks are vulnerable to adversarial input perturbations. In this paper, we present a versatile language for programmatically specifying string transformations -- e.g., insertions, deletions, substitutions, swaps, etc. -- that are relevant to the task at hand. We then…
Transformers learn functionals from distributions without losing information.
problem Lack of rigorous mathematical theory supporting Transformer performance.
method Proposed a Transformer learning framework, attention operator, and distribution regression.
result Transformers can compress distributions into function representations without loss of information.
This research examines how data transformations affect adversarial robustness in recurrent neural networks.
problem Adversarial examples reduce machine learning accuracy, especially in high-dimensional datasets.
method Analysis of feature selection, dimensionality reduction, and trend extraction techniques on recurrent neural networks.
result Data transformations may increase vulnerability to adversarial samples, but only if they approximate intrinsic dimensionality and maintain manifold coverage.
Task-agnostic data augmentation shows little benefit for pretrained transformers.
problem Evaluating the effectiveness of task-agnostic data augmentation on pretrained transformers.
method Conducted a systematic examination of two data augmentation techniques (Easy Data Augmentation and Back-Translation) across 5 tasks, 6 datasets, and 3 pretrained transformer models.
result Data augmentation techniques previously effective for non-pretrained models fail to consistently improve performance for pretrained transformers, even with limited training data.
Subspace clustering assumes that the data is sepa-rable into separate subspaces. Such a simple as-sumption, does not always hold. We assume that, even if the raw data is not separable into subspac-es, one can learn a representation (transform coef-ficients) such that the learnt representation is sep-arable into subspac…
New techniques improve channel prediction in noisy wireless systems.
problem Predicting channels in wireless communication systems from noisy observations.
method Adapted sequence-to-sequence models and transformers with reverse positional encoding and reversed encoder outputs.
result Improved robustness and relationship capture in channel prediction models.
The characteristics (or numerical patterns) of a feature vector in the transform domain of a perturbation model differ significantly from those of its corresponding feature vector in the input domain. These differences - caused by the perturbation techniques used for the transformation of feature patterns - degrade the…
We propose the first qualitative hypothesis characterizing the behavior of visual transformation based self-supervision, called the VTSS hypothesis. Given a dataset upon which a self-supervised task is performed while predicting instantiations of a transformation, the hypothesis states that if the predicted instantiati…
Lower bounds set for infinite-precision transformers.
problem Understanding limitations of infinite-precision transformers.
method Used VC dimension technique to prove lower bounds.
result First lower bounds for two tasks: function composition and SUM2. Paper transforms torse-forming vector fields into simpler forms.
problem Generalizing vector fields and their transformations.
method Present techniques to transform torse-forming vector fields into simpler cases.
result Concrete examples of transformations are provided.
This paper analyzes convergence of large-scale Transformers with weight decay.
problem Understanding optimization guarantees in large-scale Transformer training.
method Construct mean-field limit, show gradient flow convergence to PDE, demonstrate global minimum consistency.
result Gradient flow reaches global minimum in large-scale Transformers with small weight decay.
Time series forecasting with limited data is a challenging yet critical task. While transformers have achieved outstanding performances in time series forecasting, they often require many training samples due to the large number of trainable parameters. In this paper, we propose a training technique for transformers th…
The objective of this work is to improve the accuracy of building demand forecasting. This is a more challenging task than grid level forecasting. For the said purpose, we develop a new technique called recurrent transform learning (RTL). Two versions are proposed. The first one (RTL) is unsupervised; this is used as a…
FinTech uses data science and AI to transform finance.
problem Transforming finance with data science and AI.
method DSAI techniques including complex system methods, quantitative methods, etc.
result DSAI enables smart FinTech for various financial sectors.
The paper offers generalization bounds for Transformers that ignore sequence length.
problem Developing generalization bounds for Transformers that are independent of sequence length.
method Covering number approach to upper bound Rademacher complexity of bounded linear transformations.
result Theoretical bounds for Transformer generalization are independent of sequence length.
Hidformer improves stock price prediction accuracy using Transformer techniques.
problem Improving stock price prediction accuracy using machine learning.
method Adapted Transformer model (Hidformer) for stock price forecasting.
result Hidformer shows promising performance in stock price prediction.
New method learns identity-preserving transformations on data manifolds without labels.
problem Learning identity-preserving transformations on natural variations without supervision.
method Introduces a learning strategy that does not require transformation labels and learns local regions for operators.
result Trains on MNIST and Fashion MNIST, and CelebA, learning transformations without labels.
Objective: A variety of pattern analysis techniques for model training in brain interfaces exploit neural feature dimensionality reduction based on feature ranking and selection heuristics. In the light of broad evidence demonstrating the potential sub-optimality of ranking based feature selection by any criterion, we …
Dataset augmentation, the practice of applying a wide array of domain-specific transformations to synthetically expand a training set, is a standard tool in supervised learning. While effective in tasks such as visual recognition, the set of transformations must be carefully designed, implemented, and tested for every …
We establish a link between Archimedes' method of integration for calculating areas, volumes and centers of mass of segments of parabolas and quadrics of revolution by factorization via the moments of a balance and an integration technique for a particular integrable system, namely Bianchi's Bäcklund transformation for…
We define a Fourier-Mukai transform for a triple consisting of two holomorphic vector bundles over an elliptic curve and a homomorphism between them. We prove that in some cases the transform preserves the natural stability condition for a triple. We also define a Nahm transform for solutions to natural gauge-theoretic…
R2T hybrid model improves robust regression for asymmetric noise.
problem Least-squares regression fails with asymmetric structured noise.
method Transformer encoder, compression NN, fixed symbolic equation.
result Median regression MSE of 6e-6 to 3.5e-5 on synthetic data.
This paper explores the limits of Transformers in learning new patterns from scratch.
problem Understanding when Transformers can learn new patterns from scratch.
method Introducing the 'globality degree' to measure learnability and developing scratchpad techniques.
result Distributions with high globality cannot be learned efficiently by Transformers.
Paper trains a Transformer to add numbers of any length.
problem Training Transformers to handle arbitrary-length addition.
method Autoregressive generation from right to left.
result Trains a Transformer to generalize addition of numbers of any length.
The paper solves a complex financial optimization problem using a novel mathematical technique.
problem Optimizing portfolio selection in financial markets.
method Maximal monotone operator method and Riccati transformation.
result Existence and uniqueness of a solution to the transformed parabolic equation in a Sobolev space.
The Legendre transform and its generalizations, originally found in supersymmetric sigma-models, are techniques that can be used to give constructions of hyperkahler metrics. We give a twistor space interpretation to the generalizations of the Legendre transform construction. The Atiyah-Hitchin metric on the moduli spa…
Improves data normality with robust transformations.
problem Skewed data distribution.
method Modified Box-Cox and Yeo-Johnson transformations with robust parameter estimation.
result Transformed data approximates normality in the center with outliers.
New method reduces training cost by using approximate gradients.
problem Training neural networks is computationally expensive.
method Uses control variates to approximate gradients without full backward pass.
result Efficacy demonstrated on a vision transformer classification task.
This work improves Fourier pricing for multi-asset options using RQMC with domain transformation.
problem Efficiently pricing multi-asset options in high dimensions with Fourier methods.
method Randomized quasi-Monte Carlo (RQMC) with domain transformation to handle singularities.
result RQMC with domain transformation provides accurate and scalable Fourier pricing for multi-asset options.
Paper unifies propositionalization and embedding for relational learning.
problem Data fusion from diverse input formats into a single table.
method Unified framework combining propositionalization and embedding.
result New algorithms outperform existing relational learners.
SGPA calibrates transformer uncertainty for safety-critical tasks.
problem Uncertainty estimation in transformer models for safety-critical domains.
method Bayesian inference in transformer's output space using sparse Gaussian processes.
result SGPA-based Transformers improve in-distribution calibration and out-of-distribution robustness.
A new transformer model accelerates training with optimization techniques.
problem Training deep neural networks efficiently and effectively.
method Interprets transformer layers as optimization steps, applying Nesterov acceleration.
result The new model outperforms existing models on benchmark datasets.
Novel regularization for Vision Transformers improves model generalization and sparsity.
problem Improving generalization and sparsity in Vision Transformers.
method Likelihood-guided variational Ising-based regularization.
result Improved generalization and sparsity in Vision Transformers.
The paper examines how nonlinear transformations affect ridge sets in manifold learning.
problem Understanding the impact of nonlinear transformations on ridge sets in manifold learning.
method Examined the effects of nonlinear transformations on ridge sets using mathematical proofs and numerical experiments.
result The inclusion relationship $\cR(f\circ p)\subseteq \cR(p)$ holds for strictly increasing and concave transformations, and the Hausdorff distance between transformed and non-transformed ridge sets is smaller.
The paper introduces new methods for Asian option pricing using Laguerre quadrature.
problem Developing accurate pricing models for Asian options.
method Utilizes Laguerre quadrature and diffusion kernel approach.
result Demonstrates new techniques to solve complex Asian option pricing equations.
Transforms between neural networks using manifold-learning techniques.
problem Establish equivalence between different neural networks.
method Diffusion maps with a Mahalanobis-like metric to construct transformations between network outputs and internal neuron activations.
result Established equivalence classes between neural networks trained on various data types.
The study compares differencing methods for financial data and finds fractional differencing improves model performance.
problem Improving financial time series forecasting models using appropriate data transformation techniques.
method Comparative analysis of traditional logarithmic returns and fractional differencing methods, including tempered extensions.
result Fractional differencing methods improve model forecasting performance and trading strategy effectiveness.
We give a new algorithm for approximating the Discrete Fourier transform of an approximately sparse signal that has been corrupted by worst-case L0 noise, namely a bounded number of coordinates of the signal have been corrupted arbitrarily. Our techniques generalize to a wide range of linear transformations that are…
GT-PCA improves PCA for image and time series data.
problem Lack of robustness to transformations in PCA.
method GT-PCA is a neural network that estimates components invariant to specific transformations.
result GT-PCA outperforms alternative methods in synthetic and real data experiments.
The complex wave representation (CWR) converts unsigned 2D distance transforms into their corresponding wave functions. Here, the distance transform S(X) appears as the phase of the wave function φ(X)---specifically, φ(X)=exp(iS(X)/τwhere τis a free parameter. In this work, we prove a novel result using the higher-orde…
New L-functions for 3-manifolds connect to Witten invariants and relate to generalized Bernoulli polynomials.
problem Understanding L-functions for 3-manifolds and their invariants. method Using Mellin transforms and asymptotic techniques, proving entire functions and their values.
result Linear relations between L-function values at negative integers, generalizing known zeta functions. Transformers can learn spectral methods and perform unsupervised learning.
problem Learning spectral methods using unsupervised learning.
method Using multi-layered Transformers, pre-trained on a large set of instances, to learn and perform statistical estimation tasks.
result Proven that pre-trained Transformers can learn spectral methods and perform tasks like PCA and clustering.
Machine learning in high-energy physics faces challenges from nuisance parameters, which are reviewed and techniques to mitigate their impact are discussed.
problem Impact of nuisance parameters on machine learning performance in high-energy physics.
method Review and discussion of techniques including nuisance-parameterized models, modified or adversary losses, semi-supervised learning, and inference-aware techniques.
result Various methods to reduce the impact of nuisance parameters and improve model performance in high-energy physics.
New displacement technique vanishes bounded cohomology in all degrees.
problem Vanishing of bounded cohomology in all positive degrees and dual separable coefficients.
method Introducing the property of commuting cyclic conjugates as a new displacement technique.
result Vanishes bounded cohomology in all positive degrees and all dual separable coefficients.
Mixup improves model accuracy and calibration through data transformation and random perturbation.
problem Improving model accuracy and calibration in machine learning.
method Interprets Mixup as empirical risk minimization with data transformation and random perturbation.
result Mixup induces multiple known regularization schemes that prevent overfitting and overconfident predictions.
Despite the widespread adoption of Transformer models for NLP tasks, the expressive power of these models is not well-understood. In this paper, we establish that Transformer models are universal approximators of continuous permutation equivariant sequence-to-sequence functions with compact support, which is quite surp…