The paper studies how noise synchronizes tokens in deep transformer models.
problem Understanding synchronization in deep learning models with noise.
method Proves convergence to a stochastic particle system and identifies the limiting SDE.
result The limiting model displays synchronization by noise and exponential dissipation of interaction energy.
The paper explores how generative networks can transform noise distributions into other distributions.
problem Transforming noise distributions into desired distributions using generative networks.
method Developed a space-filling function for ReLU networks and provided efficient methods for univariate uniform to normal distribution transformations.
result Optimal construction for ReLU networks to increase noise dimensionality and efficient methods for distribution transformations.
New insights show stochastic initialization prevents token clustering in deep Transformers.
problem Understanding token dynamics in deep stochastic Transformers.
method Analysis of deep Transformers with random initialization noise, proving convergence to an interacting-particle system on the sphere.
result Initialization noise prevents token clustering, leading to antipodal formations.
Noise stability improves understanding of Transformer models.
problem Lack of robustness metrics for real-valued domains and junta-like input dependence in modern LLMs.
method Proposed noise stability as a new metric and developed a practical regularization method.
result Noise stability regularization method accelerates training by 35-75%.
Transformer model removes noise from light curves efficiently.
problem Challenges in processing astrophysical light curves due to noise.
method Denoising Time Series Transformer (DTST) model trained with masked objective.
result DTST model excels at removing noise and outliers in time series datasets.
R2T hybrid model improves robust regression for asymmetric noise.
problem Least-squares regression fails with asymmetric structured noise.
method Transformer encoder, compression NN, fixed symbolic equation.
result Median regression MSE of 6e-6 to 3.5e-5 on synthetic data.
New auto-encoder handles varying noise levels without retraining.
problem Auto-encoders degrade in noisy conditions.
method Formalized auto-encoders as transform learning, derived new architecture.
result Models generalize well to different noise levels.
New algorithm learns halfspaces with noise using Forster decomposition.
problem Learning halfspaces in noisy data.
method Forster decomposition and efficient mixture of distributions.
result First polynomial-time algorithm with strongly polynomial sample complexity.
Improves detection of low-rank signals from noisy data matrices.
problem Statistical detection of low-rank signals in noisy data matrices.
method Entrywise pre-transforming data matrix for non-Gaussian noise, sharp phase transition thresholds, central limit theorem for linear spectral statistics, hypothesis test.
result Improves detection of low-rank signals from noisy data matrices, generalizing known results.
Paper proposes a new e-exponentiated transformation to make convex loss functions more robust to outliers.
problem Making convex loss functions robust to outliers in the presence of label noise.
method Introduces a novel e-exponentiated transformation for loss functions and proves its effectiveness through theoretical and empirical analysis. result The transformed loss function achieves tighter generalization error bounds and higher accuracy in noisy datasets.
Class2Simi reduces noise in noisy label learning by transforming noisy class labels into noisy similarity labels.
problem Learning with noisy labels in supervised and unsupervised settings.
method Transforming noisy class labels into noisy similarity labels, training DNNs from noisy data pairs.
result The noise rate reduction is theoretically guaranteed, making it easier to handle noisy similarity labels.
The alignment of a set of objects by means of transformations plays an important role in computer vision. Whilst the case for only two objects can be solved globally, when multiple objects are considered usually iterative methods are used. In practice the iterative methods perform well if the relative transformations b…
Improves signal detection in non-Gaussian noise using transformed data.
problem Signal detection in rank-one signal-plus-noise data matrices.
method Pre-transforming matrix entries and using linear spectral statistics for hypothesis testing.
result Sharp phase transition of largest eigenvalues in spiked rectangular matrices.
Guarantees recovery of compressible signals from adversarial noise.
problem Recovering compressible signals from noise and adversarial attacks.
method Extends adversarial defense framework to ℓ0, ℓ2, and ℓ∞ norms. result Recovery guarantees for various signal recovery methods under different noise types.
New algorithm defends against adversarial examples in image classification.
problem Defending against adversarial examples in image classification.
method Approximates Discrete Fourier transform of sparse signals corrupted by L0 noise. result Successfully defends against L0 adversaries in image classification. Enhances DSN with multi-family wavelet transforms and sparsity.
problem Improving interpretability and expressive power of DSN.
method Multi-family wavelet transforms and optimal thresholding.
result Enhanced robustness and diversity of scattering coefficients.
Strongly polynomial algorithm for approximate Forster transforms and halfspace learning.
problem Computing approximate Forster transforms and halfspace learning.
method Strongly polynomial time algorithm for approximate Forster transforms and halfspace learning.
result First strongly polynomial time algorithm for distribution-free PAC learning of halfspaces.
Develops statistical confidence sets for multidimensional scaling.
problem Statistical uncertainty in multidimensional scaling of noisy data.
method Formal statistical framework, distributional convergence results, uniform confidence sets, bootstrap procedures.
result Construction of reliable confidence sets for latent configurations in multidimensional scaling.
Generative Adversarial Nets (GANs) and Variational Auto-Encoders (VAEs) provide impressive image generations from Gaussian white noise, but the underlying mathematics are not well understood. We compute deep convolutional network generators by inverting a fixed embedding operator. Therefore, they do not require to be o…
We study the theoretical advantages of active learning over passive learning. Specifically, we prove that, in noise-free classifier learning for VC classes, any passive learning algorithm can be transformed into an active learning algorithm with asymptotically strictly superior label complexity for all nontrivial targe…
A method predicts GNS of transformer layers using normalization layer norms.
problem Estimating gradient noise scale with minimal variance.
method Simultaneously compute per-example gradient norms and parameter gradients.
result Total GNS is predicted well by normalization layer GNS.
We study the property of the Fused Lasso Signal Approximator (FLSA) for estimating a blocky signal sequence with additive noise. We transform the FLSA to an ordinary Lasso problem. By studying the property of the design matrix in the transformed Lasso problem, we find that the irrepresentable condition might not hold, …
TSLANet improves time series models by capturing long-term and short-term interactions.
problem Noise sensitivity, computational efficiency, and overfitting in Transformer-based models for time series data.
method Adaptive Spectral Block and Interactive Convolution Block for robust feature representation and noise mitigation.
result TSLANet outperforms state-of-the-art models in various time series tasks.
Paper proposes integrating wavelet transform, channel attention, and LSTM for better stock price prediction.
problem Inherently difficult stock price prediction due to low signal-to-noise ratio.
method Wavelet transform convolution, channel attention, and LSTM integration.
result Robust performance in post-pandemic market conditions.
New algorithm solves Schrödinger bridge problem with mismatched channels.
problem Solving Schrödinger bridge problem with input and noise channel mismatch.
method Design of a Sinkhorn recursion with memory for nonlinear PDEs.
result Demonstrates solving control-affine Schrödinger bridge problem.
Levy processes, which have stationary independent increments, are ideal for modelling the various types of noise that can arise in communication channels. If a Levy process admits exponential moments, then there exists a parametric family of measure changes called Esscher transformations. If the parameter is replaced w…
Method generates audio attacks resistant to reverberation and noise.
problem Physical attacks on speech recognition models.
method Simulates playback and recording transformations to generate robust adversarial examples.
result Adversarial examples can attack speech recognition without being noticed by humans.
Proposes a method to estimate SDE noise from a single trajectory.
problem Estimating SDE noise from a single data trajectory without ergodicity or stationarity.
method Combining Taylor expansions, Girsanov transformations, and drift function's initial value for drift and noise estimation.
result First SSISDE algorithm capable of identifying SDE dynamics from a single trajectory.
Computing accurate estimates of the Fourier transform of analog signals from discrete data points is important in many fields of science and engineering. The conventional approach of performing the discrete Fourier transform of the data implicitly assumes periodicity and bandlimitedness of the signal. In this paper, we…
Generative model for financial time series using structured noise and signature learning.
problem Creating synthetic financial data to reflect real-world market dynamics.
method Structured noise, moving average model, signature transform, reinforcement learning.
result Model effectively captures key financial characteristics and outperforms existing methods.
We develop a classification algorithm for estimating posterior distributions from positive-unlabeled data, that is robust to noise in the positive labels and effective for high-dimensional data. In recent years, several algorithms have been proposed to learn from positive-unlabeled data; however, many of these contribu…
In nonlinear latent variable models or dynamic models, if we consider the latent variables as confounders (common causes), the noise dependencies imply further relations between the observed variables. Such models are then closely related to causal discovery in the presence of nonlinear confounders, which is a challeng…
Paper develops privacy-preserving mechanisms for machine learning using wavelet transforms.
problem Improper data privacy methods compromise user data even with small preliminary knowledge.
method Three privacy-preserving mechanisms with discrete M-band wavelet transform.
result Successfully retains both differential privacy and learnability in various machine learning environments.
Optimal test for detecting signal in noisy matrix model.
problem Signal detection in noisy matrix models with unknown rank.
method Hypothesis test based on linear spectral statistics, optimal under Gaussian noise.
result Optimal test under Gaussian noise, improved with non-Gaussian noise.
Echo noise improves compression in autoencoders with exact rate-distortion.
problem Limitations of Gaussian noise in lossy compression and VAEs.
method Introduces Echo noise, a data-driven noise channel with exact mutual information.
result Echo noise leads to improved bounds on log-likelihood and dominates VAEs.
Attention models can overfit without harming test performance.
problem Understanding benign overfitting in single-head attention models.
method Analyzing conditions for benign overfitting in a single-head softmax attention model.
result A single-head attention model can overfit without harming test performance under certain conditions.
We develop a flexible framework for low-rank matrix estimation that allows us to transform noise models into regularization schemes via a simple bootstrap algorithm. Effectively, our procedure seeks an autoencoding basis for the observed matrix that is stable with respect to the specified noise model; we call the resul…
Adaptive denoising models adjust the number of steps based on noise level.
problem Generating data with lower intrinsic dimensions.
method Adaptive diffusion models using Doob's h-transform to terminate at a random time.
result Adaptive models simplify termination to a first-hitting rule, enhancing adaptability.
Unified diffusion framework enhances generative models flexibility.
problem Improving generative models' design freedom and efficiency.
method Unified framework incorporating choice of representation, prior distribution, and noise scheduling.
result Enhanced flexibility leading to more efficient training and data generation.
Proposes ENVAR for causal discovery in structural VAR models with equal noise variance.
problem Challenges in causal discovery from multivariate time series with contemporaneous effects.
method Introduces observational equivalence and the observational alignment discrepancy for structural VAR models with equal noise variance.
result Shows that multiple structural VAR parameterizations can induce the same stationary observed process law.
Study on signal-plus-noise decomposition in nonlinear spiked random matrices.
problem Nonlinear spiked random matrix models with rank-one signal and noise.
method Signal-plus-noise decomposition and phase transition analysis.
result Identified precise phase transitions in signal components at critical thresholds.
Shot-Noise processes constitute a useful tool in various areas, in particular in finance. They allow to model abrupt changes in a more flexible way than processes with jumps and hence are an ideal tool for modelling stock prices, credit portfolio risk, systemic risk, or electricity markets. Here we consider a general f…
This work shows how transformers use multi-concept word semantics for efficient in-context learning.
problem Understanding the connection between transformer-based LLMs' multi-concept semantic representation and their innovative in-context learning abilities.
method A concept-based low-noise sparse coding prompt model, leveraging advanced techniques to analyze the exponential convergence of 0-1 loss over non-convex training dynamics.
result Transformers leverage multi-concept word semantics to enable powerful and excellent out-of-distribution in-context learning.
The problem of Poisson denoising appears in various imaging applications, such as low-light photography, medical imaging and microscopy. In cases of high SNR, several transformations exist so as to convert the Poisson noise into an additive i.i.d. Gaussian noise, for which many effective algorithms are available. Howev…
NR-GANs learn clean images from noisy data.
problem Learning clean images from noisy training data.
method Introduced a noise generator and distribution/transformation constraints.
result NR-GANs can generate clean images from noisy data.
Kernel estimator optimally recovers function from noisy exponential Radon transform.
problem Inverting noisy exponential Radon transform of a function.
method Proposed a kernel estimator to estimate the true function.
result The estimator converges to the true function at minimax optimal rate.
Paper proposes PMformer for better cryptocurrency price forecasting.
problem Huge volatility and trade-off between univariate and multivariate models.
method Partial-multivariate approach using PMformer.
result PMformer achieves significant statistical accuracy in forecasting.
New method improves few-shot learning with noisy labels.
problem Robustness to label noise in few-shot learning.
method Feature aggregation and Transformer model for noisy samples.
result TraNFS outperforms other methods in noisy conditions.