The article recovers tensor fields from partial data using weighted divergent ray transforms.
problem Recovering tensor fields from partial data.
method Weighted divergent ray transforms, unique continuation property of fractional Laplacian, explicit reconstruction formulas.
result Recovery of symmetric m-tensor fields and unique continuation for vector fields and symmetric 2-tensor fields. New Transformers maintain Lipschitz continuity for robustness.
problem Ensuring robustness in Transformers for safety-sensitive applications.
method Introducing gradient-descent-type in-context Transformers with explicit Euler steps of negative gradient flows.
result Universal approximation theorem for Lipschitz continuous Transformers.
We analyze Darboux transformations in very general settings for multidimensional linear partial differential operators. We consider all known types of Darboux transformations, and present a new type. We obtain a full classification of all operators that admit Wronskian type Darboux transformations of first order and a …
New geometric transformations link discrete and continuous curve motions.
problem Establishing a connection between discrete and continuous curve motions.
method Infinitesimal Darboux transformations of smooth curves.
result Alternate geometric interpretation for semi-discrete mKdV equation.
New method detects symmetries beyond affine transformations.
problem Current methods limit symmetry detection to affine transformations.
method Framework for discovering continuous symmetry beyond affine transformations.
result Method is competitive for large sample sizes and superior for small sample sizes.
Kendall transformation converts continuous data into categorical vectors for robust information theory.
problem Handling small number of observations and preserving ranking in continuous data.
method Kendall transformation converts ordered features into categorical vectors of pairwise order relations.
result Kendall transformation makes information theory methods applicable to continuous data robustly.
Proves conditions for Fourier transforms in rank 1 symmetric spaces.
problem Understanding Fourier transform bounds in symmetric spaces.
method Proves sufficient and necessary conditions using Lipschitz and Fourier type integral conditions.
result Establishes bounds for Fourier transforms in rank 1 symmetric spaces with specific moduli of continuity.
CAM-GAN improves GANs for continual learning with efficient feature map transformations.
problem Efficient continual learning for GANs with reduced parameter growth.
method Designing and leveraging parameter-efficient feature map transformations, including global and task-specific parameters, residual bias, and Fisher information matrix.
result Significantly improved model performance and high-quality samples with fewer parameters.
Study Transformer layers under cross-entropy training using mean field control.
problem Understanding the behavior of Transformer layers in cross-entropy training.
method Continuous-depth mean field control analysis, treating depth as time and layer parameters as controls.
result Derivation of a Pontryagin condition for the limiting population problem, involving the softmax residual.
In this paper we investigate continuous speech recognition using electroencephalography (EEG) features using recently introduced end-to-end transformer based automatic speech recognition (ASR) model. Our results demonstrate that transformer based model demonstrate faster training compared to recurrent neural network (R…
This research studies affine invariance in continuous-domain convolutional neural networks.
problem Recognizing patterns and features under affine transformations in continuous domains.
method Introduces a new criterion for assessing affine invariance, embeds images into the affine Lie group, and analyzes convolution over this group.
result Extends the scope of geometrical transformations that deep-learning pipelines can handle.
Cai, Song and Kou (2015) [Cai, N., Y. Song, S. Kou (2015) A general framework for pricing Asian options under Markov processes. Oper. Res. 63(3): 540-554] made a breakthrough by proposing a general framework for pricing both discretely and continuously monitored Asian options under one-dimensional Markov processes. In …
Proposes a new feature preprocessing method using kernel density integral transformation.
problem Feature preprocessing for tabular data in machine learning and statistics.
method Kernel density integral transformation as a drop-in replacement or improved alternative to min-max scaling and quantile transformation.
result Frequently outperforms min-max scaling and quantile transformation with hyperparameter tuning.
Graded Transformers embed algebraic structure in neural networks through graded transformations.
problem Efficiently modeling hierarchical and structured data in neural networks.
method Introduces Linearly Graded Transformer (LGT) and Exponentially Graded Transformer (EGT) with graded scaling operators.
result Establishes rigorous guarantees and improved efficiency for structured data.
PPT optimizes transformer behavior by steering its latent posterior using prior samples.
problem Eliciting desired behavior from transformers without backpropagation.
method Posterior Prefix Tuning (PPT) uses predictive Monte Carlo (PMC) samples and importance sampling to optimize the latent posterior.
result PPT optimizes transformer behavior without backpropagation, achieving high utility across different utility functions.
Transformers can interpolate between arbitrary measures.
problem Understanding the expressive power of Transformers as measure-to-measure maps.
method Provided an explicit choice of parameters for a single Transformer to match N arbitrary input measures to N arbitrary target measures.
result A single Transformer can interpolate between arbitrary measures.
Rough Transformers improve time series modeling with lower costs and better performance.
problem Inefficient modeling of irregularly sampled time series data.
method Signature patching for continuous-time representations, reducing computational costs.
result Rough Transformers outperform vanilla Transformers and Neural ODE models.
Paper tackles continual learning with Transformers, maintaining performance and speed.
problem Catastrophic forgetting in continual learning with limited resources.
method Incremental training with pre-trained Transformers and Adapters.
result Maintains good predictive performance and inference speed without retraining.
New analysis shows RPE-based Transformers can't approximate all functions.
problem Understanding the limitations of RPE-based Transformers in approximating continuous functions.
method Mathematical analysis and development of a novel attention module (URPE) to overcome limitations.
result RPE-based Transformers can't approximate all continuous sequence-to-sequence functions, even with depth and width.
Training of discrete latent variable models remains challenging because passing gradient information through discrete units is difficult. We propose a new class of smoothing transformations based on a mixture of two overlapping distributions, and show that the proposed transformation can be used for training binary lat…
Despite the widespread adoption of Transformer models for NLP tasks, the expressive power of these models is not well-understood. In this paper, we establish that Transformer models are universal approximators of continuous permutation equivariant sequence-to-sequence functions with compact support, which is quite surp…
ACSSM models irregular time series with continuous dynamics.
problem Modeling irregular time series data.
method ACSSM uses a multi-marginal Doob's h-transform and variational inference with stochastic optimal control.
result ACSSM outperforms in tasks like classification, regression, interpolation, and extrapolation.
ABHT boosts regression by filtering regions with different smoothness.
problem Improving regression performance through local adaptivity.
method Gradient boosting with adaptive histogram transform.
result ABHT converges faster than PEHT in Hölder continuous spaces.
Unique continuation for X-ray transforms of one-forms with partial data.
problem Proving unique continuation for X-ray transforms of one-forms with limited data.
method Proved unique continuation for the normal operator of X-ray transforms of one-forms, leading to partial data results.
result Unique continuation for X-ray transforms of one-forms with partial data.
Recently, we proposed short-time Fourier transform (STFT)-based loss functions for training a neural speech waveform model. In this paper, we generalize the above framework and propose a training scheme for such models based on spectral amplitude and phase losses obtained by either STFT or continuous wavelet transform …
Transformers can solve complex filtering problems for non-Gaussian signals.
problem Non-linear and non-Markovian filtering problems for conditionally Gaussian signals.
method Continuous-time transformer models called filterformers.
result Filterformers can approximate the conditional law of non-Markovian and conditionally Gaussian signal processes.
Analyzes non-Markovian environments in stochastic approximation.
problem Understanding learning mechanisms in non-ergodic, non-Markovian settings.
method Analytic framework for transformer learning and continual learning.
result Proposes a new approach to transformer and continual learning.
We show injectivity of the geodesic X-ray transform on piecewise constant functions when the transform is weighted by a continuous matrix weight. The manifold is assumed to be compact and nontrapping of any dimension, and in dimension three and higher we assume a foliation condition. We make no assumption regarding con…
New model learns symmetry transformations from complex data.
problem Learning symmetry transformations in complex domains like chemical space.
method Two latent subspaces, deep information bottleneck, continuous mutual information regularizer.
result Model outperforms state-of-the-art methods on artificial and molecular datasets.
Continuous phase transitions identified in Doi-Onsager, noisy transformer, and Hegselmann-Krause models.
problem Phase transitions in multimodal models and their properties.
method Sharp coercivity estimate and constrained Lebedev--Milin inequality.
result Continuous phase transitions at critical coupling strengths for Doi-Onsager, noisy transformer, and Hegselmann-Krause models.
Transformers can predict new tokens based on any number of context tokens, approximating continuous mappings with fixed resources.
problem Handling an arbitrarily large number of context tokens in transformers.
method Mathematical analysis of transformer's expressivity using Wasserstein distance and continuous mappings.
result Deep transformers are universal and can approximate continuous in-context mappings to arbitrary precision, uniformly over compact token domains.
Transformers preserve support and can approximate any continuous map.
problem Understanding the mathematical properties of transformers.
method Characterizing maps between measures that can be represented as transformers and proving their properties.
result Transformers preserve support and have uniformly continuous Fréchet derivatives.
New framework discovers non-affine continuous symmetries in neural networks.
problem Lack of efficient methods for detecting non-affine continuous symmetries in neural networks.
method Computational framework for discovering infinitesimal generators of multi-parameter group actions.
result Framework can discover non-affine continuous symmetries in neural networks.
We present a general theory of fractal transformations and show how it leads to a new type of method for filtering and transforming digital images. This work substantially generalizes earlier work on fractal tops. The approach involves fractal geometry, chaotic dynamics, and an interplay between discrete and continuous…
As early as 1972, Penrose - in a purely formal way - introduced a "discontinuous coordinate transformation", which relates a continuous representation of the metric of impulsive pp-waves to a discontinuous one. On the basis of the invertibility concept for generalized functions developed recently by the first author, w…
Shapelet transform improves time series classification for earthquake, wind, and wave events.
problem Autonomous detection of specific events from large time series datasets in civil engineering.
method Shapelet transform for local similarity in time series subsequences, combined with machine learning.
result Shapelet transform yields a new feature representation for time series signals in civil engineering.
Overparameterized models improve performance in sequential learning tasks.
problem Catastrophic forgetting in overparameterized neural networks.
method Two-task linear regression problem with random orthogonal transformations.
result Overparameterization mitigates catastrophic forgetting in sequential learning tasks.
We study the dynamics of the discrete bicycle (Darboux, Backlund) transformation of polygons in n-dimensional Euclidean space. This transformation is a discretization of the continuous bicycle transformation, recently studied by Foote, Levi, and Tabachnikov. We prove that the respective monodromy is a Moebius transform…
Rough Transformers improve efficiency for medical time-series data.
problem Efficiently modeling irregularly sampled, long-range time-series data.
method Introducing Rough Transformers, a Transformer variant with continuous-time representations and multi-view signature attention.
result Rough Transformers outperform vanilla Transformers while using less computational resources.
The paper proposes a method to learn the structure of continuous-action games with non-parametric utilities using a limited number of samples.
problem Learning the exact structure of continuous-action games with non-parametric utility functions.
method An ℓ1 regularized method that encourages sparsity of the Fourier transform coefficients of the utility functions, accessed via a few Nash equilibria and their noisy utilities. result The method recovers the exact structure of the utility functions and the game structure with provable theoretical guarantees.
CL methods improve monolingual ASR models across new tasks without forgetting past data.
problem Catastrophic Forgetting in monolingual ASR models when adapting to new domains or accents.
method Implement and compare various Continual Learning methods for monolingual ASR.
result Best CL method reduces performance gap by over 40% with minimal past data.
Analytical pricing formulas and Greeks are obtained for European and American basket put options using Mellin transforms. We assume assets are driven by geometric Brownian motion which exhibit correlation and pay a continuous dividend rate. A novel approach to numerical Mellin inversion is achieved via the fast Fourier…
In this paper, we continue studying the 6-dimensional pseudo-Riemannian space V^6(g_{ij}) with signature [++--], which admits projective motions, i. e. continuous transformation groups preserving geodesics. In particular, we determine a necessary and sufficient condition that the 6-dimensional rigid h-spaces have const…
Improved robustness of 1D CNNs for heart arrhythmia classification.
problem Improving the robustness of 1D CNNs for classification tasks.
method Parameterization using Cayley transform and controllability Gramian for Lipschitz-bounded CNNs.
result Improved robustness of trained Lipschitz-bounded 1D CNNs for heart arrhythmia classification.
New model recognizes emotions with missing modalities, improving accuracy.
problem Handling missing modalities in emotion recognition.
method Transformer-based architecture with cross-attention and self-attention mechanisms.
result Improvement of 37% in predicting arousal values and 30% in valence values compared to baseline.
This paper describes some applications of an incremental implementation of the principal component analysis (PCA). The algorithm updates the transformation coefficients matrix on-line for each new sample, without the need to keep all the samples in memory. The algorithm is formally equivalent to the usual batch version…
How can prior knowledge on the transformation invariances of a domain be incorporated into the architecture of a neural network? We propose Equivariant Transformers (ETs), a family of differentiable image-to-image mappings that improve the robustness of models towards pre-defined continuous transformation groups. Throu…
Proposes a differentiable STFT for more efficient optimization of hop length.
problem Efficient optimization of hop length in STFT for better temporal control.
method Introduces a differentiable version of STFT with continuous hop length.
result Improves optimization methods like gradient descent for STFT.