TransGCN combines GCNs with transformation assumptions for better link prediction in KGs.
problem Link prediction in knowledge graphs for understanding graph structure.
method Unified GCN framework with simultaneous learning of entity and relation embeddings, using transformation assumptions.
result TransGCN outperforms state-of-the-art models on FB15K-237 and WN18RR.
Under a convexity assumption on the boundary we solve a local inverse problem, namely we show that the geodesic X-ray transform can be inverted locally in a stable manner; one even has a reconstruction formula. We also show that under an assumption on the existence of a global foliation by strictly convex hypersurfaces…
Backlund transformations of admissible curves in the Galilean 3-space and pseudo-Galilean 3-space and also spatial Backlund transformations of space curves in Galilean 4-space preserve the torsions under certain assumptions.
Study on DiTs' rates of approximation and estimation under various data assumptions.
problem Investigating statistical rates of conditional diffusion transformers.
method Discretization and Taylor expansion of conditional diffusion score function under Hölder smooth data assumption.
result Establishes statistical limits for conditional and unconditional DiTs, offering practical guidance.
We investigate basic features of Bianchi's Bäcklund transformation of quadrics to see if it can be obtained under weaker assumptions and if it can be generalized to deformations of other surfaces.
We show injectivity of the X-ray transform and the d-plane Radon transform for distributions on the n-torus, lowering the regularity assumption in the recent work by Abouelaz and Rouvière. We also show solenoidal injectivity of the X-ray transform on the n-torus for tensor fields of any order, allowing the tensor…
We show injectivity of the geodesic X-ray transform on piecewise constant functions when the transform is weighted by a continuous matrix weight. The manifold is assumed to be compact and nontrapping of any dimension, and in dimension three and higher we assume a foliation condition. We make no assumption regarding con…
FAT-GAN simulates electron-proton scattering without theoretical assumptions.
problem Efficiently training GANs to simulate complex particle distributions.
method Developed FAT-GAN using transformed and augmented features to improve GAN performance.
result FAT-GAN accurately reproduces electron momenta distributions in electron-proton scattering.
Study X-ray transform on conic spaces, proving injectivity under certain conditions.
problem Injectivity of geodesic X-ray transform on conic metrics.
method Injectivity under non-trapping and no conjugate point assumptions.
result Injectivity of geodesic X-ray transform for asymptotically conic metrics.
New PAC-Bayes bounds derived using Legendre transform and f-divergences.
problem Deriving PAC-Bayes bounds under various assumptions.
method Combining Legendre transform and Fenchel--Young inequality to derive change-of-measure inequalities.
result Extended PAC-Bayesian guarantees under tailored assumptions.
Under the assumption that the X-ray transform over symmetric solenoidal 2-tensors is injective, we prove that smooth compact connected manifolds with strictly convex boundary, no conjugate points and a hyperbolic trapped set are locally marked boundary rigid.
This work studies clustering in transformer models, proving exponential convergence to a single token state.
problem Understanding the long-term behavior of tokens in transformer models.
method Investigates mean-field transformer models under specific conditions to prove exponential convergence to a single state.
result Transformer models synchronize exponentially fast to a single token state with explicit rates.
Proposes a measure to predict generalization in non-matching environments.
problem Characterizing and comparing generalization of machine learning models in non-matching environments.
method Neighborhood invariance measure, calculating invariance as the largest fraction of transformed points classified into the same class.
result Strong and robust correlation between neighborhood invariance and actual out-of-domain generalization.
Transformers improve with Fourier integral attentions.
problem Inefficiency of dot-product attention in capturing feature dependencies.
method Interpreted attention as kernel regression, proposed FourierFormer with generalized Fourier integral kernels.
result FourierFormer achieves better accuracy and reduces redundancy.
Given two compact hyperkähler surfaces X and Y and a holomorphic vector bundle Q on X×Y, which is a generalized instanton, one can define a Fourier-Mukai transform, which, under suitable assumptions, maps vector bundles on X to vector bundles on Y. If X and Y are dual complex tori, this transform …
Transformer RL optimizes A/B testing for time series experiments.
problem Challenges in applying A/B testing to time series experiments, especially with limited history and strong assumptions.
method Transformer reinforcement learning approach that conditions allocation on full history and optimizes MSE without restrictive assumptions.
result Consistently outperforms existing designs in synthetic, simulator, and real-world data.
Normalizing flows are shown to be equivalent to Bayesian networks, revealing new insights.
problem Understanding the limitations and capabilities of normalizing flows.
method Revisiting normalizing flows as probabilistic graphical models and analyzing their structure.
result Normalizing flows can be reduced to Bayesian networks, revealing new insights into their structure and capabilities.
Transformer-based method improves causal discovery from observational data.
problem Causal discovery from observational data requires explicit assumptions.
method CSIvA transformer architecture trained on synthetic data.
result Transformer-based methods adhere to identifiability theory.
TraCeR uses transformers to analyze survival data with longitudinal covariates.
problem Handling longitudinal covariates and assessing model calibration in survival analysis.
method Transformer-based survival analysis framework with factorized self-attention architecture.
result TraCeR achieves significant performance improvements over state-of-the-art methods.
Transformers converge linearly to optimal models for Gaussian mixtures classification.
problem Theoretical understanding of transformers' in-context classification.
method Gradient descent training of a single-layer transformer for Gaussian mixtures classification.
result Transformers converge linearly to globally optimal models for Gaussian mixtures classification.
For smooth compact connected manifolds with strictly convex boundary, no conjugate points and a hyperbolic trapped set, we prove an equivalence principle concerning the injectivity of the X-ray transform Im on symmetric solenoidal tensors and the surjectivity of an operator πm∗ on the set of solenoidal tensors…
The problem of inhomogeneous cluster densities has been a long-standing issue for distance-based and density-based algorithms in clustering and anomaly detection. These algorithms implicitly assume that all clusters have approximately the same density. As a result, they often exhibit a bias towards dense clusters in th…
A framework uses free probability to analyze Transformer models.
problem Understanding the dynamics and complexity of Transformer-based language models.
method Formal operator-theoretic analysis using free probability theory.
result Entropy-based generalization bounds derived under freeness assumptions.
Guillarmou extends X-ray transform to magnetic and thermostat flows.
problem Stability of magnetic X-ray transforms.
method Generalizes normal operator to thermostat and magnetic flows, proving ellipticity.
result Elliptic pseudodifferential operators of order -1 for generalized normal operators.
The paper provides convergence guarantees for ODE-based generative models using transformers.
problem Theoretical guarantees for ODE-based generative models.
method A pre-trained autoencoder maps inputs to a latent space, and a transformer predicts the velocity field.
result The distribution of samples generated via estimated ODE flow converges to the target distribution in Wasserstein-2 distance.
Current multi-view factorization methods make assumptions that are not acceptable for many kinds of data, and in particular, for graphical data with hierarchical structure. At the same time, current hierarchical methods work only in the single-view setting. We generalize the Treelet Transform to the Multi-View Treelet …
Develops European power option pricing under correlated interest rate and asset processes.
problem Pricing European power options under correlated interest rate and asset processes.
method Martingale method and Girsannov transform.
result Derives European power option pricing formulae under two market assumptions.
New method detects symmetries beyond affine transformations.
problem Current methods limit symmetry detection to affine transformations.
method Framework for discovering continuous symmetry beyond affine transformations.
result Method is competitive for large sample sizes and superior for small sample sizes.
Transformers can interpolate between arbitrary measures.
problem Understanding the expressive power of Transformers as measure-to-measure maps.
method Provided an explicit choice of parameters for a single Transformer to match N arbitrary input measures to N arbitrary target measures.
result A single Transformer can interpolate between arbitrary measures.
This paper proposes a new RV prediction model using neural distributional transformation and co-training.
problem Predicting skewed and fat-tailed realized volatility (RV) is challenging.
method The paper uses a neural distributional transformation and co-training to predict RV. It jointly trains the transformation and prediction model using a maximum-likelihood objective function.
result The proposed method significantly outperforms other methods on a dataset of 100 stocks.
Let (M,g) be a simple Riemannian manifold. Under the assumption that the metric g is real-analytic, it is shown that if the geodesic ray transform of a function f∈L2(M) vanishes on an appropriate open set of geodesics, then f=0 on the set of points lying on these geodesics. The approach is based on a micr…
Maps asymptotically embed conic transforms from circle bundles.
problem Embedding conic transforms from circle bundles.
method Asymptotic embeddings using equivariant Szegő projectors.
result Maps embed conic transforms from circle bundles.
GOAD improves anomaly detection across various data types.
problem Finding anomalies in diverse data types.
method GOAD combines classification and transformation-based methods.
result GOAD achieves state-of-the-art accuracy on multiple datasets.
Proposes a faster Transformer decoding method by truncating target-side self-attention windows.
problem Efficiency in Transformer decoding with minimal BLEU score loss.
method N-gram assumption to truncate target-side self-attention windows.
result N-gram masked self-attention model maintains BLEU score for N values from 4 to 8. A novel transformer model improves classification of partially ordered sequences.
problem Classification of partially ordered sequences with uncertainty in timestamps.
method Developed a transformer-based model for partially ordered sequences, benchmarked against set models.
result Transformer-based model outperforms set models on three datasets.
In machine learning and data mining, linear models have been widely used to model the response as parametric linear functions of the predictors. To relax such stringent assumptions made by parametric linear models, additive models consider the response to be a summation of unknown transformations applied on the predict…
Vision transformers benefit from non-smooth components in adaptation.
problem Understanding the role of non-smoothness in vision transformer adaptation.
method Theoretical analysis and extensive experiments on large-scale vision transformers.
result High plasticity of attention modules and feedforward layers leads to better finetuning performance.
New bounds on NTK's smallest eigenvalue for arbitrary data without distributional assumptions.
problem Existing bounds on NTK's smallest eigenvalue require distributional assumptions and high-dimensional data.
method Novel application of the hemisphere transform.
result Bounds on NTK's smallest eigenvalue hold with high probability even for constant input dimension.
Solves wave equation on non-flat harmonic manifolds using Abel transform and Fourier analysis.
problem Wave equation on non-flat harmonic manifolds with specific curvature conditions.
method Explicit representation using inverse dual Abel transform and Fourier transform.
result Shows asymptotic Huygens principle and equidistribution of energy.
Minimal token perturbations reveal how Transformer models process information.
problem Understanding information propagation in Transformer models for interpretability.
method Study of minimal token perturbations on embedding space.
result Rare tokens cause larger shifts, and input information mixes deeper.
Transformers learn to cluster Gaussian mixtures as well as the EM algorithm.
problem Learning guarantees of Transformers in multi-class clustering of Gaussian mixtures.
method Developed a theory connecting Transformer's Softmax Attention layers to the EM algorithm's workflow.
result Transformers achieve minimax optimal rate for clustering Gaussian mixtures with sufficient training samples and initialization.
A new model avoids the PH assumption for right-censored survival data.
problem Inflexibility of Cox model when PH assumption fails.
method Deep partially linear transformation model (DPLTM) for right-censored data.
result The DPLTM avoids the curse of dimensionality and retains interpretability.
Here, by extending the definition of circle to Finsler geometry, we show that, every circle-preserving local diffeomorphism is conformal. This result implies that in Finsler geometry, the definition of concircular change of metrics, a priori, does not require the conformal assumption.
A theorem transforms Lorentzian to signature-changing metrics.
problem Signature-changing manifolds and their initial conditions.
method Transformation prescription to change metrics.
result Transformation Theorem linking Lorentzian to signature-changing metrics.
Latent feature models are attractive for image modeling, since images generally contain multiple objects. However, many latent feature models ignore that objects can appear at different locations or require pre-segmentation of images. While the transformed Indian buffet process (tIBP) provides a method for modeling tra…
Unified view of GNNs as graph signal denoising.
problem Understanding and improving GNNs for graph data.
method Established GNNs as graph denoising problems with smoothness assumptions.
result Unified framework UGNN for adaptive smoothness graphs.
SDPM models survival analysis without parametric assumptions, achieving competitive performance.
problem Estimating survival distributions from censored data with flexibility and accuracy.
method Generative model using denoising diffusion, avoiding parametric assumptions and discretization.
result SDPM achieves competitive predictive performance across various metrics.
We characterize compact locally conformally Kähler (l.c.K.) manifolds under the assumption of a purely conformal, holomorphic circle action. As an application, we determine the structure of the compact l.c.K. manifolds with parallel Lee form. We introduce the Lee-Cauchy-Riemann (LCR) transformations as a class of diffe…