Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

58115173230 · Jun 202019922001200920172026
48 results for Uniform Transformation

UT module refines VAE latent space, improving disentanglement and interpretability.

problem Irregular latent distributions cause posterior collapse and misalignment in VAEs.
method UT module uses G-KDE clustering, GM modeling, and PIT to transform latent space into uniform distribution.
result UT module enhances disentanglement and interpretability of latent representations.

The paper studies dynamical properties in semigroups modulo ideals.

problem Analyzing shadowing, expansivity, and stability in semigroups with ideals.
method Investigates shadowing, expansivity, and stability properties in uniform transformation semigroups modulo an ideal.
result Establishes that if a semigroup exhibits shadowing and expansivity modulo an ideal, it is also topologically stable modulo that ideal.

We prove that if YY is the Gromov-Hausdorff limit of a sequence of compact manifolds, MinM^n_i, with a uniform lower bound on Ricci curvature and a uniform upper bound on diameter, then YY has a universal cover. We then show that, for ii sufficiently large, the fundamental group of MiM_i has a surjective homeomorphis…

2000-08-29abs ↗pdf ↗

Uniform scaling limits in AdamW-trained transformers converge to ODEs.

problem Understanding the dynamics of large-depth transformers trained with AdamW.
method Modeling transformer dynamics as an interacting particle system coupled through attention, proving convergence to ODEs.
result The joint dynamics of hidden states and backpropagated variables converge uniformly to an ODE system.

The paper explores uniform perfectness and centers in Morse boundaries.

problem Detecting κκ-center exhaustivity in uniformly perfect Morse boundaries.
method Analyzes CAT(0) and geodesic spaces, using visual boundary data and metric transforms.
result Fixed-basepoint uniform perfectness is insufficient for κκ-center exhaustivity.

In this paper we deduce a local deformation lemma for uniform embeddings in a metric covering space over a compact manifold from the deformation lemma for embeddings of a compact subspace in a manifold. This implies the local contractibility of the group of uniform homeomorphisms of such a metric covering space under t…

2012-03-19abs ↗pdf ↗

Study Transformer layers under cross-entropy training using mean field control.

problem Understanding the behavior of Transformer layers in cross-entropy training.
method Continuous-depth mean field control analysis, treating depth as time and layer parameters as controls.
result Derivation of a Pontryagin condition for the limiting population problem, involving the softmax residual.

Efficiently transforms samples from various statistical models.

problem Approximately transforming samples from one statistical model to another without knowing the source model's parameters.
method Constructs computationally efficient procedures to reduce uniform, Erlang, and Laplace models to general target families.
result Establishes nonasymptotic reductions between canonical high-dimensional problems, such as mixtures of experts, phase retrieval, and signal denoising.

We review the construction known as the Nahm transform in a generalized context, which includes all the examples of this construction already described in the literature. The Nahm transform for translation invariant instantons on 4\real^4 is presented in an uniform manner. We also analyze two new examples, the first o…

2003-09-18abs ↗pdf ↗

In this paper we prove mixed norm estimates for Riesz transforms related to Laplace--Beltrami operators on compact Riemannian symmetric spaces of rank one. These operators are closely related to the Riesz transforms for Jacobi polynomials expansions. The key point is to obtain sharp estimates for the kernel of the Jaco…

2013-08-29abs ↗pdf ↗

The parametrization theorem is derived in a flat nD pseudo-complex affine space. The pseudo-complex hyperbolic space accomodates n-number of uncompactified time-like extra dimensions with sugnature (s,r), where s and r are the numbers of minus and plus signs associated with the diagonalized metric matrix. The main resu…

2010-03-01abs ↗pdf ↗

Enhances Fourier estimator performance for asynchronous event-data.

problem Improving correlation and covariance estimation on event-data.
method Implement and test NUFFT methods with different averaging kernels.
result Demonstrates improved performance and relationship between averaging scales.

The paper extends Strichartz's conjecture to spinor bundles over real hyperbolic spaces.

problem Extending Strichartz's conjecture to spinor bundles.
method Characterization of Poisson transform for spinor bundles and uniform L2L^2 estimates.
result Strichartz's conjecture is extended to spinor bundles over real hyperbolic spaces.

Subsampled Randomized Hadamard Transform (SRHT), a popular random projection method that can efficiently project a dd-dimensional data into rr-dimensional space (rdr \ll d) in O(dlog(d))O(dlog(d)) time, has been widely used to address the challenge of high-dimensionality in machine learning. SRHT works by rotating the input …

2020-02-05abs ↗pdf ↗

An approach to the modelling of volatile time series using a class of uniformity-preserving transforms for uniform random variables is proposed. V-transforms describe the relationship between quantiles of the stationary distribution of the time series and quantiles of the distribution of a predictable volatility proxy …

2020-02-24abs ↗pdf ↗

Uniform approximations for RHTs improve kernel approximation and distance estimation.

problem Theoretical guarantees for RHTs in low-dimensional applications.
method Proved uniform convergence of average of function over RHTs entries.
result Improved guarantees for kernel approximation and distance estimation.

New loss function equivalence reveals PER's uniform sampling can be improved.

problem Improving Prioritized Experience Replay (PER) for better learning efficiency.
method Transforming non-uniformly sampled data loss functions into uniformly sampled ones.
result Some environments can replace PER with a new loss function without performance loss.

Study phase transitions in noisy transformer dynamics on spheres.

problem Understanding phase transitions in noisy transformer dynamics on spheres.
method Sharp Beckner--Onofri/logarithmic HLS inequality, Funk--Hecke/Bessel coefficients, degree-two quartic obstruction.
result Sharp global-minimizer dichotomy and phase transitions in noisy transformer dynamics in arbitrary dimension.

Continuous phase transitions identified in Doi-Onsager, noisy transformer, and Hegselmann-Krause models.

problem Phase transitions in multimodal models and their properties.
method Sharp coercivity estimate and constrained Lebedev--Milin inequality.
result Continuous phase transitions at critical coupling strengths for Doi-Onsager, noisy transformer, and Hegselmann-Krause models.

Scalable kernel methods for large datasets using Fourier representations and NUFFT.

problem Cubic complexity in kernel methods limits their use on large-scale datasets.
method Fourier representation of kernels combined with NUFFT for O(n log n) complexity.
result Achieves minimax convergence rates and processes up to tens of billions of samples.

Driven by the need for parallelizable hyperparameter optimization methods, this paper studies \emph{open loop} search methods: sequences that are predetermined and can be generated before a single configuration is evaluated. Examples include grid search, uniform random search, low discrepancy sequences, and other sampl…

2017-06-06abs ↗pdf ↗

ZeroS improves Transformers by adding negative weights, matching or beating softmax attention.

problem Limited performance of linear attention methods, especially in long context sequences.
method Proposes Zero-Sum Linear Attention (ZeroS) that removes the zero-order term and reweights zero-sum softmax residuals.
result ZeroS matches or exceeds standard softmax attention across various benchmarks, theoretically expanding representable functions.

Paper converts deep networks to flat, equivalent kernel machines.

problem Capacity control and uniform convergence in deep learning.
method Push-forward transformation from deep networks to indefinite kernel machines.
result Flat network weights are Lp-norm regularized (0<p<1).

New algorithm improves sparse-view tomography without needing ground-truth data.

problem Poor image reconstructions with sparse projections and non-uniform sensors.
method Unsupervised deep learning with CNN and STN modules.
result Significantly outperforms filtered backprojection in sparse-view scenarios.

Enhanced Hopfield model boosts memory retrieval capacity.

problem Memory retrieval in modern Hopfield models with limited capacity.
method Introduces a learnable feature map transforming energy function into kernel space, minimizing separation loss for uniform memory distribution.
result Significant reduction in metastable states, enhancing memory capacity and retrieval accuracy.

This paper precisely estimates transformer derivatives for explicit learning guarantees.

problem Computing fully-explicit generalization bounds for transformers with precise higher-order derivative estimates.
method Analyzes and estimates all higher-order derivatives of transformers with multiple attention heads and layer normalization.
result Obtains explicit pathwise generalization bounds for transformers learning from non-i.i.d. samples.

It is well known that the collection of uniformizations of a closed Riemann surface SS is partially ordered; the lowest ones are the Schottky unformizations, that is, tuples (Ω,Γ,P:ΩS)(Ω,Γ,P:Ω\to S), where ΓΓ is a Schottky group with region of discontinuity ΩΩ and P:ΩSP:Ω\to S is a regular holomorphic cover map with ΓΓ as it…

2013-07-09abs ↗pdf ↗

New algorithm for estimating multivariate quantiles using stochastic optimal transport.

problem Estimating multivariate quantiles from data.
method Stochastic algorithm for entropic optimal transport in Banach spaces, using Fourier coefficients.
result Almost sure convergence of the stochastic algorithm in infinite-dimensional Banach spaces.

Deep Graph Neural Networks (GNNs) are useful models for graph classification and graph-based regression tasks. In these tasks, graph pooling is a critical ingredient by which GNNs adapt to input graphs of varying size and structure. We propose a new graph pooling operation based on compressive Haar transforms -- HaarPo…

2019-09-25abs ↗pdf ↗

Rough Transformers improve efficiency for medical time-series data.

problem Efficiently modeling irregularly sampled, long-range time-series data.
method Introducing Rough Transformers, a Transformer variant with continuous-time representations and multi-view signature attention.
result Rough Transformers outperform vanilla Transformers while using less computational resources.

Proposes a new latent variable model for hyperspherical latent spaces.

problem Efficiently modeling heavy-tailed distributions in hyperspherical latent spaces.
method Introduces spherical Cauchy (spCauchy) latent variables and applies Möbius transformations.
result Shows spCauchy recovers vMF geometry in high-concentration limits and avoids complex evaluations.

Transformers with linear space and time complexity for accurate attention estimation.

problem Efficiently estimating attention in large-scale tasks without relying on priors.
method Performers use Fast Attention Via positive Orthogonal Random features (FAVOR+) for linear approximation of softmax attention.
result Performers achieve competitive results on various tasks, demonstrating the effectiveness of their attention-learning approach.

Paper studies Transformer learning theory for Euclidean and Riemannian domains.

problem Understanding and optimizing Transformer networks for regression tasks.
method Constructive approximation framework using softmax partition of unity and attention mechanism.
result Transformer can achieve uniform ε-approximation error with minimal parameters.

Accelerated coordinate descent is widely used in optimization due to its cheap per-iteration cost and scalability to large-scale problems. Up to a primal-dual transformation, it is also the same as accelerated stochastic gradient descent that is one of the central methods used in machine learning. In this paper, we imp…

2015-12-30abs ↗pdf ↗

New Performer model tackles long-sequence protein modeling.

problem Challenges of training complex Transformer models for long sequences.
method Linearly scalable long-context Transformer architecture, Performer.
result Performer provides strong theoretical guarantees and is effective for protein sequence modeling.

Rough Transformers improve time series modeling with lower costs and better performance.

problem Inefficient modeling of irregularly sampled time series data.
method Signature patching for continuous-time representations, reducing computational costs.
result Rough Transformers outperform vanilla Transformers and Neural ODE models.

Paper explores generalization of AID-based bi-level optimization methods.

problem Uncertainty in generalization properties of AID-based bi-level optimization methods.
method Uniform stability analysis and convergence study of AID-based methods.
result AID-based methods can achieve similar generalization as single-level nonconvex problems.

Vision Transformers show different internal representations compared to CNNs.

problem Understanding how Vision Transformers solve image classification tasks.
method Comparative analysis of ViT and CNN architectures on image classification benchmarks.
result ViT has more uniform representations across all layers, while CNNs have more varied representations.