In this paper, we propose several dictionary learning algorithms for sparse representations that also impose specific structures on the learned dictionaries such that they are numerically efficient to use: reduced number of addition/multiplications and even avoiding multiplications altogether. We base our work on facto…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Hybrid approach combines transformer and Bayesian filtering for robust multiple particle tracking.
Study very stable Higgs bundles on Riemann surfaces, linking to multiplicity and mirror symmetry.
MPP trains a transformer to predict multiple physical systems, improving accuracy across various tasks.
We develop a deep autoencoder architecture that can be used to find a coordinate transformation which turns a nonlinear PDE into a linear PDE. Our architecture is motivated by the linearizing transformations provided by the Cole-Hopf transform for Burgers equation and the inverse scattering transform for completely int…
We describe a method for training accurate Transformer machine-translation models to run inference using 8-bit integer (INT8) hardware matrix multipliers, as opposed to the more costly single-precision floating-point (FP32) hardware. Unlike previous work, which converted only 85 Transformer matrix multiplications to IN…
HFformer outperforms LSTM in high-frequency trading with multiple signals.
Study multiplicity-free covering of graded manifolds, proving equivalence of categories.
New CMC surfaces with dihedral symmetry constructed from Darboux transforms.
SurvTRACE uses transformers to analyze survival times with competing events.
Learning multiple tasks across heterogeneous domains is a challenging problem since the feature space may not be the same for different tasks. We assume the data in multiple tasks are generated from a latent common domain via sparse domain transforms and propose a latent probit model (LPM) to jointly learn the domain t…
Transformers use ReLUs to approximate softmax efficiently.
Fast linear transforms are ubiquitous in machine learning, including the discrete Fourier transform, discrete cosine transform, and other structured transformations such as convolutions. All of these transforms can be represented by dense matrix-vector multiplication, yet each has a specialized and highly efficient (su…
BoostTransformer uses boosting to improve transformer efficiency and accuracy.
Fusion of transformer networks using optimal transport for improved performance.
Paper proposes a robust framework for detecting multiple periodic components in time series.
MultiRocket boosts TSC speed and accuracy with pooling and transformations.
We consider an application involving a financial quadratic portfolio of options, when the joint underlying log-returns changes with multivariate elliptic distribution. This motivates the needs for methods for the approximation of multiple integrals over hyperboloids. A transformation is used to reduce the hyperboloid i…
The classical Liouville Theorem on conformal transformations determines local conformal transformations on the Euclidean space of dimension . Its natural adaptation to the general framework of Riemannian structures is the 2-rigidity of conformal transformations, that is such a transformation is fully determined…
Improved signal classification using multiple wavelets and their smooth coefficients.
LOFT separates subspace rotation and transformation for orthogonal fine-tuning.
We show that for any complete connected Kähler manifold the index of the group of complex affine transformations in the group of c-projective transformations is at most two unless the Kähler manifold is isometric to complex projective space equipped with a positive constant multiple of the Fubini-Study metric. This est…
Improves contrastive learning invariance with novel training objectives and feature averaging.
MIMONets speed up neural network inference by processing multiple inputs in parallel.
Outlier detection aims to identify unusual data instances that deviate from expected patterns. The outlier detection is particularly challenging when outliers are context dependent and when they are defined by unusual combinations of multiple outcome variable values. In this paper, we develop and study a new conditiona…
We analyze a simple asset transfer model in which the transfer amount is a fixed fraction of the giver's wealth. The model is analyzed in a new way by Laplace transforming the master equation, solving it analytically and numerically for the steady-state distribution, and exploring the solutions for various values o…
Transformers learn to recall with non-orthogonal embeddings in realistic settings.
Extends SW and GSW to compare heterogeneous joint distributions.
Transformers can learn optimal variable selection in group-sparse classification.
Transformer models show robustness across domains with domain adversarial training.
Probabilistic STNs improve image classification and robustness.
Transformers simplify modeling of small longitudinal cohort data by reducing parameters and incorporating attention mechanisms.
The alignment of a set of objects by means of transformations plays an important role in computer vision. Whilst the case for only two objects can be solved globally, when multiple objects are considered usually iterative methods are used. In practice the iterative methods perform well if the relative transformations b…
We incorporate Tensor-Product Representations within the Transformer in order to better support the explicit representation of relation structure. Our Tensor-Product Transformer (TP-Transformer) sets a new state of the art on the recently-introduced Mathematics Dataset containing 56 categories of free-form math word-pr…
Many-to-Many VTN improves voice conversion across multiple speakers.
Current multi-view factorization methods make assumptions that are not acceptable for many kinds of data, and in particular, for graphical data with hierarchical structure. At the same time, current hierarchical methods work only in the single-view setting. We generalize the Treelet Transform to the Multi-View Treelet …
Sparse coding is a common approach to learning local features for object recognition. Recently, there has been an increasing interest in learning features from spatio-temporal, binocular, or other multi-observation data, where the goal is to encode the relationship between images rather than the content of a single ima…
Investigates the fundamental components of attention mechanisms.
Segmentation maps of medical images annotated by medical experts contain rich spatial information. In this paper, we propose to decompose annotation maps to learn disentangled and richer feature transforms for segmentation problems in medical images. Our new scheme consists of two main stages: decompose and integrate. …
Unified framework for learning function representations using INRs and Transformers.
Most existing fingerprints-based indoor localization approaches are based on some single fingerprints, such as received signal strength (RSS), channel impulse response (CIR), and signal subspace. However, the localization accuracy obtained by the single fingerprint approach is rather susceptible to the changing environ…
Although nonstationary data are more common in the real world, most existing causal discovery methods do not take nonstationarity into consideration. In this letter, we propose a kernel embedding-based approach, ENCI, for nonstationary causal model inference where data are collected from multiple domains with varying d…
Transformer-based method for causal discovery with prior knowledge integration.
GOAD improves anomaly detection across various data types.
X-ray transform on H-type groups solved, revealing function injectivity.
Transforms ensemble predictions to maintain interpretability.
We discuss several aspects of Mellin transform, including distributional Mellin transform and inversion of multiple Mellin-Barnes integrals in and its connection to residue expansion or evaluation of Laplace integrals. These mathematical concepts are demonstrated on several option-pricing models. This in…
Improved SVMs learn from few samples with composition and multiple scales.