Random projections are able to perform dimension reduction efficiently for datasets with nonlinear low-dimensional structures. One well-known example is that random matrices embed sparse vectors into a low-dimensional subspace nearly isometrically, known as the restricted isometric property in compressed sensing. In th…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Generates low-dimensional node vectors for graphs with privacy while preserving structural preferences.
Classifies geodesic vectors in low-dimensional Lie algebras.
The local linear embedding algorithm (LLE) is a non-linear dimension-reducing technique, widely used due to its computational simplicity and intuitive approach. LLE first linearly reconstructs each input point from its nearest neighbors and then preserves these neighborhood relations in the low-dimensional embedding. W…
Agents collaborate to reduce regret in a multi-agent linear bandit problem with side information.
LIT-LVM improves linear predictors by estimating interaction terms with latent vectors.
We introduce {\em vector diffusion maps} (VDM), a new mathematical framework for organizing and analyzing massive high dimensional data sets, images and shapes. VDM is a mathematical and algorithmic generalization of diffusion maps and other non-linear dimensionality reduction methods, such as LLE, ISOMAP and Laplacian…
Enhances multi-tag classification using low-dimensional vector representations and virtual data.
The study lists low-dimensional stratified groups and their properties.
Embodied cognition states that semantics is encoded in the brain as firing patterns of neural circuits, which are learned according to the statistical structure of human multimodal experience. However, each human brain is idiosyncratically biased, according to its subjective experience history, making this biological s…
We introduce the problem of reconstructing a sequence of multidimensional real vectors where some of the data are missing. This problem contains regression and mapping inversion as particular cases where the pattern of missing data is independent of the sequence index. The problem is hard because it involves possibly m…
The paper analyzes side effects of learning from low-dimensional data embedded in a Euclidean space.
Proposes MR-SNE for multimodal data visualization.
Consider a dataset of vector-valued observations that consists of noisy inliers, which are explained well by a low-dimensional subspace, along with some number of outliers. This work describes a convex optimization problem, called REAPER, that can reliably fit a low-dimensional model to this type of data. This approach…
In this paper we prove that both complete and vertical lifts of a Poisson vector field from a Poisson manifold to its tangent bundle are also Poisson. We use this fact to describe the infinitesimal deformations of Poisson tensor . We study some of their properties and present a extensive…
We apply the concept of castling transform of prehomogeneous vector spaces to produce new examples of minimal homogeneous Lagrangian submanifolds in the complex projective space. Furthermore we verify the Hamiltonian stability of a low dimensional example that can be obtained in this way.
Proposes a novel approach using vector cross product to preserve directional edges in directed graphs.
Extends dimension reduction to data-driven settings without gradients.
Shifu2 discovers advisor-advisee relationships in collaboration networks.
Recently a variety of methods have been developed to encode graphs into low-dimensional vectors that can be easily exploited by machine learning algorithms. The majority of these methods start by embedding the graph nodes into a low-dimensional vector space, followed by using some scheme to aggregate the node embedding…
Principal component analysis (PCA) is an unsupervised method for learning low-dimensional features with orthogonal projections. Multilinear PCA methods extend PCA to deal with multidimensional data (tensors) directly via tensor-to-tensor projection or tensor-to-vector projection (TVP). However, under the TVP setting, i…
Word embedding maps words into a low-dimensional continuous embedding space by exploiting the local word collocation patterns in a small context window. On the other hand, topic modeling maps documents onto a low-dimensional topic space, by utilizing the global word collocation patterns in the same document. These two …
This paper considers the problem of embedding directed graphs in Euclidean space while retaining directional information. We model a directed graph as a finite set of observations from a diffusion on a manifold endowed with a vector field. This is the first generative model of its kind for directed graphs. We introduce…
The paper studies how norms of random vectors are preserved by random projections.
We review localization techniques for functional integrals which have recently been used to perform calculations in and gain insight into the structure of certain topological field theories and low-dimensional gauge theories. These are the functional integral counterparts of the Mathai-Quillen formalism, the Duistermaa…
Paper reviews advances in solving sparsest vector problem in subspaces.
Many machine learning problems, especially multi-modal learning problems, have two sets of distinct features (e.g., image and text features in news story classification, or neuroimaging data and neurocognitive data in cognitive science research). This paper addresses the joint dimensionality reduction of two feature ve…
The incredible variety of galaxy shapes cannot be summarized by human defined discrete classes of shapes without causing a possibly large loss of information. Dictionary learning and sparse coding allow us to reduce the high dimensional space of shapes into a manageable low dimensional continuous vector space. Statisti…
Training task diversity improves ICL with linear attention.
We consider dynamic pricing with many products under an evolving but low-dimensional demand model. Assuming the temporal variation in cross-elasticities exhibits low-rank structure based on fixed (latent) features of the products, we show that the revenue maximization problem reduces to an online bandit convex optimiza…
Graph representation converts complex networks into vectors for easier analysis.
New algorithms reduce computational burden for principal support vector machines.
In this paper, we propose the distributed tree kernels (DTK) as a novel method to reduce time and space complexity of tree kernels. Using a linear complexity algorithm to compute vectors for trees, we embed feature spaces of tree fragments in low-dimensional spaces where the kernel computation is directly done with dot…
Multi-label learning is concerned with the classification of data with multiple class labels. This is in contrast to the traditional classification problem where every data instance has a single label. Due to the exponential size of output space, exploiting intrinsic information in feature and label spaces has been the…
Contrastive learning adapts to data intrinsic dimensions, learning low-dimensional representations.
We discuss Poincaré duality complexes X and the question whether or not their Spivak normal fibration admits a reduction to a vector bundle in the case where the dimension of X is at most 4. We show that in dimensions less than 4 such a reduction always exists, and in dimension 4 such a reduction exists provided X is o…
Given labeled points in a high-dimensional vector space, we seek a low-dimensional subspace such that projecting onto this subspace maintains some prescribed distance between points of differing labels. Intended applications include compressive classification. Taking inspiration from large margin nearest neighbor class…
This paper tackles multilabel classification by exploiting label sparsity and hierarchy.
Multi-criteria recommender systems have been increasingly valuable for helping consumers identify the most relevant items based on different dimensions of user experiences. However, previously proposed multi-criteria models did not take into account latent embeddings generated from user reviews, which capture latent se…
Word embedding models offer continuous vector representations that can capture rich contextual semantics based on their word co-occurrence patterns. While these word vectors can provide very effective features used in many NLP tasks such as clustering similar words and inferring learning relationships, many challenges …
A new DDR framework learns low-dimensional data representations using dynamical systems.
TOFU-POV tackles partially observed linear bandits, achieving sublinear regret with low-dimensional action vectors.
A new method for analyzing product competition using low-dimensional embeddings.
Network representation learning in low dimensional vector space has attracted considerable attention in both academic and industrial domains. Most real-world networks are dynamic with addition/deletion of nodes and edges. The existing graph embedding methods are designed for static networks and they cannot capture evol…
We propose the application of a high-speed maximum likelihood clustering algorithm to detect temporal financial market states, using correlation matrices estimated from intraday market microstructure features. We first determine the ex-ante intraday temporal cluster configurations to identify market states, and then st…
Linear classifiers in product space forms improve scRNA-seq data classification.
A new method for one-class classification using ellipsoidal encapsulation.
A new framework enhances generative modeling by learning local flows over complex manifolds.