Optimizes embedding accuracy for data variance and error.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Just as semantic hashing can accelerate information retrieval, binary valued embeddings can significantly reduce latency in the retrieval of graphical data. We introduce a simple but effective model for learning such binary vectors for nodes in a graph. By imagining the embeddings as independent coin flips of varying b…
Paper generalizes kernel mean embedding to von Neumann-algebra-valued measures.
Complex embeddings handle non-metric proximity data better than traditional methods.
Learning knowledge representation is an increasingly important technology that supports a variety of machine learning related applications. However, the choice of hyperparameters is seldom justified and usually relies on exhaustive search. Understanding the effect of hyperparameter combinations on embedding quality is …
Captures data influence changes during training.
Representing entities and relations in an embedding space is a well-studied approach for machine learning on relational data. Existing approaches, however, primarily focus on simple link structure between a finite set of entities, ignoring the variety of data types that are often used in knowledge bases, such as text, …
Probabilistic hash embeddings improve online learning of categorical features.
A key obstacle in automated analytics and meta-learning is the inability to recognize when different datasets contain measurements of the same variable. Because provided attribute labels are often uninformative in practice, this task may be more robustly addressed by leveraging the data values themselves rather than ju…
Recommendation problems with large numbers of discrete items, such as products, webpages, or videos, are ubiquitous in the technology industry. Deep neural networks are being increasingly used for these recommendation problems. These models use embeddings to represent discrete items as continuous vectors, and the vocab…
Paper develops heavy-tailed embeddings for better text classification and augmentation.
We approximate derivatives of functions on manifolds by embedding them and applying vector-valued operators.
This paper introduces an approach for detecting differences in the first-order structures of spatial point patterns. The proposed approach leverages the kernel mean embedding in a novel way by introducing its approximate version tailored to spatial point processes. While the original embedding is infinite-dimensional a…
New tensors capture intrinsic embedding data of conformal hypersurfaces.
Bayesian sparsification improves complex-valued neural networks by 50-100x with minimal performance loss.
New method learns state embeddings from demonstrations for improved reinforcement learning.
This paper finds a linear relationship between t-SNE perplexity and data set size.
The recent proliferation of publicly available graph-structured data has sparked an interest in machine learning algorithms for graph data. Since most traditional machine learning algorithms assume data to be tabular, embedding algorithms for mapping graph data to real-valued vector spaces has become an active area of …
MCE reduces embedding instability in nonlinear dimensionality reduction.
Unified framework for complex-valued eigenfunctions on Riemannian symmetric spaces.
Word embeddings are a powerful approach for capturing semantic similarity among terms in a vocabulary. In this paper, we develop exponential family embeddings, a class of methods that extends the idea of word embeddings to other types of high-dimensional data. As examples, we studied neural data with real-valued observ…
The study shows how to accurately estimate embedding vectors in high dimensions.
The study analyzes XRP transaction networks to understand market dynamics.
A novel geometric algebra-based KG embedding framework improves link prediction.
Study estimates gaps in semigroup products, proving embedding properties.
Quantum kernel methods can lead to trivial models due to exponential concentration of kernel values.
The paper defines invariants for almost graph embeddings and explores their properties.
We demonstrate an equivalence between reproducing kernel Hilbert space (RKHS) embeddings of conditional distributions and vector-valued regressors. This connection introduces a natural regularized loss function which the RKHS embeddings minimise, providing an intuitive understanding of the embeddings and a justificatio…
In statistical relational learning, knowledge graph completion deals with automatically understanding the structure of large knowledge graphs---labeled directed graphs---and predicting missing relationships---labeled edges. State-of-the-art embedding models propose different trade-offs between modeling expressiveness, …
VICE embeds concepts in a vector space using human data.
Methods based on vector embeddings of knowledge graphs have been actively pursued as a promising approach to knowledge graph completion.However, embedding models generate storage-inefficient representations, particularly when the number of entities and relations, and the dimensionality of the real-valued embedding vect…
A new method estimates multi-dimensional value distributions using Hilbert space embeddings.
The universal order 1 invariant f^U of immersions of a closed orientable surface into R^3, whose existence has been established in [N3], takes values in the group G_U = K \oplus Z/2 \oplus Z/2 where K is a countably generated free Abelian group. The projections of f^U to K and to the first and second Z/2 factors are de…
A nonparametric approach for policy learning for POMDPs is proposed. The approach represents distributions over the states, observations, and actions as embeddings in feature spaces, which are reproducing kernel Hilbert spaces. Distributions over states given the observations are obtained by applying the kernel Bayes' …
We describe the Customer LifeTime Value (CLTV) prediction system deployed at ASOS.com, a global online fashion retailer. CLTV prediction is an important problem in e-commerce where an accurate estimate of future value allows retailers to effectively allocate marketing spend, identify and nurture high value customers an…
The study of Seifert linking forms for punctured n-manifolds in (2n-1)-space.
An explicit global and unique isometric embedding into hyperbolic 3-space, H^3, of an axi-symmetric 2-surface with Gaussian curvature bounded below is given. In particular, this allows the embedding into H^3 of surfaces of revolution having negative, but finite, Gaussian curvature at smooth fixed points of the U(1) iso…
We unify subsampling methods for network embeddings and prove their asymptotic distribution.
Nonlinear embedding manifold learning methods provide invaluable visual insights into the structure of high-dimensional data. However, due to a complicated nonconvex objective function, these methods can easily get stuck in local minima and their embedding quality can be poor. We propose a natural extension to several …
Investigates the rotating Kepler problem for energy values ≤ -3/2.
Paper introduces RKHM and KME for richer data analysis.
In this paper we prove that an embedded and simply connected constant mean curvature surface with curvature large at a point contains a multi-valued graph around that point on the scale of , where is the norm squared of the second fundamental form. This generalizes Colding and Minicozzi's result for mini…
Skip-gram with negative sampling, a popular variant of Word2vec originally designed and tuned to create word embeddings for Natural Language Processing, has been used to create item embeddings with successful applications in recommendation. While these fields do not share the same type of data, neither evaluate on the …
With automobiles becoming increasingly reliant on sensors to perform various driving tasks, it is important to encode the relevant CAN bus sensor data in a way that captures the general state of the vehicle in a compact form. In this paper, we develop a deep learning-based method, called Drive2Vec, for embedding such s…
LLMs can memorize economic data and recall exact values before their training cutoff.
Conventional prior for Variational Auto-Encoder (VAE) is a Gaussian distribution. Recent works demonstrated that choice of prior distribution affects learning capacity of VAE models. We propose a general technique (embedding-reparameterization procedure, or ER) for introducing arbitrary manifold-valued variables in VAE…
Paper establishes identifiability conditions for a model with two latent vectors and auxiliary data.
M2VN forecasts financial volatility by fusing time series data with news embeddings.