We present the first provably sublinear time algorithm for approximate \emph{Maximum Inner Product Search} (MIPS). Our proposal is also the first hashing algorithm for searching with (un-normalized) inner product as the underlying similarity measure. Finding hashing schemes for MIPS was considered hard. We formally sho…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We propose a quantization based approach for fast approximate Maximum Inner Product Search (MIPS). Each database vector is quantized in multiple subspaces via a set of codebooks, learned directly by minimizing the inner product quantization error. Then, the inner product of a query to a database vector is approximated …
New vector quantization method reduces relevance of parallel components in database points.
This paper addresses the nearest neighbor search problem under inner product similarity and introduces a compact code-based approach. The idea is to approximate a vector using the composition of several elements selected from a source dictionary and to represent this vector by a short code composed of the indices of th…
Inverted file and asymmetric distance computation (IVFADC) have been successfully applied to approximate nearest neighbor search and subsequently maximum inner product search. In such a framework, vector quantization is used for coarse partitioning while product quantization is used for quantizing residuals. In the ori…
Recently it was shown that the problem of Maximum Inner Product Search (MIPS) is efficient and it admits provably sub-linear hashing algorithms. Asymmetric transformations before hashing were the key in solving MIPS which was otherwise hard. In the prior work, the authors use asymmetric transformations which convert th…
Sublinear LSVI via LSH reduces runtime to sublinear in actions.
Efficient Maximum Inner Product Search (MIPS) is an important task that has a wide applicability in recommendation systems and classification with a large number of classes. Solutions based on locality-sensitive hashing (LSH) as well as tree-based solutions have been investigated in the recent literature, to perform ap…
Amortizes MIPS by training neural networks to predict optimal keys.
There has been substantial research on sub-linear time approximate algorithms for Maximum Inner Product Search (MIPS). To achieve fast query time, state-of-the-art techniques require significant preprocessing, which can be a burden when the number of subsequent queries is not sufficiently large to amortize the cost. Fu…
We consider the problem of designing locality sensitive hashes (LSH) for inner product similarity, and of the power of asymmetric hashes in this context. Shrivastava and Li argue that there is no symmetric LSH for the problem and propose an asymmetric LSH based on different mappings for query and database points. Howev…
Scalable model for slate recommendation learns reward probabilities.
Neyshabur and Srebro proposed Simple-LSH, which is the state-of-the-art hashing method for maximum inner product search (MIPS) with performance guarantee. We found that the performance of Simple-LSH, in both theory and practice, suffers from long tails in the 2-norm distribution of real datasets. We propose Norm-rangin…
Many emerging use cases of data mining and machine learning operate on large datasets with data from heterogeneous sources, specifically with both sparse and dense components. For example, dense deep neural network embedding vectors are often used in conjunction with sparse textual features to provide high dimensional …
In this paper, we propose a stochastic optimization method that adaptively controls the sample size used in the computation of gradient approximations. Unlike other variance reduction techniques that either require additional storage or the regular computation of full gradients, the proposed method reduces variance by …
Study higher rank inner products and their tilings to describe tori degenerations.
Paper speeds up policy optimization for large recommendation systems.
Researchers prove inner product recovery is impossible in latent space models.
Paper proposes a new method to optimize feature coordinates for better image classification.
Study of Gaussian distributions using entropic Gromov-Wasserstein and inner product Gromov-Wasserstein.
Inference in log-linear models scales linearly in the size of output space in the worst-case. This is often a bottleneck in natural language processing and computer vision tasks when the output space is feasibly enumerable but very large. We propose a method to perform inference in log-linear models with sublinear amor…
Study on kernel regression risk in high dimensions using Pinsker bound.
We point out that the Homfly polynomial (that is to say, Ocneanu's trace functional) contains two polynomial-valued inner products on the Hecke algebra representation of Artin's braid group. These bear a close connection to the Morton-Franks-Williams inequality. In these structures, the sets of positive, respectively n…
We study one extremal problem on the product of power of generalized inner radii of non-overlapping domains in .
Convex learning for diverse invariances in semi-inner-product space.
Bayesian optimization with Gaussian processes speeds up searches for stationary points.
Vectors of data are at the heart of machine learning and data mining. Recently, vector quantization methods have shown great promise in reducing both the time and space costs of operating on vectors. We introduce a vector quantization algorithm that can compress vectors over 12x faster than existing techniques while al…
We study estimation of (semi-)inner products between two nonparametric probability distributions, given IID samples from each distribution. These products include relatively well-studied classical and Sobolev inner products, as well as those induced by translation-invariant reproducing kernels, for whic…
Estimates latent inner products from an anisotropic Gaussian graph with improved spectral method.
Being E a vector space with inner product and S the sphere of E, will be given a demonstration that every application of the sphere S itself it such that preserve inner product is the restriction of a linear isometry in E.
This study approximates neural network features for modeling relations and attention mechanisms.
The Bergman kernels of holomorphic vector bundles are studied to extend the Fubini-Study map.
A new method for finding efficient neural interaction functions in collaborative filtering.
Classifies metrics on specific Lie groups.
We propose a robust elastic net (REN) model for high-dimensional sparse regression and give its performance guarantees (both the statistical error bound and the optimization bound). A simple idea of trimming the inner product is applied to the elastic net model. Specifically, we robustify the covariance matrix by trimm…
We show that every unimodular Lie algebra, of dimension at most 4, equipped with an inner product, possesses an orthonormal basis comprised of geodesic elements. On the other hand, we give an example of a solvable unimodular Lie algebra of dimension 5 that has no orthonormal geodesic basis, for any inner product.
Signature Isolation Forest removes constraints from FIF by using rough path theory's signature transform.
We describe a construction of fibrewise inner products on the cotangent bundle of the smooth free loop space of a Riemannian manifold. Using this inner product, we construct an operator over the loop space of a string manifold which is directly analogous to the Dirac operator of a spin manifold.
IENs reduce neural network variance without increasing complexity.
We propose (WIPS) for neural network-based graph embedding. In addition to the parameters of neural networks, we optimize the weights of the inner product by allowing positive and negative values. Despite its simplicity, WIPS can approximate arbitrary general similarities in…
Equivalent tests for SGD batch size selection found.
We give a condition for a function to produce a Möbius invariant weighted inner product on the tangent space of the space of knots, and show that some kind of Möbius invariant knot energies can produce Möbius invariant and parametrization invariant weighted inner products. They would give a natural way to study the evo…
The paper classifies hypersurfaces in Riemannian manifolds with constant inner product and torse-forming axes.
We formulate a quantization commutes with reduction principle in the setting where the Lie group , the symplectic manifold it acts on, and the orbit space of the action may all be noncompact. It is assumed that the action is proper, and the zero set of a deformation vector field, associated to the momentum map and a…
Model-based methods for recommender systems have been studied extensively in recent years. In systems with large corpus, however, the calculation cost for the learnt model to predict all user-item preferences is tremendous, which makes full corpus retrieval extremely difficult. To overcome the calculation barriers, mod…
H-type Lie algebras were introduced by Kaplan as a class of real Lie algebras generalizing the familiar Heisenberg Lie algebra . The H-type property depends on a choice of inner product on the Lie algebra . Among the H-type Lie algebras are the complex Heisenberg Lie algebras $\mathfrak{h}…
In this paper we develop several algebraic structures on the simplicial cochains of a triangulated manifold that are analogues of objects in differential geometry. We study a cochain product and prove several statements about its convergence to the wedge product on differential forms. Also, for cochains with an inner p…
Minwise hashing (Minhash) is a widely popular indexing scheme in practice. Minhash is designed for estimating set resemblance and is known to be suboptimal in many applications where the desired measure is set overlap (i.e., inner product between binary vectors) or set containment. Minhash has inherent bias towards sma…