We simplify word embeddings by removing sigmoid in SGNS, revealing connections to hyperbolic spaces.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper, we propose a new kernel-based co-occurrence measure that can be applied to sparse linguistic expressions (e.g., sentences) with a very short learning time, as an alternative to pointwise mutual information (PMI). As well as deriving PMI from mutual information, we derive this new measure from the Hilbert…
Considered an important macroeconomic indicator, the Purchasing Managers' Index (PMI) on Manufacturing generally assumes that PMI announcements will produce an impact on stock markets. International experience suggests that stock markets react to negative PMI news. In this research, we empirically investigate the stock…
PMI-Masking improves MLM pretraining by masking correlated spans efficiently.
Enhances CLIP's similarity computation using PMI's linear structure.
This paper takes a step towards theoretical analysis of the relationship between word embeddings and context embeddings in models such as word2vec. We start from basic probabilistic assumptions on the nature of word vectors, context vectors, and text generation. These assumptions are well supported either empirically o…
Semantic word embeddings represent the meaning of a word via a vector, and are created by diverse methods. Many use nonlinear operations on co-occurrence statistics, and have hand-tuned hyperparameters and reweighting methods. This paper proposes a new generative model, a dynamic version of the log-linear topic model o…
Attention improves edge prediction in e-commerce graphs.
Continuous representation of words is a standard component in deep learning-based NLP models. However, representing a large vocabulary requires significant memory, which can cause problems, particularly on resource-constrained platforms. Therefore, in this paper we propose an isotropic iterative quantization (IIQ) appr…
Word embedding is a powerful tool in natural language processing. In this paper we consider the problem of word embedding composition \--- given vector representations of two words, compute a vector for the entire phrase. We give a generative model that can capture specific syntactic relations between words. Under our …
Word2Vec (W2V) and GloVe are popular, fast and efficient word embedding algorithms. Their embeddings are widely used and perform well on a variety of natural language processing tasks. Moreover, W2V has recently been adopted in the field of graph embedding, where it underpins several leading algorithms. However, despit…
Variance-Calibrated Modulation (VCM) addresses the likelihood trap in LLMs by reshaping the probability distribution before truncation.
This paper proposes CSADA to make DNNs cost-sensitive.
Proposes a framework to quantify uncertainty in multi-step decision-making by LLMs.
DeepTMR reorders matrices without prior knowledge of structural patterns.
The paper constructs Goeritz matrices from Dehn colorings.
New matrix reveals cluster info in sparse directed graphs.
The CN matrix of a pure braid projection is characterized and applied.
Generalised matrix-matrix multiplication forms the kernel of many mathematical algorithms. A faster matrix-matrix multiply immediately benefits these algorithms. In this paper we implement efficient matrix multiplication for large matrices using the floating point Intel Pentium SIMD (Single Instruction Multiple Data) a…
Characterizes the OU matrix for up to 5 strands in braids.
Classifies SL(n) covariant matrix-valued valuations on Lp-spaces.
Unified approach for robust low rank matrix estimation with adversaries.
Most recent results in matrix completion assume that the matrix under consideration is low-rank or that the columns are in a union of low-rank subspaces. In real-world settings, however, the linear structure underlying these models is distorted by a (typically unknown) nonlinear transformation. This paper addresses the…
A new algorithm speeds up matrix operations in Neural Networks.
New NMF algorithm uses Toeplitz matrix for facial recognition.
Recommender systems are widely used to recommend the most appealing items to users. These recommendations can be generated by applying collaborative filtering methods. The low-rank matrix completion method is the state-of-the-art collaborative filtering method. In this work, we show that the skewed distribution of rati…
The paper defines the OU matrix for braid diagrams and finds determinant relationships.
New method improves robust low-rank matrix completion for computer vision.
Matrix completion is a modern missing data problem where both the missing structure and the underlying parameter are high dimensional. Although missing structure is a key component to any missing data problems, existing matrix completion methods often assume a simple uniform missing mechanism. In this work, we study ma…
Matrix approximation is a common tool in machine learning for building accurate prediction models for recommendation systems, text mining, and computer vision. A prevalent assumption in constructing matrix approximations is that the partially observed matrix is of low-rank. We propose a new matrix approximation model w…
Unified framework for nonconvex matrix completion with linearly parameterized factors.
The problem of low rank matrix completion is considered in this paper. To exploit the underlying low-rank structure of the data matrix, we propose a hierarchical Gaussian prior model, where columns of the low-rank matrix are assumed to follow a Gaussian distribution with zero mean and a common precision matrix, and a W…
Matrix completion is a problem that arises in many data-analysis settings where the input consists of a partially-observed matrix (e.g., recommender systems, traffic matrix analysis etc.). Classical approaches to matrix completion assume that the input partially-observed matrix is low rank. The success of these methods…
Proposes a robust factor analysis for matrix data.
3-manifold triangulation can be reconstructed from its intersection matrix.
We give the first algorithm for Matrix Completion whose running time and sample complexity is polynomial in the rank of the unknown target matrix, linear in the dimension of the matrix, and logarithmic in the condition number of the matrix. To the best of our knowledge, all previous algorithms either incurred a quadrat…
Paper presents a new framework for covariance matrix estimation with geometric insights.
Study improves fractional posterior for 1-bit matrix completion.
In this paper, we propose an online algorithm to compute matrix factorizations. Proposed algorithm updates the dictionary matrix and associated coefficients using a single observation at each time. The algorithm performs low-rank updates to dictionary matrix. We derive the algorithm by defining a simple objective funct…
Incorporates matrix exponential into generative flows for improved performance.
Consider a movie recommendation system where apart from the ratings information, side information such as user's age or movie's genre is also available. Unlike standard matrix completion, in this setting one should be able to predict inductively on new users/movies. In this paper, we study the problem of inductive matr…
Matrix SMD converges to unique solution minimizing Bregman divergence.
Paper derives matrix formulae and proves skein relations for non-orientable surfaces in quasi-cluster algebras.
Gradient descent proves global convergence for 4-layer matrix factorization.
Proposes a new matrix factorization model for interval-valued matrices.
New method clusters matrix-variate data with outliers.
New method for hyperparameter tuning in sparse matrix factorization.
The warping matrix has been defined for knot projections and knot diagrams by using warping degrees. In particular, the warping matrix of a knot diagram represents the knot diagram uniquely. In this paper we show that the rank of the warping matrix is one greater than the crossing number. We also discuss the linearly i…