Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

0111 · Jun 200419922001200920172026
20 results for SGN

SGNs use Hamiltonian mechanics for invertible deep generative modeling.

problem Efficient and exact likelihood evaluation for deep generative models.
method Symplectic structure in latent space, Hamiltonian dynamics for data generation.
result Exact likelihood evaluation without Jacobian calculations.

We revisit skip-gram negative sampling (SGNS), one of the most popular neural-network based approaches to learning distributed word representation. We first point out the ambiguity issue undermining the SGNS model, in the sense that the word vectors can be entirely distorted without changing the objective value. To res…

2018-04-01abs ↗pdf ↗

We study the following 11-Yamabe equation on a connected finite graph Δ1u+gSgn(u)=huα1Sgn(u),Δ_1u+g\mathrm{Sgn}(u)=h|u|^{α-1}\mathrm{Sgn}(u), where Δ1Δ_1 is the discrete 11-Laplacian, α>1α>1 and g,h>0g, h>0 are known. We show that the above 11-Yamabe equation always has a nontrivial solution u0u\geq0, u0u\neq0.

2017-09-28abs ↗pdf ↗

We simplify word embeddings by removing sigmoid in SGNS, revealing connections to hyperbolic spaces.

problem Improving word embeddings quality and understanding their relationship with hyperbolic spaces.
method Analyzing squashed shifted PMI matrix and its relation to graph properties and hyperbolic geometry.
result Word embeddings can be connected to hyperbolic spaces through squashed shifted PMI matrix.

The Global Vectors for word representation (GloVe), introduced by Jeffrey Pennington et al. is reported to be an efficient and effective method for learning vector representations of words. State-of-the-art performance is also provided by skip-gram with negative-sampling (SGNS) implemented in the word2vec tool. In this…

2014-11-20abs ↗pdf ↗

New method improves generalization in deep learning models.

problem Improving generalization in overparameterized deep neural networks.
method Stochastic Gauss-Newton method with Levenberg-Marquardt damping and mini-batch sampling.
result Established finite-time convergence and non-asymptotic generalization bounds.

This work shows dimension regularization can replace skip-gram negative sampling for graph embeddings, improving efficiency and performance.

problem Efficiently enforcing dissimilarity among node embeddings in graph learning.
method Dimension regularization as an alternative to skip-gram negative sampling.
result Dimension regularization is a more efficient approach to enforcing dissimilarity in graph embeddings.

New perspective on SGD reveals short-range memory effects in deep learning.

problem Understanding the efficacy of stochastic gradient descent (SGD) in deep learning.
method Proposed that SGD is a discretization of an SDE driven by fractional Brownian motion (FBM).
result SGD stays longer in flat minima, favoring generalization.

In this note, we investigate the well-known Yau rigidity theorem for minimal submanifolds in spheres. Using the parameter method of Yau and the DDVV inequality verified by Lu, Ge and Tang, we prove that if MM is an nn-dimensional oriented compact minimal submanifold in the unit sphere Sn+p(1)S^{n+p}(1), and if $K_{M}\geq\…

2011-02-28abs ↗pdf ↗

Let MM be an n(4)n(\geq 4)-dimensional compact submanifold in the simply connected space form Fn+p(c)F^{n+p}(c) with constant curvature c0c\geq 0, where HH is the mean curvature of MM. We verify that if the scalar curvature of MM satisfies R>n(n2)(c+H2)R>n(n-2)(c+H^2), and if RicM(n22σn2nσn)(c+H2)Ric_M\geq (n-2-\frac{2σ_n}{2n-σ_n})(c+H^2), then MM is…

2019-03-01abs ↗pdf ↗

Unified framework for word embedding models using noise examples.

problem Improving word embedding models with negative sampling.
method Formulated a Word-Context Classification (WCC) framework that generalizes SkipGram word embedding models.
result The best noise distribution is the data distribution, improving both performance and training speed.

What enables Stochastic Gradient Descent (SGD) to achieve better generalization than Gradient Descent (GD) in Neural Network training? This question has attracted much attention. In this paper, we study the distribution of the Stochastic Gradient Noise (SGN) vectors during the training. We observe that for batch sizes …

2019-10-21abs ↗pdf ↗

Note on subgaussian bounds for sign-quantized linear maps.

problem Understanding subgaussian behavior of sign-quantized linear maps.
method Developed a dimension-independent subgaussian concentration bound for Gaussian vectors under nonlinear mappings.
result Answered a question about sign-quantized linear maps using a new subgaussian bound.

This paper takes a step towards theoretical analysis of the relationship between word embeddings and context embeddings in models such as word2vec. We start from basic probabilistic assumptions on the nature of word vectors, context vectors, and text generation. These assumptions are well supported either empirically o…

2019-02-26abs ↗pdf ↗

We present a comprehensive study of utility function of the minority game in its efficient regime. We develop an effective description of state of the game. For the payoff function $g(x)=\sgn (x)$ we explicitly represent the game as the Markov process and prove the finitness of number of states. We also demonstrate bou…

2009-07-18abs ↗pdf ↗

Previous studies indicate that nonlinear properties of Gaussian time series with long-range correlations, uiu_i, can be detected and quantified by studying the correlations in the magnitude series ui|u_i|, i.e., the ``volatility''. However, the origin for this empirical observation still remains unclear, and the exact …

2004-06-14abs ↗pdf ↗

Word embeddings, i.e., low-dimensional vector representations such as GloVe and SGNS, encode word "meaning" in the sense that distances between words' vectors correspond to their semantic proximity. This enables transfer learning of semantics for a variety of natural language processing tasks. Word embeddings are typic…

2020-01-14abs ↗pdf ↗