Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

20416181 · Oct 201919922001200920172026
48 results for word analogy

Recent work has demonstrated that embeddings of tree-like graphs in hyperbolic space surpass their Euclidean counterparts in performance by a large margin. Inspired by these results and scale-free structure in the word co-occurrence graph, we present an algorithm for learning word embeddings in hyperbolic space from fr…

2018-08-30abs ↗pdf ↗

Word embeddings generated by neural network methods such as word2vec (W2V) are well known to exhibit seemingly linear behaviour, e.g. the embeddings of analogy "woman is to queen as man is to king" approximately describe a parallelogram. This property is particularly intriguing since the embeddings are not trained to a…

2019-01-28abs ↗pdf ↗

LLMs learn new tasks from unstructured data, but it depends on word co-occurrence and positional information.

problem Understanding how LLMs can learn new tasks from unstructured data without explicit training.
method Examined the capabilities of LLMs trained on unstructured data, focusing on sequence model requirements and training data structure.
result Many ICL capabilities can emerge from word co-occurrence in unstructured data, but positional information is crucial for certain tasks.

We propose a new method for learning word representations using hierarchical regularization in sparse coding inspired by the linguistic study of word meanings. We show an efficient learning algorithm based on stochastic proximal methods that is significantly faster than previous approaches, making it possible to perfor…

2014-06-08abs ↗pdf ↗

Latent topic models have been successfully applied as an unsupervised topic discovery technique in large document collections. With the proliferation of hypertext document collection such as the Internet, there has also been great interest in extending these approaches to hypertext [6, 9]. These approaches typically mo…

2012-06-13abs ↗pdf ↗

Bidirectional attention is shown to be equivalent to a continuous bag of words model with mixture-of-experts.

problem Understanding the statistical underpinnings of bidirectional attention.
method Exploring bidirectional attention as a mixture-of-experts model and reparameterizing it.
result Bidirectional attention can be viewed as a continuous bag of words model with mixture-of-experts weights.

Semantic word embeddings represent the meaning of a word via a vector, and are created by diverse methods. Many use nonlinear operations on co-occurrence statistics, and have hand-tuned hyperparameters and reweighting methods. This paper proposes a new generative model, a dynamic version of the log-linear topic model o…

2015-02-12abs ↗pdf ↗

We describe a method for learning word embeddings with data-dependent dimensionality. Our Stochastic Dimensionality Skip-Gram (SD-SG) and Stochastic Dimensionality Continuous Bag-of-Words (SD-CBOW) are nonparametric analogs of Mikolov et al.'s (2013) well-known 'word2vec' models. Vector dimensionality is made dynamic b…

2015-11-17abs ↗pdf ↗

Word embeddings learnt from large corpora have been adopted in various applications in natural language processing and served as the general input representations to learning systems. Recently, a series of post-processing methods have been proposed to boost the performance of word embeddings on similarity comparison an…

2019-05-27abs ↗pdf ↗

Embedding words in a vector space has gained a lot of attention in recent years. While state-of-the-art methods provide efficient computation of word similarities via a low-dimensional matrix embedding, their motivation is often left unclear. In this paper, we argue that word embedding can be naturally viewed as a rank…

2015-06-09abs ↗pdf ↗

Recent progress in applying machine learning for jet physics has been built upon an analogy between calorimeters and images. In this work, we present a novel class of recursive neural networks built instead upon an analogy between QCD and natural languages. In the analogy, four-momenta are like words and the clustering…

2017-02-02abs ↗pdf ↗

Continuous vector representations of words and objects appear to carry surprisingly rich semantic content. In this paper, we advance both the conceptual and theoretical understanding of word embeddings in three ways. First, we ground embeddings in semantic spaces studied in cognitive-psychometric literature and introdu…

2015-09-18abs ↗pdf ↗

Anosov representations of word hyperbolic groups into higher-rank semisimple Lie groups are representations with finite kernel and discrete image that have strong analogies with convex cocompact representations into rank-one Lie groups. However, the most naive analogy fails: generically, Anosov representations do not a…

2017-01-31abs ↗pdf ↗

The paper finds formulas for word lengths and conjugacy classes in surface groups.

problem Finding formulas for word lengths and conjugacy classes in surface groups.
method Investigating symmetric presentations and normal forms of conjugacy classes.
result Derives three formulae for word lengths and provides efficient algorithms for conjugacy problems.

Factor complexity bφ(n)b_φ(n) for a vertex coloring φφ of a regular tree is the number of colored nn-balls up to color-preserving automorphisms. Sturmian colorings are colorings of minimal unbounded factor complexity bφ(n)=n+2b_φ(n) = n+2. In this article, we prove an induction algorithm for Sturmian colorings using colored ba…

2016-09-20abs ↗pdf ↗

We develop a family of techniques to align word embeddings which are derived from different source datasets or created using different mechanisms (e.g., GloVe or word2vec). Our methods are simple and have a closed form to optimally rotate, translate, and scale to minimize root mean squared errors or maximize the averag…

2018-06-04abs ↗pdf ↗

In recent years, the Word2Vec model trained with the Negative Sampling loss function has shown state-of-the-art results in a number of machine learning tasks, including language modeling tasks, such as word analogy and word similarity, and in recommendation tasks, through Prod2Vec, an extension that applies to modeling…

2018-05-22abs ↗pdf ↗

Machine learning algorithms are optimized to model statistical properties of the training data. If the input data reflects stereotypes and biases of the broader society, then the output of the learning algorithm also captures these stereotypes. In this paper, we initiate the study of gender stereotypes in {\em word emb…

2016-06-20abs ↗pdf ↗

We prove that the exponential growth rate of the regular language of penetration sequences is smaller than the growth rate of the regular language of normal form words, if the acceptor of the regular language of normal form words is strongly connected. Moreover, we show that the latter property is satisfied for all irr…

2014-03-11abs ↗pdf ↗

In this paper we introduce flat grafting as a deformation of quadratic differentials on a surface of finite type that is analogous to the grafting map on hyperbolic surfaces. Flat grafting maps are generic in the strata structure and preserve parallel measured foliations. We use flat grafting to construct paths connect…

2018-03-27abs ↗pdf ↗

The aim of this paper is to present a short introduction to supergeometry on pure odd supermanifolds. (Pseudo)differential forms, Cartan calculus (DeRham differential, Lie derivative, "inner" product), metric, inner product, Killing's vector fields, Hodge star operator, integral forms, co-differential and connection on…

2003-09-23abs ↗pdf ↗

Let ww be a word in the free group on rr generators. The expected value of the trace of the word in rr independent Haar elements of O(n)\mathrm{O}(n) gives a function TrwO(n){\cal T}r_{w}^{\mathrm{O}}(n) of nn. We show that TrwO(n){\cal T}r_{w}^{\mathrm{O}}(n) has a convergent Laurent expansion at n=n=\infty involving maps on …

2019-04-30abs ↗pdf ↗

Analyzes word2vec-like models revealing linear subspaces learned during training.

problem Understanding representation learning in word embeddings.
method Analytical solution of word2vec loss dynamics and final embeddings.
result Models learn orthogonal linear subspaces incrementally, representing interpretable concepts.

Given a function f ⁣:XYf\colon X\to Y of metric spaces, its {\it asymptotic dimension} $\asdim(f)$ is the supremum of $\asdim(A)$ such that AXA\subset X and $\asdim(f(A))=0$. Our main result is \begin{Thm} \label{ThmAInAbstract} $\asdim(X)\leq \asdim(f)+\asdim(Y)$ for any large scale uniform function f ⁣:XYf\colon X\to Y. \end…

2006-05-16abs ↗pdf ↗

Separable Non-negative Matrix Factorization (SNMF) is an important method for topic modeling, where "separable" assumes every topic contains at least one anchor word, defined as a word that has non-zero probability only on that topic. SNMF focuses on the word co-occurrence patterns to reveal topics by two steps: anchor…

2019-05-10abs ↗pdf ↗

This paper continues a geometric study of Harvey's Complex of Curves, whose ultimate goal is to apply the theory of hyperbolic spaces and groups to algorithmic questions for the Mapping Class Group and geometric properties of Kleinian representations. The authors' previous result that the complex is delta-hyperbolic wa…

1998-07-27abs ↗pdf ↗

Most existing word embedding approaches do not distinguish the same words in different contexts, therefore ignoring their contextual meanings. As a result, the learned embeddings of these words are usually a mixture of multiple meanings. In this paper, we acknowledge multiple identities of the same word in different co…

2016-11-29abs ↗pdf ↗

There are certain families of words and word sequences (words in the generators of a two-generator group) that arise frequently in the Teichm{ü}ller theory of hyperbolic three-manifolds and Kleinian and Fuchsian groups and in the discreteness problem for two generator matrix groups. We survey some of the families of su…

2007-01-20abs ↗pdf ↗

Urban2Vec combines street view imagery and POIs for better urban neighborhood embeddings.

problem Lack of comprehensive representation of urban neighborhoods using heterogeneous data.
method Unsupervised multi-modal framework using CNN for visual features and bag-of-words for POI data.
result Urban2Vec achieves better performance than baseline models and comparable to fully-supervised methods.

Study on representations of four-punctured sphere group in hyperbolic spaces.

problem Understanding representations of the four-punctured sphere group.
method Investigation into simple-stable and Bowditch representations in Gromov-hyperbolic spaces.
result Simple-stable representations and Bowditch representations are equivalent.

We discuss a topological approach to words introduced by the author. Words on an arbitrary alphabet are approximated by Gauss words and then studied up to natural modifications inspired by the Reidemeister moves on knot diagrams. This leads us to a notion of homotopy for words. We introduce several homotopy invariants …

2006-09-19abs ↗pdf ↗

If GG is a finite group, is a function f:GCf:G\to\mathbb C determined by its sums over all cosets of cyclic subgroups of GG? In other words, is the Radon transform on GG injective? This inverse problem is a discrete analogue of asking whether a function on a compact Lie group is determined by its integrals over all ge…

2014-11-14abs ↗pdf ↗

Continuous word representation (aka word embedding) is a basic building block in many neural network-based models used in natural language processing tasks. Although it is widely accepted that words with similar semantics should be close to each other in the embedding space, we find that word embeddings learned in seve…

2018-09-18abs ↗pdf ↗

The abstract explains how word and relation representations capture semantic meaning.

problem Understanding how word and relation representations capture semantic meaning.
method Theoretical justification and extension of geometric relationships between word embeddings and knowledge graph representations.
result The geometric relationships between word embeddings correspond to semantic relations between words and entities in knowledge graphs.

The study introduces Cayley--Abels--Rosendal graphs for Polish groups.

problem Understanding the structure of Polish groups through graph theory.
method Developing Cayley--Abels--Rosendal graphs and applying them to Polish groups.
result Groups with Cayley--Abels--Rosendal graphs are topological analogues of finitely generated groups.