Classifies complex hyperbolic triangle groups by types.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
BBM models short texts using biterms to improve coherence.
A deep learning approach generates math word problems in multiple languages.
Unified framework for word embedding models using noise examples.
While the celebrated Word2Vec technique yields semantically rich representations for individual words, there has been relatively less success in extending to generate unsupervised sentences or documents embeddings. Recent work has demonstrated that a distance measure between documents called \emph{Word Mover's Distance…
The article introduces a new set of Polish word embeddings, built using KGR10 corpus, which contains more than 4 billion words. These embeddings are evaluated in the problem of recognition of temporal expressions (timexes) for the Polish language. We described the process of KGR10 corpus creation and a new approach to …
We simplify word embeddings by removing sigmoid in SGNS, revealing connections to hyperbolic spaces.
As the emergence and the thriving development of social networks, a huge number of short texts are accumulated and need to be processed. Inferring latent topics of collected short texts is useful for understanding its hidden structure and predicting new contents. Unlike conventional topic models such as latent Dirichle…
Proves involutions on Right-angled Coxeter groups without fixed points.
Given any generating set of any pseudo-Anosov-containing subgroup of the mapping class group of a surface, we construct a pseudo-Anosov with word length bounded by a constant depending only on the surface. More generally, in any subgroup G we find an element f with the property that the minimal subsurface supporting a …
Interventional cancer clinical trials are generally too restrictive, and some patients are often excluded on the basis of comorbidity, past or concomitant treatments, or the fact that they are over a certain age. The efficacy and safety of new treatments for patients with these characteristics are, therefore, not defin…
We investigate the fundamental group of Griffiths' space, and the first singular homology group of this space and of the Hawaiian Earring by using (countable) reduced tame words. We prove that two such words represent the same element in the corresponding group if and only if they can be carried to the same tame word b…
Intense volatility in financial markets affect humans worldwide. Therefore, relatively accurate prediction of volatility is critical. We suggest that massive data sources resulting from human interaction with the Internet may offer a new perspective on the behavior of market participants in periods of large market move…
In text mining, information retrieval, and machine learning, text documents are commonly represented through variants of sparse Bag of Words (sBoW) vectors (e.g. TF-IDF). Although simple and intuitive, sBoW style representations suffer from their inherent over-sparsity and fail to capture word-level synonymy and polyse…
Due to recent technical and scientific advances, we have a wealth of information hidden in unstructured text data such as offline/online narratives, research articles, and clinical reports. To mine these data properly, attributable to their innate ambiguity, a Word Sense Disambiguation (WSD) algorithm can avoid numbers…
In this paper we use theory of embedded graphs on oriented and compact -surfaces to construct minimal realizations of signed Gauss paragraphs. We prove that the genus of the ambient surface of these minimal realizations can be seen as a function of the maximum number of Carter's circles. For the case of signed Gaus…
One of the most interesting questions about a group is if its word problem can be solved and how. The word problem in the braid group is of particular interest to topologists, algebraists and geometers, and is the target of intensive current research. We look at the braid group from a topological point of view (rather …
This paper improves topic modeling by embedding words and topics together.
We use a first-order energy quantity to prove a strengthened statement of uniqueness for the Ricci flow. One consequence of this statement is that if a complete solution on a noncompact manifold has uniformly bounded Ricci curvature, then its sectional curvature will remain bounded for a short time if it is bounded ini…
A new neural topic model using optimal transport improves document representation and topic coherence.
We present a novel method for hierarchical topic detection where topics are obtained by clustering documents in multiple ways. Specifically, we model document collections using a class of graphical models called hierarchical latent tree models (HLTMs). The variables at the bottom level of an HLTM are observed binary va…
System identifies language of transliterated text.
WEEND uses a neural network to recognize speech and assign speakers to words.
Enhanced word embedding creates new consumer-friendly health terms.
The paper revisits the investment simulation based on strategies exhibited by Generalized (m,2)-Zipf law to present an interesting characterization of the wildness in financial time series. The investigations of dominant strategies on each specific time series shows that longer words dominant in larger time scale exhib…
Using Quillen's superconnection formalism we give a new "twisted" approach to the rational Gromov-Lawson-Rosenberg (GLR) conjecture on topological obstructions to the existence of Riemannian metrics of positive scalar curvature on compact spin manifolds. In particular, we present a short proof of the rational GLR conje…
We describe a simple neural language model that relies only on character-level inputs. Predictions are still made at the word-level. Our model employs a convolutional neural network (CNN) and a highway network over characters, whose output is given to a long short-term memory (LSTM) recurrent neural network language mo…
An important part of the information gathering and data analysis is to find out what people think about, either a product or an entity. Twitter is an opinion rich social networking site. The posts or tweets from this data can be used for mining people's opinions. The recent surge of activity in this area can be attribu…
We apply text analysis approaches for a specialized search engine for 3D CAD models and associated products. The main goals are to distinguish between actual product descriptions and other text on a website, as well as to decide whether a given text is or contains a product name. For this we use paragraph vectors for t…
This thesis focuses on gaining linguistic insights into textual discussions on a word level. It was of special interest to distinguish messages that constructively contribute to a discussion from those that are detrimental to them. Thereby, we wanted to determine whether "I"- and "You"-messages are indicators for eithe…
We introduce a new approach for topic modeling that is supervised by survival analysis. Specifically, we build on recent work on unsupervised topic modeling with so-called anchor words by providing supervision through an elastic-net regularized Cox proportional hazards model. In short, an anchor word being present in a…
Monitoring the biomedical literature for cases of Adverse Drug Reactions (ADRs) is a critically important and time consuming task in pharmacovigilance. The development of computer assisted approaches to aid this process in different forms has been the subject of many recent works. One particular area that has shown pro…
Contractible Vietoris-Rips complexes for integer n proved using discrete Morse theory.
SoulMate links short texts through multi-aspect embeddings.
Model clusters authors and topics in short texts like social media posts.
Neural machine translation is a relatively new approach to statistical machine translation based purely on neural networks. The neural machine translation models often consist of an encoder and a decoder. The encoder extracts a fixed-length representation from a variable-length input sentence, and the decoder generates…
The Generative Adversarial Network (GAN) has recently been applied to generate synthetic images from text. Despite significant advances, most current state-of-the-art algorithms are regular-grid region based; when attention is used, it is mainly applied between individual regular-grid regions and a word. These approach…
This short note gives a geometric interpretation of the Atiyah class of a Lie pair. It proves that it vanishes if the subalgebroid is the kernel of a fibration of Lie algebroids. In other words, the Atiyah class of a Lie pair vanishes if the subalgebroid is the fiber of an ideal system in the Lie algebroid. In order to…
In this paper we consider the problem of clustering collections of very short texts using subspace clustering. This problem arises in many applications such as product categorisation, fraud detection, and sentiment analysis. The main challenge lies in the fact that the vectorial representation of short texts is both hi…
This paper addresses the robust speech recognition problem as an adaptation task. Specifically, we investigate the cumulative application of adaptation methods. A bidirectional Long Short-Term Memory (BLSTM) based neural network, capable of learning temporal relationships and translation invariant representations, is u…
Recurrent neural network (RNN) language models (LMs) and Long Short Term Memory (LSTM) LMs, a variant of RNN LMs, have been shown to outperform traditional N-gram LMs on speech recognition tasks. However, these models are computationally more expensive than N-gram LMs for decoding, and thus, challenging to integrate in…
We introduce multiplicative LSTM (mLSTM), a recurrent neural network architecture for sequence modelling that combines the long short-term memory (LSTM) and multiplicative recurrent neural network architectures. mLSTM is characterised by its ability to have different recurrent transition functions for each possible inp…
Although deep learning models have proven effective at solving problems in natural language processing, the mechanism by which they come to their conclusions is often unclear. As a result, these models are generally treated as black boxes, yielding no insight of the underlying learned patterns. In this paper we conside…
Random groups prove length constraints on product of conjugates.
Study examines how cluster number affects short-text clustering, introducing a stability metric.
Motivated by the need to automate medical information extraction from free-text radiological reports, we present a bi-directional long short-term memory (BiLSTM) neural network architecture for modelling radiological language. The model has been used to address two NLP tasks: medical named-entity recognition (NER) and …
Survey of geometry developments, including complex structures on surfaces.
In this short note, we prove that a bi-invariant Riemannian metric on is uniquely determined by the spectrum of its Laplace-Beltrami operator within the class of left-invariant metrics on . In other words, on any of these compact simple Lie groups, every left-invariant metric which is n…