A graph model improves short text classification by integrating sentence relationships.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper introduces new indicators for forecasting crude oil prices using short news headlines.
AOBTM adapts online topic modeling for short app reviews, revealing coherent topics over time.
POTA improves short text clustering by generating reliable pseudo-labels.
Recent approaches based on artificial neural networks (ANNs) have shown promising results for short-text classification. However, many short texts occur in sequences (e.g., sentences in a document or utterances in a dialog), and most existing ANN-based systems do not leverage the preceding short texts when classifying …
Paper introduces a new text clustering model using Beta-Liouville priors.
BBM models short texts using biterms to improve coherence.
New Gamma-Poisson model improves topic selection for short text.
Linking authors of short-text contents has important usages in many applications, including Named Entity Recognition (NER) and human community detection. However, certain challenges lie ahead. Firstly, the input short-text contents are noisy, ambiguous, and do not follow the grammatical rules. Secondly, traditional tex…
System identifies language of transliterated text.
Paper improves short text clustering by integrating semantic relationships into Optimal Transport.
We give a short proof of a theorem of Handel and Mosher stating that any finitely generated subgroup of either contains a fully irreducible automorphism, or virtually fixes the conjugacy class of a proper free factor of , and we extend their result to non finitely generated subgroups of $\text{Ou…
In this paper we consider the problem of clustering collections of very short texts using subspace clustering. This problem arises in many applications such as product categorisation, fraud detection, and sentiment analysis. The main challenge lies in the fact that the vectorial representation of short texts is both hi…
Study examines how cluster number affects short-text clustering, introducing a stability metric.
Model clusters authors and topics in short texts like social media posts.
Proposes a flexible neural recommendation framework for better prediction performance.
Improves video search by balancing text and visual modalities.
Social media is increasingly used by humans to express their feelings and opinions in the form of short text messages. Detecting sentiments in the text has a wide range of applications including identifying anxiety or depression of individuals and measuring well-being or mood of a community. Sentiments can be expressed…
The Generative Adversarial Network (GAN) has achieved great success in generating realistic (real-valued) synthetic data. However, convergence issues and difficulties dealing with discrete data hinder the applicability of GAN to text. We propose a framework for generating realistic text via adversarial training. We emp…
In this paper, we prove that there exists a dimensional constant such that given any background Kähler metric , the Calabi flow with initial data satisfying \begin{equation*} \partial \bar \partial u_0 \in L^\infty (M) \text{ and } (1- δ)ω< ω_{u_0} < (1+δ)ω, \end{equation*} admits a unique short time so…
As the emergence and the thriving development of social networks, a huge number of short texts are accumulated and need to be processed. Inferring latent topics of collected short texts is useful for understanding its hidden structure and predicting new contents. Unlike conventional topic models such as latent Dirichle…
Study shows text-based news veracity models don't generalize across U.S. and U.K.
Framework incorporates prior knowledge into Bayesian models for data streams.
Ontology learning is a critical task in industry, dealing with identifying and extracting concepts captured in text data such that these concepts can be used in different tasks, e.g. information retrieval. Ontology learning is non-trivial due to several reasons with limited amount of prior research work that automatica…
This study reviews text-based stock market analysis methods.
Study proposes a multimodal model for cardiovascular risk prediction using EHRs.
Study enhances cryptocurrency sentiment analysis using TikTok and Twitter data.
Study integrates climate and text data to improve credit default prediction.
SCROLLS benchmarks long text NLP tasks, improving existing models.
The increasing volume of short texts generated on social media sites, such as Twitter or Facebook, creates a great demand for effective and efficient topic modeling approaches. While latent Dirichlet allocation (LDA) can be applied, it is not optimal due to its weakness in handling short texts with fast-changing topics…
One-hot CNN (convolutional neural network) has been shown to be effective for text categorization (Johnson & Zhang, 2015). We view it as a special case of a general framework which jointly trains a linear model with a non-linear feature generator consisting of `text region embedding + pooling'. Under this framework, we…
We apply text analysis approaches for a specialized search engine for 3D CAD models and associated products. The main goals are to distinguish between actual product descriptions and other text on a website, as well as to decide whether a given text is or contains a product name. For this we use paragraph vectors for t…
The paper predicts financial markets using news text and semantic network analysis.
SHMM models human mobility from GPS and text data, overcoming text sparsity.
A latent-variable model is introduced for text matching, inferring sentence representations by jointly optimizing generative and discriminative objectives. To alleviate typical optimization challenges in latent-variable models for text, we employ deconvolutional networks as the sequence decoder (generator), providing l…
Interventional cancer clinical trials are generally too restrictive, and some patients are often excluded on the basis of comorbidity, past or concomitant treatments, or the fact that they are over a certain age. The efficacy and safety of new treatments for patients with these characteristics are, therefore, not defin…
Entity linking is the task of mapping potentially ambiguous terms in text to their constituent entities in a knowledge base like Wikipedia. This is useful for organizing content, extracting structured data from textual documents, and in machine learning relevance applications like semantic search, knowledge graph const…
AudioPaLM combines text and speech models to improve speech processing and translation.
Let us denote by the hyperspace of all convex bodies of equipped with the Hausdorff distance topology. An affine invariant point is a continuous and Aff(n)-equivariant map , where Aff(n) denotes the group of all nonsingular affine maps of . Fo…
GCTM integrates GCN into topic models for better topic learning from data streams.
Bayesian Neural Networks detect gravitational wave events with high accuracy and real-time potential.
ICP improves text infilling and POS tagging with valid confidence sets.
Efficiently distills pretrained text-to-image models without real data, improving FID and CLIP scores.
We consider various notions of strains; quantitative measures for the deviation of a linear transformation from an isometry. The main approach, which is motivated by physical applications and follows the work of Patrizio Neff and co-workers , is to select a Riemannian metric on , and use its induced geodes…
We give a short proof of Masbaum and Reid's result that mapping class groups involve any finite group, appealing to free quotients of surface groups and a result of Gilman, following Dunfield-Thurston.
Many methods have been used to recognize author personality traits from text, typically combining linguistic feature engineering with shallow learning models, e.g. linear regression or Support Vector Machines. This work uses deep-learning-based models and atomic features of text, the characters, to build hierarchical, …
This is a short, elementary survey article about taut submanifolds. In order to simplify the exposition, we restrict to the case of compact smooth submanifolds of Euclidean or spherical spaces. Some new, partial results concerning taut 4-manifolds are discussed at the end of the text.
In text mining, information retrieval, and machine learning, text documents are commonly represented through variants of sparse Bag of Words (sBoW) vectors (e.g. TF-IDF). Although simple and intuitive, sBoW style representations suffer from their inherent over-sparsity and fail to capture word-level synonymy and polyse…