Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

15294458 · Feb 202019922001200920182026
48 results for Topic categorization

Local-HDP learns independent topics for each 3D object category in real-time.

problem Learning independent topics for each 3D object category in real-time.
method Local-Hierarchical Dirichlet Process (Local-HDP) with online variational inference.
result Local-HDP outperforms other approaches in accuracy, scalability, and memory efficiency.

The study analyzes stock price reactions to diverse financial disclosures.

problem Understanding stock price reactions to varied financial disclosures.
method Topic modeling using Latent Dirichlet Allocation (LDA) for categorizing filings.
result Significant abnormal returns in response to specific types of disclosures.

Supervised topic models utilize document's side information for discovering predictive low dimensional representations of documents. Existing models apply the likelihood-based estimation. In this paper, we present a general framework of max-margin supervised topic models for both continuous and categorical response var…

2009-12-30abs ↗pdf ↗

A new topic model uses word embeddings on a sphere for better topic coherence.

problem Traditional topic models ignore semantic word correlations.
method Proposes von Mises-Fisher distribution for word density on a unit sphere, using Hierarchical Dirichlet Process and Stochastic Variational Inference.
result The model outperforms existing methods in topic coherence.

Researchers use LRP to explain CNN predictions in NLP tasks.

problem Explaining predictions of complex non-linear classifiers in NLP.
method Layer-wise relevance propagation (LRP) applied to a CNN for topic categorization.
result LRP highlights relevant words for CNN predictions, validating its suitability for NLP.

Improved bibliographic model for author, topic, and document clustering.

problem Modeling research publications using authors, categorical labels, and citation networks.
method Citation Network Topic Model (CNTM) combining Poisson mixed-topic and author-topic models with a novel inference algorithm.
result Improved performance in model fitting and document clustering compared to baselines.

LR-Robot automates SLRs with AI, expert oversight, and multidimensional analysis.

problem Efficient but contextually limited outputs from existing SLR frameworks.
method Human-in-the-loop process, structured knowledge sources, retrieval-augmented generation.
result Empirical demonstration of AI-driven literature synthesis in option pricing.

Report on enhancing trust in ML models with visualizations.

problem Understanding and trusting the results of black box ML models.
method Categorization of visualization techniques, statistical overview, topic analyses, and interactive web-based survey.
result Expanded categorization of trust against different facets of interactive ML.

Paper proposes PP-GCN for fine-grained social event categorization.

problem Challenges in mining social events due to heterogeneous event elements and social network structures.
method Design an event meta-schema, build an HIN, propose PP-GCN, and use KIES.
result PP-GCN outperforms other techniques in social event detection and clustering.

Paper proposes a new estimator for generic discrete distributions.

problem Estimating gradients for stochastic nodes in deep generative models.
method Generalized Gumbel-Softmax estimator using truncation, Gumbel-Softmax trick, and linear transformation.
result Efficacy and practical value demonstrated in synthetic examples and topic models.

Survey on reproducibility and distortion issues in text clustering and topic modeling.

problem Reproducibility and misleading cluster geometry in unsupervised learning for text categorization.
method Systematic literature review of text clustering and topic modeling from 2011-2022.
result Outliers and initialization issues are significant factors in text clustering and topic modeling.

Paper proposes a semi-supervised learning approach for automated topic detection in support and feedback data.

problem Manual analysis of large textual feedback and support data is time-consuming and labor-intensive.
method Combines BERT-based multiclassification with a novel PSHTI model for automated topic and sub-topic inference.
result Automated system for better understanding user voice and engagement patterns in enterprise data.

Discusses new probabilistic morphisms and geometric methods in machine and statistical learning.

problem Addressing challenges in statistical, machine, and manifold learning.
method Introduces category of probabilistic morphisms and geometric methods.
result New insights and applications in various learning fields.

Study categorizes knots and links as rigid or shaky based on Reidemeister moves.

problem Classifying knots and links as rigid or shaky based on adaptability to Reidemeister moves.
method Categorization of hard diagrams as rigid or shaky, investigation of rigid and shaky hard diagrams for specific knots and links.
result Every link has a rigid hard diagram, and there is an upper limit for the number of crossings in such diagrams.

Following Roe and others (see, e.g., [MR1451755]), we (re)develop coarse geometry from the foundations, taking a categorical point of view. In this paper, we concentrate on the discrete case in which topology plays no role. Our theory is particularly suited to the development of the_Roe (C*-)algebras_ C*(X) and their K…

2007-08-29abs ↗pdf ↗

Machine learning identifies types of alterations in historical manuscripts.

problem Understanding and categorizing alterations in historical manuscripts.
method Alteration Latent Dirichlet Allocation (alterLDA) model.
result High performance in recognizing alterations on labelled data, and interesting insights on unlabelled data.

Machine learning explains text document categorization decisions.

problem Understanding how text documents are categorized by machine learning models.
method Layer-wise relevance propagation (LRP) to trace predictions back to individual words.
result Word-based ML models can be made more comprehensible through LRP.

Study uses machine learning to analyze Twitter sentiments about COVID-19.

problem Examining public concerns and sentiments about COVID-19 from Twitter.
method Machine learning (Latent Dirichlet Allocation) to identify topics and sentiments.
result Identified 13 topics and categorized into five themes, revealing dominant fears and mixed feelings.

Experimental life sciences like biology or chemistry have seen in the recent decades an explosion of the data available from experiments. Laboratory instruments become more and more complex and report hundreds or thousands measurements for a single experiment and therefore the statistical methods face challenging tasks…

2014-03-12abs ↗pdf ↗

New models automate support group formation in online health communities.

problem Challenges in traditional support group formation methods for scalability, static categorization, and insufficient personalization.
method Two novel machine learning models: gDMR and gSTM, integrating user content, demographics, and network data.
result Models outperform baselines in predictive accuracy, semantic coherence, and internal group consistency.

Survey on modeling event sequences through temporal processes.

problem Modeling phenomena with sequences of events over continuous time.
method Probabilistic models based on point processes, categorized into simple, marked, and spatio-temporal.
result Analysis of existing approaches and their applicability to prediction and modeling.

StructureBoost improves gradient boosting for complex categorical variables efficiently.

problem Efficiently handling complex categorical variables with known structure.
method Two methods to overcome computational obstacles in SCDT enumeration for structured categorical variables.
result StructureBoost outperforms existing packages on complex categorical problems.

Categorical bundles provide a natural framework for gauge theories involving multiple gauge groups. Unlike the case of traditional bundles there are distinct notions of triviality, and hence also of local triviality, for categorical bundles. We study categorical principal bundles that are product bundles in the categor…

2015-06-14abs ↗pdf ↗

Bayesian model improves categorization of explosions from sparse data.

problem Challenges in categorizing explosions from limited data.
method Bayesian update to Event Categorization Matrix model with Bayesian Decision Theory.
result Consistent gains in overall accuracy and lower false negative rates.

A new method for backpropagating through categorical distributions.

problem Difficulty in backpropagating through categorical latent variables in neural networks.
method Introducing Gumbel-Softmax distribution for differentiable sampling.
result Gumbel-Softmax estimator outperforms existing methods on tasks with categorical latent variables.

microTC is a minimalistic text classifier for various tasks.

problem Text categorization tasks across different domains and languages.
method Minimalistic approach with text transformations, representations, and supervised learning.
result microTC outperformed state-of-the-art methods in 20 out of 30 datasets.

UNTIE learns representations of coupled categorical data.

problem Challenges in learning from unlabeled categorical data with complex couplings.
method UNTIE approach for unsupervised representation learning of heterogeneous couplings.
result UNTIE significantly improves categorical data representations on 25 diverse datasets.

Paper introduces Categorical Normalizing Flows for better handling of categorical data.

problem Limited application of normalizing flows on categorical data due to lack of intrinsic order.
method Categorical Normalizing Flows use continuous transformations to model latent relations in categorical data, optimizing both continuous representation and model likelihood.
result GraphCNF, a permutation-invariant generative model, outperforms state-of-the-art on molecule generation.

This paper proposes a method to reduce complexity in GLMs with categorical predictors.

problem Wasteful, hard-to-interpret, and prone to overfitting of traditional one-hot encoding for high-cardinality categorical predictors.
method Clustering categories of categorical predictors through a numerical method that preserves or improves accuracy while reducing the number of coefficients.
result Clustering categories of categorical predictors reduces complexity substantially without harming accuracy.

Gaussian-Dirichlet posterior dominance proven for sequential categorical data.

problem Sequential learning from categorical observations bounded in [0,1]
method Establishing an ordering between Dirichlet and Gaussian posteriors under N(0,1) noise
result Posterior mean of categorical distribution stochastically dominates Gaussian distribution

A new method optimises problems with both continuous and categorical inputs.

problem Optimising black-box problems with mixed continuous and categorical inputs.
method Continuous and Categorical Bayesian Optimisation (CoCaBO) combining multi-armed bandits and Bayesian optimisation.
result CoCaBO outperforms existing methods on synthetic and real-world tasks.

The paper shows how integrating categorical semantics can enhance unsupervised domain translation.

problem Improving unsupervised domain translation between perceptually different domains.
method Learning invariant categorical semantic features in an unsupervised manner and conditioning them on the style encoder.
result Conditioning the style encoder on learned categorical semantics improves translation and stylization.