Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

8152330 · Oct 201919922001200920182026
48 results for Semi-Structured Documents

Bayesian method for semi-structured models accounts for both types of uncertainty.

problem Lack of work on epistemic uncertainty in semi-structured regression models.
method Bayesian approximation with subspace inference for joint posterior sampling.
result Validated approach recovers structured effect posteriors and approaches full-space posterior.

This paper extends semi-structured networks to functional data.

problem Maintaining interpretability in functional data analysis while capturing non-linearities and interactions.
method Proposes a functional SSN method that scales well and improves predictive performance.
result The functional SSN method accurately recovers underlying signals and performs favorably compared to competing methods.

Proposes a new model to analyze mortgage delinquency transitions.

problem Analyzing mortgage delinquency transitions in a flexible yet identifiable way.
method Combines structured additive predictor with neural network for complex interactions, orthogonalising components for identifiability.
result The semi-structured model provides modest gains in discrimination compared to a structured model, especially in the early prediction spans.

Study reveals issues with neural autoregressive models and proposes mode recovery cost.

problem Unreasonable affinity of neural autoregressive models to short and long sequences.
method Investigates modes of ground-truth, empirical, and decoding-induced distributions via mode recovery cost.
result Mode recovery cost varies depending on ground-truth distribution and impacts decoding-induced distribution.

Polytopic Matrix Factorization models data as latent vectors from a polytope, maximizing determinant for identifiability.

problem Data decomposition with semi-structured latent vectors and polytope constraints.
method Model input data as latent vectors from a polytope, using determinant maximization for identifiability.
result Identifiability condition for polytopes with specific symmetry restrictions.

Framework uses deep learning to analyze large documents and identify their logical structure.

problem Analyzing large, multi-themed documents with diverse topics.
method Deep learning techniques to model and extract logical and semantic structure.
result Framework effectively identifies and classifies different sections of documents.

Entity-GCN model answers multi-document questions by reasoning across documents.

problem Answering questions based on multiple documents and cross-document relations.
method Graph Convolutional Networks (GCNs) applied to a graph of mentions and their relations.
result Achieves state-of-the-art results on WikiHop dataset.

A new method for document network embedding interprets and generalizes well.

problem Lack of interpretability and generalization to new documents in existing methods.
method Introduces Topic-Word Attention (TWA) and Inductive Document Network Embedding (IDNE) to generate document representations.
result Achieves state-of-the-art performance on various networks and produces meaningful representations.

Deep learning captures semantic structure of large documents.

problem Understanding complex, structured documents like scholarly articles and business reports.
method Deep learning-based document ontology to capture semantic structure and domain-specific concepts.
result The ontology enhances semantic indexing for better understanding by humans and machines.

MarlRank uses multi-agent reinforcement learning to improve document ranking.

problem Neglecting mutual information among documents in ranking models.
method Formulated as a multi-agent Markov Decision Process (MDP), each document predicts relevance considering its own and similar documents features and actions.
result Significant performance gains over state-of-the-art baselines on LETOR benchmark datasets.

Generative topic embedding combines local and global patterns for document representation.

problem Representing documents in a continuous space using both local and global patterns.
method Proposes a variational inference model to generate topic embeddings and document representations.
result Performs better than existing methods in document classification tasks.

A new nonparametric model discovers hidden topics in document networks without knowing the number of topics in advance.

problem Discovering hidden topics in document networks without knowing the number of topics in advance.
method Proposes a nonparametric relational topic model using dependent gamma processes to represent topic interests of documents and their network structure.
result Can discover hidden topics and their number simultaneously using a posterior inference algorithm.

Paper develops a new unsupervised scoring function for cross-lingual document alignment.

problem Aligning documents across different languages for NLP tasks.
method Uses cross-lingual sentence embeddings to compute semantic distances and guides document alignment.
result The proposed scoring function outperforms current methods by 7-22% on various language pairs.

A scalable topic model for large document collections using MapReduce.

problem Scalability issues in topic modeling for large document collections.
method Correlated Topic Model with variational Expectation-Maximization in MapReduce framework.
result Comparable topic coherences with LDA in MapReduce framework.

A new method for measuring document similarity using hierarchical optimal transport.

problem Inability to measure semantic similarities and scalability issues in past document similarity measures.
method Model documents as distributions over topics, topics as distributions over words, solve optimal transport problem on topics.
result Hierarchical optimal transport provides better interpretability and scalability with comparable performance.

We develop a nested hierarchical Dirichlet process (nHDP) for hierarchical topic modeling. The nHDP is a generalization of the nested Chinese restaurant process (nCRP) that allows each word to follow its own path to a topic node according to a document-specific distribution on a shared tree. This alleviates the rigid, …

2012-10-25abs ↗pdf ↗

Study improves document processing in banking with multimodal analytics.

problem Raising operational efficiency in banking through document-intensive processes.
method Comparative analysis of text classifiers and multimodal model (LayoutXLM) on company register extracts.
result Incorporating layout information in a model substantially increases performance.

Document clustering and topic modeling are two closely related tasks which can mutually benefit each other. Topic modeling can project documents into a topic space which facilitates effective document clustering. Cluster labels discovered by document clustering can be incorporated into topic models to extract local top…

2013-09-26abs ↗pdf ↗

This paper improves topic modeling by embedding words and topics together.

problem Topic models struggle with short documents and approximate inference.
method Model each document as a mixture of word embeddings and each topic as a mixture of topic embeddings.
result The method optimizes topic embeddings to minimize semantic differences between words and topics.

Improved bio-surveillance through automated document classification.

problem Tracking infectious diseases across global news alerts.
method Recurrent neural networks, TF-IDF, Naive Bayes, logistic regression.
result 97% recall and 93.3% accuracy in bio-surveillance event classification.

Develops scalable autoencoder for document networks.

problem Sparse and skewed latent node representations in document relational networks.
method Combines graph Poisson factor analysis with Weibull-based graph inference networks.
result Extracts high-quality hierarchical latent document representations.

Paper proposes LAHA to improve XMTC by integrating document content and label correlation.

problem Challenges in tagging documents with most relevant labels from a large label set.
method Hybrid attention deep neural network model (LAHA) that combines multi-label self-attention and adaptive fusion strategies.
result LAHA outperforms state-of-the-art methods, especially on tail labels.

Top2Vec finds topic vectors from documents and words without needing stop words or custom settings.

problem Topic modeling weaknesses, including needing known topics, stop words, and custom settings.
method Joint document and word semantic embedding to find topic vectors automatically.
result Top2Vec finds more informative and representative topics than probabilistic models.

Proposes new listwise learning-to-rank models to address rating ties and document relevance.

problem Rating ties and document relevance in existing listwise learning-to-rank models.
method Models ranking as selecting documents from a candidate set based on unique rating levels. Uses a new loss function and adapted RNN model for refining prediction scores.
result Models notably outperform state-of-the-art learning-to-rank models on four public datasets.

List-CVAE optimizes document slates directly, improving recommendation performance.

problem Optimizing document slates for better user interest capture.
method List-CVAE learns document distributions conditioned on user responses to generate optimal slates.
result List-CVAE outperforms traditional ranking methods consistently across various document corpora.

Improves retrieval accuracy for hierarchical documents, especially for distant matches.

problem Limited expressive power of dual encoder models in hierarchical retrieval.
method Proves feasibility of DEs for HR, introduces pretrain-finetune recipe to improve long-distance retrieval.
result Pretrain-finetune boosts recall on long-distance pairs from 19% to 76%.

This paper speeds up WMD computation for multiple queries efficiently.

problem Efficiently computing the semantic dissimilarity between text documents.
method Adapting the Sinkhorn-Knopp algorithm to compute WMD of one document against many targets in parallel.
result 67x speedup on 96 cores compared to sequential and naive parallel methods.