Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

8162331 · Oct 201919922001200920182026
48 results for document ontology

Deep learning captures semantic structure of large documents.

problem Understanding complex, structured documents like scholarly articles and business reports.
method Deep learning-based document ontology to capture semantic structure and domain-specific concepts.
result The ontology enhances semantic indexing for better understanding by humans and machines.

Enhanced ontology learning from text improves question-answering systems.

problem Improving ontology learning from unstructured text for better question-answering systems.
method Heuristically modified FP-Tree with DFA for concept extraction and frequent pattern mining for ontology learning.
result Our approach significantly improves question-answering system performance, answering 80% of questions compared to 28.4% with Text2Onto.

The FAIRnets Ontology makes neural networks findable, accessible, interoperable, and reusable.

problem The resource-intensive training of neural networks and the lack of training data availability.
method Development of FAIRnets Ontology to model neural networks on a meta-level and creation of a knowledge graph (FAIRnets) of over 18,400 neural networks.
result The FAIRnets Ontology and knowledge graph enable the reuse and recommendation of neural networks to data scientists.

Prototype learns automotive industry ontology from unstructured data.

problem Automatic learning of domain-specific ontologies from unstructured text data.
method Two-stage classification system: first classifier for concepts and irrelevant collocates, second classifier for concept types.
result Prototype validated with automotive industry complaint and repair data.

Improved predictions for rare labels using neural networks and ontologies.

problem Long-tailed frequency distribution in multi-label prediction problems.
method Modified neural network output layer with a Bayesian network of sigmoids leveraging ontology relationships.
result Significant improvements in per-label AUROC and average precision for less common labels.

A new model embeds word and label hierarchies in hyperbolic space for HMLC.

problem Learning mappings from word hierarchies to label hierarchies in hierarchical multi-label classification.
method Proposes a Hyperbolic Interaction Model (HyperIM) to learn label-aware document representations in hyperbolic space.
result Demonstrates improved performance for HMLC compared to state-of-the-art methods.

We explain the meaning of local symmetries in physics.

problem Understanding the meaning of local symmetries in physics.
method We argue that general covariance and gauge principles are principles of epistemic access to physical laws, leading to ontological insights.
result Relationality is a core notion in gauge field theory, encoded by local symmetries.

ODVICE augments EHR cohorts using ontology to improve analysis robustness.

problem Limited records in cohorts for rare diseases hamper robust analysis.
method Ontology-driven Monte-Carlo graph spanning algorithm for data augmentation.
result ODVICE augmented cohorts show ~30% improvement in AUC over non-augmented datasets.

ML-Schema offers a standardized format for machine learning components.

problem Lack of interoperability and interpretability in machine learning experiments.
method Developed a top-level ontology with classes, properties, and restrictions.
result ML-Schema enables better interoperability and interpretability of machine learning experiments.

Convolutional neural networks (CNNs) have recently emerged as a popular building block for natural language processing (NLP). Despite their success, most existing CNN models employed in NLP share the same learned (and static) set of filters for all input sentences. In this paper, we consider an approach of using a smal…

2017-09-25abs ↗pdf ↗

PEHRT harmonizes EHR data for translational research.

problem Barriers in using EHR data for translational research.
method Common pipeline including open-source code, visualization tools, and detailed documentation.
result PEHRT harmonizes EHR data to standardized ontologies and generates robust embeddings.

We applied machine learning to predict whether a gene is involved in axon regeneration. We extracted 31 features from different databases and trained five machine learning models. Our optimal model, a Random Forest Classifier with 50 submodels, yielded a test score of 85.71%, which is 4.1% higher than the baseline scor…

2017-10-30abs ↗pdf ↗

Deep learning method improves gene ontology classification of neural images.

problem Classifying gene expression in neural in situ hybridization images.
method End-to-end deep learning using convolutional denoising autoencoders (CDAE).
result Significant improvement in classification accuracy (96% reduction in error rate).

Recent work in learning ontologies (hierarchical and partially-ordered structures) has leveraged the intrinsic geometry of spaces of learned representations to make predictions that automatically obey complex structural constraints. We explore two extensions of one such model, the order-embedding model for hierarchical…

2017-08-01abs ↗pdf ↗

Machine learning's data-centric philosophy conflicts with natural sciences' standards.

problem Conflict between machine learning's ontology and epistemology and natural sciences' practices.
method Identifying and analyzing contexts where ML can be beneficial or harmful in natural sciences.
result ML can enhance trustworthiness in causal inference but introduces biases in emulation and labeling.

Research creates a taxonomy to bridge AI security and regulatory gaps.

problem Disciplinary disconnect between technical and legal teams in AI risk assessment.
method Developed an AI System Threat Vector Taxonomy with 9 domains and 53 sub-threats.
result Empirically validated and aligned with ISO/IEC 42001 controls and NIST AI RMF functions.

Enhanced neural network framework improves constraint satisfaction with topological conditioning.

problem Maintaining semantic coherence while satisfying physical and logical constraints in neuro-symbolic reasoning.
method Integrates topological conditioning with gradient stabilization mechanisms using Forman-Ricci curvature, Deep Delta Learning, and Covariance Matrix Adaptation Evolution Strategy.
result Achieves mean energy reduction to 1.15 compared to baseline values of 11.68, with 95 percent success rate.

Paper proposes a graph network for EHR data that learns robust representations.

problem Learning robust representations for EHR data with implicit connections.
method Variationally regularized encoder-decoder graph network.
result Model outperforms existing methods in various EHR predictive tasks.

MuLan links music audio to natural language tags.

problem Traditional music tagging systems use rigid attributes; MuLan aims to link audio directly to natural language.
method Joint audio-text embedding model trained on 44 million music recordings and text annotations.
result MuLan's embeddings enable zero-shot functionalities and transfer learning.

The problem of multilabel classification when the labels are related through a hierarchical categorization scheme occurs in many application domains such as computational biology. For example, this problem arises naturally when trying to automatically assign gene function using a controlled vocabularies like Gene Ontol…

2012-05-09abs ↗pdf ↗

Framework uses deep learning to analyze large documents and identify their logical structure.

problem Analyzing large, multi-themed documents with diverse topics.
method Deep learning techniques to model and extract logical and semantic structure.
result Framework effectively identifies and classifies different sections of documents.

Imaging neuroscience links brain activation maps to behavior and cognition via correlational studies. Due to the nature of the individual experiments, based on eliciting neural response from a small number of stimuli, this link is incomplete, and unidirectional from the causal point of view. To come to conclusions on t…

2013-11-15abs ↗pdf ↗

Entity-GCN model answers multi-document questions by reasoning across documents.

problem Answering questions based on multiple documents and cross-document relations.
method Graph Convolutional Networks (GCNs) applied to a graph of mentions and their relations.
result Achieves state-of-the-art results on WikiHop dataset.

A new method for document network embedding interprets and generalizes well.

problem Lack of interpretability and generalization to new documents in existing methods.
method Introduces Topic-Word Attention (TWA) and Inductive Document Network Embedding (IDNE) to generate document representations.
result Achieves state-of-the-art performance on various networks and produces meaningful representations.

We show that, under mild assumptions, some unimaginable events - which we refer to as Black Swan events - must necessarily occur. It follows as a corollary of our theorem that any computational model of decision-making under uncertainty is incomplete in the sense that not all events that occur can be taken into account…

2018-03-07abs ↗pdf ↗

MarlRank uses multi-agent reinforcement learning to improve document ranking.

problem Neglecting mutual information among documents in ranking models.
method Formulated as a multi-agent Markov Decision Process (MDP), each document predicts relevance considering its own and similar documents features and actions.
result Significant performance gains over state-of-the-art baselines on LETOR benchmark datasets.

Study develops a diagnostic tool for rare diseases using reinforcement learning.

problem Minimize medical tests while reducing diagnostic uncertainty for rare diseases.
method Investigated reinforcement learning algorithms, combined expert knowledge with clinical data, integrated ontological information.
result Demonstrated a feasible decision support tool for rare diseases.

SL2MF predicts synthetic lethality using logistic matrix factorization.

problem Predicting synthetic lethality in human cancers from limited experimental data.
method Logistic matrix factorization incorporating biological knowledge.
result SL2MF effectively predicts known and unknown SL interactions.

Paper develops a new unsupervised scoring function for cross-lingual document alignment.

problem Aligning documents across different languages for NLP tasks.
method Uses cross-lingual sentence embeddings to compute semantic distances and guides document alignment.
result The proposed scoring function outperforms current methods by 7-22% on various language pairs.