Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

8.3%16.7%25.0%33.3% · Jul 199219922001200920182026
48 results for Author name disambiguation

This study examines how imbalanced training data affects author name disambiguation.

problem The impact of imbalanced training data on machine learning for author name disambiguation.
method Training three classifiers (Logistic Regression, Naïve Bayes, Random Forest) on multiple labeled datasets with various positive-negative training data ratios.
result Increasing negative training data can improve disambiguation performance but with diminishing returns.

A new method for name disambiguation in academic networks using multi-view attention and recurrent neural networks.

problem Disambiguating authors with the same name in large-scale academic networks.
method Multi-view Attention-based Pairwise Recurrent Neural Network (MA-PairRNN) that divides papers into blocks based on author attributes and merges blocks of the same author.
result MA-PairRNN significantly improves name disambiguation performance on real-world datasets.

This study improves author disambiguation without supervision using feature overlap.

problem Author name homonymy in the Web of Science.
method Probabilistic similarity measure based on feature overlap for agglomerative clustering.
result Our approach outperforms the trivial baseline and is state-of-the-art.

Paper develops neural network for Mandarin polyphone disambiguation.

problem Homograph problem in Mandarin Chinese text-to-speech.
method Bidirectional RNN for context, prediction network for mapping embeddings to pronunciations.
result Achieves 94.69% accuracy on polyphonic character dataset.

EviTrack improves sequential prediction in delayed disambiguation scenarios.

problem Challenges in sequential prediction with delayed disambiguation where early observations are ambiguous.
method EviTrack operates over latent trajectories, applying evidence- and likelihood-ratio-based selection to delay commitment until supported by data.
result EviTrack outperforms sampling-based baselines in a controlled synthetic benchmark, achieving faster post-disambiguation recovery.

Paper proposes an algorithm to recover full supervision from weakly labeled data.

problem Machine learning requires expensive data annotation, motivating the use of weak supervision.
method The paper introduces a disambiguation principle and an empirical disambiguation algorithm for partial labelling.
result The algorithm achieves exponential convergence rates under learnability assumptions.

Model learns to discover and disambiguate entities and relations in text streams.

problem Learning to follow and resolve mentions in a continuous text stream.
method End-to-end trainable memory network for online, one-shot learning.
result Improves disambiguation and discovery skills with minimal supervision.

Improves medical note processing by training model on related concepts and global context.

problem Scarce and imbalanced labeled training data limits generalizability of automated abbreviation disambiguation models.
method Data augmentation using related medical concepts and global context information within medical notes.
result Model accuracy improved by almost 14% on CASI dataset and 4% on i2b2 dataset.

A coloring scheme improves graph neural networks for node disambiguation.

problem Improving graph neural networks' ability to distinguish identical node attributes.
method Introducing a graph neural network called Colored Local Iterative Procedure (CLIP) that uses colors to disambiguate node attributes.
result CLIP is a universal approximator of continuous functions on graphs with node attributes.

System solves author name ambiguity in e-commerce catalogs.

problem Finding correct author names in e-commerce catalogs with abbreviations and spelling variants.
method Composite system using open data sources and machine learning techniques for natural language processing.
result Top proposal of the system is the normalized author name with 72% accuracy.

Active inference selects actions to maximize information gain, aiding structure learning.

problem Learning the structure of underlying world models.
method Active inference selects actions based on expected free energy, which includes information gain and value.
result Actions that maximize information gain help disambiguate among alternative models.

DivDis learns diverse hypotheses from underspecified data to improve robustness.

problem Learning from underspecified datasets leads to multiple equally viable solutions, causing out-of-distribution issues.
method DivDis framework: 1) learns diverse hypotheses using unlabeled test data, 2) selects one hypothesis with minimal additional supervision.
result DivDis finds robust features in image and natural language processing problems.

Vector representations of words have heralded a transformational approach to classical problems in NLP; the most popular example is word2vec. However, a single vector does not suffice to model the polysemous nature of many (frequent) words, i.e., words with multiple meanings. In this paper, we propose a three-fold appr…

2016-10-24abs ↗pdf ↗

PML-GAN tackles noisy multi-label annotations using adversarial learning.

problem Learning multi-label models from noisy, overcomplete annotations.
method PML-GAN uses a disambiguation network and a generative adversarial network to map noisy labels to clean labels and data samples.
result PML-GAN achieves state-of-the-art performance on partial multi-label learning datasets.

Efficient autoregressive entity linking with correction for faster, more accurate results.

problem High computational cost and non-parallelizable decoding in autoregressive entity linking.
method Parallelizes autoregressive linking across all mentions, uses a shallow decoder, and adds a discriminative correction term.
result 70 times faster and more accurate than previous methods, outperforming state-of-the-art approaches.

Enhanced word embeddings boost multiclass text classification accuracy.

problem Improving multiclass text classification accuracy using pre-trained embeddings.
method Proposed word-class embeddings (WCEs) to enhance pre-trained word embeddings.
result WCEs significantly improve multiclass text classification accuracy.

Researchers prove rigidity of first conformal Steklov eigenvalue on specific shapes.

problem Rigidity of the first conformal Steklov eigenvalue on annuli and Möbius bands.
method Proof relies on uniqueness results, compactness theorem, and asymptotic control of Steklov eigenvalues.
result Rigidity of the first conformal Steklov eigenvalue on annuli and Möbius bands proved.

The study quantifies uncertainty to improve model calibration and disambiguate annotator and data bias in emotion recognition.

problem Improving model interpretability and disambiguating bias in complex tasks like emotion recognition.
method Used a modified Monte Carlo dropout approach to quantify epistemic and aleatoric uncertainty.
result Identified a significant correlation between aleatoric uncertainty and human annotator disagreement.

Constructs minimal surfaces in balls, maximizing eigenvalues.

problem Finding minimal surfaces in Euclidean balls with controlled topology.
method Maximizing the first non-trivial Steklov eigenvalue for isoperimetric problems.
result Constructs free boundary minimal immersions with controlled topology.

Extends complex manifold structures to line bundles, revealing new projective manifolds.

problem Generalizing scalar-valued holomorphic structures to line bundles.
method Study of holomorphic pp-contact and ss-symplectic structures on complex manifolds with line bundles.
result Holomorphic pp-contact and ss-symplectic manifolds can be projective.

We construct prime amphicheiral knots that have free period 2. This settles an open question raised by the second named author, who proved that amphicheiral hyperbolic knots cannot admit free periods and that prime amphicheiral knots cannot admit free periods of order >2.

2018-04-09abs ↗pdf ↗

2-dimensional knots and links are studied in the article. The notion of parity is introduced via techniques similar to the ones used by the second named author in 1-dimensional case. By using parity new invariants are constructed and known invariants are refined.

2016-06-22abs ↗pdf ↗

Motivated by a previous work of Zheng and the second named author, we study pinching constants of compact Kähler manifolds with positive holomorphic sectional curvature. In particular we prove a gap theorem following the work of Petersen and Tao on Riemannian manifolds with almost quarter-pinched sectional curvature.

2017-09-08abs ↗pdf ↗

This paper completes the classification of certain nilpotent Lie groups with specific geometric structures.

problem Classifying nilpotent Lie groups with purely coclosed G2-structures.
method Analyzing seven-dimensional nilpotent Lie groups of various steps.
result Classification of indecomposable 5- and 6-step nilpotent Lie groups with these structures.

We construct a weak 2-functor from the bicategory of oriented tangles to a bicategory of Lagrangian cospans. This functor simultaneously extends the Burau representation of the braid groups, its generalization to tangles due to Turaev and the first-named author, and the Alexander module of 1 and 2-dimensional links.

2016-06-16abs ↗pdf ↗