This study examines how imbalanced training data affects author name disambiguation.
problem The impact of imbalanced training data on machine learning for author name disambiguation.
method Training three classifiers (Logistic Regression, Naïve Bayes, Random Forest) on multiple labeled datasets with various positive-negative training data ratios.
result Increasing negative training data can improve disambiguation performance but with diminishing returns.
A new method for name disambiguation in academic networks using multi-view attention and recurrent neural networks.
problem Disambiguating authors with the same name in large-scale academic networks.
method Multi-view Attention-based Pairwise Recurrent Neural Network (MA-PairRNN) that divides papers into blocks based on author attributes and merges blocks of the same author.
result MA-PairRNN significantly improves name disambiguation performance on real-world datasets.
This study improves author disambiguation without supervision using feature overlap.
problem Author name homonymy in the Web of Science.
method Probabilistic similarity measure based on feature overlap for agglomerative clustering.
result Our approach outperforms the trivial baseline and is state-of-the-art.
Author name disambiguation in bibliographic databases is the problem of grouping together scientific publications written by the same person, accounting for potential homonyms and/or synonyms. Among solutions to this problem, digital libraries are increasingly offering tools for authors to manually curate their publica…
A new system learns entity representations to improve local entity disambiguation.
problem Local entity disambiguation in text.
method Entity-ELMo (E-ELMo) approach for contextual entity representation.
result Outperforms state-of-the-art models by 0.5% on AIDA test-b.
Framework tracks evolving news stories across multiple sources.
problem Tracking evolving news stories across diverse sources and formats.
method Cross-domain story tracking approach using entity graphs and learning-to-rank.
result Outperforms state-of-the-art methods for real-time story tracking.
Paper develops neural network for Mandarin polyphone disambiguation.
problem Homograph problem in Mandarin Chinese text-to-speech.
method Bidirectional RNN for context, prediction network for mapping embeddings to pronunciations.
result Achieves 94.69% accuracy on polyphonic character dataset.
Improves biomedical entity linking with latent type modeling.
problem Lack of fine-grained type information for entity disambiguation.
method Jointly models entity disambiguation and latent type learning without direct supervision.
result Significant performance improvements over state-of-the-art techniques.
EviTrack improves sequential prediction in delayed disambiguation scenarios.
problem Challenges in sequential prediction with delayed disambiguation where early observations are ambiguous.
method EviTrack operates over latent trajectories, applying evidence- and likelihood-ratio-based selection to delay commitment until supported by data.
result EviTrack outperforms sampling-based baselines in a controlled synthetic benchmark, achieving faster post-disambiguation recovery.
Paper proposes an algorithm to recover full supervision from weakly labeled data.
problem Machine learning requires expensive data annotation, motivating the use of weak supervision.
method The paper introduces a disambiguation principle and an empirical disambiguation algorithm for partial labelling.
result The algorithm achieves exponential convergence rates under learnability assumptions.
Model learns to discover and disambiguate entities and relations in text streams.
problem Learning to follow and resolve mentions in a continuous text stream.
method End-to-end trainable memory network for online, one-shot learning.
result Improves disambiguation and discovery skills with minimal supervision.
Improves medical note processing by training model on related concepts and global context.
problem Scarce and imbalanced labeled training data limits generalizability of automated abbreviation disambiguation models.
method Data augmentation using related medical concepts and global context information within medical notes.
result Model accuracy improved by almost 14% on CASI dataset and 4% on i2b2 dataset.
A coloring scheme improves graph neural networks for node disambiguation.
problem Improving graph neural networks' ability to distinguish identical node attributes.
method Introducing a graph neural network called Colored Local Iterative Procedure (CLIP) that uses colors to disambiguate node attributes.
result CLIP is a universal approximator of continuous functions on graphs with node attributes.
System solves author name ambiguity in e-commerce catalogs.
problem Finding correct author names in e-commerce catalogs with abbreviations and spelling variants.
method Composite system using open data sources and machine learning techniques for natural language processing.
result Top proposal of the system is the normalized author name with 72% accuracy.
Active inference selects actions to maximize information gain, aiding structure learning.
problem Learning the structure of underlying world models.
method Active inference selects actions based on expected free energy, which includes information gain and value.
result Actions that maximize information gain help disambiguate among alternative models.
DivDis learns diverse hypotheses from underspecified data to improve robustness.
problem Learning from underspecified datasets leads to multiple equally viable solutions, causing out-of-distribution issues.
method DivDis framework: 1) learns diverse hypotheses using unlabeled test data, 2) selects one hypothesis with minimal additional supervision.
result DivDis finds robust features in image and natural language processing problems.
New method clusters unknown music artists using audio metrics.
problem Disambiguating large catalogs of unknown artists.
method Metric learning from audio data with negative sampling.
result Our method outperforms a classifier-based approach when audio data is available.
Vector representations of words have heralded a transformational approach to classical problems in NLP; the most popular example is word2vec. However, a single vector does not suffice to model the polysemous nature of many (frequent) words, i.e., words with multiple meanings. In this paper, we propose a three-fold appr…
Sparse-mode DMD disambiguates local and global modes in spatiotemporal data.
problem Disambiguating local and global modes in spatiotemporal data.
method Sparse-mode DMD with sparsity-promoting regularization.
result Explicitly constructs discrete and continuous spectra.
Paper proposes FOFE for efficient WSD.
problem Word sense disambiguation (WSD) problem.
method Fixed-size ordinally forgetting encoding (FOFE) combined with FFNN.
result FOFE-based FFNN achieves comparable performance to state-of-the-art at lower cost.
PML-GAN tackles noisy multi-label annotations using adversarial learning.
problem Learning multi-label models from noisy, overcomplete annotations.
method PML-GAN uses a disambiguation network and a generative adversarial network to map noisy labels to clean labels and data samples.
result PML-GAN achieves state-of-the-art performance on partial multi-label learning datasets.
Efficient autoregressive entity linking with correction for faster, more accurate results.
problem High computational cost and non-parallelizable decoding in autoregressive entity linking.
method Parallelizes autoregressive linking across all mentions, uses a shallow decoder, and adds a discriminative correction term.
result 70 times faster and more accurate than previous methods, outperforming state-of-the-art approaches.
DKPCA improves WSD accuracy with scarce labeled data.
problem Word sense disambiguation in natural language processing.
method DKPCA combines Kernel PCA and Semantic Diffusion Kernel.
result DKPCA outperforms SVM and KPCA on SensEval data.
Authors derive McShane identity for super tori.
problem Deriving McShane identity for super tori.
method Super Teichmüller theory and supergeometry.
result Asymptotic growth rate of length spectra established.
Pangloss improves entity linking in noisy text environments.
problem Entity linking in non-grammatical, loosely-structured text.
method Combines probabilistic key phrase identification and semantic similarity engine.
result Better than state-of-the-art results (>5% in F1).
FONDUE identifies ambiguous nodes in networks for better analysis.
problem Ambiguous nodes in network data.
method Network embedding for node disambiguation.
result FONDUE outperforms existing methods in ambiguous node identification.
Enhanced word embeddings boost multiclass text classification accuracy.
problem Improving multiclass text classification accuracy using pre-trained embeddings.
method Proposed word-class embeddings (WCEs) to enhance pre-trained word embeddings.
result WCEs significantly improve multiclass text classification accuracy.
Survey on complex Monge-Ampère equations in Hermitian contexts.
problem Complex Monge-Ampère equations in Hermitian settings.
method Survey of results from many authors over 15 years.
result Overview of known results on complex Monge-Ampère equations.
This article reviews entity resolution methods and their applications.
problem Integrating information from multiple sources to clean and accurately link records.
method Clustering, semi- and fully supervised methods, canonicalization.
result Modern probabilistic record linkage has been foundational.
The abstract conjectures a formula for Higgs sheaves on complex surfaces.
problem Understanding moduli spaces of Higgs sheaves on surfaces with holomorphic 2-forms.
method Interpolates between K-theoretic Donaldson and Vafa-Witten invariants.
result Verlinde formula for moduli space of Higgs sheaves verified in examples.
We complete the classification, initiated by the second named author, of homogeneous singular Riemannian foliations of spheres that are lifts of foliations produced from Clifford systems.
Review of conformal geometry in irrational rotation algebra.
problem Understanding conformal geometry of irrational rotation algebra.
method Review of recent progress by Connes and others.
result Recent progress in understanding conformal geometry.
The author proposes a finance trading strategy named Entropy Oriented Trading and apply thermodynamics on the strategy. The state variables are chosen so that the strategy satisfies the second law of thermodynamics. Using the law, the author proves that the rate of investment (ROI) of the strategy is equal to or more t…
In this paper we investigate the differential geometric and algebro-geometric properties of the noncollapsing limit in the continuity method that was introduced by the first two named authors in \cite{LaTi14}.
Due to recent technical and scientific advances, we have a wealth of information hidden in unstructured text data such as offline/online narratives, research articles, and clinical reports. To mine these data properly, attributable to their innate ambiguity, a Word Sense Disambiguation (WSD) algorithm can avoid numbers…
Researchers prove rigidity of first conformal Steklov eigenvalue on specific shapes.
problem Rigidity of the first conformal Steklov eigenvalue on annuli and Möbius bands.
method Proof relies on uniqueness results, compactness theorem, and asymptotic control of Steklov eigenvalues.
result Rigidity of the first conformal Steklov eigenvalue on annuli and Möbius bands proved.
The study quantifies uncertainty to improve model calibration and disambiguate annotator and data bias in emotion recognition.
problem Improving model interpretability and disambiguating bias in complex tasks like emotion recognition.
method Used a modified Monte Carlo dropout approach to quantify epistemic and aleatoric uncertainty.
result Identified a significant correlation between aleatoric uncertainty and human annotator disagreement.
The paper proves a unique orbit for a specific genus 3 curve.
problem Proving the uniqueness of a closed orbit in genus 3.
method Understanding the Forni subspace and solving the jump problem.
result The Eierlegende Wollmilchsau orbit is the only one with zero Lyapunov exponent.
Constructs minimal surfaces in balls, maximizing eigenvalues.
problem Finding minimal surfaces in Euclidean balls with controlled topology.
method Maximizing the first non-trivial Steklov eigenvalue for isoperimetric problems.
result Constructs free boundary minimal immersions with controlled topology.
Extends complex manifold structures to line bundles, revealing new projective manifolds.
problem Generalizing scalar-valued holomorphic structures to line bundles.
method Study of holomorphic p-contact and s-symplectic structures on complex manifolds with line bundles. result Holomorphic p-contact and s-symplectic manifolds can be projective. GENRE retrieves entities autoregressively, improving efficiency and accuracy.
problem Retrieving entities from queries efficiently and accurately.
method Autoregressive generation of entity names, reducing memory footprint and improving context encoding.
result Significantly improved performance on entity disambiguation, linking, and retrieval tasks.
Study convolution of invariant valuations on Lie groups.
problem Understanding convolution of valuations on Lie groups.
method Explicit formula for left-invariant valuations, showing existence of smooth bi-invariant valuations, defining convolution on arbitrary Lie groups.
result Unified convolution operations on Lie groups.
We construct prime amphicheiral knots that have free period 2. This settles an open question raised by the second named author, who proved that amphicheiral hyperbolic knots cannot admit free periods and that prime amphicheiral knots cannot admit free periods of order >2.
2-dimensional knots and links are studied in the article. The notion of parity is introduced via techniques similar to the ones used by the second named author in 1-dimensional case. By using parity new invariants are constructed and known invariants are refined.
Completes classification of G2-structures on specific nilpotent Lie groups.
problem Classifying seven-dimensional nilpotent Lie groups with purely coclosed G2-structures.
method Analyzing nilpotent Lie groups of various steps and dimensions.
result Classification of indecomposable 5- and 6-step nilpotent Lie groups.
Motivated by a previous work of Zheng and the second named author, we study pinching constants of compact Kähler manifolds with positive holomorphic sectional curvature. In particular we prove a gap theorem following the work of Petersen and Tao on Riemannian manifolds with almost quarter-pinched sectional curvature.
This paper completes the classification of certain nilpotent Lie groups with specific geometric structures.
problem Classifying nilpotent Lie groups with purely coclosed G2-structures.
method Analyzing seven-dimensional nilpotent Lie groups of various steps.
result Classification of indecomposable 5- and 6-step nilpotent Lie groups with these structures.
We construct a weak 2-functor from the bicategory of oriented tangles to a bicategory of Lagrangian cospans. This functor simultaneously extends the Burau representation of the braid groups, its generalization to tangles due to Turaev and the first-named author, and the Alexander module of 1 and 2-dimensional links.