Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

1122 · Apr 201519922001200920182026
33 results for synonyms

SAFER method certifies robustness to word substitutions without model structure.

problem Certified robustness against synonymous word substitutions in NLP models.
method Randomized smoothing with stochastic ensemble of randomized inputs.
result Significantly outperforms state-of-the-art methods for certified robustness.

Custom NLP system extracts clinical data for breast cancer analysis.

problem Manual extraction of information from text-based medical records is tedious and requires specialized knowledge.
method Combines standard text mining techniques with advanced synonym detection for global analysis.
result Achieved good extraction accuracy for various concepts of interest without requiring existing corpora or ontologies.

Improves text clustering by incorporating sequential features and word embeddings.

problem Lack of sequential information and synonym handling in current text clustering methods.
method SiDPMM model that models documents as joint of bags of words, sequential features, and word embeddings.
result Significant improvement in performance and accurate inference of cluster numbers.

This work improves neural network robustness to symbol substitutions using formal verification.

problem Neural networks' vulnerability to adversarial attacks, especially under discrete text perturbations.
method Formal verification using Interval Bound Propagation on a simplex model of input perturbations.
result Models show improved verified accuracy under perturbations with formal guarantees.

As we show using the notion of equilibrium in the theory of infinite sequential games, bubbles and escalations are rational for economic and environmental agents, who believe in an infinite world. This goes against a vision of a self regulating, wise and pacific economy in equilibrium. In other words, in this context, …

2013-05-01abs ↗pdf ↗

Convolutional neural networks win SemEval-2017 for scientific relation extraction.

problem Extracting relations between scientific concepts from scholarly articles.
method Convolutional neural network model for relation extraction.
result Ranked first in SemEval-2017 Task 10 for relation extraction in scientific articles.

We seek to better understand the difference in quality of the several publicly released embeddings. We propose several tasks that help to distinguish the characteristics of different embeddings. Our evaluation of sentiment polarity and synonym/antonym relations shows that embeddings are able to capture surprisingly nua…

2013-01-15abs ↗pdf ↗

The paper analyzes word embeddings and their failure to distinguish polarized terms.

problem Word embeddings fail to correctly distinguish terms with opposite polarities.
method Mathematical analysis of word2vec model, synthetic corpus generation, empirical assessment.
result Word embeddings treat antonyms as frequentist synonyms, leading to mixed polarity terms.

Lemma on smooth maps from Azumaya/matrix manifolds to smooth manifolds lays groundwork for D-brane symplectic and calibrated geometry.

problem Understanding smooth maps from Azumaya/matrix manifolds to smooth manifolds.
method Laying down a fundamental lemma concerning the algebraicness property of smooth maps.
result Provides a starting point for synthetic (synonymous with CC^{\infty}-algebraic) symplectic and calibrated geometry.

Paper improves relation extraction in clinical texts with limited data.

problem Relation extraction in narrow knowledge domains with scarce annotated data.
method Introduces a bag-of-concepts (BoC) model and compares it with window-bounded co-occurrence (WBC).
result BoC model outperforms baseline and other complex methods on small dataset.

New method improves consistency in preference learning for neural networks.

problem Inconsistent surrogate losses in preference learning for neural networks.
method Formulated a margin-shifted ranking framework and introduced Structure-Aware HH-consistency.
result Proved superior consistency guarantees for capacity-bounded models using heavy-tailed surrogates.

New analysis shows low volatility can be unstable in financial markets.

problem Understanding the relationship between volatility and market stability.
method Using mean first hitting time as a stability indicator and comparing to standard volatility measures.
result Low volatility can be associated with higher instability in financial markets.

This paper tackles rare word problem in low-resource language pairs using NMT.

problem Rare word problem in neural machine translation, especially for low-resource languages.
method Three solutions: enhanced source context, morphology learning, and wordnet synonyms.
result Significant improvements in BLEU scores (+1.0 points) on English-Vietnamese and Japanese-Vietnamese.

Hyperbolic embeddings reduce dimensions for hierarchical data with high precision.

problem Embedding hierarchical data structures like synonym or type hierarchies efficiently.
method Combinatorial construction and hyperbolic multidimensional scaling (h-MDS) for metric spaces.
result Hyperbolic embeddings achieve high precision with few dimensions, e.g., 0.989 MAP with only 2 dimensions on WordNet.

New study shows tradeoffs between compression quality, distortion, and perception.

problem Optimizing compression for low distortion often sacrifices perceptual quality.
method Adopted Blau & Michaeli's perceptual quality definition and studied the rate-distortion-perception tradeoff.
result Restricting perceptual quality to high generally requires a trade-off between rate and distortion.

Automated author disambiguation using crowdsourced data and semi-supervised learning.

problem Grouping scientific publications by the same author, accounting for homonyms and synonyms.
method Exploits crowdsourced annotations for training an accurate classifier and clustering publications semi-supervisedly.
result Improves recall and tailors disambiguation to non-Western author names.

This paper explores how LLMs can improve pipeline-based conversational agents.

problem Limitations of pipeline-based conversational agents in human-like conversations.
method Investigated LLMs' capabilities in two phases: design and development, and operations.
result LLMs can enhance pipeline-based agents in various tasks like data generation, intent classification, and auto-correction.

In this Part II of D(11), we introduce new objects: super-CkC^k-schemes and Azumaya super-CkC^k-manifolds with a fundamental module (or, synonymously, matrix super-CkC^k-manifolds with a fundamental module), and extend the study in D(11.1) ([L-Y3], arXiv:1406.0929 [math.DG]) to define the notion of `differentiable maps…

2014-12-02abs ↗pdf ↗

The paper assesses text classification robustness through maximal safe radius computation.

problem Vulnerability of neural network models to small input modifications.
method Maximal safe radius computation, Monte Carlo Tree Search, syntactic filtering, linear bounding techniques.
result Approximation methods for computing upper and lower bounds of maximal safe radius.

Study classifies Persian speech acts for better understanding of text intent.

problem Understanding the intended function of Persian texts.
method Dictionary-based statistical technique using WordNet for SA recognition.
result Proposed method achieved state-of-the-art accuracy of 0.95 for Persian SA classification.

ExCIR provides efficient, consistent, and scalable explainability for complex models.

problem Complex models lack transparency and require efficient, stable, and scalable explainability methods.
method ExCIR uses correlation-aware feature attribution with robust centering and groupwise aggregation.
result ExCIR delivers trustworthy agreement with global baselines and full model rankings, reduces runtime, and scales to large datasets.

Proposes a method to derive knowledge graphs from EHR data.

problem Challenges in deriving generalizable knowledge from EHR data.
method Infer conditional dependency structure via a latent graphical block model (LGBM).
result Perfect recovery of block structure demonstrated.

The study predicts clinical significance of BRCA1 and BRCA2 nsSNPs using neural networks.

problem Predicting clinical significance of BRCA1 and BRCA2 nsSNPs with unknown clinical significance.
method Hydrophobicity/hydrophilicity scores, Probabilistic Neural Network (PNN), and Deep Neural Network-Stacked AutoEncoder (DNN).
result The methods achieve high prediction accuracy (87.97% for BRCA1, 82.17% for BRCA2 using PNN, and 95.41% for BRCA1, 92.80% for BRCA2 using DNN).