Enhances BLEU for better SMT evaluation.
problem Improving BLEU for more human-like evaluation of machine translations.
method Adapts BLEU to consider synonyms, word order, and style variations in human translations.
result Improves SMT evaluation metrics and correlates with human translation quality.
SAFER method certifies robustness to word substitutions without model structure.
problem Certified robustness against synonymous word substitutions in NLP models.
method Randomized smoothing with stochastic ensemble of randomized inputs.
result Significantly outperforms state-of-the-art methods for certified robustness.
Custom NLP system extracts clinical data for breast cancer analysis.
problem Manual extraction of information from text-based medical records is tedious and requires specialized knowledge.
method Combines standard text mining techniques with advanced synonym detection for global analysis.
result Achieved good extraction accuracy for various concepts of interest without requiring existing corpora or ontologies.
Isothermic parameterizations} are synonyms of isothermal curvature line parameterizations, for surfaces immersed in Euclidean spaces. We provide a method of constructing isothermic coordinate charts on surfaces which admit them, starting from an arbitrary chart. One of the primary applications of this work consists of …
Improves text clustering by incorporating sequential features and word embeddings.
problem Lack of sequential information and synonym handling in current text clustering methods.
method SiDPMM model that models documents as joint of bags of words, sequential features, and word embeddings.
result Significant improvement in performance and accurate inference of cluster numbers.
This work improves neural network robustness to symbol substitutions using formal verification.
problem Neural networks' vulnerability to adversarial attacks, especially under discrete text perturbations.
method Formal verification using Interval Bound Propagation on a simplex model of input perturbations.
result Models show improved verified accuracy under perturbations with formal guarantees.
As we show using the notion of equilibrium in the theory of infinite sequential games, bubbles and escalations are rational for economic and environmental agents, who believe in an infinite world. This goes against a vision of a self regulating, wise and pacific economy in equilibrium. In other words, in this context, …
Proposes BU-SPO method to improve text classification robustness.
problem Vulnerability of deep models in text classification.
method Bigram and unigram based adaptive Semantic Preservation Optimization (BU-SPO) method.
result Achieves highest attack success rates and semantic similarity by changing the smallest number of words.
Convolutional neural networks win SemEval-2017 for scientific relation extraction.
problem Extracting relations between scientific concepts from scholarly articles.
method Convolutional neural network model for relation extraction.
result Ranked first in SemEval-2017 Task 10 for relation extraction in scientific articles.
We seek to better understand the difference in quality of the several publicly released embeddings. We propose several tasks that help to distinguish the characteristics of different embeddings. Our evaluation of sentiment polarity and synonym/antonym relations shows that embeddings are able to capture surprisingly nua…
The paper analyzes word embeddings and their failure to distinguish polarized terms.
problem Word embeddings fail to correctly distinguish terms with opposite polarities.
method Mathematical analysis of word2vec model, synthetic corpus generation, empirical assessment.
result Word embeddings treat antonyms as frequentist synonyms, leading to mixed polarity terms.
Project classifies Hinglish social content on platforms like Twitter, Reddit.
problem Classifying abusive and hate-inducing content in Hinglish on social media.
method Used deep learning with bi-directional sequence models and text augmentation techniques.
result Produced a state-of-the-art classifier that outperforms previous work.
Lemma on smooth maps from Azumaya/matrix manifolds to smooth manifolds lays groundwork for D-brane symplectic and calibrated geometry.
problem Understanding smooth maps from Azumaya/matrix manifolds to smooth manifolds.
method Laying down a fundamental lemma concerning the algebraicness property of smooth maps.
result Provides a starting point for synthetic (synonymous with C∞-algebraic) symplectic and calibrated geometry. Develops techniques to align word embeddings from different sources.
problem Aligning word embeddings from diverse datasets or methods.
method Simple closed-form techniques for optimal rotation, translation, and scaling.
result Maximizes cosine similarity and minimizes root mean squared errors.
This paper introduces the distinction between aleatoric and epistemic uncertainty in machine learning.
problem The need to distinguish between types of uncertainty in machine learning.
method Introduction and overview of existing methods.
result The distinction between aleatoric and epistemic uncertainty.
Paper improves relation extraction in clinical texts with limited data.
problem Relation extraction in narrow knowledge domains with scarce annotated data.
method Introduces a bag-of-concepts (BoC) model and compares it with window-bounded co-occurrence (WBC).
result BoC model outperforms baseline and other complex methods on small dataset.
Sparse PCA selects variables with FDR control for improved performance.
problem Sparse PCA selects irrelevant variables when maximizing explained variance.
method Proposes FDR-controlled selection using T-Rex selector.
result Significant performance improvement over traditional sparse PCA.
New method improves consistency in preference learning for neural networks.
problem Inconsistent surrogate losses in preference learning for neural networks.
method Formulated a margin-shifted ranking framework and introduced Structure-Aware H-consistency. result Proved superior consistency guarantees for capacity-bounded models using heavy-tailed surrogates.
New analysis shows low volatility can be unstable in financial markets.
problem Understanding the relationship between volatility and market stability.
method Using mean first hitting time as a stability indicator and comparing to standard volatility measures.
result Low volatility can be associated with higher instability in financial markets.
This paper tackles rare word problem in low-resource language pairs using NMT.
problem Rare word problem in neural machine translation, especially for low-resource languages.
method Three solutions: enhanced source context, morphology learning, and wordnet synonyms.
result Significant improvements in BLEU scores (+1.0 points) on English-Vietnamese and Japanese-Vietnamese.
Hyperbolic embeddings reduce dimensions for hierarchical data with high precision.
problem Embedding hierarchical data structures like synonym or type hierarchies efficiently.
method Combinatorial construction and hyperbolic multidimensional scaling (h-MDS) for metric spaces.
result Hyperbolic embeddings achieve high precision with few dimensions, e.g., 0.989 MAP with only 2 dimensions on WordNet.
New method improves KB completion by predicting unseen relations.
problem KB completion with improved multi-hop inference.
method Recursive neural network (RNN) for composing multi-hop relations.
result Improves KB completion by 11% over traditional methods.
Proposes a method for imputing missing data and estimating class uncertainties.
problem Missing data and uncertainties in target class assignments.
method Generative imputation and stochastic prediction using generator, discriminator, and predictor networks.
result Effectiveness in generating imputations and estimating class uncertainties.
New study shows tradeoffs between compression quality, distortion, and perception.
problem Optimizing compression for low distortion often sacrifices perceptual quality.
method Adopted Blau & Michaeli's perceptual quality definition and studied the rate-distortion-perception tradeoff.
result Restricting perceptual quality to high generally requires a trade-off between rate and distortion.
Automated author disambiguation using crowdsourced data and semi-supervised learning.
problem Grouping scientific publications by the same author, accounting for homonyms and synonyms.
method Exploits crowdsourced annotations for training an accurate classifier and clustering publications semi-supervisedly.
result Improves recall and tailors disambiguation to non-Western author names.
This paper explores how LLMs can improve pipeline-based conversational agents.
problem Limitations of pipeline-based conversational agents in human-like conversations.
method Investigated LLMs' capabilities in two phases: design and development, and operations.
result LLMs can enhance pipeline-based agents in various tasks like data generation, intent classification, and auto-correction.
In this Part II of D(11), we introduce new objects: super-Ck-schemes and Azumaya super-Ck-manifolds with a fundamental module (or, synonymously, matrix super-Ck-manifolds with a fundamental module), and extend the study in D(11.1) ([L-Y3], arXiv:1406.0929 [math.DG]) to define the notion of `differentiable maps…
The paper assesses text classification robustness through maximal safe radius computation.
problem Vulnerability of neural network models to small input modifications.
method Maximal safe radius computation, Monte Carlo Tree Search, syntactic filtering, linear bounding techniques.
result Approximation methods for computing upper and lower bounds of maximal safe radius.
Study classifies Persian speech acts for better understanding of text intent.
problem Understanding the intended function of Persian texts.
method Dictionary-based statistical technique using WordNet for SA recognition.
result Proposed method achieved state-of-the-art accuracy of 0.95 for Persian SA classification.
ExCIR provides efficient, consistent, and scalable explainability for complex models.
problem Complex models lack transparency and require efficient, stable, and scalable explainability methods.
method ExCIR uses correlation-aware feature attribution with robust centering and groupwise aggregation.
result ExCIR delivers trustworthy agreement with global baselines and full model rankings, reduces runtime, and scales to large datasets.
Tests neural networks for compositional generalization in language.
problem Understanding how neural networks generalize in a compositional manner.
method Developed five tests bridging linguistic and philosophical compositionality with neural models.
result Revealed strengths and weaknesses of different neural architectures.
Proposes a method to derive knowledge graphs from EHR data.
problem Challenges in deriving generalizable knowledge from EHR data.
method Infer conditional dependency structure via a latent graphical block model (LGBM).
result Perfect recovery of block structure demonstrated.
The study predicts clinical significance of BRCA1 and BRCA2 nsSNPs using neural networks.
problem Predicting clinical significance of BRCA1 and BRCA2 nsSNPs with unknown clinical significance.
method Hydrophobicity/hydrophilicity scores, Probabilistic Neural Network (PNN), and Deep Neural Network-Stacked AutoEncoder (DNN).
result The methods achieve high prediction accuracy (87.97% for BRCA1, 82.17% for BRCA2 using PNN, and 95.41% for BRCA1, 92.80% for BRCA2 using DNN).