Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

2795588361,115 · Jun 202019922001200920182026
48 results for Monolingual Data

Extracts high-quality monolingual datasets from web crawl data.

problem Improving text representation quality through larger corpora.
method Automated pipeline using deduplication and language identification, augmented with filtering for high-quality documents.
result Extracted massive high-quality monolingual datasets from Common Crawl.

CL methods improve monolingual ASR models across new tasks without forgetting past data.

problem Catastrophic Forgetting in monolingual ASR models when adapting to new domains or accents.
method Implement and compare various Continual Learning methods for monolingual ASR.
result Best CL method reduces performance gap by over 40% with minimal past data.

Improved multilingual speech recognition with low latency for nine Indic languages.

problem Imbalance in training data across languages and low latency in interactive applications.
method Conditioning on a language vector and training language-specific adapter layers.
result Lower word error rate than monolingual E2E models and conventional systems.

This study improves NMT using reinforcement learning, overcoming its instability.

problem Stability issues in reinforcement learning for neural machine translation.
method Systematic study on reinforcement learning factors and a new method for monolingual data.
result Competitive results on WMT17 Chinese-English translation task, setting a state-of-the-art performance.

Forward translation improves neural machine translation for sentences originally in source language.

problem Improving neural machine translation quality using synthetic data.
method Case study with French-English news translation, separating test sets by original language, analyzing domains, translationese, and noise.
result Forward translation delivers superior gains on sentences originally in source language, complementing back-translation on target language sentences.

Geometric approach learns bilingual mappings from monolingual embeddings.

problem Bilingual lexicon induction and cross-lingual word similarity.
method Decouples learning into rotations and a metric, modeled as optimization on Riemannian manifolds.
result Outperforms previous approaches on bilingual lexicon induction and cross-lingual word similarity tasks.

Introduces challenges and techniques for creating machine translation for indigenous languages.

problem Limited data for machine translation of indigenous languages.
method Introduction to challenges, concepts, and techniques for creating MT systems.
result Discussion of recent advances and open questions in NLP for these languages.

Improved unsupervised word translation using adversarial autoencoder with cycle consistency and input reconstruction.

problem Challenging language pairs and lack of parallel data for unsupervised word translation.
method Adversarial autoencoder with cycle consistency and input reconstruction regularization.
result More stable and better performance than recent approaches.

Improves cross-lingual NER by projecting entities from one language to another.

problem Limited annotated corpora for named entity recognition in many languages.
method Uses machine translation twice: first for sentences, then for entities; matches based on ortho- and phonetic similarity; identifies matches using distributional statistics.
result Improves cross-lingual NER by an average of 4.1 points on 5 diverse languages.

A simple modification enables a universal NMT model with language-specific parameters.

problem Creating a universal NMT model that can adapt to different languages and domains.
method Introducing a contextual parameter generator (CPG) that dynamically adjusts model parameters based on source and target language embeddings.
result The system achieves state-of-the-art performance and zero-shot translation, demonstrating the effectiveness of the CPG.

Unsupervised MT struggles with morphologically rich languages.

problem Limitations of unsupervised machine translation on morphologically rich languages.
method Adversarial unsupervised alignment of word embedding spaces for bilingual dictionary induction.
result A simple trick exploiting weak supervision from identical words improves unsupervised bilingual dictionary induction performance.

Simple framework decouples word alignment and multilingual embedding mapping.

problem Learning multilingual embeddings without supervision.
method Two-stage approach: 1) unsupervised word alignment, 2) mapping embeddings to shared space.
result Robust performance across various multilingual tasks, including distant languages.

End-to-end neural network improves QbE-STD in multilingual speech search.

problem Query by example spoken term detection in zero-resource scenarios.
method Use of multilingual bottleneck features and CNN for pattern matching, integrated into a fully neural network framework.
result CNN-based matching outperforms DTW-based matching using bottleneck features.

GIRNet tackles multi-tasking with mixed domain sequences, improving sentiment and tagging tasks.

problem Labeling sequences with mixed domain data and position-specific inference.
method Unified position-sensitive multi-task RNN architecture with gated state sequences from auxiliary data.
result GIRNet achieves new state-of-the-art performance in sentiment classification, POS tagging, and target position-sensitive annotation.

This thesis evaluates text-based vs audio-based classification of mental health interviews.

problem Classifying psychiatric illness using text-based methods.
method Design and evaluate a text classification network on mental health interviews, using belabBERT.
result Text-based classification is a strong alternative to audio-based methods.

This paper uses LLMs and cycle consistency for better machine translation evaluation.

problem Evaluating translation quality and LLM capabilities without ground truth.
method Generate translation candidates, back-translate, and evaluate cycle consistency.
result Larger LLMs or more inference passes improve cycle consistency.

Prevents sensitive data generation in diffusion models using labeled and unlabeled data.

problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.

Study reveals Data Shapley's inconsistent performance in data selection tasks.

problem Inconsistency of Data Shapley's performance in data selection across different settings.
method Hypothesis testing framework and identification of utility functions.
result Data Shapley's performance is no better than random selection without specific constraints.

PRRO generates synthetic tabular data that improves SL performance and class distribution.

problem Low SL utility of synthetic data due to class imbalance and overlooked data relationships.
method Data pruning and column reordering to optimize SL utility.
result Synthetic data generated with PRRO enhances predictive performance and class distribution.

Defines data science as a natural ecosystem with challenges and missions.

problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.

Data science redefines causal inference from observational data, classifying tasks into description, prediction, and counterfactual prediction.

problem Widespread misunderstandings about data science's role in causal inference from observational data.
method Organizing data science tasks into three classes: Description, prediction, and counterfactual prediction (including causal inference).
result The necessity of subject-matter expert knowledge for causal analyses in data science.

This paper evaluates how dirty data affects data mining and machine learning results.

problem Negative impacts of dirty data on data mining and machine learning results.
method Experimental comparison of missing, inconsistent, and conflicting data on classification and clustering algorithms.
result Guidelines for algorithm selection and data cleaning based on experimental findings.

This paper introduces C-DSL to improve data mining outcomes by considering context.

problem Data collection ambiguities, data imbalance, hidden biases, lack of domain info, and data incompleteness.
method Developed Context-Driven Data Science Lifecycle (C-DSL) to address data quality issues.
result Tangible improvements to data mining outcomes were achieved through C-DSL.

Proposes using probabilistic models for privacy-preserving synthetic data.

problem Designing high-quality synthetic data for privacy preservation.
method Formulate the problem through probabilistic modelling, choosing a model for the data.
result Statistical discoveries can be reliably reproduced from synthetic data.

Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.

problem Ensuring fair predictions across sensitive attributes in synthetic data.
method Equalizing target probability distributions across sensitive attributes in synthetic data generation.
result Synthetic data provides strong fair predictions, equal across all thresholds.

A new method classifies multiple correlated data streams simultaneously.

problem Classifying multiple correlated data streams in practical scenarios.
method Double-Coupling Support Vector Machines (DC-SVM) considers both internal and external correlations.
result The proposed method outperforms traditional methods on artificial and real-world data streams.

This paper improves neural machine translation training by selecting and denoising data.

problem Reduces negative impact of noisy data on neural machine translation training.
method Measures and selects domain data, applies denoising curriculum using online data selection.
result Significant effectiveness for training on noisy data.