ATD measures language distance using neural models, recovering linguistic groupings.
problem Lack of a unified quantitative measure for cross-linguistic distance.
method Pretrained multilingual language models, attention mechanisms, optimal transport.
result ATD quantifies representational distance between languages, recovering linguistic groupings.
Study maps Spanish dialects on Twitter, revealing urban vs regional variations.
problem Understanding linguistic diversity across Spanish-speaking regions.
method Geographically tagged Twitter corpus, machine learning for dialect analysis.
result Urban dialects show international characteristics, rural dialects regional uniformity.
Paper introduces sampling-based speech synthesis with natural variation.
problem Synthetic speech lacks natural inter-utterance variation.
method Moment-matching networks trained to match moments of generated speech parameters to natural speech parameters.
result Sampling-based generation does not degrade synthetic speech quality.
Computational model uncovers linguistic universals.
problem Manual processing of linguistic typology by linguists is time-consuming and leaves key universals unexplored.
method Presented a computational model to identify known and new linguistic universals.
result The model successfully identifies known universals and uncovers new ones.
Enhances speech quality in noisy environments using symbolic sequential modeling.
problem Improving speech quality in noisy conditions.
method Incorporates symbolic sequential modeling into speech enhancement framework.
result Significant improvement in speech quality metrics (PESQ, STOI) on TIMIT dataset.
Paper proposes a method to align word embedding models in a joint latent space.
problem Challenges in aligning variations of word embedding models.
method Generative process using synthetic data points based on linguistic relationships.
result Substantial improvements in recovering embeddings of local neighborhoods.
Bayesian algorithm improves word representations using semantic taxonomy.
problem Improving word representations in semantic taxonomy.
method Bayesian Hierarchical Words Representation (BHWR) learning algorithm combining Variational Bayes and semantic taxonomy modeling.
result BHWR produces better representations for rare words.
Study shows adding noise to training data improves speech synthesis system's performance under noisy test conditions.
problem Impact of noisy linguistic features on neural network-based speech synthesis systems.
method Comparison of systems using ideal and corrupted linguistic features in training and test sets.
result Adding noise to training data can regularize the model and improve performance under noisy test conditions.
A new method extracts linguistic objects from text using CNNs.
problem Lack of interpretability in deep learning models for text.
method Weighted extension of Text Deconvolution Saliency (wTDS) measure.
result Extracts interpretable linguistic objects from text.
GTI network learns linguistic features for multi-task sequence tagging.
problem Improving neural model performance on multi-task sequence tagging without explicit features.
method GTI network with neural gate modules to learn relations between tasks.
result GTI network outperforms baselines on chunking and NER tasks.
Proposes a method to interpret linguistic data models using parse trees and least-squares scores.
problem Interpreting trained classification models in linguistic data sets.
method Assigns least-squares based importance scores to words in a sentence using syntactic constituency structure and relates them to the Banzhaf value in coalitional game theory.
result Demonstrates the effectiveness of the proposed method in aiding interpretability and diagnostics for language models.
Model predicts upcoming discourse referents using linguistic and script knowledge.
problem Predicting upcoming discourse referents based on linguistic knowledge.
method Built a computational model that predicts referents using linguistic knowledge and scripts.
result Script knowledge significantly improves model estimates of human predictions.
The paper analyzes how CNNs interpret NLP tasks and identify linguistic features.
problem Understanding how CNNs capture linguistic features in NLP tasks.
method Visualization techniques and error analysis to interpret CNNs.
result Identified how CNNs capture different linguistic features and their impact on model performance.
The starting point of this article is the question "How to retrieve fingerprints of rhythm in written texts?" We address this problem in the case of Brazilian and European Portuguese. These two dialects of Modern Portuguese share the same lexicon and most of the sentences they produce are superficially identical. Yet t…
Study investigates how simple speech sounds can form abstract categories.
problem How do abstract categories like phonemes emerge from speech exposure?
method Used modeling techniques to test Memory-Based Learning and Error-Correction Learning.
result Error-Correction Learning models can learn abstractions, identifying phone inventory and grouping.
Study shows how socioeconomic status influences language use on Twitter.
problem Global variability of linguistic patterns due to socioeconomic factors.
method Multivariate analysis of French Twitter corpus and socioeconomic data.
result People with higher socioeconomic status use more standard language.
New model combines stats and grammar, proving key property.
problem Creating a new model for linguistics and beyond.
method Introducing Markov substitute processes and proving their exponential family property.
result Markov substitute processes with a given support form an exponential family.
Linguistic calibration improves long-form text confidence.
problem LMs hallucinate, leading to suboptimal decisions.
method Defining linguistic calibration, training framework, reinforcement learning.
result Llama 2 7B is significantly more calibrated than baselines.
Model learns disentangled, interpretable representations from sequential data without supervision.
problem Learning disentangled and interpretable representations from sequential data without supervision.
method Factorized hierarchical variational autoencoder with multi-scale priors.
result Model outperforms i-vector baseline in speaker verification and reduces word error rate by 35% in mismatched scenarios.
The abstract discusses parallels between Galois theory and Stone-Weierstrass theorem in various fields.
problem Connecting distinguishing power and expressive power in different fields.
method Elementary theorem connecting distinguishing power and expressive power.
result Foundational principle in linguistics linking distinguishing power and expressive power.
BERT captures linguistic features in separate semantic and syntactic subspaces.
problem Understanding how transformer models like BERT represent linguistic features internally.
method Qualitative and quantitative investigations of BERT's internal representations.
result Evidence of a fine-grained geometric representation of word senses and syntactic representations.
Study evaluates natural language models' ability to generalize across tasks.
problem Natural language models struggle with generalizing to new tasks.
method Empirical evaluation of state-of-the-art models using new metrics.
result Models require extensive in-domain training and are prone to forgetting.
ContextBench benchmarks methods for generating linguistically fluent inputs that activate specific latent features in language models.
problem Identifying inputs that trigger specific behaviours or latent features in language models.
method Context modification and benchmarking methods like Evolutionary Prompt Optimisation (EPO) with LLM-assistance and diffusion model inpainting.
result Enhanced methods achieve state-of-the-art performance in balancing elicitation effectiveness and fluency.
A new TTS method uses diffusion and VAE for better speech synthesis.
problem Improving text-to-speech synthesis for better speech quality and robustness.
method Combines diffusion probabilistic model and variational autoencoder for latent variable conversion.
result The method is robust to poor orthography and alignment errors.
Investigates neural TTS systems for Japanese and English.
problem Improving neural TTS systems for high-quality speech synthesis.
method Comparative study of neural sequence-to-sequence TTS vs. DNN pipeline TTS, varying model architecture, parameter size, and language.
result A neural sequence-to-sequence TTS system requires sufficient model parameters and a powerful encoder for high-quality speech synthesis.
A fast kernel-based measure for sparse linguistic expressions.
problem Efficiently measuring co-occurrence in sparse linguistic data.
method Derives PHSIC from HSIC, estimates it linearly, and uses various kernels.
result Empirically, PHSIC outperforms PMI in accuracy and learning speed.
Paper tackles deconfounding age effects in dementia detection models.
problem Dementia detection models are affected by age, leading to potential non-generalizable accuracies.
method Proposes fair representation learning to learn age-invariant representations.
result Best models compromise accuracy by only 2.56% and 1.54% on clinical datasets.
Predicts whether online discussions will be productive or not.
problem Identifying wasteful discussions in group settings.
method Analyze conversational dynamics and linguistic cues.
result Linguistic cues and conversational patterns predict discussion productivity.
Expanding spoken language understanding to handle complex entities and intents.
problem Handling compound entities and intents in spoken language understanding.
method Introducing a domain-agnostic shallow parser that handles linguistic coordination, learning domain-independent and slot-independent features.
result The model learns to segment conjunct boundaries of various phrasal categories and improves generalization across different slot types using adversarial training.
Qwant Research improves clinical case matching and information retrieval.
problem Matching and retrieving relevant clinical cases and discussions.
method Approach based on language models and preprocessings, information extraction system using neural networks and linguistic analysis.
result Very encouraging results in information extraction accuracy.
Paper proposes MMD-Sense-Analysis for detecting word sense shifts.
problem Detecting and interpreting shifts in word meanings over time.
method Leverages Maximum Mean Discrepancy (MMD) to identify and explain word sense changes.
result Demonstrates effectiveness of MMD-Sense-Analysis through empirical results.
Interpersonal relations are fickle, with close friendships often dissolving into enmity. In this work, we explore linguistic cues that presage such transitions by studying dyadic interactions in an online strategy game where players form alliances and break those alliances through betrayal. We characterize friendships …
Analysts use vague language in reports to convey useful information about future payoffs.
problem Lack of precise numerical forecasts in analyst reports.
method Empirical analysis of analyst reports to assess the predictive power of linguistic tone.
result The textual tone of analyst reports has predictive power for forecast errors and subsequent revisions, especially when language is vague and uncertainty is high.
Study shows mutual information can reward structure learning agents without expert systems.
problem Designing rewards for structure learning agents in natural language environments.
method Revisited Information Theory of unsupervised induction of phrase-structure grammars, using random sets of linguistic samples.
result Empirical evidence that simulated semantic structures can be distinguished from random ones by mutual information among their constituents.
Study uses social media to analyze COVID-19 impact.
problem Understanding the global impact of COVID-19.
method Machine learning and linguistic tools to analyze social media posts.
result Automatic detection of positive reports of COVID-19.
Analyzes Twitter users' opinions on self-driving cars.
problem Understanding public perception of self-driving cars.
method Annotated Twitter dataset, topic modeling, sentiment classification using Twitter features.
result People are generally optimistic but also concerned about self-driving cars.
We present the Bayesian Echo Chamber, a new Bayesian generative model for social interaction data. By modeling the evolution of people's language usage over time, this model discovers latent influence relationships between them. Unlike previous work on inferring influence, which has primarily focused on simple temporal…
Proposes a new voice conversion model that preserves pitch patterns.
problem Preserving pitch patterns while changing speaker identity.
method Variational-autoencoder-based model with an auxiliary network.
result Ensures the conversion result correctly reflects specified F0/timbre information.
This work decouples language from math problems to enable cross-language learning.
problem Current machine learning representations are language dependent.
method Inspired by linguistics, the work learns language agnostic representations.
result Models trained on one language achieve similar accuracies in other languages.
OT domain adaptation improves aphasia detection across languages.
problem Detecting aphasia in low-resource languages with limited data.
method Utilized OT domain adaptation to map linguistic features across multiple languages.
result OT domain adaptation significantly improved F1 scores for French and Mandarin aphasia detection.
Preserves content while changing style in text.
problem Text style transfer often fails to preserve content.
method Leverages linguistic information in structured supervisions.
result Significant improvement in content preservation and style transfer.
EigenNoise provides a competitive word vector initialization scheme without pre-training data.
problem Improving word vector initialization without pre-training data.
method EigenNoise uses a dense, independent co-occurrence model to initialize word vectors.
result EigenNoise can approach GloVe performance without pre-training data.
New RL framework learns task completion without prior knowledge.
problem Learning task completion without linguistic or perceptual knowledge.
method Sequentially imagining visual goals and choosing actions.
result Framework outperforms flat and hierarchical architectures.
Paper predicts Indian stocks using news psycholinguistic features.
problem Predicting Indian stock market performance using financial news.
method Hybrid intelligent models using psycholinguistic variables (LIWC and TAALES) from news articles.
result GMDH and GRNN are statistically the best techniques for prediction.
New distress dictionary improves bankruptcy prediction from disclosure text.
problem Bankruptcy prediction from financial disclosures.
method Proposes a distress dictionary based on managers' sentences, quantifies linguistic features, and builds predictive models.
result Predictive models based on the distress dictionary outperform existing methods.
Large language models predict human sensory judgments across multiple modalities.
problem Determining the extent of perceptual information in language.
method State-of-the-art large language models were used to predict sensory judgments across six psychophysical datasets.
result Large language models can predict human sensory judgments across multiple modalities with significant correlation to human data.
This paper analyzes text in financial disclosures to improve financial analysis.
problem Insufficient analysis of unstructured text in financial disclosures.
method Reviews and explores methods in computational linguistics and NLP.
result Highlights limitations of sentiment metrics and suggests future research areas.
This paper proposes a new method to connect language and physical actions in reinforcement learning.
problem Connecting linguistic representations to the physical world in embodied agents.
method Language-conditioned goal generators to decouple sensorimotor learning from language acquisition.
result Agents can demonstrate a diversity of behaviors for any given instruction.