Efficient multimodal fusion reduces complexity and improves performance.
problem Multimodal data fusion with tensor transformations.
method Low-rank Multimodal Fusion using tensors.
result Significant reduction in computational complexity with competitive performance.
Paper proposes RMFN for multimodal language analysis.
problem Modeling interactions between language, visual, and acoustic modalities.
method Recurrent Multistage Fusion Network (RMFN) decomposes fusion into stages focusing on subsets of multimodal signals.
result RMFN achieves state-of-the-art performance across multimodal sentiment analysis, emotion recognition, and speaker traits recognition datasets.
Although highly correlated, speech and speaker recognition have been regarded as two independent tasks and studied by two communities. This is certainly not the way that people behave: we decipher both speech content and speaker traits at the same time. This paper presents a unified model to perform speech and speaker …
Privacy-preserving method protects user speech data from cloud services.
problem Privacy compromise in cloud-based speech analysis.
method Collects and sanitizes speech data before sharing, using transformation functions and voice conversion.
result Identification of sensitive emotional state reduced by ~96%.
Study uses GMM-UBM and i-vectors to assess Parkinson's patients via speech, handwriting, and gait.
problem Assessing neurological state of Parkinson's disease patients using speech, handwriting, and gait signals.
method GMM-UBM and i-vectors applied to speech, handwriting, and gait signals.
result Different feature sets from each signal are crucial for assessing Parkinson's patients.
Many complex disease syndromes such as asthma consist of a large number of highly related, rather than independent, clinical phenotypes, raising a new technical challenge in identifying genetic variations associated simultaneously with correlated traits. In this study, we propose a new statistical framework called grap…
New models extrapolate false alarms in ASV without new data.
problem Reliable extrapolation of false alarm rates in ASV without new speaker data.
method Generative models in ASV score space for arbitrary systems.
result Models accurately extrapolate false alarm rates for large speaker populations.
End-to-end speaker recognition method using neural networks.
problem Speaker and session variability in speaker verification.
method Joint Factor Analysis with tied hidden variables, MAP adaptation, two-step backpropagation.
result Improved likelihood ratios and robust performance on RSR2015 database.
This paper presents a novel approach to speaker subspace modelling based on Gaussian-Binary Restricted Boltzmann Machines (GRBM). The proposed model is based on the idea of shared factors as in the Probabilistic Linear Discriminant Analysis (PLDA). GRBM hidden layer is divided into speaker and channel factors, herein t…
Advances of modern sensing and sequencing technologies generate a deluge of high dimensional space-temporal physiological and next-generation sequencing (NGS) data. Physiological traits are observed either as continuous random functions, or on a dense grid and referred to as function-valued traits. Both physiological a…
The NL score optimizes speaker recognition tasks.
problem Improving speaker recognition accuracy.
method Established the theory of optimal scores based on normalized likelihood.
result NL score is equivalent to PLDA likelihood ratio under certain conditions.
A genome-wide association study (GWAS) correlates marker variation with trait variation in a sample of individuals. Each study subject is genotyped at a multitude of SNPs (single nucleotide polymorphisms) spanning the genome. Here we assume that subjects are unrelated and collected at random and that trait values are n…
SepIt improves speech separation for multiple speakers.
problem Improving speech separation for multiple speakers in single channel recordings.
method SepIt uses a deep neural network that iteratively improves estimates of different speakers based on mutual information.
result SepIt outperforms state-of-the-art methods for 2, 3, 5, and 10 speakers.
PerSense assesses personality traits from text for commonsense reasoning.
problem Estimating human personality traits from text for mental health analysis.
method Aggregated Probability Density Functions (PDF) and Machine Learning (ML) models.
result PerSense algorithms achieve comparable results to ground truth data, with high accuracy for personality assessment and commonsense prediction.
Model predicts personality traits from Facebook Likes.
problem Predicting personality traits from social media data.
method Mapped Facebook Likes to page categories, used machine learning.
result 83% accuracy distinguishing religious vs non-religious.
Study shows emotion affects speaker recognition and vice versa.
problem Dependencies between emotion and speaker recognition.
method Transfer learning and fine-tuning for emotion classification.
result Fine-tuning improves emotion recognition performance by 30.40% on IEMOCAP, 7.99% on MSP-Podcast, and 8.61% on Crema-D.
Paper proposes LSTM for speaker similarity measurement and improves diarization accuracy.
problem Improving speaker diarization accuracy using neural networks.
method Uses Bi-LSTM for similarity measurement and spectral clustering for clustering.
result Significantly outperforms state-of-the-art methods with diarization error rate of 6.63%.
In this paper we describe the recent advancements made in the IBM i-vector speaker recognition system for conversational speech. In particular, we identify key techniques that contribute to significant improvements in performance of our system, and quantify their contributions. The techniques include: 1) a nearest-neig…
Standard probabilistic linear discriminant analysis (PLDA) for speaker recognition assumes that the sample's features (usually, i-vectors) are given by a sum of three terms: a term that depends on the speaker identity, a term that models the within-speaker variability and is assumed independent across samples, and a fi…
Improves TTS accuracy by correcting context-dependent units.
problem Improves text-to-speech accuracy through speaker adaptation.
method Statistical model predicting context-dependent phonetic unit classes and their mean error values.
result Corrected boundaries of units improve TTS accuracy compared to HMM segmentation.
Many biological characteristics of evolutionary interest are not scalar variables but continuous functions. Here we use phylogenetic Gaussian process regression to model the evolution of simulated function-valued traits. Given function-valued data only from the tips of an evolutionary tree and utilising independent pri…
Paper proposes D3M to improve anti-spoofing detection by balancing loss function and using complementary features.
problem Improving automatic speaker verification systems against high-quality playback attacks.
method D3M uses a balanced focal loss function to dynamically scale loss based on sample traits, and combines three feature types for robust detection.
result D3M systems outperform conventional methods significantly, achieving min-tDCF of 0.0124 and EER of 0.55%.
Proposes DNN-based speaker embedding correlated with subjective inter-speaker similarity for speech synthesis.
problem Inadequate speaker representation for open speakers not in training data.
method Two training algorithms using inter-speaker similarity matrices: similarity vector embedding and similarity matrix embedding.
result Proposed algorithms learn speaker embedding highly correlated with subjective inter-speaker similarity.
A new method for speaker recognition on hyperspheres improves on PLDA's limitations.
problem Improving speaker recognition on hyperspheres with PLDA's limitations.
method Probabilistic Spherical Discriminant Analysis (PSDA) using Von Mises-Fisher distributions.
result PSDA scores are closed-form and can handle various trials, improving over PLDA.
VoiceFilter separates target speaker from multi-speaker signals.
problem Speech recognition in multi-speaker environments.
method Speaker recognition network and spectrogram masking network trained together.
result Significant reduction in speech recognition WER on multi-speaker signals.
Develops a new method to analyze brain networks for cognitive traits.
problem Challenges in summarizing and relating brain connectomes to human traits.
method Graph Auto-Encoding (GATE) model using deep learning.
result GATE improves prediction accuracy and efficiency over existing methods.
An important step in speaker verification is extracting features that best characterize the speaker voice. This paper investigates a front-end processing that aims at improving the performance of speaker verification based on the SVMs classifier, in text independent mode. This approach combines features based on conven…
Combining data from multiple speakers improves neural TTS quality, especially with imbalanced data.
problem Training high-quality TTS systems with imbalanced speaker data.
method Combine data from multiple speakers, train multi-speaker models, and use ensemble methods.
result Ensemble multi-speaker models improve synthetic speech quality for underrepresented speakers.
Improved autoencoder for F0-consistent voice conversion.
problem Non-parallel many-to-many voice conversion with prosodic information leakage.
method Conditional autoencoder with information-constraining bottlenecks.
result Controlled F0 contour and improved speech quality.
Study identifies personality traits from dance movements in music.
problem Predicting individual differences from music-induced movement.
method Identified Big Five personality traits and EQ/SQ scores from dance movements.
result Successfully explored unseen space for personality and EQ/SQ.
We present the recent advances along with an error analysis of the IBM speaker recognition system for conversational speech. Some of the key advancements that contribute to our system include: a nearest-neighbor discriminant analysis (NDA) approach (as opposed to LDA) for intersession variability compensation in the i-…
T-PSDA improves speaker recognition accuracy on toroidal submanifolds.
problem Improving speaker recognition accuracy on hypersphere embeddings.
method Extends PSDA to model within and between-speaker variabilities in toroidal submanifolds of the hypersphere.
result T-PSDA achieves accuracy on par with cosine scoring on VoxCeleb and large accuracy gains on NIST SRE'21.
Study largest maize SNP dataset for heterosis traits prediction.
problem Predicting improved biological qualities in maize hybrids.
method Developed linear and non-linear models considering hybrid relationships.
result Specially designed model efficiently predicts maize traits.
Improved speaker recognition with deep metric learning.
problem Performance gap between training and unseen speakers.
method Optimized speaker embedding model with prototypical network loss (PNL).
result Outperforms state-of-the-art models in speaker verification and identification.
New unsupervised speaker adaptation method for speech synthesis.
problem Adapting speech synthesis to new speakers with minimal data.
method Concatenating audio and text inputs, proposing new training schemes.
result Improves adaptation to unseen speakers and multi-speaker modeling.
Method distinguishes genetic correlations from causation in GWAS.
problem Identifying causal relationships among genetically correlated traits.
method Mixed fourth moments to quantify causal relationships.
result Identified 30 putative genetically causal relationships across 52 traits.
The paper proposes deep normalization to improve speaker recognition performance.
problem Non-Gaussian and non-homogeneous distributions of deep speaker vectors negatively impact speaker recognition.
method Proposes a deep normalization approach based on a novel discriminative normalization flow (DNF) model.
result DNF-based normalization delivers substantial performance gains and strong generalization capability.
End-to-end speaker verification framework reduces text dependency.
problem Improving text-independent speaker verification.
method Jointly trains SE and ASR networks with triplet loss and adversarial gradient.
result Lower equal error rate and better text-independency compared to other approaches.
Study presents MMC model for better fitting multiple choice data.
problem Improving accuracy of latent trait estimates in IRT models.
method Fit autoencoders to MMC model, demonstrating better fit than nominal response model.
result MMC model outperforms traditional IRT models in fit.
FinBERT model identifies key speakers in earnings calls, boosting stock returns.
problem Unequal impact of all speakers in earnings call transcripts on stock returns.
method Utilized FinBERT, a domain-specific transformer model, to parse transcripts and weight speakers' sentiment.
result FinBERT section-weighted sentiment generates significant long-short alpha of 2.03%.
Graph neural networks refine speaker embeddings for better session-level diarization.
problem Local speaker distinction in meeting sessions using deep embeddings.
method Graph Neural Networks (GNNs) refine speaker embeddings using session-level structural information.
result Spectral clustering on refined embeddings outperforms original embeddings significantly.
Two methods use BART to model missing data in leaf photosynthetic trait data.
problem Handling missing data in multivariate outcomes with non-ignorable mechanisms.
method Bayesian Additive Regression Trees (BART) for joint modeling of data and missingness indicators.
result Both methods effectively recover various missingness mechanisms and outperform existing approaches.
Improved neural speaker embeddings enhance ASR performance.
problem Few studies have explored neural speaker embeddings for ASR.
method Integrating improved neural speaker embeddings into a conformer-based hybrid HMM ASR system.
result Improved neural embeddings achieve on-par performance with i-vectors.
Generative x-vectors improve SV performance.
problem Improving text-independent speaker verification.
method Proposes a novel method to combine i-vectors and x-vectors using a transformation model derived from canonical correlation analysis.
result Generative x-vectors provide better performance than baseline systems, especially for long-duration utterances.
Paper improves speaker verification with federated learning and differential privacy.
problem Improving speaker verification accuracy using private data.
method Combining federated learning and differential privacy to train an auxiliary model that predicts vocal characteristics.
result 6% relative improvement in equal error rate over a baseline system.
Anonymizes speech data to protect privacy.
problem Protecting personal speech data from misuse.
method Extracts features, uses x-vectors, and neural models to synthesize anonymized speech.
result Effective in concealing speaker identities without compromising speech quality.
VoxCeleb 2019 challenge assesses speaker recognition in uncontrolled settings.
problem Evaluate speaker recognition technology in unconstrained data.
method Public dataset, challenge, and workshop at Interspeech 2019.
result Baseline results and discussions provided.
Study reveals LLM personality patterns but lacks behavioral consistency.
problem Understanding and validating personality traits in LLMs.
method Characterized LLM personality across three dimensions: training dynamics, self-report validity, and intervention effects.
result Self-reported traits do not reliably predict behavior, and instructional alignment affects trait expression but not behavior.