TristouNet improves speaker comparison using neural networks and triplet loss.
problem Speaker comparison and change detection in short speech turns.
method Triplet loss for training neural network to project speech sequences into fixed-dimensional space.
result Significant improvements over state-of-the-art techniques for speaker comparison and change detection.
Bayesian approach improves speech recognition with limited speaker data.
problem Reduces mismatch between training and evaluation data due to speaker differences.
method Bayesian learning for DNN adaptation models with limited speaker data.
result Bayesian adaptation consistently outperforms deterministic methods, reducing word error rates up to 1.4%.
Seq2seq ASR adapts to speakers, improving performance by 25%.
problem Speaker adaptation for seq2seq ASR systems to match conventional methods.
method Applied Kullback-Leibler divergence and Linear Hidden Network adaptation to seq2seq models.
result 25% relative word error rate improvement with seq2seq model adaptation.
Study compares metric learning loss functions for speaker verification.
problem Comparing metric learning loss functions for end-to-end speaker verification.
method Cross entropy loss, cosine loss, angular margin loss, center loss, contrastive loss, triplet loss.
result Additive angular margin loss outperforms other loss functions.
The study compares adversarial and multi-task learning for speech recognition, finding invariant representations are key.
problem Improving speech recognition performance with speaker information.
method Investigated multi-task learning and adversarial learning for speech recognition, comparing their effects on error rates.
result Deep models already develop speaker-invariant representations, and adversarial learning has a minor impact.
GPU acceleration speeds up i-vector extraction 3000x, enabling new research.
problem Speeding up i-vector extraction for speaker verification.
method GPU acceleration for i-vector extraction, including re-computing UBM and frame alignments.
result Significant speed-up allows rigorous study of i-vector variations.
Paper compares different spoofing detection methods for speech verification.
problem Detecting audio replay attacks in speech verification systems.
method GMM based methods, high level features extraction with simple classifier, deep learning frameworks.
result Deep learning approaches are efficient in changing acoustic conditions.
Proposes DNN-based speaker embedding correlated with subjective inter-speaker similarity for speech synthesis.
problem Inadequate speaker representation for open speakers not in training data.
method Two training algorithms using inter-speaker similarity matrices: similarity vector embedding and similarity matrix embedding.
result Proposed algorithms learn speaker embedding highly correlated with subjective inter-speaker similarity.
VoiceFilter separates target speaker from multi-speaker signals.
problem Speech recognition in multi-speaker environments.
method Speaker recognition network and spectrogram masking network trained together.
result Significant reduction in speech recognition WER on multi-speaker signals.
Combining data from multiple speakers improves neural TTS quality, especially with imbalanced data.
problem Training high-quality TTS systems with imbalanced speaker data.
method Combine data from multiple speakers, train multi-speaker models, and use ensemble methods.
result Ensemble multi-speaker models improve synthetic speech quality for underrepresented speakers.
Meta-embeddings improve PLDA model's accuracy for speaker recognition.
problem Improving PLDA model's accuracy for speaker recognition.
method Proposed Gaussian meta-embeddings (GMEs) for heavy-tailed PLDA models.
result GMEs with variable precisions propagate uncertainty, leading to up to 20% more accurate results.
Improved speaker recognition with deep metric learning.
problem Performance gap between training and unseen speakers.
method Optimized speaker embedding model with prototypical network loss (PNL).
result Outperforms state-of-the-art models in speaker verification and identification.
New unsupervised speaker adaptation method for speech synthesis.
problem Adapting speech synthesis to new speakers with minimal data.
method Concatenating audio and text inputs, proposing new training schemes.
result Improves adaptation to unseen speakers and multi-speaker modeling.
The paper proposes deep normalization to improve speaker recognition performance.
problem Non-Gaussian and non-homogeneous distributions of deep speaker vectors negatively impact speaker recognition.
method Proposes a deep normalization approach based on a novel discriminative normalization flow (DNF) model.
result DNF-based normalization delivers substantial performance gains and strong generalization capability.
End-to-end speaker verification framework reduces text dependency.
problem Improving text-independent speaker verification.
method Jointly trains SE and ASR networks with triplet loss and adversarial gradient.
result Lower equal error rate and better text-independency compared to other approaches.
Graph neural networks refine speaker embeddings for better session-level diarization.
problem Local speaker distinction in meeting sessions using deep embeddings.
method Graph Neural Networks (GNNs) refine speaker embeddings using session-level structural information.
result Spectral clustering on refined embeddings outperforms original embeddings significantly.
Improved neural speaker embeddings enhance ASR performance.
problem Few studies have explored neural speaker embeddings for ASR.
method Integrating improved neural speaker embeddings into a conformer-based hybrid HMM ASR system.
result Improved neural embeddings achieve on-par performance with i-vectors.
End-to-end system improves speaker verification using attention mechanism.
problem Improving text-dependent speaker verification accuracy.
method Speaker discriminative CNNs extract features, attention mechanism combines them, end-to-end training optimizes system.
result The proposed system achieves better performance on Windows 10 speaker verification task.
Paper improves speaker verification with federated learning and differential privacy.
problem Improving speaker verification accuracy using private data.
method Combining federated learning and differential privacy to train an auxiliary model that predicts vocal characteristics.
result 6% relative improvement in equal error rate over a baseline system.
Anonymizes speech data to protect privacy.
problem Protecting personal speech data from misuse.
method Extracts features, uses x-vectors, and neural models to synthesize anonymized speech.
result Effective in concealing speaker identities without compromising speech quality.
VoxCeleb 2019 challenge assesses speaker recognition in uncontrolled settings.
problem Evaluate speaker recognition technology in unconstrained data.
method Public dataset, challenge, and workshop at Interspeech 2019.
result Baseline results and discussions provided.
Paper explores using EEG for better speaker identification, even in noisy environments.
problem Speaker identification performance degrades in background noise.
method Uses EEG signals to enhance speaker identification systems, comparing with acoustic features.
result Speaker identification system using only EEG features outperforms one using only acoustic features in high background noise.
Personal VAD detects target speaker voice activity efficiently.
problem Efficiently detect target speaker voice activity for reduced computational cost and battery usage.
method Trains a neural network conditioned on speaker embedding or verification score, outputs probabilities for three speech classes.
result Trained model with 130K parameters outperforms combined standard VAD and speaker recognition networks.
Improved speaker diarization with LSTM and d-vectors.
problem Improving speaker diarization performance.
method Combining LSTM-based d-vector audio embeddings with non-parametric clustering.
result Achieved state-of-the-art diarization error rate of 12.0% on NIST SRE 2000 CALLHOME.
New models extrapolate false alarms in ASV without new data.
problem Reliable extrapolation of false alarm rates in ASV without new speaker data.
method Generative models in ASV score space for arbitrary systems.
result Models accurately extrapolate false alarm rates for large speaker populations.
Unified framework for speaker-adaptive models using scaling and bias codes.
problem Improving speaker-adaptive neural-network based speech synthesis systems.
method Unified framework representation and generalized scaling and bias codes.
result Improves performance of speaker adaptation compared to conventional methods.
Improves speaker verification for variable-duration utterances using a feature pyramid module.
problem Improving robustness for variable-duration utterances in speaker verification.
method Integrates a feature pyramid module into multi-scale aggregation to enhance speaker-discriminative information from multiple layers.
result Improves performance for both short and long utterances compared to state-of-the-art approaches.
Paper proposes UIS-RNN for fully supervised speaker diarization.
problem Speaker diarization with unknown number of speakers and time-stamped labels.
method Parameter-sharing RNNs with interleaved states, ddCRP for clustering.
result 7.6% diarization error rate on NIST SRE 2000 CALLHOME.
The paper uses neural networks to convert one speaker's voice to another.
problem Identifying speakers uniquely based on their voice.
method Uses convolutional neural networks to manipulate pitch and timbre.
result Preliminary results show encouraging voice conversion.
Proposes SPE for robust speaker verification.
problem Improving text-independent speaker verification accuracy.
method Spatial pyramid encoding and deep length normalization.
result Proposed system outperforms i-vector and d-vector baselines.
EEG signals enhance speaker verification system robustness.
problem Improving speaker verification in noisy environments.
method Used end-to-end deep learning model with EEG and speech features.
result EEG signals improve speaker verification robustness, especially in noisy conditions.
BOFFIN TTS optimizes hyper-parameters for new speaker adaptation.
problem Fine-tuning a pre-trained TTS model for a new speaker with limited data.
method Bayesian optimization to efficiently find optimal hyper-parameters.
result Average 30% improvement in speaker similarity over standard techniques.
New method separates multiple voices in mixed audio.
problem Separating multiple simultaneous speakers in audio.
method Gated neural networks trained at multiple steps, selecting actual number of speakers.
result Outperforms current state of the art for more than two speakers.
Improved multi-user VoiceFilter-Lite model for speech recognition.
problem Limited performance of multi-user VoiceFilter-Lite models.
method Dual learning rate schedule and FiLM for feature conditioning.
result Closed the performance gap between multi-user and single-user VoiceFilter-Lite models.
Improved deep neural networks for text-independent speaker recognition.
problem Text-independent speaker recognition using deep neural networks.
method Angular softmax activation, residual frame level connections, cosine similarity, discriminative similarity metric learning.
result Improved speaker recognition accuracy on real-life conditions.
Improved far-field speaker verification for short utterances in noisy conditions.
problem Challenges in speaker verification on short utterances in uncontrolled noisy environments.
method Used deep neural network architectures (TDNN and ResNet) and experimented with various embedding extractors and training procedures.
result ResNet architectures outperform x-vector approach in speaker verification quality for both long and short utterances.
CNN trained with noise improves multi-speaker localization accuracy.
problem Multi-speaker localization in noisy environments.
method Convolutional Neural Network (CNN) trained with synthesized noise.
result The CNN-based method outperforms the steered response power method.
One-shot VC model converts voices without parallel data.
problem Limited VC applicability due to training data restriction.
method Disentangles speaker and content representations with instance normalization.
result Model converts voices from unseen speakers with high similarity.
Meta-learning approach for adaptive TTS with few data.
problem Adapting TTS systems to new speakers with minimal data.
method Meta-learning with shared WaveNet core and independent speaker embeddings, using three training strategies.
result Successful adaptation of multi-speaker neural network to new speakers with minimal data.
Study shows emotion affects speaker recognition and vice versa.
problem Dependencies between emotion and speaker recognition.
method Transfer learning and fine-tuning for emotion classification.
result Fine-tuning improves emotion recognition performance by 30.40% on IEMOCAP, 7.99% on MSP-Podcast, and 8.61% on Crema-D.
Adversarial ASV improves speaker verification robustness.
problem Mismatches in training, enrollment, and test conditions degrade deep speaker embeddings.
method Adversarial multi-task training to learn condition-invariant embeddings.
result 8.8% and 14.5% relative EER improvements for known and unknown conditions.
Algorithm separates multiple speakers from a single audio input.
problem Separating multiple speakers from a single microphone input.
method Deep recurrent neural networks regression to a speaker characteristic vector space.
result Better empirical performance compared to other techniques.
Method converts facial expressions and voice of a source speaker into a target speaker.
problem Separate conversion of facial and acoustic features leads to unnatural results.
method Uses three neural networks: conversion, waveform generation, and image reconstruction.
result Significantly higher naturalness achieved when converting both features together.
Speech enhancement improved by adapting to unknown speakers without auxiliary signals.
problem Improving speech enhancement accuracy for unknown speakers.
method Adopting multi-task learning for speech enhancement and speaker identification, using multi-head self-attention.
result Achieved state-of-the-art performance and improved subjective quality.
Study speaker verification security using hierarchical Bayesian modeling.
problem Estimating false alarm rate in ASV systems for large speaker databases.
method Hierarchical Bayesian modeling of ASV scores to assess security against closest impostors.
result Neither i-vector nor x-vector systems are secure against increased impostor database size.
Study shows neural networks outperform traditional methods in speaker identification.
problem Open-set speaker identification with large populations.
method Discriminative neural networks compared to Gaussian mixture models.
result Multi-class neural networks outperform traditional methods for large speaker populations.
Study uses GMM-UBM and i-vectors to assess Parkinson's patients via speech, handwriting, and gait.
problem Assessing neurological state of Parkinson's disease patients using speech, handwriting, and gait signals.
method GMM-UBM and i-vectors applied to speech, handwriting, and gait signals.
result Different feature sets from each signal are crucial for assessing Parkinson's patients.
WEEND uses a neural network to recognize speech and assign speakers to words.
problem End-to-end neural diarization without additional ASR and orchestration.
method Multi-task learning with an auxiliary network for ASR and speaker diarization.
result WEEND outperforms turn-based diarization and can handle 5-minute audio.