Speaker verification (SV) systems using deep neural network embeddings, so-called the x-vector systems, are becoming popular due to its good performance superior to the i-vector systems. The fusion of these systems provides improved performance benefiting both from the discriminatively trained x-vectors and generative …
Probabilistic embeddings improve speaker diarization accuracy.
problem Improving speaker diarization accuracy using embeddings.
method Extracting x-vectors and precision matrices from speech segments, interfacing with PLDA model, applying agglomerative clustering, joint training of PLDA and extractor.
result Joint training of PLDA and probabilistic x-vector extractor yields accuracy gains.
The standard state-of-the-art backend for text-independent speaker recognizers that use i-vectors or x-vectors, is Gaussian PLDA (G-PLDA), assisted by a Gaussianization step involving length normalization. G-PLDA can be trained with both generative or discriminative methods. It has long been known that heavy-tailed PLD…
Anonymizes speech data to protect privacy.
problem Protecting personal speech data from misuse.
method Extracts features, uses x-vectors, and neural models to synthesize anonymized speech.
result Effective in concealing speaker identities without compromising speech quality.
Study shows emotion affects speaker recognition and vice versa.
problem Dependencies between emotion and speaker recognition.
method Transfer learning and fine-tuning for emotion classification.
result Fine-tuning improves emotion recognition performance by 30.40% on IEMOCAP, 7.99% on MSP-Podcast, and 8.61% on Crema-D.
We consider technology-assisted mimicry attacks in the context of automatic speaker verification (ASV). We use ASV itself to select targeted speakers to be attacked by human-based mimicry. We recorded 6 naive mimics for whom we select target celebrities from VoxCeleb1 and VoxCeleb2 corpora (7,365 potential targets) usi…
Deep learning improves speaker recognition verification and identification.
problem Limited progress in speaker recognition for 5-6 years.
method Applied deep learning techniques in speaker verification and identification.
result Deep learning becomes the state-of-the-art solution for speaker recognition.
Study compares metric learning loss functions for speaker verification.
problem Comparing metric learning loss functions for end-to-end speaker verification.
method Cross entropy loss, cosine loss, angular margin loss, center loss, contrastive loss, triplet loss.
result Additive angular margin loss outperforms other loss functions.
Neural network framework for language recognition considers sequence information and improves accuracy.
problem Challenging task of automatic language identification in noisy conditions.
method Proposes a neural network framework with bidirectional LSTM and attention modeling for relevance weighting.
result Significant improvements over conventional methods in noisy conditions and multi-speaker speech.
Hybrid diarization framework handles overlapped speech and long recordings.
problem Challenges in clustering-based and end-to-end neural diarization approaches.
method Proposes a hybrid framework combining clustering and end-to-end neural diarization.
result Significantly better performance on long recordings with overlapped speech.
Study speaker verification security using hierarchical Bayesian modeling.
problem Estimating false alarm rate in ASV systems for large speaker databases.
method Hierarchical Bayesian modeling of ASV scores to assess security against closest impostors.
result Neither i-vector nor x-vector systems are secure against increased impostor database size.
State-of-the-art speaker recognition systems comprise an x-vector (or i-vector) speaker embedding front-end followed by a probabilistic linear discriminant analysis (PLDA) backend. The effectiveness of these components relies on the availability of a large collection of labeled training data. In practice, it is common …
New method predicts y distributions from imperfect data.
problem Predicting y from imperfect data (discrete, truncated, censored).
method Optimal transformations to estimate p(y|x).
result Estimates location, scale, and shape of y distribution.
Improved neural speaker embeddings enhance ASR performance.
problem Few studies have explored neural speaker embeddings for ASR.
method Integrating improved neural speaker embeddings into a conformer-based hybrid HMM ASR system.
result Improved neural embeddings achieve on-par performance with i-vectors.
Extends BV functions and divergence-measure fields to metric spaces.
problem Defining BV functions and divergence-measure fields in metric spaces.
method Employing differential structure developed by N. Gigli, extending BV functions and divergence-measure fields to metric spaces.
result Gauss-Green formulas established for BV functions and divergence-measure fields in metric spaces.
System diagnoses Alzheimer's disease from spoken language using multi-modal features.
problem Early diagnosis of Alzheimer's disease from spoken language.
method Classification system based on spoken language using three approaches (N-gram, i-vector, x-vector).
result Accuracy of 83.6% on the cookie picture description task from Pitt Corpus dementia bank.
GPU acceleration speeds up i-vector extraction 3000x, enabling new research.
problem Speeding up i-vector extraction for speaker verification.
method GPU acceleration for i-vector extraction, including re-computing UBM and frame alignments.
result Significant speed-up allows rigorous study of i-vector variations.
Improved far-field speaker verification for short utterances in noisy conditions.
problem Challenges in speaker verification on short utterances in uncontrolled noisy environments.
method Used deep neural network architectures (TDNN and ResNet) and experimented with various embedding extractors and training procedures.
result ResNet architectures outperform x-vector approach in speaker verification quality for both long and short utterances.
Non-Gaussian component analysis (NGCA) is a problem in multidimensional data analysis which, since its formulation in 2006, has attracted considerable attention in statistics and machine learning. In this problem, we have a random variable X in n-dimensional Euclidean space. There is an unknown subspace Γ of the …