Unified framework for speaker-adaptive models using scaling and bias codes.
problem Improving speaker-adaptive neural-network based speech synthesis systems.
method Unified framework representation and generalized scaling and bias codes.
result Improves performance of speaker adaptation compared to conventional methods.
New unsupervised speaker adaptation method for speech synthesis.
problem Adapting speech synthesis to new speakers with minimal data.
method Concatenating audio and text inputs, proposing new training schemes.
result Improves adaptation to unseen speakers and multi-speaker modeling.
BOFFIN TTS optimizes hyper-parameters for new speaker adaptation.
problem Fine-tuning a pre-trained TTS model for a new speaker with limited data.
method Bayesian optimization to efficiently find optimal hyper-parameters.
result Average 30% improvement in speaker similarity over standard techniques.
Seq2seq ASR adapts to speakers, improving performance by 25%.
problem Speaker adaptation for seq2seq ASR systems to match conventional methods.
method Applied Kullback-Leibler divergence and Linear Hidden Network adaptation to seq2seq models.
result 25% relative word error rate improvement with seq2seq model adaptation.
Improved robustness in ASR systems with speaker adaptation.
problem Improving robustness in automatic speech recognition systems.
method Weighted-Simple-Add method for adding weighted speaker information vectors to the conformer-based acoustic model.
result Achieved 11% relative improvement in WER on Switchboard 300h Hub5'00 dataset.
Bayesian approach improves speech recognition with limited speaker data.
problem Reduces mismatch between training and evaluation data due to speaker differences.
method Bayesian learning for DNN adaptation models with limited speaker data.
result Bayesian adaptation consistently outperforms deterministic methods, reducing word error rates up to 1.4%.
Improves TTS accuracy by correcting context-dependent units.
problem Improves text-to-speech accuracy through speaker adaptation.
method Statistical model predicting context-dependent phonetic unit classes and their mean error values.
result Corrected boundaries of units improve TTS accuracy compared to HMM segmentation.
ASA improves ASR by adapting SD models to SI model's deep feature distribution.
problem Improving ASR performance on new speakers with limited data.
method Adversarial learning to regularize SD model's deep features to match SI model's.
result ASA achieves significant word error rate improvements over SI models.
Conditional T/S learning improves student model performance by selectively learning from teacher or ground truth.
problem Teacher's occasional wrong guidance leads to suboptimal student model performance.
method Proposes a conditional T/S learning scheme where the student selectively chooses between teacher and ground truth based on teacher correctness.
result The conditional learning achieves significant performance improvements over traditional T/S learning.
CAT is a new ASR toolkit using CRF and CTC for state-of-the-art speech recognition.
problem Improving automatic speech recognition systems.
method CRF-based discriminative training with CTC-inspired state topology.
result CAT achieves state-of-the-art results with fewer parameters and is competitive with hybrid models.
Improved speech recognition with cumulative adaptation methods.
problem Robust speech recognition in varying environments and speakers.
method Used a bidirectional LSTM neural network and i-vectors for adaptation.
result Achieved 13% relative improvement in word error rate.
Meta-learning approach for adaptive TTS with few data.
problem Adapting TTS systems to new speakers with minimal data.
method Meta-learning with shared WaveNet core and independent speaker embeddings, using three training strategies.
result Successful adaptation of multi-speaker neural network to new speakers with minimal data.
Deep Convolutional Neural Networks (CNNs) are more powerful than Deep Neural Networks (DNN), as they are able to better reduce spectral variation in the input signal. This has also been confirmed experimentally, with CNNs showing improvements in word error rate (WER) between 4-12% relative compared to DNNs across a var…
Speech enhancement improved by adapting to unknown speakers without auxiliary signals.
problem Improving speech enhancement accuracy for unknown speakers.
method Adopting multi-task learning for speech enhancement and speaker identification, using multi-head self-attention.
result Achieved state-of-the-art performance and improved subjective quality.
End-to-end speaker recognition method using neural networks.
problem Speaker and session variability in speaker verification.
method Joint Factor Analysis with tied hidden variables, MAP adaptation, two-step backpropagation.
result Improved likelihood ratios and robust performance on RSR2015 database.
End-to-end voice conversion without vocoder.
problem Speech conversion without vocoder.
method Transformer network for raw spectrum conversion.
result Transformer model converts real voices efficiently.
In this paper we describe the recent advancements made in the IBM i-vector speaker recognition system for conversational speech. In particular, we identify key techniques that contribute to significant improvements in performance of our system, and quantify their contributions. The techniques include: 1) a nearest-neig…