WaveCycleGAN converts synthetic speech to natural speech using cycle-consistent adversarial networks.
problem Over-smoothing effect in synthetic speech, leading to quality degradation.
method Cycle-consistent adversarial networks for waveform-level modification.
result Improves naturalness of generated speech sounds.
Enhanced transformer converts whispered speech to natural speech.
problem Machine recognition of whispered speech is challenging.
method Proposes an enhanced transformer architecture trained end-to-end using supervised learning.
result Similar formant distributions of converted speech to groundtruth.
This paper presents sampling-based speech parameter generation using moment-matching networks for Deep Neural Network (DNN)-based speech synthesis. Although people never produce exactly the same speech even if we try to express the same linguistic and para-linguistic information, typical statistical speech synthesis pr…
WaveCycleGAN2 improves speech synthesis quality by reducing aliasing.
problem Human ear can still distinguish synthesized speech from natural speech.
method WaveCycleGAN2 uses generators without down/up-sampling modules and combines discriminators from waveform and acoustic parameter domains.
result WaveCycleGAN2 achieves high-quality speech synthesis with comparable mean opinion scores to natural speech.
This paper improves speech synthesis using a DGP with SRU for naturalness.
problem Improving naturalness in synthetic speech.
method Deep Gaussian process with a recurrent architecture using SRU.
result SRU-DGP outperforms other models in naturalness of synthetic speech.
Natural speech can be easily manipulated to alter perception.
problem The susceptibility of natural speech to manipulation.
method Investigation of the McGurk effect and Yanny or Laurel illusion.
result A significant fraction of natural speech is illusionable.
An ability to model a generative process and learn a latent representation for speech in an unsupervised fashion will be crucial to process vast quantities of unlabelled speech data. Recently, deep probabilistic generative models such as Variational Autoencoders (VAEs) have achieved tremendous success in modeling natur…
RUSLAN is a large Russian speech corpus for text-to-speech.
problem Lack of high-quality annotated Russian speech data for text-to-speech.
method Developed a large annotated Russian speech corpus and trained a neural network for text-to-speech synthesis.
result Synthesized speech quality evaluated with MOS scores: 4.05 for naturalness, 3.78 for intelligibility.
Improves naturalness in TTS samples using quantized VAE and auto-regressive prosody.
problem Discontinuous and unnatural speech from standard VAE priors.
method Discretized latent features using vector quantization (VQ), and separately trained autoregressive (AR) prior model.
result Significantly improves naturalness in random sample generation.
Survey reviews code-switched speech and language processing.
problem Processing code-switched text and speech for multilingual communities.
method Reviews computational approaches and lists available resources.
result Essential for building intelligent agents that interact in multilingual settings.
Improved speech emotion recognition using pre-trained language models.
problem Challenging task of speech emotion recognition for natural human-machine interaction.
method Fine-tuning pre-trained language models for text emotion recognition, combining with speech emotion recognition.
result 73.5% accuracy in speech emotion recognition on a subset of IEMOCAP dataset.
Enhanced Tacotron for Japanese speech synthesis improves naturalness.
problem Challenges in end-to-end Japanese speech synthesis due to pitch accents.
method Extended Tacotron with self-attention to capture pitch accent dependencies.
result Proposed systems show improvements but still lag behind traditional pipeline methods.
Generative models produce varied intonation in speech synthesis.
problem Typical TTS systems lack the ability to produce multiple distinct renditions of a sentence.
method Use variational autoencoders (VAEs) to capture a distribution over multiple renditions and produce varied intonation.
result Sampling from the tails of the VAE prior produces more varied intonation than traditional approaches, while maintaining naturalness.
Enhanced text-to-speech synthesizes expressive speech from a single example.
problem Creating a new expressive speech style from a single example of speech.
method Combines VAE and Normalizing Flows to improve disentanglement and naturalness.
result Reduces KL-divergence by 22% and improves perceptual metrics.
This paper benchmarks speech LVMs against deterministic models and adapts a video model to speech.
problem Speech generation models are inferior to deterministic models.
method Developed a speech benchmark of LVMs and compared them against deterministic models.
result The Clockwork VAE outperforms previous LVMs and reduces the gap to deterministic models.
Novel BCI system classifies imagined speech with high accuracy.
problem Classifying imagined speech from brain signals.
method Hierarchical deep learning with CNN and autoencoder.
result Achieved 83.42% average accuracy across six phonological tasks.
This research improves neural synthesizers for music sounds from speech data.
problem Applying speech synthesis techniques to musical instrument sounds.
method Comparison of three neural synthesizers in three scenarios: training, zero-shot learning, and fine-tuning.
result Neural synthesizers trained on speech data and fine-tuned on music data perform better.
Recent neural networks such as WaveNet and sampleRNN that learn directly from speech waveform samples have achieved very high-quality synthetic speech in terms of both naturalness and speaker similarity even in multi-speaker text-to-speech synthesis systems. Such neural networks are being used as an alternative to voco…
New speech recognition method uses hypergraphs for better label prediction.
problem Missed information in pairwise relationships between speech samples.
method Hypergraph Laplacian based semi-supervised learning methods applied to speech recognition.
result Sensitivity performance measures of hypergraph methods are better than state-of-the-art methods.
Models predict Alzheimer's Dementia from spontaneous speech with high accuracy.
problem Early diagnosis of Alzheimer's Dementia (AD) through spontaneous speech analysis.
method Compared natural language processing techniques including SVM, GBDT, CRFs, and Transformer-based models.
result Top models achieve 0.81-0.82 test set scores for AD vs controls and 4.58 RMSE for Mental Mini State Exam scores.
Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep recurrent neural network whose inputs consist solely of a raw waveform. Our best models …
Activates speech DNNs to generate understandable examples.
problem Difficulty in understanding DNN classifications for speech.
method Activation maximization to generate speech samples.
result Activation maximization can generate understandable speech samples.
Learning the latent representation of data in unsupervised fashion is a very interesting process that provides relevant features for enhancing the performance of a classifier. For speech emotion recognition tasks, generating effective features is crucial. Currently, handcrafted features are mostly used for speech emoti…
New method uses Multiple Choice Learning for speech separation.
problem Ambiguous task of assigning model predictions to ground truth signals.
method Uses Multiple Choice Learning (MCL) instead of Permutation Invariant Training (PIT).
result MCL matches PIT performance but is computationally advantageous.
SepIt improves speech separation for multiple speakers.
problem Improving speech separation for multiple speakers in single channel recordings.
method SepIt uses a deep neural network that iteratively improves estimates of different speakers based on mutual information.
result SepIt outperforms state-of-the-art methods for 2, 3, 5, and 10 speakers.
Improved anger detection in speech using transfer learning from SoundNet.
problem Detecting anger in speech with limited emotion datasets.
method Transfer learning from SoundNet, a multimodal audio classifier trained on video data.
result Improved performance and generalization on various speech emotion datasets.
This research improves emotion detection from speech, enhancing CCC by 30%.
problem Improving emotion detection from speech for categorical emotions.
method Used LSTM and TC-LSTM networks, trained with multiple datasets and robust features.
result Improved CCC for valence by 30% compared to baseline.
End-to-end voice conversion without vocoder.
problem Speech conversion without vocoder.
method Transformer network for raw spectrum conversion.
result Transformer model converts real voices efficiently.
Recent advances in speech synthesis suggest that limitations such as the lossy nature of the amplitude spectrum with minimum phase approximation and the over-smoothing effect in acoustic modeling can be overcome by using advanced machine learning approaches. In this paper, we build a framework in which we can fairly co…
Study shows integrating acoustic features in financial forecasting models can degrade performance.
problem Predicting stock market volatility from corporate earnings calls using speech features.
method Empirical investigation of acoustic feature extraction in teleconference environments using a two-stream late-fusion architecture.
result Integrating acoustic features via late fusion significantly degraded performance, reducing recall to 47.08%.
End-to-end Sanskrit TTS developed with limited data, achieving good quality.
problem Developing natural-sounding speech for Sanskrit with scarce data.
method Fine-tuning Tacotron2 model with WaveGlow and transfer learning.
result Achieved an overall MOS of 3.38 from 37 evaluators.
Thanks to the growing availability of spoofing databases and rapid advances in using them, systems for detecting voice spoofing attacks are becoming more and more capable, and error rates close to zero are being reached for the ASVspoof2015 database. However, speech synthesis and voice conversion paradigms that are not…
Improved method learns robust speech representations from multiple tasks.
problem Learning good speech representations without supervision.
method Single encoder with multiple self-supervised tasks.
result Transferable, robust, problem-agnostic features learned.
Bayesian method improves reliability of BERT for hate speech detection.
problem Reliable detection of hate speech in user-generated content.
method Bayesian approach using Monte Carlo dropout in transformer models.
result Monte Carlo dropout provides well-calibrated reliability estimates.
Unsupervised model generates distinct intonation codes for speech synthesis.
problem Lack of understanding of what prosodic variations are controlled in speech synthesis.
method Phrase-level variational autoencoder with multi-modal prior, using mode centres as intonation codes.
result Generated intonation codes are perceptually distinct and carry various affect-related styles.
This paper examines the speaker identification potential of breath sounds in continuous speech. Speech is largely produced during exhalation. In order to replenish air in the lungs, speakers must periodically inhale. When inhalation occurs in the midst of continuous speech, it is generally through the mouth. Intra-spee…
Paper tackles toxic comment detection using deep learning.
problem Automatic detection of toxic comments on the internet.
method Designs binary classification and regression-based approaches using DNN.
result BERT fine-tuning outperforms other methods.
Dr.VOT measures both positive and negative VOTs accurately in natural speech.
problem Accurate measurement of VOT in natural speech.
method Deep-learning model based on RNNs for structured prediction.
result Dr.VOT improves over state-of-the-art performance on VOT estimation.
How to model distribution of sequential data, including but not limited to speech and human motions, is an important ongoing research problem. It has been demonstrated that model capacity can be significantly enhanced by introducing stochastic latent variables in the hidden states of recurrent neural networks. Simultan…
We investigated the impact of noisy linguistic features on the performance of a Japanese speech synthesis system based on neural network that uses WaveNet vocoder. We compared an ideal system that uses manually corrected linguistic features including phoneme and prosodic information in training and test sets against a …
Lipper synthesizes speech from silent videos, improving over single-view methods.
problem Lipreading as text classification is limited; multi-view approach needed.
method Multi-view lipreading as a regression task, producing speech from silent videos.
result Improvement in speech reconstruction with multi-view silent videos.
Speech Translation has always been about giving source text or audio input and waiting for system to give translated output in desired form. In this paper, we present the Acoustic Dialect Decoder (ADD) - a voice to voice ear-piece translation device. We introduce and survey the recent advances made in the field of Spee…
New tool detects weak and strong Islamophobic hate speech on social media.
problem Detecting Islamophobic hate speech on social media is challenging due to its varied nature.
method Built a multi-class classifier distinguishing between non-Islamophobic, weak Islamophobic, and strong Islamophobic content using GloVe word embeddings.
result Accuracy of 77.6% and balanced accuracy of 83% on a dataset of 109,488 tweets.
Telephonetic improves robustness of neural language models to ASR errors.
problem Handling ASR errors and semantic noise in neural language models trained on written text.
method Data augmentation framework using character-level and embedding space perturbations.
result Achieves state-of-the-art perplexity of 37.49 on Penn Treebank corpus.
Deep learning system classifies phonological categories from EEG data.
problem Speech-related BCI for people with speaking disabilities.
method Hierarchical deep learning approach using CNN, LSTM, and autoencoder.
result Average accuracy of 77.9% across five binary classification tasks.
End-to-end speech recognition using EEG without speech input.
problem Speech recognition without direct speech input.
method Implemented attention model and CTC-based ASR systems for EEG signals; fused EEG with noisy speech features.
result Demonstrated end-to-end speech recognition using EEG signals.
Study classifies Persian speech acts for better understanding of text intent.
problem Understanding the intended function of Persian texts.
method Dictionary-based statistical technique using WordNet for SA recognition.
result Proposed method achieved state-of-the-art accuracy of 0.95 for Persian SA classification.
Foreign policy analysis has been struggling to find ways to measure policy preferences and paradigm shifts in international political systems. This paper presents a novel, potential solution to this challenge, through the application of a neural word embedding (Word2vec) model on a dataset featuring speeches by heads o…