Method converts emotions in nonparallel speech data.
problem Lack of parallel data for speech emotion conversion.
method Unsupervised style transfer technique for nonparallel training.
result Effectiveness demonstrated on nonparallel corpora with four emotions.
End-to-end deep learning detects emotions in real-life emergency calls.
problem Recognizing emotions in real-life emergency call center recordings.
method Used an end-to-end deep learning architecture trained on IEMOCAP and CEMO datasets.
result Obtained 45.6% Unweighted Accuracy Recall on CEMO with 4 classes, 76.9% on 2 classes (Anger, Neutral).
Privacy-preserving method protects user speech data from cloud services.
problem Privacy compromise in cloud-based speech analysis.
method Collects and sanitizes speech data before sharing, using transformation functions and voice conversion.
result Identification of sensitive emotional state reduced by ~96%.
Speech emotion recognition system using features and text.
problem Improving accuracy in emotion recognition from speech.
method Used speech features (Spectrogram, MFCC) and text, trained Deep Neural Networks.
result Combined MFCC-Text CNN model achieved highest accuracy.
Deep architecture learns transferable features for robust speech emotion recognition.
problem Robust and discriminative features for diverse speech emotion domains.
method Jointly uses CNN for domain-shared features and LSTM for domain-specific emotion classification.
result Transferable features provide gains up to 18.4% in speech emotion recognition.
This research improves emotion detection from speech, enhancing CCC by 30%.
problem Improving emotion detection from speech for categorical emotions.
method Used LSTM and TC-LSTM networks, trained with multiple datasets and robust features.
result Improved CCC for valence by 30% compared to baseline.
Improved speech emotion recognition using pre-trained language models.
problem Challenging task of speech emotion recognition for natural human-machine interaction.
method Fine-tuning pre-trained language models for text emotion recognition, combining with speech emotion recognition.
result 73.5% accuracy in speech emotion recognition on a subset of IEMOCAP dataset.
Study on mother-infant affect communication using audio recordings.
problem Lack of accurate emotional speech databases for real-life settings.
method Used RAVDESS database and trained a Convolutional Neural Nets model.
result Dominant emotions in mother-infant speech were angry and sad.
ResNet with Focal Loss improves speech emotion recognition.
problem Speech emotion recognition using plain text features is insufficient.
method Residual Convolutional Neural Network (ResNet) trained with Focal Loss.
result Focal Loss enhances model's focus on hard examples.
Paper proposes a CNN for speech emotion recognition using center loss and reconstruction.
problem Speech emotion recognition (SER) in audio signals.
method Convolutional Neural Network (CNN) with center loss and reconstruction as regularizers.
result Proposed method achieves highly discriminative features for SER.
Learning the latent representation of data in unsupervised fashion is a very interesting process that provides relevant features for enhancing the performance of a classifier. For speech emotion recognition tasks, generating effective features is crucial. Currently, handcrafted features are mostly used for speech emoti…
This work tackles domain shift in speech emotion recognition by proposing class-wise adversarial domain adaptation.
problem Domain shift between corpora poses a challenge for speech emotion recognition, especially for positive/negative emotions.
method Class-wise adversarial domain adaptation to reduce shift between different corpora.
result Our method is effective even with limited target labeled examples, as demonstrated on EMODB and Aibo corpora.
Enhances speech emotion recognition using transfer learning.
problem Improving accuracy in speech emotion recognition.
method Transformer-based Predictive Coding with transfer learning.
result Significantly improved emotion recognition accuracy.
Improved anger detection in speech using transfer learning from SoundNet.
problem Detecting anger in speech with limited emotion datasets.
method Transfer learning from SoundNet, a multimodal audio classifier trained on video data.
result Improved performance and generalization on various speech emotion datasets.
Paper explores deep learning features for complex emotion recognition.
problem Improving emotion recognition accuracy in complex emotions.
method Used pretrained networks (AudioSet Net, VoxCeleb Net, Deep Speech Net) and their deep layer features for emotion recognition.
result Achieved highest F1 score of 0.85 on EmoReact dataset.
Enhances speech emotion recognition by adapting to varying time scales.
problem Robust emotion recognition from speech audio with temporal variations.
method Introduces multi-time-scale (MTS) convolutional layers to CNNs.
result MTS layers improve generalization, especially on smaller datasets.
Speech emotion recognition improved with simpler machine learning models.
problem Identifying emotions from speech is ambiguous and challenging.
method Feature-engineering approach using hand-crafted audio features and text features. Comparison of traditional machine learning and deep learning models.
result Lighter machine learning models outperform deep learning models for emotion recognition.
Enhances speech emotion recognition by integrating visual data with attention mechanisms.
problem Improving emotion detection accuracy by combining multiple modalities.
method Introducing an attention mechanism to combine audio, text, and video modalities.
result Significant improvement of 3.65% in weighted accuracy.
ADDoG improves cross-dataset speech emotion recognition.
problem Cross-dataset speech emotion recognition failure.
method ADDoG uses an iterative approach to move representations closer together across datasets.
result ADDoG and MADDoG improve cross-dataset speech emotion recognition.
Novel bio-inspired masking for robust speech emotion recognition.
problem Noise degradation in speech emotion recognition.
method Cochlear cepstrogram-based contrastive learning with temporal and frequency masking.
result Improved speech emotion recognition performance on K-EmoCon benchmark.
Bardo Composer generates tabletop RPG music based on player speech.
problem Creating immersive background music for tabletop RPGs.
method Speech recognition, emotion classification, and music generation using a novel beam search algorithm.
result Generated music pieces can be accurately identified by human subjects as conveying the intended emotion.
Enhanced speech emotion recognition using nonlinear recurrence dynamics.
problem Improving speech emotion recognition accuracy.
method Phase space reconstruction, Recurrence Plot, Recurrence Quantification Analysis, statistical functionals, feature fusion, Bidirectional Recurrent Neural Network.
result State-of-the-art performance on IEMOCAP with up to 10.7% improvement in accuracy.
Automatically assessing emotional valence in human speech has historically been a difficult task for machine learning algorithms. The subtle changes in the voice of the speaker that are indicative of positive or negative emotional states are often "overshadowed" by voice characteristics relating to emotional intensity …
Study evaluates adversarial attacks on speech emotion recognition systems.
problem Adversarial examples challenge the robustness of speech emotion recognition systems.
method Proposes adversarial training and GAN as defenses.
result Demonstrates effective use of adversarial examples for robustness.
This paper proposes an approach to detect emotion from human speech employing majority voting technique over several machine learning techniques. The contribution of this work is in two folds: firstly it selects those features of speech which is most promising for classification and secondly it uses the majority voting…
End-to-end network predicts and aligns continuous emotion labels.
problem Inconsistent alignment of continuous emotion labels with speech signals.
method Convolutional neural network with a multi-delay sinc layer.
result State-of-the-art results in predicting and aligning continuous emotion labels.
ConvS2S-VC converts voice characteristics and pitch contour using a fully convolutional seq2seq model.
problem Voice conversion with preservation of pitch contour and duration.
method Fully convolutional seq2seq architecture with conditional batch normalization.
result ConvS2S-VC outperforms baseline methods in sound quality and speaker similarity.
Improved speech emotion recognition using pitch-synchronous single frequency filtering spectrogram.
problem Uncertainty principles in STFT spectrogram limit time and frequency resolutions.
method Modified SFF spectrogram by averaging amplitudes between GCI locations, named pitch-synchronous SFF spectrogram.
result Improved SER accuracy (63.95% to 70.4%) on IEMOCAP dataset.
Generative Adversarial Networks improve affective speech feature generation.
problem Improving feature representation for emotion recognition.
method Experimented with GAN architectures to generate feature vectors corresponding to emotions.
result GANs generate realistic synthetic samples for emotion recognition.
Improved TTS style transfer across disjoint datasets with adversarial cycle consistency.
problem Suboptimal TTS style transfer on disjoint datasets with underrepresented styles.
method Adversarial cycle consistency training with paired and unpaired triplets.
result 78% improvement in style transfer with minimal reduction in fidelity and naturalness.
FINs enhance performance in diverse datasets like finance, speech, and health.
problem Improving neural network performance across various domains.
method Feature Imitating Networks (FINs) initialize weights to approximate specific statistical features.
result FINs significantly improve performance in Bitcoin price prediction, speech emotion recognition, and chronic neck pain detection.
Self-attentive network improves emotion recognition in conversations.
problem Emotion recognition in dyadic conversations using deep learning.
method Introduces a novel self-attention mechanism for capturing temporal dynamics without a decoder.
result Outperforms state-of-the-art alternatives on the IEMOCAP benchmark.
Survey on DNNs for speech processing, focusing on limited data challenges.
problem Challenges in training DNNs for speech tasks with limited data.
method Overview of techniques for few-shot speech processing.
result Promising few-shot techniques for speech processing.
Generating versatile and appropriate synthetic speech requires control over the output expression separate from the spoken text. Important non-textual speech variation is seldom annotated, in which case output control must be learned in an unsupervised fashion. In this paper, we perform an in-depth study of methods for…
In this work we explored building automatic speech recognition models for transcribing doctor patient conversation. We collected a large scale dataset of clinical conversations (14,000 hr), designed the task to represent the real word scenario, and explored several alignment approaches to iteratively improve data qua…
WaveCycleGAN converts synthetic speech to natural speech using cycle-consistent adversarial networks.
problem Over-smoothing effect in synthetic speech, leading to quality degradation.
method Cycle-consistent adversarial networks for waveform-level modification.
result Improves naturalness of generated speech sounds.
Paper improves emotion expression in AI chatbots.
problem AI chatbots generate responses that lack emotion.
method Developed neural models to express specific emotions in generated responses.
result An encoder-decoder model with multiple attention layers performs best in expressing required emotion.
Enhanced transformer converts whispered speech to natural speech.
problem Machine recognition of whispered speech is challenging.
method Proposes an enhanced transformer architecture trained end-to-end using supervised learning.
result Similar formant distributions of converted speech to groundtruth.
Study evaluates feature selection methods for emotion recognition in resource-constrained settings.
problem Reducing memory and computational requirements for emotion recognition in low-resource settings.
method Evaluation of three feature selection methods: ILFS, ReliefF, Fisher, and AFS.
result Smaller feature sets can achieve similar or better accuracy, reducing resource usage.
Gait patterns reveal emotions, offering a non-invasive method for automated recognition.
problem Automated emotion recognition from gait patterns.
method Data collection, preprocessing, and classification techniques.
result Gait patterns can indicate different emotion states, making them a promising source for emotion detection.
End-to-end voice conversion without vocoder.
problem Speech conversion without vocoder.
method Transformer network for raw spectrum conversion.
result Transformer model converts real voices efficiently.
Improved autoencoder for F0-consistent voice conversion.
problem Non-parallel many-to-many voice conversion with prosodic information leakage.
method Conditional autoencoder with information-constraining bottlenecks.
result Controlled F0 contour and improved speech quality.
Method converts facial expressions and voice of a source speaker into a target speaker.
problem Separate conversion of facial and acoustic features leads to unnatural results.
method Uses three neural networks: conversion, waveform generation, and image reconstruction.
result Significantly higher naturalness achieved when converting both features together.
Deep learning predicts mental disorders from audio and text samples.
problem Predicting mental disorders from speech samples.
method Multimodal deep learning structure using various pre-trained models for audio and text embeddings, transfer learning, and auxiliary corpora.
result Acceptable accuracy in predicting mental disorders through multimodal analysis.
Method converts speech with attention and context preservation.
problem Voice conversion with improved stability and efficiency.
method Sequence-to-Sequence learning with attention and context preservation.
result Synthesized speech quality comparable to advanced methods.
A new TTS method uses diffusion and VAE for better speech synthesis.
problem Improving text-to-speech synthesis for better speech quality and robustness.
method Combines diffusion probabilistic model and variational autoencoder for latent variable conversion.
result The method is robust to poor orthography and alignment errors.
Improved SAR in asynchronous conversations using neural models and unlabeled data.
problem Lack of labeled data for SAR in asynchronous conversations.
method Hierarchical LSTM-CRF model, semi-supervised learning with word embeddings, adversarial training.
result Adversarial training improves SAR performance by leveraging labeled data from synchronous domains.
Joint training model for TTS and VC tasks using Tacotron and WaveNet.
problem Training a shared model for text-to-speech and voice conversion.
method Extended Tacotron model with dual attention mechanism for shared tasks, WaveNet for waveform generation.
result Joint training of a shared model achieves both TTS and VC tasks efficiently.