Voice-controlled e-commerce app enhances user experience for visually impaired.
problem Limited user experience for visually impaired in e-commerce applications.
method Proposes a voice-controlled e-commerce application using IBM Watson speech-to-text.
result Demonstrates enhanced usability for visually impaired users.
Arcades uses deep learning to adapt smart-home behavior based on context.
problem Adaptive decision making in voice-controlled smart-home environments.
method Deep reinforcement learning with graphical representation of home automation system.
result Arcades promises long-term context-aware control of smart-home systems.
We propose a reinforcement learning (RL) based closed loop power control algorithm for the downlink of the voice over LTE (VoLTE) radio bearer for an indoor environment served by small cells. The main contributions of our paper are to 1) use RL to solve performance tuning problems in an indoor cellular network for voic…
Study improves voice conversion model with Mel-spectrogram augmentation.
problem Insufficient speech pairs data for training sequence-to-sequence voice conversion models.
method Experimented with Mel-spectrogram augmentation using SpecAugment policies and proposed new augmentation policies.
result Time axis warping policies showed better performance in training the voice conversion model.
Improved autoencoder for F0-consistent voice conversion.
problem Non-parallel many-to-many voice conversion with prosodic information leakage.
method Conditional autoencoder with information-constraining bottlenecks.
result Controlled F0 contour and improved speech quality.
Proposes a new voice conversion model that preserves pitch patterns.
problem Preserving pitch patterns while changing speaker identity.
method Variational-autoencoder-based model with an auxiliary network.
result Ensures the conversion result correctly reflects specified F0/timbre information.
The paper enhances a virtual assistant's humor to improve user satisfaction.
problem Improving a virtual assistant's ability to deliver humorous responses.
method Combines traditional NLP techniques with self-attentional networks and multi-task learning, using implicit feedback for labeling.
result Deep-learning models outperform heuristic methods in real-world user satisfaction.
Study develops sign recognition system for DHH users.
problem Accessibility of voice-controlled devices for Deaf and Hard-of-Hearing users.
method Multimodal data (RGB video and skeletal data) for sign language recognition using deep learning.
result Validation on GMUASL51 dataset of 12 users and 13107 samples across 51 signs.
Paper proposes singing voice conversion without parallel data.
problem Convert singing voices without parallel data.
method Phonetic posterior feature, DBLSTM, vocoder.
result Successfully converts singing voices without parallel data.
Paper proposes ACVAE-VC for non-parallel voice conversion.
problem Non-parallel many-to-many voice conversion with attribute class label retention.
method Uses ACVAE with fully convolutional networks, information-theoretic regularization, and auxiliary classifier.
result Successfully retains attribute class labels and avoids buzzy speech.
This paper converts speech to match a face image and vice versa.
problem Matching speech to a face image and vice versa.
method Proposes a model with speech converter, face encoder/decoder, and voice encoder.
result Trained model converts speech to match a face image and generates a face image that matches the voice of input speech.
Solution for voice conversion with limited data using hierarchical seq2seq and attention models.
problem Voice conversion between speakers with limited parallel audio pairs.
method Hierarchical sequence to sequence model with attention-based decoder, trained on single speaker dataset.
result Improved voice conversion quality using mel spectrograms and wavenet vocoder.
Convolutional network converts speaker voices without text.
problem Speaker conversion without text-based methods.
method Fully convolutional wav-to-wav network with ASR pre-training.
result Successfully converts TTS robot's voice to narrated audiobook voices.
Data augmentation improves keyword spotting accuracy in noisy conditions.
problem Maintaining low false reject rates in far-field KWS with playback interference.
method Artificially corrupted training data with mixed music and TV audio.
result 30-45% reduction in false reject rates under audio playback.
LSTM model detects voice disorders with high accuracy.
problem Automated detection of voice disorders is challenging due to continuous audio data.
method Used Long Short Term Memory (LSTM) model for feature extraction and classification of voice disorders.
result 22% sensitivity, 97% specificity, 56% unweighted average recall.
Deep learning converts one singer's voice to another without supervision.
problem Unsupervised singing voice conversion.
method Deep learning network with a single CNN encoder, WaveNet decoder, and classifier.
result Natural, recognizable singing voices converted without supervision.
A fast voice conversion method using diffusion models.
problem One-shot many-to-many voice conversion.
method Diffusion probabilistic modeling with Fast Maximum Likelihood Sampling Scheme.
result Superior quality compared to state-of-the-art approaches.
Study improves voice disorder detection system robust to channel effects.
problem Voice signals are sensitive to recording devices.
method Bidirectional LSTM network with domain adversarial training (DAT).
result Increased PR-AUC from 0.8448 to 0.9455 (and 0.9522 with labels).
End-to-end voice conversion without vocoder.
problem Speech conversion without vocoder.
method Transformer network for raw spectrum conversion.
result Transformer model converts real voices efficiently.
Improved voice conversion with semi-supervised learning.
problem Voice conversion with limited parallel data.
method Amortized variational inference with parallel and non-parallel utterances.
result Semi-supervised training improves voice conversion performance.
One-shot VC model converts voices without parallel data.
problem Limited VC applicability due to training data restriction.
method Disentangles speaker and content representations with instance normalization.
result Model converts voices from unseen speakers with high similarity.
AUTOVC converts voices without parallel data, achieving state-of-the-art results.
problem Non-parallel many-to-many voice conversion and zero-shot voice conversion.
method Only an autoencoder with a carefully designed bottleneck is used, training on a self-reconstruction loss.
result AUTOVC achieves state-of-the-art results in many-to-many voice conversion with non-parallel data and performs zero-shot voice conversion.
Semi-supervised singing voice separation using synthetic mixtures.
problem Singing voice separation with limited labeled data.
method Trains a single mapping function g on synthetic mixtures of singing and instrumental music.
result Performance comparable to fully supervised methods, better than semi-supervised alternatives.
New method separates multiple voices in mixed audio.
problem Separating multiple simultaneous speakers in audio.
method Gated neural networks trained at multiple steps, selecting actual number of speakers.
result Outperforms current state of the art for more than two speakers.
A new neural network separates singing voices more effectively.
problem Separating singing voices from mixed signals with high accuracy.
method MBR-FCN that processes different frequency bands with varying resolutions and filters.
result The MBR-FCN achieves better performance with fewer parameters.
Improved U-Nets with various intermediate blocks enhance singing voice separation.
problem Improving singing voice separation accuracy using U-Net architectures.
method Implemented and compared U-Nets with different intermediate spectrogram transformation blocks.
result A specific block type achieves state-of-the-art SDR by 0.9 dB.
Privacy-preserving method protects user speech data from cloud services.
problem Privacy compromise in cloud-based speech analysis.
method Collects and sanitizes speech data before sharing, using transformation functions and voice conversion.
result Identification of sensitive emotional state reduced by ~96%.
Wave-U-Net with MHE regularization improves singing voice separation.
problem Singing voice separation from mixed music recordings.
method Wave-U-Net architecture with MHE regularization applied to 1D filters.
result Adding MHE regularization to the loss function consistently improves singing voice separation.
DONUT spots custom wakewords from voice recordings.
problem Spotting personalized wakewords for hands-free devices.
method CTC-based algorithm using training examples and hypothesis aggregation.
result DONUT enables custom wakewords without private data upload.
Blow converts non-parallel raw audio voices efficiently.
problem Voice conversion with non-parallel data.
method Single-scale normalizing flow with hypernetwork conditioning.
result Blow outperforms existing flow-based architectures in voice conversion.
Improved speech recognition for voice assistants by analyzing speech data.
problem Reducing false triggers in speech-enabled assistants.
method Post-processing LVCSR hypothesis lattice with a Bidirectional Lattice Recurrent Neural Network (LatticeRNN).
result LatticeRNN significantly improves detection accuracy over traditional methods.
WeSinger improves singing voice synthesis with data augmentation and specialized modules.
problem Improving the accuracy and naturalness of synthesized singing voices.
method Developed a multi-singer Chinese neural singing voice synthesis system with deep bi-directional LSTM, Transformer, LPCNet, and data augmentation.
result WeSinger achieves state-of-the-art performance on the Opencpop corpus.
Many-to-Many VTN improves voice conversion across multiple speakers.
problem Voice conversion across multiple speakers.
method Sequence-to-sequence learning framework with many-to-many VTN architecture.
result Improved sound quality and speaker similarity compared to baseline methods.
Thanks to the growing availability of spoofing databases and rapid advances in using them, systems for detecting voice spoofing attacks are becoming more and more capable, and error rates close to zero are being reached for the ASVspoof2015 database. However, speech synthesis and voice conversion paradigms that are not…
New neural network separates singing voices from music using cross entropy loss.
problem Separating singing voices from music accompaniment.
method Deep Convolutional Neural Network (CNN) trained with Ideal Binary Mask (IBM) and cross entropy loss.
result Proposed CNN outperforms existing systems in MIREX evaluations.
VoiceFilter separates target speaker from multi-speaker signals.
problem Speech recognition in multi-speaker environments.
method Speaker recognition network and spectrogram masking network trained together.
result Significant reduction in speech recognition WER on multi-speaker signals.
Dr.VOT measures both positive and negative VOTs accurately in natural speech.
problem Accurate measurement of VOT in natural speech.
method Deep-learning model based on RNNs for structured prediction.
result Dr.VOT improves over state-of-the-art performance on VOT estimation.
Personal VAD detects target speaker voice activity efficiently.
problem Efficiently detect target speaker voice activity for reduced computational cost and battery usage.
method Trains a neural network conditioned on speaker embedding or verification score, outputs probabilities for three speech classes.
result Trained model with 130K parameters outperforms combined standard VAD and speaker recognition networks.
Paper introduces a new model for polyphonic music composition.
problem Creating music with multiple interwoven voices.
method Developed a coupled recurrent model using probabilistic factorization and neural network ideas.
result Trained models for single-voice and multi-voice composition on a large dataset.
A-StarGAN improves nonparallel voice conversion speed and realism.
problem Nonparallel voice conversion without parallel data.
method Augmented classifier StarGAN (A-StarGAN) for nonparallel voice conversion.
result A-StarGAN generates realistic-sounding speech quickly and efficiently.
Transformer models show distinct spectral fingerprints under voice changes.
problem Detecting architectural biases in transformer models.
method Spectral analysis of attention-induced token graphs.
result Clear architectural signatures in model fingerprints correlate with language-specific behavior.
Study compares neural and statistical models for Parkinson's disease progression from voice data.
problem Difficult statistical analysis of longitudinal voice biomarkers due to subject correlation, small cohorts, and varied disease trajectories.
method Evaluated Neural Mixed Effects (NME), Generalized Neural Network Mixed Models (GNMMs), and semi-parametric Generalized Additive Mixed Models (GAMMs).
result GAMMs achieve stronger predictive performance and retain interpretable smooth effects and subject-level structure.
MaskCycleGAN-VC improves voice conversion without parallel data.
problem Limited ability to convert mel-spectrogram data without parallel data.
method Integrates a novel auxiliary task called filling in frames (FIF) to learn time-frequency structures.
result MaskCycleGAN-VC outperforms existing methods with similar model size.
Joint training model for TTS and VC tasks using Tacotron and WaveNet.
problem Training a shared model for text-to-speech and voice conversion.
method Extended Tacotron model with dual attention mechanism for shared tasks, WaveNet for waveform generation.
result Joint training of a shared model achieves both TTS and VC tasks efficiently.
The human auditory system is able to distinguish the vocal source of thousands of speakers, yet not much is known about what features the auditory system uses to do this. Fourier Transforms are capable of capturing the pitch and harmonic structure of the speaker but this alone proves insufficient at identifying speaker…
Study combines speaker verification and voice trigger detection in a single network.
problem Separate training for speaker verification and voice trigger detection.
method Multi-task learning with a single network trained on both tasks.
result Single network achieves comparable accuracy to independent models for each task.
Singing voice separation attempts to separate the vocal and instrumental parts of a music recording, which is a fundamental problem in music information retrieval. Recent work on singing voice separation has shown that the low-rank representation and informed separation approaches are both able to improve separation qu…
Paper proposes a new method for robust speaker verification.
problem Improving robustness in speaker verification systems.
method Combines soft VAD and self-adaptive VAD with DNN-based VAD.
result Significant improvement in verification performance in real-world environments.