Paper proposes singing voice conversion without parallel data.
problem Convert singing voices without parallel data.
method Phonetic posterior feature, DBLSTM, vocoder.
result Successfully converts singing voices without parallel data.
This paper converts speech to match a face image and vice versa.
problem Matching speech to a face image and vice versa.
method Proposes a model with speech converter, face encoder/decoder, and voice encoder.
result Trained model converts speech to match a face image and generates a face image that matches the voice of input speech.
Solution for voice conversion with limited data using hierarchical seq2seq and attention models.
problem Voice conversion between speakers with limited parallel audio pairs.
method Hierarchical sequence to sequence model with attention-based decoder, trained on single speaker dataset.
result Improved voice conversion quality using mel spectrograms and wavenet vocoder.
Convolutional network converts speaker voices without text.
problem Speaker conversion without text-based methods.
method Fully convolutional wav-to-wav network with ASR pre-training.
result Successfully converts TTS robot's voice to narrated audiobook voices.
LSTM model detects voice disorders with high accuracy.
problem Automated detection of voice disorders is challenging due to continuous audio data.
method Used Long Short Term Memory (LSTM) model for feature extraction and classification of voice disorders.
result 22% sensitivity, 97% specificity, 56% unweighted average recall.
Deep learning converts one singer's voice to another without supervision.
problem Unsupervised singing voice conversion.
method Deep learning network with a single CNN encoder, WaveNet decoder, and classifier.
result Natural, recognizable singing voices converted without supervision.
A fast voice conversion method using diffusion models.
problem One-shot many-to-many voice conversion.
method Diffusion probabilistic modeling with Fast Maximum Likelihood Sampling Scheme.
result Superior quality compared to state-of-the-art approaches.
Voice-controlled e-commerce app enhances user experience for visually impaired.
problem Limited user experience for visually impaired in e-commerce applications.
method Proposes a voice-controlled e-commerce application using IBM Watson speech-to-text.
result Demonstrates enhanced usability for visually impaired users.
Study improves voice disorder detection system robust to channel effects.
problem Voice signals are sensitive to recording devices.
method Bidirectional LSTM network with domain adversarial training (DAT).
result Increased PR-AUC from 0.8448 to 0.9455 (and 0.9522 with labels).
End-to-end voice conversion without vocoder.
problem Speech conversion without vocoder.
method Transformer network for raw spectrum conversion.
result Transformer model converts real voices efficiently.
Improved voice conversion with semi-supervised learning.
problem Voice conversion with limited parallel data.
method Amortized variational inference with parallel and non-parallel utterances.
result Semi-supervised training improves voice conversion performance.
One-shot VC model converts voices without parallel data.
problem Limited VC applicability due to training data restriction.
method Disentangles speaker and content representations with instance normalization.
result Model converts voices from unseen speakers with high similarity.
AUTOVC converts voices without parallel data, achieving state-of-the-art results.
problem Non-parallel many-to-many voice conversion and zero-shot voice conversion.
method Only an autoencoder with a carefully designed bottleneck is used, training on a self-reconstruction loss.
result AUTOVC achieves state-of-the-art results in many-to-many voice conversion with non-parallel data and performs zero-shot voice conversion.
Semi-supervised singing voice separation using synthetic mixtures.
problem Singing voice separation with limited labeled data.
method Trains a single mapping function g on synthetic mixtures of singing and instrumental music.
result Performance comparable to fully supervised methods, better than semi-supervised alternatives.
New method separates multiple voices in mixed audio.
problem Separating multiple simultaneous speakers in audio.
method Gated neural networks trained at multiple steps, selecting actual number of speakers.
result Outperforms current state of the art for more than two speakers.
Study improves low-quality speech data for voice cloning.
problem Feasibility of training spoofing systems with low-quality data.
method Developed a GAN-based speech enhancement system, trained TTS and voice conversion models.
result Significant improvement in SNR and perceptual cleanliness of low-quality data.
A new neural network separates singing voices more effectively.
problem Separating singing voices from mixed signals with high accuracy.
method MBR-FCN that processes different frequency bands with varying resolutions and filters.
result The MBR-FCN achieves better performance with fewer parameters.
The paper uses neural networks to convert one speaker's voice to another.
problem Identifying speakers uniquely based on their voice.
method Uses convolutional neural networks to manipulate pitch and timbre.
result Preliminary results show encouraging voice conversion.
Improved U-Nets with various intermediate blocks enhance singing voice separation.
problem Improving singing voice separation accuracy using U-Net architectures.
method Implemented and compared U-Nets with different intermediate spectrogram transformation blocks.
result A specific block type achieves state-of-the-art SDR by 0.9 dB.
Privacy-preserving method protects user speech data from cloud services.
problem Privacy compromise in cloud-based speech analysis.
method Collects and sanitizes speech data before sharing, using transformation functions and voice conversion.
result Identification of sensitive emotional state reduced by ~96%.
Wave-U-Net with MHE regularization improves singing voice separation.
problem Singing voice separation from mixed music recordings.
method Wave-U-Net architecture with MHE regularization applied to 1D filters.
result Adding MHE regularization to the loss function consistently improves singing voice separation.
Improved autoencoder for F0-consistent voice conversion.
problem Non-parallel many-to-many voice conversion with prosodic information leakage.
method Conditional autoencoder with information-constraining bottlenecks.
result Controlled F0 contour and improved speech quality.
System generates speech textures and converts voices using backpropagation.
problem Generating speech textures and voice conversion from limited data.
method Approximate inversion of speech recognition network, matching neuron activations.
result System can generate realistic speech babble and reconstruct voices with limited data.
Proposes a new voice conversion model that preserves pitch patterns.
problem Preserving pitch patterns while changing speaker identity.
method Variational-autoencoder-based model with an auxiliary network.
result Ensures the conversion result correctly reflects specified F0/timbre information.
Paper surveys and introduces Acoustic Dialect Decoder for voice translation.
problem Machine understanding of natural language in speech translation.
method Recognition, Translation, and Synthesis units using HMMs, RNNs, and HTS.
result Initial successful translation of English to Tamil.
Blow converts non-parallel raw audio voices efficiently.
problem Voice conversion with non-parallel data.
method Single-scale normalizing flow with hypernetwork conditioning.
result Blow outperforms existing flow-based architectures in voice conversion.
Improved speech recognition for voice assistants by analyzing speech data.
problem Reducing false triggers in speech-enabled assistants.
method Post-processing LVCSR hypothesis lattice with a Bidirectional Lattice Recurrent Neural Network (LatticeRNN).
result LatticeRNN significantly improves detection accuracy over traditional methods.
WeSinger improves singing voice synthesis with data augmentation and specialized modules.
problem Improving the accuracy and naturalness of synthesized singing voices.
method Developed a multi-singer Chinese neural singing voice synthesis system with deep bi-directional LSTM, Transformer, LPCNet, and data augmentation.
result WeSinger achieves state-of-the-art performance on the Opencpop corpus.
Challenge evaluates voice conversion systems using parallel and non-parallel data.
problem Evaluate and compare different voice conversion systems.
method Hub and Spoke tasks with crowdsourced evaluation.
result Naturalness and similarity ratings provided for submitted systems.
Many-to-Many VTN improves voice conversion across multiple speakers.
problem Voice conversion across multiple speakers.
method Sequence-to-sequence learning framework with many-to-many VTN architecture.
result Improved sound quality and speaker similarity compared to baseline methods.
Paper proposes CycleGAN for nonparallel voice conversion.
problem Difficulty in achieving high-quality voice conversion with nonparallel data.
method Cycle-consistent adversarial network (CycleGAN) for nonparallel data-based voice conversion.
result The proposed method outperforms parallel VC methods in subjective evaluations.
Study improves voice conversion model with Mel-spectrogram augmentation.
problem Insufficient speech pairs data for training sequence-to-sequence voice conversion models.
method Experimented with Mel-spectrogram augmentation using SpecAugment policies and proposed new augmentation policies.
result Time axis warping policies showed better performance in training the voice conversion model.
Voice conversion methods degrade in noisy conditions, with BLFWAS outperforming others.
problem Robustness of voice conversion techniques under mismatched conditions.
method Comparative analysis of five VC techniques on CMU ARCTIC corpus, exploring speech enhancement techniques.
result Bilinear frequency warping with amplitude scaling (BLFWAS) outperforms other methods in noisy conditions.
New neural network separates singing voices from music using cross entropy loss.
problem Separating singing voices from music accompaniment.
method Deep Convolutional Neural Network (CNN) trained with Ideal Binary Mask (IBM) and cross entropy loss.
result Proposed CNN outperforms existing systems in MIREX evaluations.
New algorithm separates vocals from music recordings efficiently.
problem Separate vocal and instrumental parts in music recordings.
method Informed group-sparse representation for linear-time singing voice separation.
result Efficacy confirmed on iKala dataset; music accompaniment follows group-sparse structure.
VoiceFilter separates target speaker from multi-speaker signals.
problem Speech recognition in multi-speaker environments.
method Speaker recognition network and spectrogram masking network trained together.
result Significant reduction in speech recognition WER on multi-speaker signals.
Paper proposes StarGAN-VC for non-parallel voice conversion.
problem Non-parallel many-to-many voice conversion.
method Variant of GAN called StarGAN for simultaneous many-to-many mappings.
result Obtained higher sound quality and speaker similarity than state-of-the-art methods.
Dr.VOT measures both positive and negative VOTs accurately in natural speech.
problem Accurate measurement of VOT in natural speech.
method Deep-learning model based on RNNs for structured prediction.
result Dr.VOT improves over state-of-the-art performance on VOT estimation.
Personal VAD detects target speaker voice activity efficiently.
problem Efficiently detect target speaker voice activity for reduced computational cost and battery usage.
method Trains a neural network conditioned on speaker embedding or verification score, outputs probabilities for three speech classes.
result Trained model with 130K parameters outperforms combined standard VAD and speaker recognition networks.
Paper introduces a new model for polyphonic music composition.
problem Creating music with multiple interwoven voices.
method Developed a coupled recurrent model using probabilistic factorization and neural network ideas.
result Trained models for single-voice and multi-voice composition on a large dataset.
A-StarGAN improves nonparallel voice conversion speed and realism.
problem Nonparallel voice conversion without parallel data.
method Augmented classifier StarGAN (A-StarGAN) for nonparallel voice conversion.
result A-StarGAN generates realistic-sounding speech quickly and efficiently.
Transformer models show distinct spectral fingerprints under voice changes.
problem Detecting architectural biases in transformer models.
method Spectral analysis of attention-induced token graphs.
result Clear architectural signatures in model fingerprints correlate with language-specific behavior.
New algorithms separate singing voices from accompaniment using complex and quaternionic principal component pursuit.
problem Separating singing voices from instrumental accompaniment using phase information.
method Extended principal component pursuit to complex and quaternionic cases, developed new proximity operators, applied inexact augmented Lagrange multiplier algorithm.
result Phase information improves singing voice separation.
Study compares neural and statistical models for Parkinson's disease progression from voice data.
problem Difficult statistical analysis of longitudinal voice biomarkers due to subject correlation, small cohorts, and varied disease trajectories.
method Evaluated Neural Mixed Effects (NME), Generalized Neural Network Mixed Models (GNMMs), and semi-parametric Generalized Additive Mixed Models (GAMMs).
result GAMMs achieve stronger predictive performance and retain interpretable smooth effects and subject-level structure.
Paper proposes ACVAE-VC for non-parallel voice conversion.
problem Non-parallel many-to-many voice conversion with attribute class label retention.
method Uses ACVAE with fully convolutional networks, information-theoretic regularization, and auxiliary classifier.
result Successfully retains attribute class labels and avoids buzzy speech.
MaskCycleGAN-VC improves voice conversion without parallel data.
problem Limited ability to convert mel-spectrogram data without parallel data.
method Integrates a novel auxiliary task called filling in frames (FIF) to learn time-frequency structures.
result MaskCycleGAN-VC outperforms existing methods with similar model size.
Arcades uses deep learning to adapt smart-home behavior based on context.
problem Adaptive decision making in voice-controlled smart-home environments.
method Deep reinforcement learning with graphical representation of home automation system.
result Arcades promises long-term context-aware control of smart-home systems.
Joint training model for TTS and VC tasks using Tacotron and WaveNet.
problem Training a shared model for text-to-speech and voice conversion.
method Extended Tacotron model with dual attention mechanism for shared tasks, WaveNet for waveform generation.
result Joint training of a shared model achieves both TTS and VC tasks efficiently.