Universal music translation network across instruments and genres.
problem Translating music across different instruments, genres, and styles.
method Multi-domain wavenet autoencoder with a shared encoder and disentangled latent space trained end-to-end on waveforms.
result Achieves convincing translations even from domains not seen during training.
Stochastic WaveNet models sequential data with latent variables and dilated convolutions.
problem Modeling distribution of sequential data like speech and motions.
method Combines stochastic latent variables and dilated convolutions in WaveNet architecture.
result Obtains state-of-the-art performances on speech and handwriting datasets.
NSF models generate speech waveforms faster and better than WaveNet.
problem Efficiently generating speech waveforms for statistical parametric synthesis.
method Neural source-filter (NSF) models that combine sine-based excitation, non-AR filter, and conditional preprocessing.
result NSF models generate waveforms 100 times faster than WaveNet and have better quality.
Recent speech technology research has seen a growing interest in using WaveNets as statistical vocoders, i.e., generating speech waveforms from acoustic features. These models have been shown to improve the generated speech quality over classical vocoders in many tasks, such as text-to-speech synthesis and voice conver…
WaveNet reconstructs speech from brain activity, revealing acoustic features.
problem Reconstructing speech from brain activity with limited data.
method WaveNet model applied to STG intracranial recordings.
result WaveNet models reveal phoneme-level acoustic features.
Adaptive multi-domain learning reduces parameter count for efficient deep learning.
problem Different domains have varying complexity, leading to inefficient model training.
method Proposes adaptive parameterization to reduce model complexity without sacrificing performance.
result Efficient multi-domain learning solutions with far fewer parameters.
Dual adversarial co-learning improves multi-domain text classification.
problem Improving text classification across multiple domains.
method Dual adversarial co-learning with shared-private networks and dual adversarial regularizations.
result Achieves state-of-the-art performance on multi-domain sentiment classification datasets.
MuLANN tackles multi-domain learning with adversarial approach.
problem Automated microscopy data with domain bias.
method Semi-supervised multi-domain learning with MuLANN.
result Improves state of the art on image benchmarks and bioimage dataset.
A faster neural waveform model for speech synthesis.
problem Slow waveform generation in existing neural models.
method Proposes a non-autoregressive neural source-filter model.
result Generated waveforms 100 times faster than AR WaveNet.
Proposes MD-LiNA for multi-domain latent factor causal discovery.
problem Discovering causal structures among latent factors from multi-domain data.
method Multi-Domain Linear Non-Gaussian Acyclic Models (MD-LiNA) with an integrated two-phase algorithm.
result Locally consistent estimators of causal structure among shared latent factors.
New framework for multi-domain translation using autoencoders.
problem Learning probabilistic coupling between different domains.
method Learning multiple uncoupled autoencoders under shared latent distribution.
result New autoencoders can be added sequentially without retraining.
Proposes a model to improve multi-domain recommender systems.
problem Challenges in transferring knowledge between domains in recommender systems.
method Generative adversarial networks (GANs), Variational Autoencoders (VAEs), and Cycle-Consistency (CC) for weight-sharing.
result Improves performance of multi-domain recommender systems by capturing both similarities and differences among domains.
Express Wavenet reduces neural network parameters to 1% of standard networks.
problem Optical neural networks with high parameter count.
method Wavelet modulation, random shift wavelets, expressway structure.
result Express Wavenet achieves high accuracy with significantly fewer parameters.
Neural network VQ-VAE with WaveNet decodes speech at 1.6 kbps with high quality.
problem Efficiently transmitting and storing speech signals at low bit-rates.
method VQ-VAE and WaveNet architecture for speech coding.
result Speech coding at 1.6 kbps with perceptual quality between MELP and AMR-WB.
TimbreTron transfers musical timbre using CQT and WaveNet.
problem Transfer musical timbre while preserving pitch, rhythm, and loudness.
method Apply image domain style transfer to CQT representation, then generate high-quality waveform with WaveNet.
result TimbreTron recognizably transfers timbre while preserving musical content.
Joint training model for TTS and VC tasks using Tacotron and WaveNet.
problem Training a shared model for text-to-speech and voice conversion.
method Extended Tacotron model with dual attention mechanism for shared tasks, WaveNet for waveform generation.
result Joint training of a shared model achieves both TTS and VC tasks efficiently.
Neural model synthesizes music with flexible timbre controls.
problem Creating audio samples with varied timbres from musical scores.
method Recurrent neural network conditioned on learned instrument embedding followed by WaveNet vocoder.
result Learned embedding space captures diverse timbres and enables interpolation for morphing.
WaveGlow generates high-quality speech from spectrograms.
problem Speech synthesis quality and efficiency.
method WaveGlow combines Glow and WaveNet insights, using a single network and cost function for efficient, high-quality audio synthesis.
result WaveGlow produces audio samples at over 500 kHz, matching WaveNet quality.
Improved multi-speaker TTS using GANs and waveform loss.
problem Training acoustic models for neural vocoders in multi-speaker TTS systems.
method Proposed frameworks incorporating Wasserstein GAN with gradient penalty (WGAN-GP) and discretized mixture logistic loss (DML) into acoustic models trained with WaveNet.
result Acoustic models trained with WGAN-GP and DML loss achieve highest subjective evaluation scores in multi-speaker TTS.
MetalGAN synthesizes images across multiple domains without labels.
problem Synthesizing images across multiple domains using a single network.
method Combines cGAN for image generation and Meta-Learning for domain switch.
result MetalGAN successfully produces multi-domain images without hard-coded labels.
GANSynth uses GANs to efficiently synthesize high-fidelity audio.
problem Efficient and high-fidelity audio synthesis is challenging.
method Model log magnitudes and instantaneous frequencies with GANs.
result GANSynth outperforms WaveNet on automated and human evaluation metrics.
Unsupervised learning of speech representations using WaveNet autoencoders.
problem Extract meaningful latent representations of speech signals.
method Applying autoencoding neural networks to speech waveforms, using a high capacity WaveNet decoder, and comparing three variants of latent representations.
result Comparable performance with top entries in the ZeroSpeech 2017 unsupervised acoustic unit discovery task.
New method identifies stable latent variables across different domains using weak distributional invariances.
problem Learning causal representations for multi-domain datasets.
method Autoencoders incorporating weak distributional invariances.
result Autoencoders can identify stable latent variables across different domains.
This study improves text-to-speech synthesis using GANs for glottal excitation.
problem Slow inference and computational cost of WaveNet and difficulty in parallel training of GANs.
method Adopted GANs for parallel waveform generation in speech signal and glottal excitation.
result GAN-based glottal excitation model achieves quality and voice similarity on par with WaveNet.
Graph WaveNet models spatial-temporal graphs by learning hidden dependencies and long sequences.
problem Capturing hidden spatial dependencies and long-range temporal sequences in graphs.
method Graph WaveNet integrates adaptive dependency matrix learning and stacked dilated 1D convolution.
result Graph WaveNet outperforms existing methods on public traffic network datasets.
We address the problems of multi-domain and single-domain regression based on distinct and unpaired labeled training sets for each of the domains and a large unlabeled training set from all domains. We formulate these problems as a Bayesian estimation with partial knowledge of statistical relations. We propose a worst-…
Improves dialogue state tracking across multiple domains.
problem Incomplete domain ontology limits DST models' adaptability.
method Model DST as Q&A, using evolving knowledge graph.
result 5.80% and 12.21% relative improvement on datasets.
Meta-learning approach for adaptive TTS with few data.
problem Adapting TTS systems to new speakers with minimal data.
method Meta-learning with shared WaveNet core and independent speaker embeddings, using three training strategies.
result Successful adaptation of multi-speaker neural network to new speakers with minimal data.
Estimates causal effect using proxies in multi-domain settings.
problem Estimating causal effect in settings with unobserved confounders across domains.
method Proposes estimation techniques using proxy variables for discrete or categorical data.
result Proves identifiability and consistency of causal effect estimation.
Combines symbolic and raw audio models for structured, realistic-sound music generation.
problem Lack of long-range dependencies in raw audio models and unstructured music.
method Uses a Long Short Term Memory network for melodic structure and WaveNet for raw audio generation with symbolic conditioning.
result Creates structured, realistic-sounding compositions using both symbolic and raw audio models.
Unified framework for multi-domain learning and data imputation.
problem Improving performance across different domains with missing data.
method Adversarial autoencoder for domain-invariant embeddings and data imputation.
result Superior performance compared to state-of-the-art methods in various settings.
Bayesian model learns cancer subtypes from diverse NGS data.
problem Overdispersed NGS count data and limited samples for specific cancer types.
method Bayesian Multi-Domain Learning (BMDL) model using hierarchical negative binomial factorization.
result BMDL achieves reproducible cancer subtyping without negative transfer effects.
WaveCycleGAN2 improves speech synthesis quality by reducing aliasing.
problem Human ear can still distinguish synthesized speech from natural speech.
method WaveCycleGAN2 uses generators without down/up-sampling modules and combines discriminators from waveform and acoustic parameter domains.
result WaveCycleGAN2 achieves high-quality speech synthesis with comparable mean opinion scores to natural speech.
Proposes adversarial normalization for multi-domain image segmentation.
problem Current image normalization is per-dataset, limiting multi-domain segmentation.
method Adversarial training to learn common normalizing functions across multiple datasets.
result Optimal normalizer improves segmentation accuracy and realism.
Paper develops a new method to improve model calibration under distribution shifts.
problem Challenges in uncertainty quantification with different training and test distributions.
method Develops multi-domain temperature scaling to handle distribution shifts.
result Outperforms existing methods on in-distribution and out-of-distribution test sets.
Domain Fusion uses GANs to augment data for low-volume target datasets.
problem High costs in data development for deep learning applications.
method Multi-domain learning GANs to generate new samples.
result Domain Fusion achieves better classification accuracy with less data.
Learn to automatically plug domain-specific modules into a common network.
problem Learning inflexibility and computational intensiveness in multi-domain learning.
method Neural Architecture Search (NAS) for data-driven adapter plugging and structure design.
result NAS-driven MDL model achieves comparable performance to existing approaches.
Enhances performance on downstream tasks using multi-domain data.
problem Improving performance on tasks with limited downstream data.
method Deep transfer learning framework leveraging shared and domain-specific features.
result Significantly improves convergence rate for learning Lipschitz functions.
Framework integrates mental disorder measurements for personalized treatment.
problem Optimizing treatment for mental disorders with latent mental states and heterogeneity.
method Measurement theory and multi-layer neural network for complex treatment effects.
result Learned treatment policies outperform alternatives on heterogeneous treatment effects.
TESTED improves multi-domain stance detection with topic-guided sampling and contrastive learning.
problem Challenges in multi-domain stance detection due to domain-specific variations and imbalanced annotations.
method Topic-guided diversity sampling and contrastive learning objective.
result Significant improvement in F1 scores, up to 10.2 points out-of-domain.
Empirical Bayes improves causal representation learning across multiple domains.
problem Estimating causal representations from data across multiple domains.
method Developed an EB f-modeling algorithm for linearly-mixed causal representations. result Our method achieves more accurate estimation of causal variables than other methods.
Recent advances in speech synthesis suggest that limitations such as the lossy nature of the amplitude spectrum with minimum phase approximation and the over-smoothing effect in acoustic modeling can be overcome by using advanced machine learning approaches. In this paper, we build a framework in which we can fairly co…
New findings on how convolutional architectures approximate time series data.
problem Understanding the approximation properties of convolutional architectures in time series modeling.
method Mathematical analysis of convolutional architectures applied to time series modeling.
result A new definition of spectrum-based regularity for measuring temporal relationships under convolutional approximation.
Single CNN removes multiple ultrasound artifacts.
problem Efficiently remove multiple ultrasound artifacts.
method OT-driven multi-domain unsupervised deep learning.
result Single neural network removes various artifacts.
Anchor PCA improves robustness in multi-domain PCA.
problem PCA on pooled data can focus on spurious directions.
method Anchor PCA focuses on shared directions of variation.
result Anchor PCA outperforms pooling and worst-case alternatives.
Proposes a unified normalization method for multi-domain medical images.
problem Inadequate joint information across multiple datasets hinders image segmentation performance.
method Adversarial and task-driven normalization approach to learn a common normalizing function across multiple datasets.
result Jointly normalized images improve segmentation accuracy by up to 57.5%.
In this paper, we provide a new neural-network based perspective on multi-task learning (MTL) and multi-domain learning (MDL). By introducing the concept of a semantic descriptor, this framework unifies MDL and MTL as well as encompassing various classic and recent MTL/MDL algorithms by interpreting them as different w…
Robust image translation model for noisy labels.
problem Learning mappings among multiple domains with noisy labeled data.
method Proposes a novel loss function and techniques to handle noisy labeled data.
result Demonstrates robustness in various settings including synthetic and real-world noise.