WaveCycleGAN converts synthetic speech to natural speech using cycle-consistent adversarial networks.
problem Over-smoothing effect in synthetic speech, leading to quality degradation.
method Cycle-consistent adversarial networks for waveform-level modification.
result Improves naturalness of generated speech sounds.
VPFD uses vocoder features for adversarial training in VC.
problem Adversarial training on waveform data is time-consuming and memory-intensive.
method VPFD employs vocoder features for adversarial training.
result VPFD achieves VC performance comparable to waveform discriminators with reduced training time and memory.
Convolutional network converts speaker voices without text.
problem Speaker conversion without text-based methods.
method Fully convolutional wav-to-wav network with ASR pre-training.
result Successfully converts TTS robot's voice to narrated audiobook voices.
Method converts facial expressions and voice of a source speaker into a target speaker.
problem Separate conversion of facial and acoustic features leads to unnatural results.
method Uses three neural networks: conversion, waveform generation, and image reconstruction.
result Significantly higher naturalness achieved when converting both features together.
Recent speech technology research has seen a growing interest in using WaveNets as statistical vocoders, i.e., generating speech waveforms from acoustic features. These models have been shown to improve the generated speech quality over classical vocoders in many tasks, such as text-to-speech synthesis and voice conver…
A new TTS method uses diffusion and VAE for better speech synthesis.
problem Improving text-to-speech synthesis for better speech quality and robustness.
method Combines diffusion probabilistic model and variational autoencoder for latent variable conversion.
result The method is robust to poor orthography and alignment errors.
iSTFTNet speeds up mel-spectrogram vocoders without sacrificing quality.
problem Efficiently converting mel-spectrograms to speech with minimal computation.
method Replaces convolutional layers with iSTFT after frequency dimension reduction.
result Significant reduction in computational cost with comparable quality.
Joint training model for TTS and VC tasks using Tacotron and WaveNet.
problem Training a shared model for text-to-speech and voice conversion.
method Extended Tacotron model with dual attention mechanism for shared tasks, WaveNet for waveform generation.
result Joint training of a shared model achieves both TTS and VC tasks efficiently.
Inspired by recent work on neural network image generation which rely on backpropagation towards the network inputs, we present a proof-of-concept system for speech texture synthesis and voice conversion based on two mechanisms: approximate inversion of the representation learned by a speech recognition neural network,…
Paper proposes a new loss function for training neural speech models.
problem Training high-performance neural speech waveform models.
method Uses short-time Fourier transform (STFT) spectra and assumes Gaussian and von Mises distributions for amplitude and phase spectra.
result Synthesized high-quality speech waveforms.
DiffWave generates high-fidelity audio waveforms efficiently.
problem Conditional and unconditional audio waveform generation.
method Non-autoregressive diffusion model using Markov chain synthesis.
result DiffWave produces high-quality audios in various tasks.
Novel radar waveform design for autonomous vehicles.
problem Efficient radar waveform design in time-varying environments.
method Hybrid model-driven and data-driven architecture.
result Adaptive unimodular waveform design for real-time scenarios.
DYMAG uses dynamic waveforms to improve graph neural networks.
problem Improving graph neural networks for better graph understanding.
method DYMAG employs dynamical system-based waveforms for message aggregation in graph neural networks.
result DYMAG outperforms baseline models in graph recovery, property prediction, and random graph generation.
This paper improves speech recognition by using raw waveform signals in multi-span CNN acoustic models.
problem Improving speech recognition accuracy using raw waveform signals.
method Proposes a novel multi-span structure for acoustic modelling based on raw waveform signals with multiple CNN input layers.
result Multi-span acoustic models yield a lower word error rate (WER) than traditional FBANK feature-based models.
This study proposes a fully convolutional network (FCN) model for raw waveform-based speech enhancement. The proposed system performs speech enhancement in an end-to-end (i.e., waveform-in and waveform-out) manner, which dif-fers from most existing denoising methods that process the magnitude spectrum (e.g., log power …
NSF models generate speech waveforms faster and better than WaveNet.
problem Efficiently generating speech waveforms for statistical parametric synthesis.
method Neural source-filter (NSF) models that combine sine-based excitation, non-AR filter, and conditional preprocessing.
result NSF models generate waveforms 100 times faster than WaveNet and have better quality.
A faster neural waveform model for speech synthesis.
problem Slow waveform generation in existing neural models.
method Proposes a non-autoregressive neural source-filter model.
result Generated waveforms 100 times faster than AR WaveNet.
Paper presents a novel waveform-to-waveform model for music source separation.
problem Isolating individual instruments from mixed music recordings.
method Adapted Conv-Tasnet to waveform domain and developed Demucs with U-Net and bidirectional LSTM.
result Demucs achieves superior performance on music source separation tasks, surpassing existing state-of-the-art.
Develops a universal waveform selection scheme for radar tracking.
problem Optimal waveform selection for target tracking in active sensors.
method Uses reinforcement learning and universal source coding techniques.
result Achieves optimal waveform selection for any radar scene modeled as a Markov process.
Artificial neural networks infer gravitational-wave parameters from reduced-order waveforms.
problem Efficiently infer gravitational-wave parameters from noisy data.
method Represent waveforms as weighted sums over reduced bases, train neural networks to map source parameters to coefficients.
result Fast and accurate interpolation of gravitational-wave coefficients.
WaveCycleGAN2 improves speech synthesis quality by reducing aliasing.
problem Human ear can still distinguish synthesized speech from natural speech.
method WaveCycleGAN2 uses generators without down/up-sampling modules and combines discriminators from waveform and acoustic parameter domains.
result WaveCycleGAN2 achieves high-quality speech synthesis with comparable mean opinion scores to natural speech.
Real-time speech enhancement model removes various noises and reverb.
problem Real-time speech enhancement in noisy environments.
method Causal speech enhancement model using encoder-decoder architecture with skip-connections, optimized in time and frequency domains.
result The model matches state-of-the-art performance while working directly on raw waveform.
This study improves text-to-speech synthesis using GANs for glottal excitation.
problem Slow inference and computational cost of WaveNet and difficulty in parallel training of GANs.
method Adopted GANs for parallel waveform generation in speech signal and glottal excitation.
result GAN-based glottal excitation model achieves quality and voice similarity on par with WaveNet.
Thanks to the growing availability of spoofing databases and rapid advances in using them, systems for detecting voice spoofing attacks are becoming more and more capable, and error rates close to zero are being reached for the ASVspoof2015 database. However, speech synthesis and voice conversion paradigms that are not…
Adversarial attacks on spectrograms can fool audio classifiers trained on waveforms.
problem Susceptibility of audio classifiers to adversarial attacks on spectrograms.
method Applying adversarial attacks to spectrograms and reconstructing audio waveforms.
result Perturbed spectrograms can fool 2D CNNs and 1D CNNs trained on audio waveforms.
Deep neural network learns robust acoustic models from speech waveforms.
problem Robustness in speech recognition systems using standard feature extraction techniques.
method Deep convolutional neural network with stochastic variational inference and cosine modulated filters.
result Superior performance compared to baseline waveform-based models and deep CNNs with FBANK features.
AI model enhances grid monitoring with synchro-waveform tech.
problem Dynamic, stochastic, low-inertia future grids need advanced monitoring.
method AI Foundation Model with synchro-waveform tech.
result Significantly improved fault detection accuracy and speed.
Study compares geometric approaches for shape and deformation statistics.
problem Characterizing statistical models of shapes and deformations.
method Information geometry and Wasserstein geometry.
result Wasserstein estimator is robust against waveform perturbation.
Physics-consistent method improves seismic inversion accuracy.
problem Challenges in seismic full-waveform inversion (FWI) due to ill-posedness and high cost.
method Hybrid approach combining physics-based models with data-driven methodologies, incorporating physics into data augmentation.
result Physics-consistent data-driven inversion yields higher accuracy and better generalization.
Improved music source separation using unlabeled data remixing.
problem Music source separation with deep learning.
method Introduces a simple convolutional and recurrent model and a scheme to leverage unlabeled music.
result Waveform methods can now match spectrogram methods on standard benchmarks.
Machine learning classifies gravitational wave signals to test General Relativity.
problem Testing General Relativity with gravitational wave signals from binary black hole mergers.
method Convolutional Neural Networks (CNNs) trained on whitened waveforms and response function type observables.
result CNNs improve classification sensitivity by a factor of approximately 33 compared to whitened waveforms.
DeepClean detects and removes artefacts from ICU waveform data.
problem Accurate removal of artefacts from ICU waveform data reduces bias and uncertainty in clinical assessment.
method Self-supervised deep generative learning using a convolutional variational autoencoder.
result DeepClean detects artefacts with high sensitivity and specificity, significantly outperforming baseline methods.
WaveGrad generates high-fidelity audio using gradient estimation.
problem Generating high-fidelity audio efficiently.
method Conditional model using score matching and diffusion models, iteratively refining a Gaussian white noise signal.
result WaveGrad can generate high-fidelity audio samples using as few as six iterations.
SING generates musical notes from instruments in real-time.
problem Efficiently generating high-quality audio from MIDI data.
method Frame-by-frame waveform generation with a single decoder, using a new loss function.
result SING produces significantly improved audio quality compared to state-of-the-art models, with 32x faster training and 2,500x faster inference.
Hybrid model improves music source separation by 1.4 dB.
problem Improving music source separation accuracy.
method End-to-end hybrid spectrogram and waveform model, using model decision for domain choice.
result 1.4 dB improvement in Signal-to-Distortion (SDR) on MusDB HQ dataset.
Unsupervised method detects earthquakes from raw waveforms, generalizing across datasets.
problem Lack of labeled data for earthquake detection.
method Uses deep autoencoders with cross-covariance triggering at bottleneck.
result Performance comparable to supervised methods, with strong cross-dataset generalization.
Bayesian method refines surrogate models for accurate full waveform inversion.
problem Complex input/output relations in full waveform inversion make accurate surrogate models difficult.
method Iterative refinement of surrogate models using MCMC samples and progressively expanding frequency bandwidth.
result Highly accurate surrogate model across full bandwidth enables accurate final MCMC inversion.
This research improves neural synthesizers for music sounds from speech data.
problem Applying speech synthesis techniques to musical instrument sounds.
method Comparison of three neural synthesizers in three scenarios: training, zero-shot learning, and fine-tuning.
result Neural synthesizers trained on speech data and fine-tuned on music data perform better.
In this paper, a generalized multivariate Student-t mixture model is developed for classification and clustering of Low Probability of Intercept radar waveforms. A Low Probability of Intercept radar signal is characterized by a pulse compression waveform which is either frequency-modulated or phase-modulated. The propo…
The paper uses learned prototypes to explain deep learning models for time-series data.
problem Lack of explainable AI in deep learning models for high-risk decisions.
method Learned prototypes in latent space of deep learning models.
result Prototypes improve classification decisions and provide explainable insights.
GANSynth uses GANs to efficiently synthesize high-fidelity audio.
problem Efficient and high-fidelity audio synthesis is challenging.
method Model log magnitudes and instantaneous frequencies with GANs.
result GANSynth outperforms WaveNet on automated and human evaluation metrics.
Recent advances in speech synthesis suggest that limitations such as the lossy nature of the amplitude spectrum with minimum phase approximation and the over-smoothing effect in acoustic modeling can be overcome by using advanced machine learning approaches. In this paper, we build a framework in which we can fairly co…
Improves speech recognition in noisy environments using robust acoustic models.
problem Adverse environments with significant mismatch between training and test conditions.
method Theoretical analysis of data augmentation as vicinal risk minimization, using mixture of Gaussians to incorporate robust inductive bias.
result Waveform-based approach shows 150% relative improvement in out-of-distribution generalization.
Machine learning helps create accurate models of neutron star postmerger signals.
problem Creating accurate postmerger waveforms for binary neutron stars is challenging due to theoretical uncertainties and limited numerical simulations.
method Used a conditional variational autoencoder (CVAE) to construct postmerger models based on numerical-relativity simulations.
result The CVAE can accurately generate postmerger waveforms and encode the neutron star equation of state.
The paper proposes using CWT and STFT for training neural speech models.
problem Training high-quality neural speech models.
method Proposes spectral amplitude and phase losses from STFT and CWT for training.
result Shows that CWT spectral loss can train a high-quality model as good as STFT-based loss.
We make posterior sampling in FWI feasible for large surveys.
problem Uncertainty-aware subsurface models at field scale.
method Coupling diffusion-based posterior sampling with simultaneous-source FWI data.
result Lower model error and better data fit at reduced computational cost.
Realistic music generation is a challenging task. When building generative models of music that are learnt from data, typically high-level representations such as scores or MIDI are used that abstract away the idiosyncrasies of a particular performance. But these nuances are very important for our perception of musical…
Paper shows geometric frequency and Lagrange derivative equivalence for electric and fluid systems.
problem Understanding and classifying system operating conditions based on electric quantity waveform distortions.
method Demonstrates equivalence between geometric frequency and Lagrange derivative through numerical examples.
result Identifies components of Lagrange derivative that relate to geometric frequency and waveform distortions.