This paper addresses converting speech to EGG signals without hardware, improving accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this work we explored building automatic speech recognition models for transcribing doctor patient conversation. We collected a large scale dataset of clinical conversations ( hr), designed the task to represent the real word scenario, and explored several alignment approaches to iteratively improve data qua…
Enhanced transformer converts whispered speech to natural speech.
We propose a nonparallel data-driven emotional speech conversion method. It enables the transfer of emotion-related characteristics of a speech signal while preserving the speaker's identity and linguistic content. Most existing approaches require parallel data and time alignment, which is not available in most real ap…
End-to-end voice conversion without vocoder.
Improved autoencoder for F0-consistent voice conversion.
We propose a learning-based filter that allows us to directly modify a synthetic speech waveform into a natural speech waveform. Speech-processing systems using a vocoder framework such as statistical parametric speech synthesis and voice conversion are convenient especially for a limited number of data because it is p…
A new TTS method uses diffusion and VAE for better speech synthesis.
Improved SAR in asynchronous conversations using neural models and unlabeled data.
Joint training model for TTS and VC tasks using Tacotron and WaveNet.
Improved voice conversion with semi-supervised learning.
This paper proposes a non-parallel many-to-many voice conversion (VC) method using a variant of the conditional variational autoencoder (VAE) called an auxiliary classifier VAE (ACVAE). The proposed method has three key features. First, it adopts fully convolutional architectures to construct the encoder and decoder ne…
Improved speech enhancement with larger neural networks using novel embeddings and biases.
Inspired by recent work on neural network image generation which rely on backpropagation towards the network inputs, we present a proof-of-concept system for speech texture synthesis and voice conversion based on two mechanisms: approximate inversion of the representation learned by a speech recognition neural network,…
This paper proposes a method that allows non-parallel many-to-many voice conversion (VC) by using a variant of a generative adversarial network (GAN) called StarGAN. Our method, which we call StarGAN-VC, is noteworthy in that it (1) requires no parallel utterances, transcriptions, or time alignment procedures for speec…
A-StarGAN improves nonparallel voice conversion speed and realism.
This paper converts speech to match a face image and vice versa.
This paper describes a method based on a sequence-to-sequence learning (Seq2Seq) with attention and context preservation mechanism for voice conversion (VC) tasks. Seq2Seq has been outstanding at numerous tasks involving sequence modeling such as speech synthesis and recognition, machine translation, and image captioni…
Neural network converts speech from one language to multiple languages.
In this paper, we propose a dictionary update method for Nonnegative Matrix Factorization (NMF) with high dimensional data in a spectral conversion (SC) task. Voice conversion has been widely studied due to its potential applications such as personalized speech synthesis and speech enhancement. Exemplar-based NMF (ENMF…
Solution for voice conversion with limited data using hierarchical seq2seq and attention models.
This paper proposes a novel selective autoencoder approach within the framework of deep convolutional networks. The crux of the idea is to train a deep convolutional autoencoder to suppress undesired parts of an image frame while allowing the desired parts resulting in efficient object detection. The efficacy of the fr…
Voice conversion (VC) aims at conversion of speaker characteristic without altering content. Due to training data limitations and modeling imperfections, it is difficult to achieve believable speaker mimicry without introducing processing artifacts; performance assessment of VC, therefore, usually involves both speaker…
Thanks to the growing availability of spoofing databases and rapid advances in using them, systems for detecting voice spoofing attacks are becoming more and more capable, and error rates close to zero are being reached for the ASVspoof2015 database. However, speech synthesis and voice conversion paradigms that are not…
We discuss the structure of "exceptional generalised geometry" (EGG), an extension of Hitchin's generalised geometry that provides a unified geometrical description of backgrounds in eleven-dimensional supergravity. On a d-dimensional background, as first described by Hull, the action of the generalised geometrical O(d…
Most of the existing studies on voice conversion (VC) are conducted in acoustically matched conditions between source and target signal. However, the robustness of VC methods in presence of mismatch remains unknown. In this paper, we report a comparative analysis of different VC techniques under mismatched conditions. …
End-to-end model detects articulatory features from speech data.
This paper proposes a voice conversion (VC) method using sequence-to-sequence (seq2seq or S2S) learning, which flexibly converts not only the voice characteristics but also the pitch contour and duration of input speech. The proposed method, called ConvS2S-VC, has three key features. First, it uses a model with a fully…
Survey reviews code-switched speech and language processing.
Paper proposes M2H-GAN to improve speech theme identification.
Improved voice conversion model using single generator.
Many-to-Many VTN improves voice conversion across multiple speakers.
FasterVoiceGrad speeds up VC by 6-7x with novel distillation.
Large speech dataset for commercial use with 9.98% word error rate.
Recent speech technology research has seen a growing interest in using WaveNets as statistical vocoders, i.e., generating speech waveforms from acoustic features. These models have been shown to improve the generated speech quality over classical vocoders in many tasks, such as text-to-speech synthesis and voice conver…
CycleGAN-VC2 improves voice conversion without parallel data.
A fast voice conversion method using diffusion models.
Proposes a new voice conversion model that preserves pitch patterns.
Paper proposes singing voice conversion without parallel data.
Improved CSKS with limited data using novel loss functions and transfer learning.
LGD breaks the chicken-and-egg loop in adaptive SGD by using LSH sampling.
Convolutional network converts speaker voices without text.
The fundamental frequency (F0) contour of speech is a key aspect to represent speech prosody that finds use in speech and spoken language analysis such as voice conversion and speech synthesis as well as speaker and language identification. This work proposes new methods to estimate the F0 contour of speech using deep …
ACI converts call center conversations into actionable data.
Privacy-preserving method protects user speech data from cloud services.
Although voice conversion (VC) algorithms have achieved remarkable success along with the development of machine learning, superior performance is still difficult to achieve when using nonparallel data. In this paper, we propose using a cycle-consistent adversarial network (CycleGAN) for nonparallel data-based VC train…
Deep neural networks (DNNs) are now a central component of nearly all state-of-the-art speech recognition systems. Building neural network acoustic models requires several design decisions including network architecture, size, and training loss function. This paper offers an empirical investigation on which aspects of …
AUTOVC converts voices without parallel data, achieving state-of-the-art results.