Combining data from multiple speakers improves neural TTS quality, especially with imbalanced data.
problem Training high-quality TTS systems with imbalanced speaker data.
method Combine data from multiple speakers, train multi-speaker models, and use ensemble methods.
result Ensemble multi-speaker models improve synthetic speech quality for underrepresented speakers.
Convolutional network converts speaker voices without text.
problem Speaker conversion without text-based methods.
method Fully convolutional wav-to-wav network with ASR pre-training.
result Successfully converts TTS robot's voice to narrated audiobook voices.
Meta-learning approach for adaptive TTS with few data.
problem Adapting TTS systems to new speakers with minimal data.
method Meta-learning with shared WaveNet core and independent speaker embeddings, using three training strategies.
result Successful adaptation of multi-speaker neural network to new speakers with minimal data.
Improved multi-speaker TTS using GANs and waveform loss.
problem Training acoustic models for neural vocoders in multi-speaker TTS systems.
method Proposed frameworks incorporating Wasserstein GAN with gradient penalty (WGAN-GP) and discretized mixture logistic loss (DML) into acoustic models trained with WaveNet.
result Acoustic models trained with WGAN-GP and DML loss achieve highest subjective evaluation scores in multi-speaker TTS.
RoyalFlush system improves multi-speaker ASR in M2MeT challenge.
problem Improving multi-speaker automatic speech recognition in noisy environments.
method Front-end processing with WPE and beamforming, data augmentation, and fusion of two ASR models.
result 12.22% absolute CER reduction on validation set and 12.11% on test set compared to baseline.
VoiceFilter separates target speaker from multi-speaker signals.
problem Speech recognition in multi-speaker environments.
method Speaker recognition network and spectrogram masking network trained together.
result Significant reduction in speech recognition WER on multi-speaker signals.
A new framework for real-time multi-speaker diarization without prior registration.
problem Real-time multi-speaker diarization without prior registration and pretraining.
method Semi-supervised and self-supervised learning methods applied to a new benchmark.
result Robust performance in the online MiniVox framework.
Single model performs well across different speaker scenarios.
problem Limited research on multi-scenario deep learning for multi-speaker separation.
method Combining data from different speaker scenarios into a single model.
result Single model outperforms scenario-specific models.
Paper proposes CNN-LSTM model for multi-speaker speech separation.
problem Multi-speaker source separation using deep learning.
method Parallel CNN-LSTM architecture with Bayesian hyperparameter optimization.
result Parallel CNN-LSTM model outperforms LSTM-only and CNN-only models.
The problem of multi-speaker localization is formulated as a multi-class multi-label classification problem, which is solved using a convolutional neural network (CNN) based source localization method. Utilizing the common assumption of disjoint speaker activities, we propose a novel method to train the CNN using synth…
Analyzes memory time span in LSTMs for multi-speaker speech separation.
problem Understanding how long-term dependencies are handled by LSTMs in speech separation tasks.
method Leaked state variable with controlled lifetime to evaluate task performance.
result Estimates the time span LSTMs exploit in multi-speaker speech separation.
Classifies Toda-type tt*-structures and their fixed points.
problem Classifying Toda-type tt*-structures and their fixed points.
method Fixed point description and reduction of anti-symmetry conditions.
result Reduces possibilities of anti-symmetry condition to two cases.
Using nonlinear pde techniques, we construct a new family of globally smooth tt* structures. This includes tt* structures associated to the (orbifold) quantum cohomology of a finite number of complex projective spaces and weighted projective spaces. The existence of such "magical solutions" of the tt* equations, namely…
End-to-end TTS learns context features from text input.
problem Lack of understanding of context features learned by end-to-end TTS.
method Evaluated encoder outputs against context criteria derived from parametric TTS.
result Encoder outputs reflect linguistic and phonetic context features.
The paper proves an isomorphism between tt∗ structures of Landau-Ginzburg and Calabi-Yau models.
problem Establishing an isomorphism between tt∗ structures of different geometries. method Using Landau-Ginzburg models and Calabi-Yau hypersurfaces, proving the isomorphism via the big residue map.
result An isomorphism between tt∗ structures of Landau-Ginzburg and Calabi-Yau models is proven. Relates quantum cohomology to tt*-Toda equations for minuscule flag manifolds.
problem Quantum cohomology of minuscule flag manifolds.
method Combining Lie-theoretic treatments of tt*-Toda equations and quantum cohomology.
result Relates quantum cohomology to tt*-Toda equations for minuscule flag manifolds.
Analyzes tt*-structures from ADE-type Stokes data.
problem Classifying tt*-structures over C∗. method Isomonodromic deformations with upper unitriangular real Stokes matrices.
result Establishes a direct analytic realization of the ADE classification. Polynomial growth elements found in all subgroups of Out(F_n).
problem Understanding polynomial growth in subgroups of Out(F_n).
method Analyzing conjugacy classes and elements of Out(F_n).
result Polynomial growth elements exist in all subgroups of Out(F_n).
Explains quantum cohomology of Grassmannians using tt* equations.
problem Relates quantum cohomology of complex Grassmannians to projective space.
method Uses tt* equations and Lie-theoretic connections.
result Illustrates relations between tt* equations and quantum cohomology.
Proves existence and uniqueness of solutions for A_n tt*-Toda equations.
problem Existence and uniqueness of solutions for A_n tt*-Toda equations.
method Proof of existence and uniqueness for any n, new treatment of asymptotic data.
result Existence and uniqueness of global solutions for any n.
We study transverse-tracefree (TT)-tensors on conformally flat 3-manifolds (M,g). The Cotton-York tensor linearized at g maps every symmetric tracefree tensor into one which is TT. The question as to whether this is the general solution to the TT-condition is viewed as a cohomological problem within an elliptic com…
Solves constant pre-factor problem for tt*-Toda equations using asymptotic data and symplectic structures.
problem Constant pre-factor problem for the tt*-Toda equations.
method Explicit evaluation using asymptotic data and introduction of symplectic structures.
result Preservation of symplectic structures by Riemann-Hilbert correspondence for wider class of solutions.
BOFFIN TTS optimizes hyper-parameters for new speaker adaptation.
problem Fine-tuning a pre-trained TTS model for a new speaker with limited data.
method Bayesian optimization to efficiently find optimal hyper-parameters.
result Average 30% improvement in speaker similarity over standard techniques.
Study of symplectic groupoids from tt*-Toda equations.
problem Geometry of meromorphic connections with irregular singularities.
method Holomorphic symplectic groupoid structure over Steinberg cross section.
result Proves the space of tt*-Toda connections is a symplectic Lie groupoid.
End-to-end TTS framework uses hard alignment to improve accuracy.
problem End-to-end TTS systems struggle with accurate alignment between input text and output acoustic features.
method Proposes a constrained alignment scheme with hard monotonic alignments, marginalized during training.
result Improves alignment learning and prediction in end-to-end TTS systems.
Solutions of tt*-equation from SU(2)_k fusion algebra.
problem Describe solutions to the tt*-equation from a specific algebra.
method Use DPW method and representations of SU(2).
result Construct solutions corresponding to A_k minimal model.
New surfaces with conjugate points have global blow-down maps in their TT spaces.
problem Constructing global blow-down maps for surfaces with conjugate points.
method Explicit construction of a family of non-trapping Riemannian surfaces with global blow-down maps.
result Global blow-down maps exist for some non-simple surfaces with conjugate points.
New unsupervised speaker adaptation method for speech synthesis.
problem Adapting speech synthesis to new speakers with minimal data.
method Concatenating audio and text inputs, proposing new training schemes.
result Improves adaptation to unseen speakers and multi-speaker modeling.
Proposes Textual Echo Cancellation to improve speech recognition.
problem Improving speech recognition performance and user experience for smart devices.
method A novel sequence-to-sequence model with multi-source attention that processes both the microphone mixture signal and source text of TTS playback.
result Demonstrates enhanced speech recognition performance and reduced latency.
Proposes DNN-based speaker embedding correlated with subjective inter-speaker similarity for speech synthesis.
problem Inadequate speaker representation for open speakers not in training data.
method Two training algorithms using inter-speaker similarity matrices: similarity vector embedding and similarity matrix embedding.
result Proposed algorithms learn speaker embedding highly correlated with subjective inter-speaker similarity.
Representation mixing combines character and phoneme inputs for flexible TTS synthesis.
problem Limited control over pronunciation in character or phoneme-based TTS systems.
method Representation mixing combines multiple linguistic inputs in a single encoder.
result Flexibility in choosing between character, phoneme, or mixed representations during inference.
Investigates neural TTS systems for Japanese and English.
problem Improving neural TTS systems for high-quality speech synthesis.
method Comparative study of neural sequence-to-sequence TTS vs. DNN pipeline TTS, varying model architecture, parameter size, and language.
result A neural sequence-to-sequence TTS system requires sufficient model parameters and a powerful encoder for high-quality speech synthesis.
RandUCB combines UCB and TS for optimal bandit performance.
problem Optimizing decision-making in uncertain environments with limited feedback.
method Randomized UCB algorithm using confidence intervals.
result Achieves minimax-optimal regret in various bandit settings.
FPETS speeds up TTS by 600X and reduces errors.
problem High latency and errors in end-to-end TTS systems.
method Non-autoregressive, fully parallel approach with UFANS and trainable position encoding.
result Significant speed up and better quality audios with fewer errors.
Bayesian tensor train method recovers streaming data with high accuracy.
problem Recovering high-order, incomplete, and noisy streaming data.
method Bayesian tensor train decomposition using streaming variational Bayes method.
result The proposed SPTT algorithm excels in recovering streaming data compared to state-of-the-art methods.
Recent speech technology research has seen a growing interest in using WaveNets as statistical vocoders, i.e., generating speech waveforms from acoustic features. These models have been shown to improve the generated speech quality over classical vocoders in many tasks, such as text-to-speech synthesis and voice conver…
Recently, there has been growing interest in multi-speaker speech recognition, where the utterances of multiple speakers are recognized from their mixture. Promising techniques have been proposed for this task, but earlier works have required additional training data such as isolated source signals or senone alignments…
We describe all smooth solutions of the two-function tt*-Toda equations (a version of the tt* equations, or equations for harmonic maps into SL(n,R)/SO(n)) in terms of (i) asymptotic data, (ii) holomorphic data, and (iii) monodromy data. This allows us to find all solutions with integral Stokes data. These include solu…
Paper proposes efficient tensor completion method using Gaussian Process.
problem Tensor completion in high-dimensional data with unknown smooth functions.
method Gaussian Process Regression for initialization and TT-cross approximation for tensor rank selection.
result Improved reconstruction error compared to random initialization.
In "Isomonodromy aspects of the tt* equations of Cecotti and Vafa I. Stokes data" (arxiv:1209.2045) we described all smooth solutions of the two-function tt*-Toda equations in terms of asymptotic data, holomorphic data, and monodromy data. In this supplementary article we focus on the holomorphic data and its interpret…
A new TTS method uses diffusion and VAE for better speech synthesis.
problem Improving text-to-speech synthesis for better speech quality and robustness.
method Combines diffusion probabilistic model and variational autoencoder for latent variable conversion.
result The method is robust to poor orthography and alignment errors.
This research explores various sampling methods and probability distributions for hard alignment in sequence-to-sequence TTS synthesis.
problem Improving alignment accuracy in sequence-to-sequence text-to-speech synthesis.
method Investigated various sampling methods (greedy, beam, random) and probability distributions (Bernoulli, Concrete) for hard alignment.
result Deterministic search is more preferable than stochastic search for natural alignment transition.
End-to-end Sanskrit TTS developed with limited data, achieving good quality.
problem Developing natural-sounding speech for Sanskrit with scarce data.
method Fine-tuning Tacotron2 model with WaveGlow and transfer learning.
result Achieved an overall MOS of 3.38 from 37 evaluators.
In this note we prove an existence result for the Einstein conformal constraint equations for metrics with vanishing Yamabe invariant assuming that the TT-tensor is small in L2.
We give an overview on the tt*-geometry defined for isolated hypersurface singularities and tame functions via Brieskorn lattices. We discuss nilpotent orbits in this context, as well as classifying spaces of Brieskorn lattices and (limits of) period maps.
Tensor network surrogate for efficient option pricing in large portfolios.
problem Large-scale portfolio revaluation problems in market risk management.
method Tensor-train (TT) approximation for high-dimensional price surfaces, direct inference using Laplacian kernel and TT representations.
result Tensor surrogate achieves lower test error and faster evaluation times compared to standard GPR.
Characterizes kernel of linearization for minimal surfaces problem
problem Characterizing kernel of linearization for minimal surfaces problem
method Show kernel consists of potential fields and TT fields
result In whole-space Euclidean decomposition, kernel consists of potential fields and TT fields
TensorGuide improves LoRA efficiency and expressivity through joint tensor-train optimization.
problem Limited expressivity and generalization of standard LoRA.
method TensorGuide uses a unified tensor-train structure with controlled Gaussian noise to generate correlated low-rank matrices.
result TensorGuide achieves superior accuracy and scalability with fewer parameters compared to standard LoRA and TT-LoRA.