Representation mixing combines character and phoneme inputs for flexible TTS synthesis.
problem Limited control over pronunciation in character or phoneme-based TTS systems.
method Representation mixing combines multiple linguistic inputs in a single encoder.
result Flexibility in choosing between character, phoneme, or mixed representations during inference.
A new TTS method uses diffusion and VAE for better speech synthesis.
problem Improving text-to-speech synthesis for better speech quality and robustness.
method Combines diffusion probabilistic model and variational autoencoder for latent variable conversion.
result The method is robust to poor orthography and alignment errors.
End-to-end TTS learns context features from text input.
problem Lack of understanding of context features learned by end-to-end TTS.
method Evaluated encoder outputs against context criteria derived from parametric TTS.
result Encoder outputs reflect linguistic and phonetic context features.
This research explores various sampling methods and probability distributions for hard alignment in sequence-to-sequence TTS synthesis.
problem Improving alignment accuracy in sequence-to-sequence text-to-speech synthesis.
method Investigated various sampling methods (greedy, beam, random) and probability distributions (Bernoulli, Concrete) for hard alignment.
result Deterministic search is more preferable than stochastic search for natural alignment transition.
End-to-end TTS framework uses hard alignment to improve accuracy.
problem End-to-end TTS systems struggle with accurate alignment between input text and output acoustic features.
method Proposes a constrained alignment scheme with hard monotonic alignments, marginalized during training.
result Improves alignment learning and prediction in end-to-end TTS systems.
Investigates neural TTS systems for Japanese and English.
problem Improving neural TTS systems for high-quality speech synthesis.
method Comparative study of neural sequence-to-sequence TTS vs. DNN pipeline TTS, varying model architecture, parameter size, and language.
result A neural sequence-to-sequence TTS system requires sufficient model parameters and a powerful encoder for high-quality speech synthesis.
BOFFIN TTS optimizes hyper-parameters for new speaker adaptation.
problem Fine-tuning a pre-trained TTS model for a new speaker with limited data.
method Bayesian optimization to efficiently find optimal hyper-parameters.
result Average 30% improvement in speaker similarity over standard techniques.
New method improves speech synthesis quality.
problem Efficiently train parallel speech synthesis models.
method Spectral energy distance for implicit generative models.
result State-of-the-art generation quality achieved.
Grad-TTS models speech from text using diffusion probabilistic techniques.
problem Creating high-quality speech from text input.
method Score-based decoder with stochastic differential equations for noise-to-speech transformation.
result Grad-TTS produces mel-spectrograms from text input with competitive quality.
Generative models produce varied intonation in speech synthesis.
problem Typical TTS systems lack the ability to produce multiple distinct renditions of a sentence.
method Use variational autoencoders (VAEs) to capture a distribution over multiple renditions and produce varied intonation.
result Sampling from the tails of the VAE prior produces more varied intonation than traditional approaches, while maintaining naturalness.
Joint training model for TTS and VC tasks using Tacotron and WaveNet.
problem Training a shared model for text-to-speech and voice conversion.
method Extended Tacotron model with dual attention mechanism for shared tasks, WaveNet for waveform generation.
result Joint training of a shared model achieves both TTS and VC tasks efficiently.
This study improves text-to-speech synthesis using GANs for glottal excitation.
problem Slow inference and computational cost of WaveNet and difficulty in parallel training of GANs.
method Adopted GANs for parallel waveform generation in speech signal and glottal excitation.
result GAN-based glottal excitation model achieves quality and voice similarity on par with WaveNet.
Classifies Toda-type tt*-structures and their fixed points.
problem Classifying Toda-type tt*-structures and their fixed points.
method Fixed point description and reduction of anti-symmetry conditions.
result Reduces possibilities of anti-symmetry condition to two cases.
Using nonlinear pde techniques, we construct a new family of globally smooth tt* structures. This includes tt* structures associated to the (orbifold) quantum cohomology of a finite number of complex projective spaces and weighted projective spaces. The existence of such "magical solutions" of the tt* equations, namely…
The paper proves an isomorphism between tt∗ structures of Landau-Ginzburg and Calabi-Yau models.
problem Establishing an isomorphism between tt∗ structures of different geometries. method Using Landau-Ginzburg models and Calabi-Yau hypersurfaces, proving the isomorphism via the big residue map.
result An isomorphism between tt∗ structures of Landau-Ginzburg and Calabi-Yau models is proven. Relates quantum cohomology to tt*-Toda equations for minuscule flag manifolds.
problem Quantum cohomology of minuscule flag manifolds.
method Combining Lie-theoretic treatments of tt*-Toda equations and quantum cohomology.
result Relates quantum cohomology to tt*-Toda equations for minuscule flag manifolds.
Analyzes tt*-structures from ADE-type Stokes data.
problem Classifying tt*-structures over C∗. method Isomonodromic deformations with upper unitriangular real Stokes matrices.
result Establishes a direct analytic realization of the ADE classification. Polynomial growth elements found in all subgroups of Out(F_n).
problem Understanding polynomial growth in subgroups of Out(F_n).
method Analyzing conjugacy classes and elements of Out(F_n).
result Polynomial growth elements exist in all subgroups of Out(F_n).
Explains quantum cohomology of Grassmannians using tt* equations.
problem Relates quantum cohomology of complex Grassmannians to projective space.
method Uses tt* equations and Lie-theoretic connections.
result Illustrates relations between tt* equations and quantum cohomology.
Proves existence and uniqueness of solutions for A_n tt*-Toda equations.
problem Existence and uniqueness of solutions for A_n tt*-Toda equations.
method Proof of existence and uniqueness for any n, new treatment of asymptotic data.
result Existence and uniqueness of global solutions for any n.
We study transverse-tracefree (TT)-tensors on conformally flat 3-manifolds (M,g). The Cotton-York tensor linearized at g maps every symmetric tracefree tensor into one which is TT. The question as to whether this is the general solution to the TT-condition is viewed as a cohomological problem within an elliptic com…
Solves constant pre-factor problem for tt*-Toda equations using asymptotic data and symplectic structures.
problem Constant pre-factor problem for the tt*-Toda equations.
method Explicit evaluation using asymptotic data and introduction of symplectic structures.
result Preservation of symplectic structures by Riemann-Hilbert correspondence for wider class of solutions.
Study of symplectic groupoids from tt*-Toda equations.
problem Geometry of meromorphic connections with irregular singularities.
method Holomorphic symplectic groupoid structure over Steinberg cross section.
result Proves the space of tt*-Toda connections is a symplectic Lie groupoid.
Convolutional network converts speaker voices without text.
problem Speaker conversion without text-based methods.
method Fully convolutional wav-to-wav network with ASR pre-training.
result Successfully converts TTS robot's voice to narrated audiobook voices.
Solutions of tt*-equation from SU(2)_k fusion algebra.
problem Describe solutions to the tt*-equation from a specific algebra.
method Use DPW method and representations of SU(2).
result Construct solutions corresponding to A_k minimal model.
VPFD uses vocoder features for adversarial training in VC.
problem Adversarial training on waveform data is time-consuming and memory-intensive.
method VPFD employs vocoder features for adversarial training.
result VPFD achieves VC performance comparable to waveform discriminators with reduced training time and memory.
New surfaces with conjugate points have global blow-down maps in their TT spaces.
problem Constructing global blow-down maps for surfaces with conjugate points.
method Explicit construction of a family of non-trapping Riemannian surfaces with global blow-down maps.
result Global blow-down maps exist for some non-simple surfaces with conjugate points.
Proposes Textual Echo Cancellation to improve speech recognition.
problem Improving speech recognition performance and user experience for smart devices.
method A novel sequence-to-sequence model with multi-source attention that processes both the microphone mixture signal and source text of TTS playback.
result Demonstrates enhanced speech recognition performance and reduced latency.
RandUCB combines UCB and TS for optimal bandit performance.
problem Optimizing decision-making in uncertain environments with limited feedback.
method Randomized UCB algorithm using confidence intervals.
result Achieves minimax-optimal regret in various bandit settings.
FPETS speeds up TTS by 600X and reduces errors.
problem High latency and errors in end-to-end TTS systems.
method Non-autoregressive, fully parallel approach with UFANS and trainable position encoding.
result Significant speed up and better quality audios with fewer errors.
Bayesian tensor train method recovers streaming data with high accuracy.
problem Recovering high-order, incomplete, and noisy streaming data.
method Bayesian tensor train decomposition using streaming variational Bayes method.
result The proposed SPTT algorithm excels in recovering streaming data compared to state-of-the-art methods.
We describe all smooth solutions of the two-function tt*-Toda equations (a version of the tt* equations, or equations for harmonic maps into SL(n,R)/SO(n)) in terms of (i) asymptotic data, (ii) holomorphic data, and (iii) monodromy data. This allows us to find all solutions with integral Stokes data. These include solu…
Paper proposes efficient tensor completion method using Gaussian Process.
problem Tensor completion in high-dimensional data with unknown smooth functions.
method Gaussian Process Regression for initialization and TT-cross approximation for tensor rank selection.
result Improved reconstruction error compared to random initialization.
In "Isomonodromy aspects of the tt* equations of Cecotti and Vafa I. Stokes data" (arxiv:1209.2045) we described all smooth solutions of the two-function tt*-Toda equations in terms of asymptotic data, holomorphic data, and monodromy data. In this supplementary article we focus on the holomorphic data and its interpret…
End-to-end Sanskrit TTS developed with limited data, achieving good quality.
problem Developing natural-sounding speech for Sanskrit with scarce data.
method Fine-tuning Tacotron2 model with WaveGlow and transfer learning.
result Achieved an overall MOS of 3.38 from 37 evaluators.
In this note we prove an existence result for the Einstein conformal constraint equations for metrics with vanishing Yamabe invariant assuming that the TT-tensor is small in L2.
We give an overview on the tt*-geometry defined for isolated hypersurface singularities and tame functions via Brieskorn lattices. We discuss nilpotent orbits in this context, as well as classifying spaces of Brieskorn lattices and (limits of) period maps.
Combining data from multiple speakers improves neural TTS quality, especially with imbalanced data.
problem Training high-quality TTS systems with imbalanced speaker data.
method Combine data from multiple speakers, train multi-speaker models, and use ensemble methods.
result Ensemble multi-speaker models improve synthetic speech quality for underrepresented speakers.
Tensor network surrogate for efficient option pricing in large portfolios.
problem Large-scale portfolio revaluation problems in market risk management.
method Tensor-train (TT) approximation for high-dimensional price surfaces, direct inference using Laplacian kernel and TT representations.
result Tensor surrogate achieves lower test error and faster evaluation times compared to standard GPR.
Characterizes kernel of linearization for minimal surfaces problem
problem Characterizing kernel of linearization for minimal surfaces problem
method Show kernel consists of potential fields and TT fields
result In whole-space Euclidean decomposition, kernel consists of potential fields and TT fields
TensorGuide improves LoRA efficiency and expressivity through joint tensor-train optimization.
problem Limited expressivity and generalization of standard LoRA.
method TensorGuide uses a unified tensor-train structure with controlled Gaussian noise to generate correlated low-rank matrices.
result TensorGuide achieves superior accuracy and scalability with fewer parameters compared to standard LoRA and TT-LoRA.
Proposes a new kernel technique for tensor data in SVM.
problem Handling tensorial data in machine learning.
method Kernelized support tensor train machine for image classification.
result Tensorizes the standard SVM on its input structure and kernel mapping scheme.
Improved TTS style transfer across disjoint datasets with adversarial cycle consistency.
problem Suboptimal TTS style transfer on disjoint datasets with underrepresented styles.
method Adversarial cycle consistency training with paired and unpaired triplets.
result 78% improvement in style transfer with minimal reduction in fidelity and naturalness.
A new method computes Greeks for multi-asset options using tensor trains and Fourier transforms.
problem Efficient computation of Greeks for multi-asset options with high accuracy and low sample complexity.
method Tensor train (TT) representations of Fourier-based pricing functions, combined with numerical differentiation or analytical approaches.
result Significant speed-ups of up to 105imes over Monte Carlo simulations while maintaining comparable accuracy. Establishes correspondence between Calabi-Yau and Landau-Ginzburg structures.
problem Preserving real structures in the Calabi-Yau/Landau-Ginzburg correspondence.
method Detailed analysis of period integrals and modification of real structures.
result Full CY/LG correspondence for tt∗ structures established. For any triple (W,L,ρ), where W is a closed connected and oriented 3-manifold, L is a link in W and ρ is a flat principal B-bundle over W (B is the Borel subgroup of $SL(2,\mc)$), one constructs a $\Dd$-scissors congruence class $\cG_{\Dd}(W,L,ρ)$ which belongs to a (pre)-Bloch group $\Pp (\Dd)$. The class $\cG_{\D…
In this note, we prove that for a cobounded,Lipschitz path $γ:I\to\TT$, if the pull back bundle Hγ over I is a strongly relatively hyperbolic metric space then there exists a geodesic ξ in $\TT$ such that γ(I) and ξ are close to each other.
Bayesian tensor train kernel machine uses Laplace approximation for scalable GP regression.
problem Scalability limitations of Gaussian process regression.
method Bayesian tensor train kernel machine with Laplace approximation and variational inference.
result VI replaces cross-validation and offers up to 65x faster training.