Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

1.5%2.9%4.4%5.9% · Jun 199619922001200920182026
48 results for TTS synthesis

Representation mixing combines character and phoneme inputs for flexible TTS synthesis.

problem Limited control over pronunciation in character or phoneme-based TTS systems.
method Representation mixing combines multiple linguistic inputs in a single encoder.
result Flexibility in choosing between character, phoneme, or mixed representations during inference.

A new TTS method uses diffusion and VAE for better speech synthesis.

problem Improving text-to-speech synthesis for better speech quality and robustness.
method Combines diffusion probabilistic model and variational autoencoder for latent variable conversion.
result The method is robust to poor orthography and alignment errors.

This research explores various sampling methods and probability distributions for hard alignment in sequence-to-sequence TTS synthesis.

problem Improving alignment accuracy in sequence-to-sequence text-to-speech synthesis.
method Investigated various sampling methods (greedy, beam, random) and probability distributions (Bernoulli, Concrete) for hard alignment.
result Deterministic search is more preferable than stochastic search for natural alignment transition.

End-to-end TTS framework uses hard alignment to improve accuracy.

problem End-to-end TTS systems struggle with accurate alignment between input text and output acoustic features.
method Proposes a constrained alignment scheme with hard monotonic alignments, marginalized during training.
result Improves alignment learning and prediction in end-to-end TTS systems.

Investigates neural TTS systems for Japanese and English.

problem Improving neural TTS systems for high-quality speech synthesis.
method Comparative study of neural sequence-to-sequence TTS vs. DNN pipeline TTS, varying model architecture, parameter size, and language.
result A neural sequence-to-sequence TTS system requires sufficient model parameters and a powerful encoder for high-quality speech synthesis.

Generative models produce varied intonation in speech synthesis.

problem Typical TTS systems lack the ability to produce multiple distinct renditions of a sentence.
method Use variational autoencoders (VAEs) to capture a distribution over multiple renditions and produce varied intonation.
result Sampling from the tails of the VAE prior produces more varied intonation than traditional approaches, while maintaining naturalness.

Joint training model for TTS and VC tasks using Tacotron and WaveNet.

problem Training a shared model for text-to-speech and voice conversion.
method Extended Tacotron model with dual attention mechanism for shared tasks, WaveNet for waveform generation.
result Joint training of a shared model achieves both TTS and VC tasks efficiently.

This study improves text-to-speech synthesis using GANs for glottal excitation.

problem Slow inference and computational cost of WaveNet and difficulty in parallel training of GANs.
method Adopted GANs for parallel waveform generation in speech signal and glottal excitation.
result GAN-based glottal excitation model achieves quality and voice similarity on par with WaveNet.

Using nonlinear pde techniques, we construct a new family of globally smooth tt* structures. This includes tt* structures associated to the (orbifold) quantum cohomology of a finite number of complex projective spaces and weighted projective spaces. The existence of such "magical solutions" of the tt* equations, namely…

2010-10-10abs ↗pdf ↗

The paper proves an isomorphism between tttt^* structures of Landau-Ginzburg and Calabi-Yau models.

problem Establishing an isomorphism between tttt^* structures of different geometries.
method Using Landau-Ginzburg models and Calabi-Yau hypersurfaces, proving the isomorphism via the big residue map.
result An isomorphism between tttt^* structures of Landau-Ginzburg and Calabi-Yau models is proven.

We study transverse-tracefree (TT)-tensors on conformally flat 3-manifolds (M,g)(M,g). The Cotton-York tensor linearized at gg maps every symmetric tracefree tensor into one which is TT. The question as to whether this is the general solution to the TT-condition is viewed as a cohomological problem within an elliptic com…

1996-06-18abs ↗pdf ↗

Solves constant pre-factor problem for tt*-Toda equations using asymptotic data and symplectic structures.

problem Constant pre-factor problem for the tt*-Toda equations.
method Explicit evaluation using asymptotic data and introduction of symplectic structures.
result Preservation of symplectic structures by Riemann-Hilbert correspondence for wider class of solutions.

New surfaces with conjugate points have global blow-down maps in their TT spaces.

problem Constructing global blow-down maps for surfaces with conjugate points.
method Explicit construction of a family of non-trapping Riemannian surfaces with global blow-down maps.
result Global blow-down maps exist for some non-simple surfaces with conjugate points.

Proposes Textual Echo Cancellation to improve speech recognition.

problem Improving speech recognition performance and user experience for smart devices.
method A novel sequence-to-sequence model with multi-source attention that processes both the microphone mixture signal and source text of TTS playback.
result Demonstrates enhanced speech recognition performance and reduced latency.

In "Isomonodromy aspects of the tt* equations of Cecotti and Vafa I. Stokes data" (arxiv:1209.2045) we described all smooth solutions of the two-function tt*-Toda equations in terms of asymptotic data, holomorphic data, and monodromy data. In this supplementary article we focus on the holomorphic data and its interpret…

2012-09-11abs ↗pdf ↗

Combining data from multiple speakers improves neural TTS quality, especially with imbalanced data.

problem Training high-quality TTS systems with imbalanced speaker data.
method Combine data from multiple speakers, train multi-speaker models, and use ensemble methods.
result Ensemble multi-speaker models improve synthetic speech quality for underrepresented speakers.

Tensor network surrogate for efficient option pricing in large portfolios.

problem Large-scale portfolio revaluation problems in market risk management.
method Tensor-train (TT) approximation for high-dimensional price surfaces, direct inference using Laplacian kernel and TT representations.
result Tensor surrogate achieves lower test error and faster evaluation times compared to standard GPR.

TensorGuide improves LoRA efficiency and expressivity through joint tensor-train optimization.

problem Limited expressivity and generalization of standard LoRA.
method TensorGuide uses a unified tensor-train structure with controlled Gaussian noise to generate correlated low-rank matrices.
result TensorGuide achieves superior accuracy and scalability with fewer parameters compared to standard LoRA and TT-LoRA.

Improved TTS style transfer across disjoint datasets with adversarial cycle consistency.

problem Suboptimal TTS style transfer on disjoint datasets with underrepresented styles.
method Adversarial cycle consistency training with paired and unpaired triplets.
result 78% improvement in style transfer with minimal reduction in fidelity and naturalness.

A new method computes Greeks for multi-asset options using tensor trains and Fourier transforms.

problem Efficient computation of Greeks for multi-asset options with high accuracy and low sample complexity.
method Tensor train (TT) representations of Fourier-based pricing functions, combined with numerical differentiation or analytical approaches.
result Significant speed-ups of up to 105imes10^{5} imes over Monte Carlo simulations while maintaining comparable accuracy.

Establishes correspondence between Calabi-Yau and Landau-Ginzburg structures.

problem Preserving real structures in the Calabi-Yau/Landau-Ginzburg correspondence.
method Detailed analysis of period integrals and modification of real structures.
result Full CY/LG correspondence for tttt^* structures established.

In this note, we prove that for a cobounded,Lipschitz path $γ:I\to\TT$, if the pull back bundle Hγ\mathcal H_γ over II is a strongly relatively hyperbolic metric space then there exists a geodesic ξξ in $\TT$ such that γ(I)γ(I) and ξξ are close to each other.

2011-09-17abs ↗pdf ↗

Bayesian tensor train kernel machine uses Laplace approximation for scalable GP regression.

problem Scalability limitations of Gaussian process regression.
method Bayesian tensor train kernel machine with Laplace approximation and variational inference.
result VI replaces cross-validation and offers up to 65x faster training.