Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

1345 · Feb 202019922001200920182026
48 results for Multi-speaker TTS

Combining data from multiple speakers improves neural TTS quality, especially with imbalanced data.

problem Training high-quality TTS systems with imbalanced speaker data.
method Combine data from multiple speakers, train multi-speaker models, and use ensemble methods.
result Ensemble multi-speaker models improve synthetic speech quality for underrepresented speakers.

Improved multi-speaker TTS using GANs and waveform loss.

problem Training acoustic models for neural vocoders in multi-speaker TTS systems.
method Proposed frameworks incorporating Wasserstein GAN with gradient penalty (WGAN-GP) and discretized mixture logistic loss (DML) into acoustic models trained with WaveNet.
result Acoustic models trained with WGAN-GP and DML loss achieve highest subjective evaluation scores in multi-speaker TTS.

RoyalFlush system improves multi-speaker ASR in M2MeT challenge.

problem Improving multi-speaker automatic speech recognition in noisy environments.
method Front-end processing with WPE and beamforming, data augmentation, and fusion of two ASR models.
result 12.22% absolute CER reduction on validation set and 12.11% on test set compared to baseline.

Using nonlinear pde techniques, we construct a new family of globally smooth tt* structures. This includes tt* structures associated to the (orbifold) quantum cohomology of a finite number of complex projective spaces and weighted projective spaces. The existence of such "magical solutions" of the tt* equations, namely…

2010-10-10abs ↗pdf ↗

The paper proves an isomorphism between tttt^* structures of Landau-Ginzburg and Calabi-Yau models.

problem Establishing an isomorphism between tttt^* structures of different geometries.
method Using Landau-Ginzburg models and Calabi-Yau hypersurfaces, proving the isomorphism via the big residue map.
result An isomorphism between tttt^* structures of Landau-Ginzburg and Calabi-Yau models is proven.

We study transverse-tracefree (TT)-tensors on conformally flat 3-manifolds (M,g)(M,g). The Cotton-York tensor linearized at gg maps every symmetric tracefree tensor into one which is TT. The question as to whether this is the general solution to the TT-condition is viewed as a cohomological problem within an elliptic com…

1996-06-18abs ↗pdf ↗

Solves constant pre-factor problem for tt*-Toda equations using asymptotic data and symplectic structures.

problem Constant pre-factor problem for the tt*-Toda equations.
method Explicit evaluation using asymptotic data and introduction of symplectic structures.
result Preservation of symplectic structures by Riemann-Hilbert correspondence for wider class of solutions.

End-to-end TTS framework uses hard alignment to improve accuracy.

problem End-to-end TTS systems struggle with accurate alignment between input text and output acoustic features.
method Proposes a constrained alignment scheme with hard monotonic alignments, marginalized during training.
result Improves alignment learning and prediction in end-to-end TTS systems.

New surfaces with conjugate points have global blow-down maps in their TT spaces.

problem Constructing global blow-down maps for surfaces with conjugate points.
method Explicit construction of a family of non-trapping Riemannian surfaces with global blow-down maps.
result Global blow-down maps exist for some non-simple surfaces with conjugate points.

Proposes Textual Echo Cancellation to improve speech recognition.

problem Improving speech recognition performance and user experience for smart devices.
method A novel sequence-to-sequence model with multi-source attention that processes both the microphone mixture signal and source text of TTS playback.
result Demonstrates enhanced speech recognition performance and reduced latency.

Proposes DNN-based speaker embedding correlated with subjective inter-speaker similarity for speech synthesis.

problem Inadequate speaker representation for open speakers not in training data.
method Two training algorithms using inter-speaker similarity matrices: similarity vector embedding and similarity matrix embedding.
result Proposed algorithms learn speaker embedding highly correlated with subjective inter-speaker similarity.

Representation mixing combines character and phoneme inputs for flexible TTS synthesis.

problem Limited control over pronunciation in character or phoneme-based TTS systems.
method Representation mixing combines multiple linguistic inputs in a single encoder.
result Flexibility in choosing between character, phoneme, or mixed representations during inference.

Investigates neural TTS systems for Japanese and English.

problem Improving neural TTS systems for high-quality speech synthesis.
method Comparative study of neural sequence-to-sequence TTS vs. DNN pipeline TTS, varying model architecture, parameter size, and language.
result A neural sequence-to-sequence TTS system requires sufficient model parameters and a powerful encoder for high-quality speech synthesis.

Recent speech technology research has seen a growing interest in using WaveNets as statistical vocoders, i.e., generating speech waveforms from acoustic features. These models have been shown to improve the generated speech quality over classical vocoders in many tasks, such as text-to-speech synthesis and voice conver…

2018-04-25abs ↗pdf ↗

Recently, there has been growing interest in multi-speaker speech recognition, where the utterances of multiple speakers are recognized from their mixture. Promising techniques have been proposed for this task, but earlier works have required additional training data such as isolated source signals or senone alignments…

2018-05-15abs ↗pdf ↗

In "Isomonodromy aspects of the tt* equations of Cecotti and Vafa I. Stokes data" (arxiv:1209.2045) we described all smooth solutions of the two-function tt*-Toda equations in terms of asymptotic data, holomorphic data, and monodromy data. In this supplementary article we focus on the holomorphic data and its interpret…

2012-09-11abs ↗pdf ↗

A new TTS method uses diffusion and VAE for better speech synthesis.

problem Improving text-to-speech synthesis for better speech quality and robustness.
method Combines diffusion probabilistic model and variational autoencoder for latent variable conversion.
result The method is robust to poor orthography and alignment errors.

This research explores various sampling methods and probability distributions for hard alignment in sequence-to-sequence TTS synthesis.

problem Improving alignment accuracy in sequence-to-sequence text-to-speech synthesis.
method Investigated various sampling methods (greedy, beam, random) and probability distributions (Bernoulli, Concrete) for hard alignment.
result Deterministic search is more preferable than stochastic search for natural alignment transition.

Tensor network surrogate for efficient option pricing in large portfolios.

problem Large-scale portfolio revaluation problems in market risk management.
method Tensor-train (TT) approximation for high-dimensional price surfaces, direct inference using Laplacian kernel and TT representations.
result Tensor surrogate achieves lower test error and faster evaluation times compared to standard GPR.

TensorGuide improves LoRA efficiency and expressivity through joint tensor-train optimization.

problem Limited expressivity and generalization of standard LoRA.
method TensorGuide uses a unified tensor-train structure with controlled Gaussian noise to generate correlated low-rank matrices.
result TensorGuide achieves superior accuracy and scalability with fewer parameters compared to standard LoRA and TT-LoRA.