Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

18375573 · Jun 202019922001200920182026
48 results for Speech Translation

This research investigates the Statistical Machine Translation approaches to translate speech in real time automatically. Such systems can be used in a pipeline with speech recognition and synthesis software in order to produce a real-time voice communication system between foreigners. We obtained three main data sets …

2015-09-30abs ↗pdf ↗

Models improve syntactic clustering by adding more translation and part-of-speech decoders.

problem Improving syntactic saliency in hidden sentence representations.
method Training multi-task autoencoders on linguistic tasks and analyzing the learned hidden representations.
result The representation space becomes less entangled with more decoders, leading to better syntactic clustering.

This paper tackles multilingual speech processing by optimizing conflicting objectives hierarchically.

problem Training models for multilingual, multi-task speech processing is hampered by conflicting objectives.
method Investigates three multi-objective MSP formulations and introduces a lightweight layer-selection mechanism.
result A bi-level recipe outperforms standard flat optimization in state-of-the-art MSP models.

Many machine learning tasks can be expressed as the transformation---or \emph{transduction}---of input sequences into output sequences: speech recognition, machine translation, protein secondary structure prediction and text-to-speech to name but a few. One of the key challenges in sequence transduction is learning to …

2012-11-14abs ↗pdf ↗

Study uses PPG signals for detecting speech events and speaker characteristics.

problem Detecting speech events and speaker characteristics from PPG signals.
method End-to-end convolutional neural network architectures for gender and person verification.
result Promising results showing potential of PPG for speech processing tasks.

Novel fusion of autoencoders predicts sleepiness from speech.

problem Predicting sleepiness from speech recordings.
method Attention-based and recurrent sequence to sequence autoencoders for unsupervised representation learning.
result Fusion of autoencoders' representations achieves higher correlation with sleepiness scales.

Attention forcing improves sequence-to-sequence model training stability.

problem Training auto-regressive sequence-to-sequence models with attention mechanism is challenging.
method Attention forcing guides the model with generated output history and reference attention.
result Attention forcing trains models to recover from mistakes without requiring a schedule or classifier.

Compact neural networks for speech recognition with dropout training.

problem Training large neural networks for speech recognition tasks.
method Introducing a sparsity-inducing prior on dropout retention probability to prune hidden units during training.
result Achieved comparable accuracy with fewer than 50% of the hidden units, resulting in a 2.5x speedup.

Bardo Composer generates tabletop RPG music based on player speech.

problem Creating immersive background music for tabletop RPGs.
method Speech recognition, emotion classification, and music generation using a novel beam search algorithm.
result Generated music pieces can be accurately identified by human subjects as conveying the intended emotion.

Recurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks in- cluding machine translation, handwriting synthesis and image caption gen- eration. We extend the attention-mechanism with features needed for speech recognition. We sh…

2015-06-24abs ↗pdf ↗

Paper introduces sampling-based speech synthesis with natural variation.

problem Synthetic speech lacks natural inter-utterance variation.
method Moment-matching networks trained to match moments of generated speech parameters to natural speech parameters.
result Sampling-based generation does not degrade synthetic speech quality.

New method prevents entropy collapse in Transformer training, leading to more stable and robust models.

problem Training instability in Transformers, especially in attention layers.
method Spectral normalization with a learned scalar to prevent entropy collapse.
result Prevents entropy collapse, leading to more stable training.

Paper shows continuous speech recognition with EEG features, no speech input.

problem Continuous speech recognition with limited vocabulary and noisy/no speech input.
method Connectionist temporal classification (CTC) model, EEG features, new deep learning architecture.
result Continuous speech recognition achieved on limited vocabulary with noisy/no speech input.

RUSLAN is a large Russian speech corpus for text-to-speech.

problem Lack of high-quality annotated Russian speech data for text-to-speech.
method Developed a large annotated Russian speech corpus and trained a neural network for text-to-speech synthesis.
result Synthesized speech quality evaluated with MOS scores: 4.05 for naturalness, 3.78 for intelligibility.

Unified model improves speech enhancement in unseen environments.

problem Robustness against unknown environments in speech enhancement.
method Probabilistic integration of VAE and NMF.
result Outperforms conventional DNN-based method in unseen environments.

NovoGrad improves deep learning training with adaptive moments and layer-wise normalization.

problem Training deep neural networks efficiently and effectively.
method Layer-wise adaptive moments with gradient normalization and decoupled weight decay.
result NovoGrad outperforms well-tuned SGD with momentum and Adam/AdamW in various tasks.

Detects AI-synthesized speech using cepstral and bispectral analysis.

problem Validating the authenticity of speech from AI-generated content.
method Integrates cepstral and bispectral analysis for distinguishing human from AI-synthesized speech.
result Higher-order statistics show less correlation for human speech compared to AI-synthesis, and cepstral analysis reveals unique power components.

Solution for voice conversion with limited data using hierarchical seq2seq and attention models.

problem Voice conversion between speakers with limited parallel audio pairs.
method Hierarchical sequence to sequence model with attention-based decoder, trained on single speaker dataset.
result Improved voice conversion quality using mel spectrograms and wavenet vocoder.

Improved visual speech synthesis using adapted ASR acoustic models.

problem Lack of synchronized audio, video, and depth data for speaker-independent speech-driven visual speech synthesis.
method Adapted an ASR acoustic model trained on audio-only data to the visual speech synthesis domain.
result Viewers significantly prefer animations generated from the adapted ASR acoustic model.

WaveCycleGAN2 improves speech synthesis quality by reducing aliasing.

problem Human ear can still distinguish synthesized speech from natural speech.
method WaveCycleGAN2 uses generators without down/up-sampling modules and combines discriminators from waveform and acoustic parameter domains.
result WaveCycleGAN2 achieves high-quality speech synthesis with comparable mean opinion scores to natural speech.

WaveCycleGAN converts synthetic speech to natural speech using cycle-consistent adversarial networks.

problem Over-smoothing effect in synthetic speech, leading to quality degradation.
method Cycle-consistent adversarial networks for waveform-level modification.
result Improves naturalness of generated speech sounds.

Improved speech enhancement using diffusion models with MSE loss.

problem Efficient incorporation of noisy speech in generative speech enhancement.
method Augmented diffusion-based generative model with a MSE loss for enhanced speech.
result Proposed method improves speech enhancement performance compared to original diffusion model.

End-to-end FCN framework optimizes speech enhancement metrics.

problem Inconsistency between model optimization and evaluation metrics.
method End-to-end utterance-based FCN for direct optimization of STOI.
result Enhanced speech has better STOI and improved intelligibility.

This work disentangles speech and non-speech components from found data.

problem Building robust acoustic models from found data with non-standard variations.
method Latent Stochastic Models and Multinode Latent Space Variational Autoencoders (VAE).
result Speech and music can be separated in the latent space of a VAE, independent of the language.