Improved multilingual speech recognition with low latency for nine Indic languages.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper tackles multilingual speech processing by optimizing conflicting objectives hierarchically.
New ASR system handles multiple languages without needing language-specific encoding.
Neural language modeling (LM) has led to significant improvements in several applications, including Automatic Speech Recognition. However, they typically require large amounts of training data, which is not available for many domains and languages. In this study, we propose a multilingual neural language model archite…
Improved language identification accuracy through signal combination methods.
Code-switching, the alternation of languages within a conversation or utterance, is a common communicative phenomenon that occurs in multilingual communities across the world. This survey reviews computational approaches for code-switched Speech and Natural Language Processing. We motivate why processing code-switched …
Improved LID for multilingual speakers using context-aware models.
Improved ASR for English-isiZulu code-switched speech with semi-supervised training.
Paper classifies Parkinson's disease from speech in three languages using CNNs and transfer learning.
End-to-end neural network improves QbE-STD in multilingual speech search.
Improved speech recognition using EEG and video.
In this paper we demonstrate end-to-end continuous speech recognition (CSR) using electroencephalography (EEG) signals with no speech signal as input. An attention model based automatic speech recognition (ASR) and connectionist temporal classification (CTC) based ASR systems were implemented for performing recognition…
Continuous speech recognition from brain activity without vocalization.
The performance of automatic speech recognition systems(ASR) degrades in the presence of noisy speech. This paper demonstrates that using electroencephalography (EEG) can help automatic speech recognition systems overcome performance loss in the presence of noise. The paper also shows that distillation training of auto…
AV-CPL uses continuous pseudo-labels for AVSR combining labeled and unlabeled data.
Paper explores EEG-based speech recognition using transformers, showing faster training and better performance for smaller vocabularies.
Bayesian method improves reliability of BERT for hate speech detection.
In this paper we demonstrate continuous noisy speech recognition using connectionist temporal classification (CTC) model on limited Chinese vocabulary using electroencephalography (EEG) features with no speech signal as input and we further demonstrate single CTC model based continuous noisy speech recognition on limit…
VoiceFilter-Lite separates speech from background in real-time for on-device speech recognition.
Improved speech emotion recognition using pre-trained language models.
Paper improves EEG-based speech recognition using CTC and beam search.
Large receptive field CNNs improve distant speech recognition.
Deep architecture learns transferable features for robust speech emotion recognition.
New dataset for evaluating speech recognition fairness across demographics.
Survey on DNNs for speech processing, focusing on limited data challenges.
Spoken language translation (SLT) has become very important in an increasingly globalized world. Machine translation (MT) for automatic speech recognition (ASR) systems is a major challenge of great interest. This research investigates that automatic sentence segmentation of speech that is important for enriching speec…
Study proposes a decision tree for more accurate depression recognition in speech.
Mobile app improves speech recognition of names with user feedback.
Quaternion neural networks improve distant speech recognition.
Paper proposes an online speech recognition model using Transformer.
Speech recognition systems have achieved high recognition performance for several tasks. However, the performance of such systems is dependent on the tremendously costly development work of preparing vast amounts of task-matched transcribed speech data for supervised training. The key problem here is the cost of transc…
Recurrent neural network (RNN) language models (LMs) and Long Short Term Memory (LSTM) LMs, a variant of RNN LMs, have been shown to outperform traditional N-gram LMs on speech recognition tasks. However, these models are computationally more expensive than N-gram LMs for decoding, and thus, challenging to integrate in…
ADReSS Challenge at INTERSPEECH 2020 benchmarks speech recognition for Alzheimer's dementia.
In this work we explored building automatic speech recognition models for transcribing doctor patient conversation. We collected a large scale dataset of clinical conversations ( hr), designed the task to represent the real word scenario, and explored several alignment approaches to iteratively improve data qua…
Novel bio-inspired masking for robust speech emotion recognition.
Articulatory distinctive features, as well as phonetic transcription, play important role in speech-related tasks: computer-assisted pronunciation training, text-to-speech conversion (TTS), studying speech production mechanisms, speech recognition for low-resourced languages. End-to-end approaches to speech-related tas…
Introspects convolutional speech recognition models using Gradient-adjusted Neuron Activation Profiles.
We consider multilingual bottleneck features (BNFs) for nearly zero-resource keyword spotting. This forms part of a United Nations effort using keyword spotting to support humanitarian relief programmes in parts of Africa where languages are severely under-resourced. We use 1920 isolated keywords (40 types, 34 minutes)…
Mobile training improves speech recognition for users with unique speech characteristics.
This paper proposes a Convolutional Neural Network (CNN) inspired by Multitask Learning (MTL) and based on speech features trained under the joint supervision of softmax loss and center loss, a powerful metric learning strategy, for the recognition of emotion in speech. Speech features such as Spectrograms and Mel-freq…
Long Short-Term Memory (LSTM) is a recurrent neural network (RNN) architecture that has been designed to address the vanishing and exploding gradient problems of conventional RNNs. Unlike feedforward neural networks, RNNs have cyclic connections making them powerful for modeling sequences. They have been successfully u…
This paper improves speech recognition by distilling knowledge from acoustic models.
Although highly correlated, speech and speaker recognition have been regarded as two independent tasks and studied by two communities. This is certainly not the way that people behave: we decipher both speech content and speaker traits at the same time. This paper presents a unified model to perform speech and speaker …
Convolutional Neural Networks (CNNs) are effective models for reducing spectral variations and modeling spectral correlations in acoustic features for automatic speech recognition (ASR). Hybrid speech recognition systems incorporating CNNs with Hidden Markov Models/Gaussian Mixture Models (HMMs/GMMs) have achieved the …
In training speech recognition systems, labeling audio clips can be expensive, and not all data is equally valuable. Active learning aims to label only the most informative samples to reduce cost. For speech recognition, confidence scores and other likelihood-based active learning methods have been shown to be effectiv…
Enhances speech emotion recognition by adapting to varying time scales.
Deep neural networks (DNNs) are now a central component of nearly all state-of-the-art speech recognition systems. Building neural network acoustic models requires several design decisions including network architecture, size, and training loss function. This paper offers an empirical investigation on which aspects of …
AeGAN improves speech clarity in noisy environments.