Recent work on end-to-end automatic speech recognition (ASR) has shown that the connectionist temporal classification (CTC) loss can be used to convert acoustics to phone or character sequences. Such systems are used with a dictionary and separately-trained Language Model (LM) to produce word sequences. However, they a…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper improves EEG-based speech recognition using CTC and beam search.
Because no closed timelike curve (CTC) on a Lorentzian manifold can be deformed to a point, any such manifold containing a CTC must have a topological feature, to be called a timelike wormhole, that prevents the CTC from being deformed to a point. If all wormholes have horizons, which typically seems to be the case in …
We consider the region of closed timelike curves (CTC's) in three-dimensional flat Lorentz spacetimes. The interest in this global geometrical feature goes beyond the purely mathematical. Such spacetimes may be considered lower-dimensional toy models of sourceless Einstein gravity or cosmology. In particular, our inter…
CAT toolkit combines hybrid and E2E approaches for efficient speech recognition.
A new framework improves ASR alignment accuracy via optimal transport.
We report an extension of a Keras Model, called CTCModel, to perform the Connectionist Temporal Classification (CTC) in a transparent way. Combined with Recurrent Neural Networks, the Connectionist Temporal Classification is the reference method for dealing with unsegmented input sequences, i.e. with data that are a co…
Paper shows continuous speech recognition with EEG features, no speech input.
CAT is a new ASR toolkit using CRF and CTC for state-of-the-art speech recognition.
Paper proposes an online speech recognition model using Transformer.
Improved speech recognition using EEG and video.
This paper improves ASR performance by aligning frames more accurately.
Convolutional Neural Networks (CNNs) are effective models for reducing spectral variations and modeling spectral correlations in acoustic features for automatic speech recognition (ASR). Hybrid speech recognition systems incorporating CNNs with Hidden Markov Models/Gaussian Mixture Models (HMMs/GMMs) have achieved the …
Multidimensional recurrent neural networks (MDRNNs) have shown a remarkable performance in the area of speech and handwriting recognition. The performance of an MDRNN is improved by further increasing its depth, and the difficulty of learning the deeper network is overcome by using Hessian-free (HF) optimization. Given…
Continuous speech recognition from brain activity without vocalization.
This paper improves speech recognition by distilling knowledge from acoustic models.
In this work we explored building automatic speech recognition models for transcribing doctor patient conversation. We collected a large scale dataset of clinical conversations ( hr), designed the task to represent the real word scenario, and explored several alignment approaches to iteratively improve data qua…
Keyword spotting--or wakeword detection--is an essential feature for hands-free operation of modern voice-controlled devices. With such devices becoming ubiquitous, users might want to choose a personalized custom wakeword. In this work, we present DONUT, a CTC-based algorithm for online query-by-example keyword spotti…
Improved neural transducer model outperforms attention model on longer sequences.
Causality violations are typically seen as unrealistic and undesirable features of a physical model. The following points out three reasons why causality violations, which Bonnor and Steadman identified even in solutions to the Einstein equation referring to ordinary laboratory situations, are not necessarily undesirab…
CL methods improve monolingual ASR models across new tasks without forgetting past data.
Hybrid and end-to-end models compare in syllable recognition.
Recently, the connectionist temporal classification (CTC) model coupled with recurrent (RNN) or convolutional neural networks (CNN), made it easier to train speech recognition systems in an end-to-end fashion. However in real-valued models, time frame components such as mel-filter-bank energies and the cepstral coeffic…
Neural network generates music scores directly from polyphonic audio.
End-to-end speech recognition using EEG without speech input.
We have recently shown that deep Long Short-Term Memory (LSTM) recurrent neural networks (RNNs) outperform feed forward deep neural networks (DNNs) as acoustic models for speech recognition. More recently, we have shown that the performance of sequence trained context dependent (CD) hidden Markov model (HMM) acoustic m…
This study provides benchmarks for different implementations of LSTM units between the deep learning frameworks PyTorch, TensorFlow, Lasagne and Keras. The comparison includes cuDNN LSTMs, fused LSTM variants and less optimized, but more flexible LSTM implementations. The benchmarks reflect two typical scenarios for au…
We present Listen, Attend and Spell (LAS), a neural network that learns to transcribe speech utterances to characters. Unlike traditional DNN-HMM models, this model learns all the components of a speech recognizer jointly. Our system has two components: a listener and a speller. The listener is a pyramidal recurrent ne…
Sequence-to-sequence attention-based models on subword units allow simple open-vocabulary end-to-end speech recognition. In this work, we show that such models can achieve competitive results on the Switchboard 300h and LibriSpeech 1000h tasks. In particular, we report the state-of-the-art word error rates (WER) of 3.5…
Connectionist temporal classification (CTC) is widely used for maximum likelihood learning in end-to-end speech recognition models. However, there is usually a disparity between the negative maximum likelihood and the performance metric used in speech recognition, e.g., word error rate (WER). This results in a mismatch…
In this paper we deal with the offline handwriting text recognition (HTR) problem with reduced training datasets. Recent HTR solutions based on artificial neural networks exhibit remarkable solutions in referenced databases. These deep learning neural networks are composed of both convolutional (CNN) and long short-ter…
State-level minimum Bayes risk (sMBR) training has become the de facto standard for sequence-level training of speech recognition acoustic models. It has an elegant formulation using the expectation semiring, and gives large improvements in word error rate (WER) over models trained solely using cross-entropy (CE) or co…
Inspired by recent work on neural network image generation which rely on backpropagation towards the network inputs, we present a proof-of-concept system for speech texture synthesis and voice conversion based on two mechanisms: approximate inversion of the representation learned by a speech recognition neural network,…
We study a multiclass multiple instance learning (MIL) problem where the labels only suggest whether any instance of a class exists or does not exist in a training sample or example. No further information, e.g., the number of instances of each class, relative locations or orders of all instances in a training sample, …
Develops STC for sequential data with missing labels.
Study combines speaker verification and voice trigger detection in a single network.
RoyalFlush system improves multi-speaker ASR in M2MeT challenge.
Attention-based encoder decoder network uses a left-to-right beam search algorithm in the inference step. The current beam search expands hypotheses and traverses the expanded hypotheses at the next time step. This traversal is implemented using a for-loop program in general, and it leads to speed down of the recogniti…
Brain2Char decodes text from brain recordings, achieving state-of-the-art performance.
The paper introduces BCART models for aggregate claim amount, improving frequency-severity and joint modeling.
The paper uses model-based trees to create interpretable surrogate models for complex machine learning models.
Gauge Flow Models use a learnable Gauge Field in Generative Flow Models.
The study examines how model predictions hold up under model extensions.
MALC combines interpretable linear models with black-box models for better predictions and transparency.
Revises Bayesian model averaging for foundation models.
Paper introduces symmetric divergence link models for probability distributions.
New method to handle credit portfolio model uncertainties.
The paper tests stock return models and uses LSTM to predict stock returns.