Some explanations to Kaldi's PLDA implementation to make formula derivation easier to catch.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
GPU acceleration speeds up i-vector extraction 3000x, enabling new research.
Improved hybrid acoustic model using interleaved self-attention and convolution.
Hybrid and end-to-end models compare in syllable recognition.
In this paper we study the probabilistic properties of the posteriors in a speech recognition system that uses a deep neural network (DNN) for acoustic modeling. We do this by reducing Kaldi's DNN shared pdf-id posteriors to phone likelihoods, and using test set forced alignments to evaluate these using a calibration s…
We describe the neural-network training framework used in the Kaldi speech recognition toolkit, which is geared towards training DNNs with large amounts of training data using multiple GPU-equipped or multi-core machines. In order to be as hardware-agnostic as possible, we needed a way to use multiple machines without …
CAT toolkit combines hybrid and E2E approaches for efficient speech recognition.
Non-autoregressive transformer improves speech recognition speed and accuracy.
Study speaker verification security using hierarchical Bayesian modeling.