Neural model detects phoneme boundaries from speech, outperforming baselines.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Self-supervised model detects phoneme boundaries without annotations.
The dominant automatic lexical stress detection method is to split the utterance into syllable segments using phoneme sequence and their time-aligned boundaries. Then we extract features from syllable to use classification method to classify the lexical stress. However, we can't get very accurate time boundaries of eac…
We consider the problem of training speech recognition systems without using any labeled data, under the assumption that the learner can only access to the input utterances and a phoneme language model estimated from a non-overlapping corpus. We propose a fully unsupervised learning algorithm that alternates between so…
Metric functions for phoneme perception capture the similarity structure among phonemes in a given language and therefore play a central role in phonology and psycho-linguistics. Various phenomena depend on phoneme similarity, such as spoken word recognition or serial recall from verbal working memory. This study prese…
Study investigates predictive coding models for phonemic learning.
For decades, context-dependent phonemes have been the dominant sub-word unit for conventional acoustic modeling systems. This status quo has begun to be challenged recently by end-to-end models which seek to combine acoustic, pronunciation, and language model components into a single neural network. Such systems, which…
Recent character and phoneme-based parametric TTS systems using deep learning have shown strong performance in natural speech generation. However, the choice between character or phoneme input can create serious limitations for practical deployment, as direct control of pronunciation is crucial in certain cases. We dem…
-algebra consists of expressions constructed with four kinds operations, the minimum, maximum, difference and additively homogeneous generalized means. Five families of -classifiers are investigated on binary classification tasks between English phonemes. It is shown that the classifiers are able to reflect well…
The goal of this project is to develop a limited lip reading algorithm for a subset of the English language. We consider a scenario in which no audio information is available. The raw video is processed and the position of the lips in each frame is extracted. We then prepare the lip data for processing and classify the…
We replace the Hidden Markov Model (HMM) which is traditionally used in in continuous speech recognition with a bi-directional recurrent neural network encoder coupled to a recurrent neural network decoder that directly emits a stream of phonemes. The alignment between the input and output sequences is established usin…
We stabilize the activations of Recurrent Neural Networks (RNNs) by penalizing the squared distance between successive hidden states' norms. This penalty term is an effective regularizer for RNNs including LSTMs and IRNNs, improving performance on character-level language modeling and phoneme recognition, and outperfor…
We use automatic speech recognition to assess spoken English learner pronunciation based on the authentic intelligibility of the learners' spoken responses determined from support vector machine (SVM) classifier or deep learning neural network model predictions of transcription correctness. Using numeric features produ…
In this project we further investigate the idea of reducing the dimensionality of datasets using a Borel isomorphism with the purpose of subsequently applying supervised learning algorithms, as originally suggested by my supervisor V. Pestov (in 2011 Dagstuhl preprint). Any consistent learning algorithm, for example kN…
Neural model detects DD risk in 5-year-olds, predicting 2 years ahead.
Develops slope detection for 3-manifolds with torus boundaries.
This paper benchmarks speech LVMs against deterministic models and adapts a video model to speech.
One-Class Boundary Peeling detects outliers efficiently and robustly.
Study detects boundaries in unlabeled noisy images without labels.
Sharp boundaries for detecting dense subhypergraphs established.
Grapheme-based acoustic modeling has recently been shown to outperform phoneme-based approaches in both hybrid and end-to-end automatic speech recognition (ASR), even on non-phonemic languages like English. However, graphemic ASR still has problems with rare long-tail words that do not follow the standard spelling conv…
Speech-related Brain Computer Interface (BCI) technologies provide effective vocal communication strategies for controlling devices through speech commands interpreted from brain signals. In order to infer imagined speech from active thoughts, we propose a novel hierarchical deep learning BCI system for subject-indepen…
Recurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks in- cluding machine translation, handwriting synthesis and image caption gen- eration. We extend the attention-mechanism with features needed for speech recognition. We sh…
The paper improves boundary detection and density estimation on noisy data.
Enhanced neural networks detect thin boundaries between different types of anomalies.
It has been an open question whether all boundary slopes of hyperbolic knots are strongly detected by the character variety. The main result of this paper produces an infinite family of hyperbolic knots each of which has at least one strict boundary slope that is not strongly detected by the character variety.
This manuscript addresses the problem of the automatic lesion boundary detection in dermoscopy, using deep neural networks. An approach is based on the adaptation of the U-net convolutional neural network with skip connections for lesion boundary segmentation task. I hope this paper could serve, to some extent, as an e…
BDSG generates samples on distribution boundaries, improving anomaly detection.
Recently, the connectionist temporal classification (CTC) model coupled with recurrent (RNN) or convolutional neural networks (CNN), made it easier to train speech recognition systems in an end-to-end fashion. However in real-valued models, time frame components such as mel-filter-bank energies and the cepstral coeffic…
Proposes a boundary detection method inspired by LLE for high-dimensional data.
Detects exotic embeddings in 4-manifolds with boundary.
Constructs Gabor frames for curved manifolds to detect boundaries.
Paper extends ICA to ISA with auxiliary variables for better speech representation learning.
We consider the task of unsupervised extraction of meaningful latent representations of speech by applying autoencoding neural networks to speech waveforms. The goal is to learn a representation able to capture high level semantic content from the signal, e.g.\ phoneme identities, while being invariant to confounding l…
Unified framework detects shifts in climate boundaries using GP regression and MAD test.
Study investigates how simple speech sounds can form abstract categories.
MCBP detects boundaries in high-dimensional data using curvature.
Invariants from surface Khovanov-Jacobsson classes help detect knots and slices.
New method uses weighting vectors for efficient boundary and outlier detection.
The A-polynomial of a manifold whose boundary consists of a single torus is generalised to an eigenvalue variety of a manifold whose boundary consists of a finite number of tori, and the set of strongly detected boundary curves is determined by Bergman's logarithmic limit set, which describes the exponential behaviour …
CRAUM-Net improves salient object detection with context and uncertainty modeling.
Proposes a novel method for generating hard negatives near time series data boundaries.
FROB model improves robustness and reliable confidence for few-shot OoD detection.
Detects dense subhypergraphs in heterogeneous random hypergraphs.
TailGAN uses GANs to detect anomalies near data distribution tails.
TPA-AD detects axle-box bearing anomalies using pseudo anomalies near normal boundaries.
We propose a simple but efficient method termed Guided Learning for weakly-labeled semi-supervised sound event detection (SED). There are two sub-targets implied in weakly-labeled SED: audio tagging and boundary detection. Instead of designing a single model by considering a trade-off between the two sub-targets, we de…
Study slopes on knot manifolds to understand their fundamental groups.