A key barrier to making phonetic studies scalable and replicable is the need to rely on subjective, manual annotation. To help meet this challenge, a machine learning algorithm was developed for automatic measurement of a widely used phonetic measure: vowel duration. Manually-annotated data were used to train a model t…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper explores how different audio signal representations affect topological signatures and their predictive power.
A comparative study of the application of Gaussian Mixture Model (GMM) and Radial Basis Function (RBF) in biometric recognition of voice has been carried out and presented. The application of machine learning techniques to biometric authentication and recognition problems has gained a widespread acceptance. In this res…
VOWEL trains WTA-SNNs for multi-valued events, overcoming resource limitations.
Deep learning maps tongue movements to speech sounds for voiceless individuals.
Selecting important features in non-linear or kernel spaces is a difficult challenge in both classification and regression problems. When many of the features are irrelevant, kernel methods such as the support vector machine and kernel ridge regression can sometimes perform poorly. We propose weighting the features wit…
-algebra consists of expressions constructed with four kinds operations, the minimum, maximum, difference and additively homogeneous generalized means. Five families of -classifiers are investigated on binary classification tasks between English phonemes. It is shown that the classifiers are able to reflect well…
Paper aims to find joint representation between vocal tract geometry and speech sound acoustics.
Human infants can discover words directly from unsegmented speech signals without any explicitly labeled data. In this paper, we develop a novel machine learning method called nonparametric Bayesian double articulation analyzer (NPB-DAA) that can directly acquire language and acoustic models from observed continuous sp…
Recent studies have introduced end-to-end TTS, which integrates the production of context and acoustic features in statistical parametric speech synthesis. As a result, a single neural network replaced laborious feature engineering with automated feature learning. However, little is known about what types of context in…
Improved speech recognition using EEG and video.
Study shows emotion affects speaker recognition and vice versa.
With the recent renaissance of deep convolution neural networks, encouraging breakthroughs have been achieved on the supervised recognition tasks, where each class has sufficient training data and fully annotated training data. However, to scale the recognition to a large number of classes with few or now training samp…
New proof for sphere recognition algorithm.
Continuous speech recognition from brain activity without vocalization.
VoxCeleb 2019 challenge assesses speaker recognition in uncontrolled settings.
Paper proposes Roweisposes for 3D action recognition using generalized eigenvalue problem.
This is a survey article on recognition problem of frontal singularities. We specify geometrically several frontal singularities and then we solve the recognition problem of such singularities, giving explicit normal forms. We combine the recognition results by K. Saji and several arguments on openings, which was perfo…
Paper tackles zero-shot activity recognition using video features and text embeddings.
Paper explores EEG-based speech recognition using transformers, showing faster training and better performance for smaller vocabularies.
EmbraceNet fusion model for multi-sensor activity recognition.
In this paper we demonstrate end-to-end continuous speech recognition (CSR) using electroencephalography (EEG) signals with no speech signal as input. An attention model based automatic speech recognition (ASR) and connectionist temporal classification (CTC) based ASR systems were implemented for performing recognition…
TransFall uses transfer learning to improve activity recognition from mobile sensors.
AV-CPL uses continuous pseudo-labels for AVSR combining labeled and unlabeled data.
Neural network framework for language recognition considers sequence information and improves accuracy.
Fawkes protects images from unauthorized facial recognition models.
Paper proposes a framework to protect user anonymity in emotion recognition.
Learned feature representations and sub-phoneme posteriors from Deep Neural Networks (DNNs) have been used separately to produce significant performance gains for speaker and language recognition tasks. In this work we show how these gains are possible using a single DNN for both speaker and language recognition. The u…
Random forest can be adapted for open-set recognition with improved performance.
Enhances 2D face recognition with 3D features using active illumination.
This paper presents a novel method for structural data recognition using a large number of graph models. In general, prevalent methods for structural data recognition have two shortcomings: 1) Only a single model is used to capture structural variation. 2) Naive recognition methods are used, such as the nearest neighbo…
We build CSI-Net, a unified Deep Neural Network~(DNN), to learn the representation of WiFi signals. Using CSI-Net, we jointly solved two body characterization problems: biometrics estimation (including body fat, muscle, water, and bone rates) and person recognition. We also demonstrated the application of CSI-Net on tw…
Improved speech emotion recognition using pre-trained language models.
Proves NP and co-NP status for knot core recognition in solid torus.
Todays interactive devices such as smart-phone assistants and smart speakers often deal with short-duration speech segments. As a result, speaker recognition systems integrated into such devices will be much better suited with models capable of performing the recognition task with short-duration utterances. In this pap…
The performance of automatic speech recognition systems(ASR) degrades in the presence of noisy speech. This paper demonstrates that using electroencephalography (EEG) can help automatic speech recognition systems overcome performance loss in the presence of noise. The paper also shows that distillation training of auto…
NIST CTS Superset offers a large dataset for telephony speaker recognition.
Study proposes a decision tree for more accurate depression recognition in speech.
FineHand learns hand shapes for better ASL recognition.
A new approach to unsupervised learning using recognition-parametrised models.
Long Short-Term Memory (LSTM) is a recurrent neural network (RNN) architecture that has been designed to address the vanishing and exploding gradient problems of conventional RNNs. Unlike feedforward neural networks, RNNs have cyclic connections making them powerful for modeling sequences. They have been successfully u…
Personalized activity recognition improves performance for diverse users.
Current UAV-recorded datasets are mostly limited to action recognition and object tracking, whereas the gesture signals datasets were mostly recorded in indoor spaces. Currently, there is no outdoor recorded public video dataset for UAV commanding signals. Gesture signals can be effectively used with UAVs by leveraging…
VoiceFilter-Lite separates speech from background in real-time for on-device speech recognition.
We introduce an approach based on moving frames for polygon recognition and symmetry detection. We present detailed algorithms for recognition of polygons modulo the special Euclidean, Euclidean, equi-affine, skewed-affine and similarity Lie groups, and explain the procedure for a generic Lie group. The time complexity…
Survey examines public views on facial recognition technology.
Despite the remarkable progress in face recognition related technologies, reliably recognizing faces across ages still remains a big challenge. The appearance of a human face changes substantially over time, resulting in significant intra-class variations. As opposed to current techniques for age-invariant face recogni…
Transfer learning does not improve character recognition performance.