Models predict Alzheimer's Dementia from spontaneous speech with high accuracy.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
ADReSS Challenge at INTERSPEECH 2020 benchmarks speech recognition for Alzheimer's dementia.
In this paper, we propose to apply object detection methods from the vision domain on the speech recognition domain, by treating audio fragments as objects. More specifically, we present SpeechYOLO, which is inspired by the YOLO algorithm for object detection in images. The goal of SpeechYOLO is to localize boundaries …
Emotion recognition from speech is one of the key steps towards emotional intelligence in advanced human-machine interaction. Identifying emotions in human speech requires learning features that are robust and discriminative across diverse domains that differ in terms of language, spontaneity of speech, recording condi…
Study reveals AI's spontaneous topic changes in text prediction.
This paper provides an attempt to formalize Hayek's notion of spontaneous order within the framework of the Arrow-Debreu economy. Our study shows that if a competitive economy is enough fair and free, then a spontaneous economic order shall emerge in long-run competitive equilibria so that social members together occup…
Study on spontaneous symmetry breaking in financial markets using quantum mechanics.
The article explores Helfrich flow with spontaneous curvature, finding singularities and convergence behaviors.
Investigates spontaneous symmetry breaking in non-equilibrium systems.
The paper connects financial vacuum conditions to spontaneous symmetry breaking in quantum finance.
New model detects Alzheimer's and severity from speech, cognitive, and language data.
Automated prediction of public speaking performance enables novel systems for tutoring public speaking skills. We use the largest open repository---TED Talks---to predict the ratings provided by the online viewers. The dataset contains over 2200 talk transcripts and the associated meta information including over 5.5 mi…
Study on dynamic curves with elastic energy and spontaneous curvature.
Study detects SLI in children from spontaneous narrative transcripts.
Recent theoretical advances in elasticity of membranes following Helfrich's famous spontaneous curvature model are summarized in this review. The governing equations describing equilibrium configurations of lipid vesicles, lipid membranes with free edges, and chiral lipid membranes are presented. Several analytic solut…
New model suggests universe emerges from single particle quantum mechanics.
Permutation of any two hidden units yields invariant properties in typical deep generative neural networks. This permutation symmetry plays an important role in understanding the computation performance of a broad class of neural networks with two or more hidden units. However, a theoretical study of the permutation sy…
GE-autoencoder identifies spontaneous symmetry breaking in systems.
We introduce the concept of spontaneous symmetry breaking to arbitrage modeling. In the model, the arbitrage strategy is considered as being in the symmetry breaking phase and the phase transition between arbitrage mode and no-arbitrage mode is triggered by a control parameter. We estimate the control parameter for mom…
The goal of this paper is twofold. First we prove a rigidity estimate, which generalises the theorem on geometric rigidity of Friesecke, James and Müller to 1-forms with non-vanishing exterior derivative. Second we use this estimate to prove a kind of spontaneous breaking of rotational symmetry for some models of cryst…
The study optimizes cell membranes' shapes based on curvature and proves existence of minimizers.
Geometric mechanism mimics physics' symmetry breaking.
Firm foundation theory estimates a security's firm fundamental value based on four determinants: expected growth rate, expected dividend payout, the market interest rate and the degree of risk. In contrast, other views of decision-making in the stock market, using alternatives such as human psychology and behavior, bou…
This paper presents sampling-based speech parameter generation using moment-matching networks for Deep Neural Network (DNN)-based speech synthesis. Although people never produce exactly the same speech even if we try to express the same linguistic and para-linguistic information, typical statistical speech synthesis pr…
Study on surfaces minimizing elastic energy with boundary constraints.
In this paper we demonstrate end-to-end continuous speech recognition (CSR) using electroencephalography (EEG) signals with no speech signal as input. An attention model based automatic speech recognition (ASR) and connectionist temporal classification (CTC) based ASR systems were implemented for performing recognition…
Detects AI-synthesized speech using cepstral and bispectral analysis.
The performance of automatic speech recognition systems(ASR) degrades in the presence of noisy speech. This paper demonstrates that using electroencephalography (EEG) can help automatic speech recognition systems overcome performance loss in the presence of noise. The paper also shows that distillation training of auto…
This paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach to speech enhancement is to train a deep neural network (DNN) to take noisy speech as input and output clean speech. Although this supervis…
AudioPaLM combines text and speech models to improve speech processing and translation.
WaveCycleGAN has recently been proposed to bridge the gap between natural and synthesized speech waveforms in statistical parametric speech synthesis and provides fast inference with a moving average model rather than an autoregressive model and high-quality speech synthesis with the adversarial training. However, the …
Improved speech enhancement using diffusion models with MSE loss.
In this paper we demonstrate continuous noisy speech recognition using connectionist temporal classification (CTC) model on limited Chinese vocabulary using electroencephalography (EEG) features with no speech signal as input and we further demonstrate single CTC model based continuous noisy speech recognition on limit…
Continuous speech recognition from brain activity without vocalization.
AV-CPL uses continuous pseudo-labels for AVSR combining labeled and unlabeled data.
Speech-driven visual speech synthesis involves mapping features extracted from acoustic speech to the corresponding lip animation controls for a face model. This mapping can take many forms, but a powerful approach is to use deep neural networks (DNNs). However, a limitation is the lack of synchronized audio, video, an…
In this paper we demonstrate spoken speech enhancement using electroencephalography (EEG) signals using a generative adversarial network (GAN) based model, gated recurrent unit (GRU) regression based model, temporal convolutional network (TCN) regression model and finally using a mixed TCN GRU regression model. We comp…
Diffusion models enhance speech without supervision.
Speech synthesis from EEG features using RNN.
Survey on DNNs for speech processing, focusing on limited data challenges.
We present RUSLAN -- a new open Russian spoken language corpus for the text-to-speech task. RUSLAN contains 22200 audio samples with text annotations -- more than 31 hours of high-quality speech of one person -- being the largest annotated Russian corpus in terms of speech duration for a single speaker. We trained an e…
Automatic speech recognition (ASR) systems are of vital importance nowadays in commonplace tasks such as speech-to-text processing and language translation. This created the need for an ASR system that can operate in realistic crowded environments. Thus, speech enhancement is a valuable building block in ASR systems an…
Speech enhancement improved by adapting to unknown speakers without auxiliary signals.
Synthetic speech data improves keyword spotting models with fewer real examples.
VoiceFilter-Lite separates speech from background in real-time for on-device speech recognition.
This paper proposes a speech emotion recognition method based on speech features and speech transcriptions (text). Speech features such as Spectrogram and Mel-frequency Cepstral Coefficients (MFCC) help retain emotion-related low-level characteristics in speech whereas text helps capture semantic meaning, both of which…
We propose a learning-based filter that allows us to directly modify a synthetic speech waveform into a natural speech waveform. Speech-processing systems using a vocoder framework such as statistical parametric speech synthesis and voice conversion are convenient especially for a limited number of data because it is p…
Enhanced transformer converts whispered speech to natural speech.