End-to-end speech recognition system trained on GPUs and CPUs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper aims to find joint representation between vocal tract geometry and speech sound acoustics.
This paper examines the speaker identification potential of breath sounds in continuous speech. Speech is largely produced during exhalation. In order to replenish air in the lungs, speakers must periodically inhale. When inhalation occurs in the midst of continuous speech, it is generally through the mouth. Intra-spee…
Geometric framework for aligning fiber tracts across subjects.
Recent speech technology research has seen a growing interest in using WaveNets as statistical vocoders, i.e., generating speech waveforms from acoustic features. These models have been shown to improve the generated speech quality over classical vocoders in many tasks, such as text-to-speech synthesis and voice conver…
Tract-specific diffusion measures, as derived from brain diffusion MRI, have been linked to white matter tract structural integrity and neurodegeneration. As a consequence, there is a large interest in the automatic segmentation of white matter tract in diffusion tensor MRI data. Methods based on the tractography are p…
In this paper we propose a fusion approach to continuous emotion recognition that combines visual and auditory modalities in their representation spaces to predict the arousal and valence levels. The proposed approach employs a pre-trained convolution neural network and transfer learning to extract features from video …
New method clusters infant vocalizations using topological data.
Framework converts singer identity and vocal technique from non-parallel corpora.
Understanding how housing values evolve over time is important to policy makers, consumers and real estate professionals. Existing methods for constructing housing indices are computed at a coarse spatial granularity, such as metropolitan regions, which can mask or distort price dynamics apparent in local markets, such…
Study improves machine learning models for GI tract disease detection using comprehensive evaluations and cross-dataset testing.
Study examines equity in post-Snow Uri recovery, finds disparities.
Vocal disorders have affected several patients all over the world. Due to the inherent difficulty of diagnosing vocal disorders without sophisticated equipment and trained personnel, a number of patients remain undiagnosed. To alleviate the monetary cost of diagnosis, there has been a recent growth in the use of data a…
In this paper, we present our approach for the 2018 Medico Task classifying diseases in the gastrointestinal tract. We have proposed a system based on global features and deep neural networks. The best approach combines two neural networks, and the reproducible experimental results signify the efficiency of the propose…
We describe a machine-learning approach to pitch correcting a solo singing performance in a karaoke setting, where the solo voice and accompaniment are on separate tracks. The proposed approach addresses the situation where no musical score of the vocals nor the accompaniment exists: It predicts the amount of correctio…
A new model cleans vocal note event annotations in music.
In this paper, we propose a classification based glottal closure instants (GCI) detection from pathological acoustic speech signal, which finds many applications in vocal disorder analysis. Till date, GCI for pathological disorder is extracted from laryngeal (glottal source) signal recorded from Electroglottograph, a d…
In this paper, we use several techniques with conventional vocal feature extraction (MFCC, STFT), along with deep-learning approaches such as CNN, and also context-level analysis, by providing the textual data, and combining different approaches for improved emotion-level classification. We explore models that have not…
DC-SIS selects features faster than mRMR for Parkinson's vocal diagnosis.
The paper provides conditions for realizing graphs and polytopes with specified edge lengths.
Fundamental frequency (f0) estimation from polyphonic music includes the tasks of multiple-f0, melody, vocal, and bass line estimation. Historically these problems have been approached separately, and only recently, using learning-based approaches. We present a multitask deep learning architecture that jointly estimate…
Have you ever wondered how a song might sound if performed by a different artist? In this work, we propose SCM-GAN, an end-to-end non-parallel song conversion system powered by generative adversarial and transfer learning that allows users to listen to a selected target singer singing any song. SCM-GAN first separates …
Speech produced by human vocal apparatus conveys substantial non-semantic information including the gender of the speaker, voice quality, affective state, abnormalities in the vocal apparatus etc. Such information is attributed to the properties of the voice source signal, which is usually estimated from the speech sig…
Singing voice separation attempts to separate the vocal and instrumental parts of a music recording, which is a fundamental problem in music information retrieval. Recent work on singing voice separation has shown that the low-rank representation and informed separation approaches are both able to improve separation qu…
Quantum trace map defines invariants for knots and links, confirming a length conjecture.
Paper proposes efficient multivariate spatial Fay-Herriot models using variational autoencoders.
Given a state-of-the-art deep neural network text classifier, we show the existence of a universal and very small perturbation vector (in the embedding space) that causes natural text to be misclassified with high probability. Unlike images on which a single fixed-size adversarial perturbation can be found, text is of …
A new neural network separates vocals from music accompaniment.
Research into automated systems for detecting and classifying marine mammals in acoustic recordings is expanding internationally due to the necessity to analyze large collections of data for conservation purposes. In this work, we present a Convolutional Neural Network that is capable of classifying the vocalizations o…
New method provides reliable probabilistic bounds for VUR detection.
Study geodesics on graphs with random lengths, proving bi-infinite paths exist.
Improved online Lasso reduces regret in sparse linear contextual bandits.
To accurately analyze changes of anatomical structures in longitudinal imaging studies, consistent segmentation across multiple time-points is required. Existing solutions often involve independent registration and segmentation components. Registration between time-points is used either as a prior for segmentation in a…
Following \cite{citeSavelyevVirtualMorsetheoryonHam.}, we develop here a connection between Morse theory for the (positive) Hofer length functional , with Gromov-Witten/Floer theory, for monotone symplectic manifolds . This gives some immediate restrictio…
A faster Bayesian method for estimating spatial count data models.
Hierarchical CNNs improve diagnosis of GI diseases from histopathological images.
We prove a dynamical wave trace formula for asymptotically hyperbolic (n+1) dimensional manifolds with negative (but not necessarily constant) sectional curvatures which equates the renormalized wave trace to the lengths of closed geodesics. A corollary of this dynamical trace formula is a dynamical resonance-wave trac…
One way to interpret trained deep neural networks (DNNs) is by inspecting characteristics that neurons in the model respond to, such as by iteratively optimising the model input (e.g., an image) to maximally activate specific neurons. However, this requires a careful selection of hyper-parameters to generate interpreta…
We extend the well-known Denjoy-Ahlfors theorem on the number of different asymptotic tracts of holomorphic functions to subharmonic functions on arbitrary Riemannian manifolds. We obtain some new versions of the Liouville theorem for $\p$-harmonic functions without requiring the geodesic completeness requirement of a …
Continuous speech recognition from brain activity without vocalization.
Shorter adversarial prompts help protect LLMs from jailbreak attacks.
New counterexample shows curved surfaces can deform geodesics without diffeomorphism.
We prove stability and exponential convergence of the Perfectly Matched Layer (PML) method for acoustic scattering on manifolds with axial analytic quasicylindrical ends. These manifolds model long-range geometric perturbations (e.g. bending or stretching) of tubular waveguides filled with homogeneous or inhomogeneous …
New approach reduces unconstrained linear bandits to simpler optimization problems.
Fast algorithm for rescaling vectors with clipping, improving training efficiency.
In this work we show that the systems of balance equations (balance systems) of continuum thermodynamics occupy a natural place in the variational bicomplex formalism. We apply the vertical homotopy decomposition to get a local splitting (in a convenient domain) of a general balance system as the sum of a Lagrangian pa…
In this paper, we present a technique for unsupervised learning of visual representations. Specifically, we train a model for foreground and background classification task, in the process of which it learns visual representations. Foreground and background patches for training come af- ter mining for such patches from …
This study tackles offline RL with perturbed data sources, deriving a lower bound and proposing an optimal algorithm.