Improved speech emotion recognition using pitch-synchronous single frequency filtering spectrogram.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper proposes a method for instrument classification in polyphonic music using monophonic data.
Deep Autotuner corrects singing pitch using neural networks.
This paper proposes a method for generating speech from filterbank mel frequency cepstral coefficients (MFCC), which are widely used in speech applications, such as ASR, but are generally considered unusable for speech synthesis. First, we predict fundamental frequency and voicing information from MFCCs with an autoreg…
In this paper, we learn disentangled representations of timbre and pitch for musical instrument sounds. We adapt a framework based on variational autoencoders with Gaussian mixture latent distributions. Specifically, we use two separate encoders to learn distinct latent spaces for timbre and pitch, which form Gaussian …
Pitch or fundamental frequency (f0) extraction is a fundamental problem studied extensively for its potential applications in speech and clinical applications. In literature, explicit mode specific (modal speech or singing voice or emotional/ expressive speech or noisy speech) signal processing and deep learning f0 ext…
The state-of-the-art in text-to-speech synthesis has recently improved considerably due to novel neural waveform generation methods, such as WaveNet. However, these methods suffer from their slow sequential inference process, while their parallel versions are difficult to train and even more expensive computationally. …
We describe a machine-learning approach to pitch correcting a solo singing performance in a karaoke setting, where the solo voice and accompaniment are on separate tracks. The proposed approach addresses the situation where no musical score of the vocals nor the accompaniment exists: It predicts the amount of correctio…
In this paper, we investigate the relations between the pitch, the angle of pitch and drall of parallel ruled surface of a closed curve in dual Lorentzian space.
In this paper, we investigate the relations between the pitch, the angle of pitch and drall of parallel ruled surface of a closed spacelike curve with timelike binormal in dual Lorentzian space.
In this paper, we investigate the relations between the pitch, the angle of pitch and drall of parallel ruled surface of a closed spacelike curve with a spacelike binormal in dual Lorentzian space.
The task of estimating the fundamental frequency of a monophonic sound recording, also known as pitch tracking, is fundamental to audio processing with multiple applications in speech processing and music information retrieval. To date, the best performing techniques, such as the pYIN algorithm, are based on a combinat…
Proposes a new voice conversion model that preserves pitch patterns.
Music SketchNet generates missing measures in incomplete music pieces, guided by user input.
We advance the state of the art in polyphonic piano music transcription by using a deep convolutional and recurrent neural network which is trained to jointly predict onsets and frames. Our model predicts pitch onset events and then uses those predictions to condition framewise pitch predictions. During inference, we r…
Researchers create exact minimal surfaces with helical motifs in biological structures.
The fundamental frequency (F0) represents pitch in speech that determines prosodic characteristics of speech and is needed in various tasks for speech analysis and synthesis. Despite decades of research on this topic, F0 estimation at low signal-to-noise ratios (SNRs) in unexpected noise conditions remains difficult. T…
Automatic music transcription (AMT) aims to infer a latent symbolic representation of a piece of music (piano-roll), given a corresponding observed audio recording. Transcribing polyphonic music (when multiple notes are played simultaneously) is a challenging problem, due to highly structured overlapping between harmon…
Novel higher-order group synchronization for noisy local measurements on hypergraphs.
End-to-end speech synthesis is a promising approach that directly converts raw text to speech. Although it was shown that Tacotron2 outperforms classical pipeline systems with regards to naturalness in English, its applicability to other languages is still unknown. Japanese could be one of the most difficult languages …
New method uses neural networks for accurate angle estimation in noisy conditions.
The sectoral synchronization observed for the Japanese business cycle in the Indices of Industrial Production data is an example of synchronization. The stability of this synchronization under a shock, e.g., fluctuation of supply or demand, is a matter of interest in physics and economics. We consider an economic syste…
Hybrid approach for large-scale network synchronization using KF and PTP.
We analyze how an observer synchronizes to the internal state of a finite-state information source, using the epsilon-machine causal representation. Here, we treat the case of exact synchronization, when it is possible for the observer to synchronize completely after a finite number of observations. The more difficult …
In this paper, we investigate the asymptotic behavior of regular ends of flat surfaces in the hyperbolic 3-space H^3. Galvez, Martinez and Milan showed that when the singular set does not accumulate at an end, the end is asymptotic to a rotationally symmetric flat surface. As a refinement of their result, we show that …
The human auditory system is able to distinguish the vocal source of thousands of speakers, yet not much is known about what features the auditory system uses to do this. Fourier Transforms are capable of capturing the pitch and harmonic structure of the speaker but this alone proves insufficient at identifying speaker…
New algorithm uses PSO to optimize DNN training parameters in distributed systems.
ShadowSync separates background synchronization for scalable distributed training.
New method synchronizes graphs with probability measures on rotations.
We introduce MIDI-VAE, a neural network model based on Variational Autoencoders that is capable of handling polyphonic music with multiple instrument tracks, as well as modeling the dynamics of music by incorporating note durations and velocities. We show that MIDI-VAE can perform style transfer on symbolic music by au…
Study predicts synchronization state of financial time series using cross-recurrence plots.
Efficiently estimates rotations with corrupted data.
This paper proposes a general model for synchronized crowding behavior. An order parameter is introduced to quantify the level of synchronization which is shown a function of percentage of agents in reactive state. Further, synchronization is shown to be driven by the most active agents with the highest volatility. A t…
Study optimizes estimation of orthogonal and rotation matrices from noisy data.
Solves complex clustering and rotation synchronization problem.
Spectral method for joint community detection and group synchronization.
Spectral methods achieve near-optimal performance in orthogonal and permutation group synchronization.
Study financial markets using synchronization measures and clustering algorithms.
Networks of coupled dynamical systems provide a powerful way to model systems with enormously complex dynamics, such as the human brain. Control of synchronization in such networked systems has far reaching applications in many domains, including engineering and medicine. In this paper, we formulate the synchronization…
Machine learning predicts synchronization transitions in unknown systems.
New method linearizes nonlinear coupled oscillators on graphs.
New approach predicts stock price synchronization using RNNs and LSTMs.
New method solves group synchronization with cycle-edge message passing.
Adaptive synchronization improves deep reinforcement learning performance.
KuramotoGNN uses Kuramoto model to prevent over-smoothing in graph neural networks.
New MAB model for online caching costs.
Proposes ELASTICBSP for faster, more flexible distributed deep learning training.
Paper proposes a bijective approach for signal/symbol translation using variational auto-encoders.