Enhanced Tacotron for Japanese speech synthesis improves naturalness.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
GE2E-AC improves accent classification by focusing on accent embeddings.
This paper describes the data collection effort that is part of the project Sprekend Nederland (The Netherlands Talking), and discusses its potential use in Automatic Accent Location. We define Automatic Accent Location as the task to describe the accent of a speaker in terms of the location of the speaker and its hist…
MPSA-DenseNet improves accent classification accuracy.
State-of-the-art automatic speech recognition (ASR) systems struggle with the lack of data for rare accents. For sufficiently large datasets, neural engines tend to outshine statistical models in most natural language processing problems. However, a speech accent remains a challenge for both approaches. Phonologists ma…
We investigated the impact of noisy linguistic features on the performance of a Japanese speech synthesis system based on neural network that uses WaveNet vocoder. We compared an ideal system that uses manually corrected linguistic features including phoneme and prosodic information in training and test sets against a …
ConvS2S-VC converts voice characteristics and pitch contour using a fully convolutional seq2seq model.
Deep Autotuner corrects singing pitch using neural networks.
Model disentangles timbre and pitch for musical instruments.
Deep Autotuner corrects singing pitch without scores, using vocal and accompaniment spectral data.
Hybrid f0 extraction method for various speech modes with high accuracy.
In this paper, we investigate the relations between the pitch, the angle of pitch and drall of parallel ruled surface of a closed curve in dual Lorentzian space.
Improved speech emotion recognition using pitch-synchronous single frequency filtering spectrogram.
In this paper, we investigate the relations between the pitch, the angle of pitch and drall of parallel ruled surface of a closed spacelike curve with timelike binormal in dual Lorentzian space.
In this paper, we investigate the relations between the pitch, the angle of pitch and drall of parallel ruled surface of a closed spacelike curve with a spacelike binormal in dual Lorentzian space.
The task of estimating the fundamental frequency of a monophonic sound recording, also known as pitch tracking, is fundamental to audio processing with multiple applications in speech processing and music information retrieval. To date, the best performing techniques, such as the pYIN algorithm, are based on a combinat…
Proposes a new voice conversion model that preserves pitch patterns.
Improved LID for multilingual speakers using context-aware models.
Music SketchNet generates missing measures in incomplete music pieces, guided by user input.
We advance the state of the art in polyphonic piano music transcription by using a deep convolutional and recurrent neural network which is trained to jointly predict onsets and frames. Our model predicts pitch onset events and then uses those predictions to condition framewise pitch predictions. During inference, we r…
The aim of this article is to present the category of bounded Frechet manifolds in respect to which we will review the geometry of Frechet manifolds with a stronger accent on its metric aspect. An inverse function theorem in the sense of Nash and Moser in this category is proved, and some applications to Riemannian geo…
Researchers create exact minimal surfaces with helical motifs in biological structures.
Unified analysis of efficient local training methods for distributed variational inequalities.
Active learning reduces SP calculations by 90%.
The fundamental frequency (F0) represents pitch in speech that determines prosodic characteristics of speech and is needed in various tasks for speech analysis and synthesis. Despite decades of research on this topic, F0 estimation at low signal-to-noise ratios (SNRs) in unexpected noise conditions remains difficult. T…
Automatic music transcription (AMT) aims to infer a latent symbolic representation of a piece of music (piano-roll), given a corresponding observed audio recording. Transcribing polyphonic music (when multiple notes are played simultaneously) is a challenging problem, due to highly structured overlapping between harmon…
EC^2-VAE generates music analogies by disentangling pitch and rhythm representations.
In this paper, we investigate the asymptotic behavior of regular ends of flat surfaces in the hyperbolic 3-space H^3. Galvez, Martinez and Milan showed that when the singular set does not accumulate at an end, the end is asymptotic to a rotationally symmetric flat surface. As a refinement of their result, we show that …
The human auditory system is able to distinguish the vocal source of thousands of speakers, yet not much is known about what features the auditory system uses to do this. Fourier Transforms are capable of capturing the pitch and harmonic structure of the speaker but this alone proves insufficient at identifying speaker…
Paper proposes a bijective approach for signal/symbol translation using variational auto-encoders.
MIDI-VAE models music dynamics and instrumentation for style transfer.
Hidden Markov Models (HMMs) are a ubiquitous tool to model time series data, and have been widely used in two main tasks of Automatic Music Transcription (AMT): note segmentation, i.e. identifying the played notes after a multi-pitch estimation, and sequential post-processing, i.e. correcting note segmentation using tr…
TimbreTron transfers musical timbre using CQT and WaveNet.
WeSinger improves singing voice synthesis with data augmentation and specialized modules.
This study compares PRNG and QRNG in machine learning models, revealing significant differences in performance.
Geometric characterization of sub-Riemannian geodesics on frame bundles.
This paper proposes a method for instrument classification in polyphonic music using monophonic data.
The fundamental frequency (F0) contour of speech is a key aspect to represent speech prosody that finds use in speech and spoken language analysis such as voice conversion and speech synthesis as well as speaker and language identification. This work proposes new methods to estimate the F0 contour of speech using deep …
We present a supervised neural network model for polyphonic piano music transcription. The architecture of the proposed model is analogous to speech recognition systems and comprises an acoustic model and a music language model. The acoustic model is a neural network used for estimating the probabilities of pitches in …
For every genus , we prove that contains complete, properly embedded, genus- minimal surfaces whose two ends are asymptotic to helicoids of any prescribed pitch. We also show that as the radius of the tends to infinity, these examples converge smoothly to complete, properly embedded minimal s…
Study of motion control systems on Lie groups with specific geometric constraints.
SING generates musical notes from instruments in real-time.
For every genus g, we prove that S^2 x R contains complete, properly embedded, genus-g minimal surfaces whose two ends are asymptotic to helicoids of any prescribed pitch. We also show that as the radius of the S^2 tends to infinity, these examples converge smoothly to complete, properly embedded minimal surfaces in Eu…
In this paper, we present a reverberation removal approach for speaker verification, utilizing dual-label deep neural networks (DNNs). The networks perform feature mapping between the spectral features of reverberant and clean speech. Long short term memory recurrent neural networks (LSTMs) are trained to map corrupted…
We describe all possible self-similar motions of immersed hypersurfaces in Euclidean space under the mean curvature flow and derive the corresponding hypersurface equations. Then we present a new two-parameter family of immersed helicoidal surfaces that rotate/translate with constant velocity under the flow. We look at…
DDSP integrates signal processing with deep learning for high-fidelity audio synthesis.
Paper proposes a new framework for SAD using GANs.
Paper compares XGB and BPNN for music style classification.