Improved autoencoder for F0-consistent voice conversion.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The fundamental frequency (F0) represents pitch in speech that determines prosodic characteristics of speech and is needed in various tasks for speech analysis and synthesis. Despite decades of research on this topic, F0 estimation at low signal-to-noise ratios (SNRs) in unexpected noise conditions remains difficult. T…
Deep learning model estimates multiple f0s, melodies, vocals, and bass lines from music.
Proposes a new voice conversion model that preserves pitch patterns.
The fundamental frequency (F0) contour of speech is a key aspect to represent speech prosody that finds use in speech and spoken language analysis such as voice conversion and speech synthesis as well as speaker and language identification. This work proposes new methods to estimate the F0 contour of speech using deep …
Improved GEC with weakly supervised data and iterative decoding.
Hybrid f0 extraction method for various speech modes with high accuracy.
PHBench predicts Series A funding from Product Hunt launch signals with 7.8% accuracy.
We provide a combinatorial presentation of the set F of 3-dimensional generic flows, namely the set of pairs (M,v) with M a compact oriented 3-manifold and v a nowhere-zero vector field on M having generic behaviour along the boundary of M, with M viewed up to diffeomorphism and v up to homotopy on M fixed on the bound…
Let (M,g) be a compact, connected riemannian manifold that is homogeneous, i.e. each pair of points p,q in M have isometric neighborhoods. This paper is a first step towards an understanding of the extent to which it is true that for each "generic" initial condition f0, the solution to the Heat Equation is such that fo…
Real music signals are highly variable, yet they have strong statistical structure. Prior information about the underlying physical mechanisms by which sounds are generated and rules by which complex sound structure is constructed (notes, chords, a complete musical score), can be naturally unified using Bayesian modell…
Develops confidence intervals for unique elements in data streams.
Paper proposes singing voice conversion without parallel data.
Unsupervised model generates distinct intonation codes for speech synthesis.
WaveCycleGAN converts synthetic speech to natural speech using cycle-consistent adversarial networks.
DeepTrust uses NLP to quickly identify and verify financial anomalies on Twitter.