Deep Autotuner corrects singing pitch using neural networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
WeSinger improves singing voice synthesis with data augmentation and specialized modules.
Singing voice conversion is a task to convert a song sang by a source singer to the voice of a target singer. In this paper, we propose using a parallel data free, many-to-one voice conversion technique on singing voices. A phonetic posterior feature is first generated by decoding singing voices through a robust Automa…
We study the problem of semi-supervised singing voice separation, in which the training data contains a set of samples of mixed music (singing and instrumental) and an unmatched set of instrumental music. Our solution employs a single mapping function g, which, applied to a mixed sample, recovers the underlying instrum…
SING improves state inference in latent SDE models for better drift function estimation.
We study codimension one (transversally oriented) foliations $\fa$ on oriented closed manifolds having non-empty compact singular set $\sing(\fa)$ which is locally defined by Bott-Morse functions. We prove that if the transverse type of $\fa$ at each singular point is a center and $\fa$ has a compact leaf with fini…
Deep neural networks with convolutional layers usually process the entire spectrogram of an audio signal with the same time-frequency resolutions, number of filters, and dimensionality reduction scale. According to the constant-Q transform, good features can be extracted from audio signals if the low frequency bands ar…
Jukebox generates high-fidelity songs with singing in raw audio.
Singing Voice Separation (SVS) tries to separate singing voice from a given mixed musical signal. Recently, many U-Net-based models have been proposed for the SVS task, but there were no existing works that evaluate and compare various types of intermediate blocks that can be used in the U-Net architecture. In this pap…
In recent years, deep learning has surpassed traditional approaches to the problem of singing voice separation. The Wave-U-Net is a recent deep network architecture that operates directly on the time domain. The standard Wave-U-Net is trained with data augmentation and early stopping to prevent overfitting. Minimum hyp…
Singing voice separation attempts to separate the vocal and instrumental parts of a music recording, which is a fundamental problem in music information retrieval. Recent work on singing voice separation has shown that the low-rank representation and informed separation approaches are both able to improve separation qu…
Simplicial sets deformation retract onto transverse simplices.
This paper summarizes some recent advances on a set of tasks related to the processing of singing using state-of-the-art deep learning techniques. We discuss their achievements in terms of accuracy and sound quality, and the current challenges, such as availability of data and computing resources. We also discuss the i…
We describe a machine-learning approach to pitch correcting a solo singing performance in a karaoke setting, where the solo voice and accompaniment are on separate tracks. The proposed approach addresses the situation where no musical score of the vocals nor the accompaniment exists: It predicts the amount of correctio…
Separating a singing voice from its music accompaniment remains an important challenge in the field of music information retrieval. We present a unique neural network approach inspired by a technique that has revolutionized the field of vision: pixel-wise image classification, which we combine with cross entropy loss a…
Parallel transport in a fibre bundle with respect to smooth paths in the base space B have recently been extended to representations of the smooth singular simplicial set Sing_{smooth}(B). Inspired by these extensions,I revisit the development of a notion of `parallel' transport in the topological setting of fibrations…
We present a deep learning method for singing voice conversion. The proposed network is not conditioned on the text or on the notes, and it directly converts the audio of one singer to the voice of another. Training is performed without any form of supervision: no lyrics or any kind of phonetic features, no notes, and …
We prove the factoriality of the following nodal threefolds: a complete intersection of hypersurfaces and of degree and respectively, where is smooth, , ; a double cover of a smooth hypersurface $F\subset\mathbb{P}^{…
If and are homotopic embedded surfaces in a -manifold then they may be related by a regular homotopy (at the expense of introducing double points) or by a sequence of stabilisations and destabilisations (at the expense of adding genus). This naturally gives rise to two integer-valued notions of distance bet…
A new algorithm for training generative models using Sinkhorn divergence.
Recent progress in deep learning for audio synthesis opens the way to models that directly produce the waveform, shifting away from the traditional paradigm of relying on vocoders or MIDI synthesizers for speech or music generation. Despite their successes, current state-of-the-art neural audio synthesizers such as Wav…
The paper proves smoothness of almost-minimizers' boundaries near the free boundary.
{\bf Construction.} For a dominating polynomial mapping {} with an isolated critical value at 0 ( an algebraically closed field of characteristic zero) we construct a closed {\it bundle} . We restrict over the critical points of in and partiti…
The singular set of a foliation is always connected under certain conditions.
Recently, the principal component pursuit has received increasing attention in signal processing research ranging from source separation to video surveillance. So far, all existing formulations are real-valued and lack the concept of phase, which is inherent in inputs such as complex spectrograms or color images. Thus,…
Study on singularities of area-minimizing currents, focusing on frequency and branch points.
Algorithm learns non-Gaussian graphical models via Hessian scores and triangular transport.
Let be a compact and irreducible Hermitian complex space. This paper is devoted to various questions concerning the analytic K-homology of . In the fist part, assuming either or , we show that the rolled-up operator of the minimal -$\overline{\pa…
Establishes functoriality of Baum-Bott residues under specific conditions.
We introduce a combinatorial argument to study closed minimal hypersurfaces of bounded area and high Morse index. Let be a closed Riemannian manifold and be a closed embedded minimal hypersurface with area at most and with a singular set of Hausdorff dimension at most . We show the…
Study quantifies properties of PMC hypersurfaces with area bounds.
Sound source separation has attracted attention from Music Information Retrieval(MIR) researchers, since it is related to many MIR tasks such as automatic lyric transcription, singer identification, and voice conversion. In this paper, we propose an intuitive spectrogram-based model for source separation by adapting U-…
In this work we prove a Baum-Bott type formula for non-compact complex manifold of the form , where is a complex compact manifold and is a normal crossing divisor on . As applications, we provide a Poincaré-Hopf type Theorem and an optimal description for a smooth hypersur…
One way to interpret trained deep neural networks (DNNs) is by inspecting characteristics that neurons in the model respond to, such as by iteratively optimising the model input (e.g., an image) to maximally activate specific neurons. However, this requires a careful selection of hyper-parameters to generate interpreta…
Neural networks approximate Calabi-Yau metrics and curvature.
Defines axial curvatures for corank 1 singular manifolds in higher dimensions.
We propose a flexible framework that deals with both singer conversion and singers vocal technique conversion. The proposed model is trained on non-parallel corpora, accommodates many-to-many conversion, and leverages recent advances of variational autoencoders. It employs separate encoders to learn disentangled latent…
Vector-valued neural learning has emerged as a promising direction in deep learning recently. Traditionally, training data for neural networks (NNs) are formulated as a vector of scalars; however, its performance may not be optimal since associations among adjacent scalars are not modeled. In this paper, we propose a n…
Let be a smooth manifold and let $\F$ be a codimension one, foliation on , with isolated singularities of Morse type. The study and classification of pairs $(M,\F)$ is a challenging (and difficult) problem. In this setting, a classical result due to Reeb \cite{Reeb} states that a manifold admitting a …
Deep clustering is the first method to handle general audio separation scenarios with multiple sources of the same type and an arbitrary number of sources, performing impressively in speaker-independent speech separation tasks. However, little is known about its effectiveness in other challenging situations such as mus…
Let be a -dimensional complex manifold and two distinct holomorphic self-maps. Suppose that and coincide on a globally irreducible compact hypersurface . We show that if one of the two maps is a local biholomorphism around and, if needed, sits into …
Models for audio source separation usually operate on the magnitude spectrum, which ignores phase information and makes separation performance dependant on hyper-parameters for the spectral front-end. Therefore, we investigate end-to-end source separation in the time-domain, which allows modelling phase information and…
We extend the results of our recent preprint [arXiv: 1811.00515] into higher dimensions . For minimizing harmonic maps from -dimensional domains into the two dimensional sphere we prove: (1) An extension of Almgren and Lieb's linear law, namely \[\mathcal{H}^{n-3}(\textrm{sin…
Optimal regularity theory for stable minimal hypersurfaces with small singular set.
This research investigates reliable local explanations for machine listening models.
In this paper, we propose a provably correct algorithm for convolutive nonnegative matrix factorization (CNMF) under separability assumptions. CNMF is a convolutive variant of nonnegative matrix factorization (NMF), which functions as an NMF with additional sequential structure. This model is useful in a number of appl…
Have you ever wondered how a song might sound if performed by a different artist? In this work, we propose SCM-GAN, an end-to-end non-parallel song conversion system powered by generative adversarial and transfer learning that allows users to listen to a selected target singer singing any song. SCM-GAN first separates …
Global singularities propagate in magnetic mechanical systems on Riemannian manifolds.