Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

2579 · Apr 201919922001200920182026
48 results for singing voice

Semi-supervised singing voice separation using synthetic mixtures.

problem Singing voice separation with limited labeled data.
method Trains a single mapping function g on synthetic mixtures of singing and instrumental music.
result Performance comparable to fully supervised methods, better than semi-supervised alternatives.

Improved U-Nets with various intermediate blocks enhance singing voice separation.

problem Improving singing voice separation accuracy using U-Net architectures.
method Implemented and compared U-Nets with different intermediate spectrogram transformation blocks.
result A specific block type achieves state-of-the-art SDR by 0.9 dB.

WeSinger improves singing voice synthesis with data augmentation and specialized modules.

problem Improving the accuracy and naturalness of synthesized singing voices.
method Developed a multi-singer Chinese neural singing voice synthesis system with deep bi-directional LSTM, Transformer, LPCNet, and data augmentation.
result WeSinger achieves state-of-the-art performance on the Opencpop corpus.

Wave-U-Net with MHE regularization improves singing voice separation.

problem Singing voice separation from mixed music recordings.
method Wave-U-Net architecture with MHE regularization applied to 1D filters.
result Adding MHE regularization to the loss function consistently improves singing voice separation.

New neural network separates singing voices from music using cross entropy loss.

problem Separating singing voices from music accompaniment.
method Deep Convolutional Neural Network (CNN) trained with Ideal Binary Mask (IBM) and cross entropy loss.
result Proposed CNN outperforms existing systems in MIREX evaluations.

New algorithm separates vocals from music recordings efficiently.

problem Separate vocal and instrumental parts in music recordings.
method Informed group-sparse representation for linear-time singing voice separation.
result Efficacy confirmed on iKala dataset; music accompaniment follows group-sparse structure.

New algorithms separate singing voices from accompaniment using complex and quaternionic principal component pursuit.

problem Separating singing voices from instrumental accompaniment using phase information.
method Extended principal component pursuit to complex and quaternionic cases, developed new proximity operators, applied inexact augmented Lagrange multiplier algorithm.
result Phase information improves singing voice separation.

Deep Autotuner corrects singing pitch without scores, using vocal and accompaniment spectral data.

problem Automatic pitch correction without musical scores for singing performances.
method Convolutional Gated Recurrent Unit (CGRU) model trained on karaoke data.
result The model predicts pitch correction from vocal and accompaniment spectral contents, making the voice sound in tune with the accompaniment.

ABIPNN improves neural network performance by processing vectors in each neuron.

problem Traditional neural networks fail to model associations among adjacent scalars.
method ABIPNN uses arbitrary bilinear products to process vector-valued neurons.
result ABIPNN outperforms conventional neural networks in multispectral image denoising and singing voice separation.

Spectrogram-Channels U-Net separates sounds by treating each channel as a source's spectrogram.

problem Sound source separation in music information retrieval.
method Adapting U-Net to treat each channel of the output as a source's spectrogram, balancing volumes between sources.
result State-of-the-art performance on singing voice and multi-instrument separation.

Efficiently generates and selects explanations for neural networks using GANs and FID.

problem Manual selection of hyper-parameters for generating interpretable neural network explanations is slow and requires qualitative evaluation.
method Proposes a novel metric using Fréchet Inception Distance (FID) and a GAN-based method for efficient search and realistic output generation.
result Successfully selects hyper-parameters leading to interpretable examples, avoiding manual evaluation.

Framework converts singer identity and vocal technique from non-parallel corpora.

problem Converts singer identity and vocal technique from non-parallel corpora.
method Uses variational autoencoders with separate encoders for singer identity and vocal technique.
result Successfully disentangles and converts singer identity and vocal technique.

We study codimension one (transversally oriented) foliations $\fa$ on oriented closed manifolds MM having non-empty compact singular set $\sing(\fa)$ which is locally defined by Bott-Morse functions. We prove that if the transverse type of $\fa$ at each singular point is a center and $\fa$ has a compact leaf with fini…

2006-08-23abs ↗pdf ↗

Unified approach to DP problems using Gumbel distribution and variational Bayesian inference.

problem Solving classical optimal path problems in a probabilistic framework.
method Gumbel distribution and variational Bayesian inference for latent optimal paths.
result Unified approach transforms DP problems into directed acyclic graphs with Gibbs distribution.

A system is presented that segments, clusters and predicts musical audio in an unsupervised manner, adjusting the number of (timbre) clusters instantaneously to the audio input. A sequence learning algorithm adapts its structure to a dynamically changing clustering tree. The flow of the system is as follows: 1) segment…

2015-02-02abs ↗pdf ↗

This research investigates reliable local explanations for machine listening models.

problem Generating reliable local explanations for machine listening models.
method Investigates the sensitivity of SoundLIME explanations to input perturbations and proposes a novel method for identifying suitable content types.
result SoundLIME explanations are sensitive to the content in occluded input regions, and the average magnitude of input mel-spectrogram bins is the most suitable content type for temporal explanations.

Solution for voice conversion with limited data using hierarchical seq2seq and attention models.

problem Voice conversion between speakers with limited parallel audio pairs.
method Hierarchical sequence to sequence model with attention-based decoder, trained on single speaker dataset.
result Improved voice conversion quality using mel spectrograms and wavenet vocoder.

Parallel transport in a fibre bundle with respect to smooth paths in the base space B have recently been extended to representations of the smooth singular simplicial set Sing_{smooth}(B). Inspired by these extensions,I revisit the development of a notion of `parallel' transport in the topological setting of fibrations…

2011-11-23abs ↗pdf ↗

AUTOVC converts voices without parallel data, achieving state-of-the-art results.

problem Non-parallel many-to-many voice conversion and zero-shot voice conversion.
method Only an autoencoder with a carefully designed bottleneck is used, training on a self-reconstruction loss.
result AUTOVC achieves state-of-the-art results in many-to-many voice conversion with non-parallel data and performs zero-shot voice conversion.

SING generates musical notes from instruments in real-time.

problem Efficiently generating high-quality audio from MIDI data.
method Frame-by-frame waveform generation with a single decoder, using a new loss function.
result SING produces significantly improved audio quality compared to state-of-the-art models, with 32x faster training and 2,500x faster inference.

Wave-U-Net improves audio source separation by modeling phase information.

problem Fixed spectral transformations and high sampling rates limit audio source separation performance.
method Wave-U-Net adapts U-Net to time-domain, using repeated resampling to capture different time scales.
result Wave-U-Net achieves comparable performance to spectrogram-based U-Net on singing voice separation.

We prove the factoriality of the following nodal threefolds: a complete intersection of hypersurfaces FF and GP5G\subset\mathbb{P}^{5} of degree nn and kk respectively, where GG is smooth, Sing(FG)(n+k2)(n1)/5|\mathrm{Sing}(F\cap G)|\leqslant(n+k-2)(n-1)/5, nkn\geqslant k; a double cover of a smooth hypersurface $F\subset\mathbb{P}^{…

2004-10-10abs ↗pdf ↗

Study improves low-quality speech data for voice cloning.

problem Feasibility of training spoofing systems with low-quality data.
method Developed a GAN-based speech enhancement system, trained TTS and voice conversion models.
result Significant improvement in SNR and perceptual cleanliness of low-quality data.