Improved U-Nets with various intermediate blocks enhance singing voice separation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
CLCNet improves noise reduction in hearing aids with deep learning.
Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of clean speech. To improve speech enhancement performance, we tackle the phase estima…
MASnet enhances speech on mobile devices with low latency.
iSTFTNet speeds up mel-spectrogram vocoders without sacrificing quality.
This paper aims to address two issues existing in the current speech enhancement methods: 1) the difficulty of phase estimations; 2) a single objective function cannot consider multiple metrics simultaneously. To solve the first problem, we propose a novel convolutional neural network (CNN) model for complex spectrogra…
This paper shows the susceptibility of spectrogram-based audio classifiers to adversarial attacks and the transferability of such attacks to audio waveforms. Some commonly used adversarial attacks to images have been applied to Mel-frequency and short-time Fourier transform spectrograms, and such perturbed spectrograms…
Convolutional neural networks (CNN) are widely used for speech emotion recognition (SER). In such cases, the short time fourier transform (STFT) spectrogram is the most popular choice for representing speech, which is fed as input to the CNN. However, the uncertainty principles of the short-time Fourier transform preve…
Study improves voice conversion model with Mel-spectrogram augmentation.
Sound source separation has attracted attention from Music Information Retrieval(MIR) researchers, since it is related to many MIR tasks such as automatic lyric transcription, singer identification, and voice conversion. In this paper, we propose an intuitive spectrogram-based model for source separation by adapting U-…
CycleGAN-VC3 improves CycleGAN-VCs for mel-spectrogram conversion.
New method creates minimal submanifolds using complex-valued eigenfunctions.
Complex-valued neural networks avoid spurious local minima.
A complex-valued convolutional network (convnet) implements the repeated application of the following composition of three operations, recursively applying the composition to an input vector of nonnegative real numbers: (1) convolution with complex-valued vectors followed by (2) taking the absolute value of every entry…
This paper proposes a multichannel source separation technique called the multichannel variational autoencoder (MVAE) method, which uses a conditional VAE (CVAE) to model and estimate the power spectrograms of the sources in a mixture. By training the CVAE using the spectrograms of training examples with source-class l…
Complex-valued (p,q)-harmonic morphisms defined and studied.
In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker. We achieve this by training two separate neural networks: (1) A speaker recognition network that produces speaker-discriminative embeddings; (2) …
In this paper, we address the problem of reconstructing a time-domain signal (or a phase spectrogram) solely from a magnitude spectrogram. Since magnitude spectrograms do not contain phase information, we must restore or infer phase information to reconstruct a time-domain signal. One widely used approach for dealing w…
The paper uses complex-valued functions to simplify plane differential geometry and kinematics.
A new RBM model handles both linear and log-amplitude spectrograms.
In the spirit of Ray and Singer we define a complex valued analytic torsion using non-selfadjoint Laplacians. We establish an anomaly formula which permits to turn this into a topological invariant. Conjecturally this analytically defined invariant computes the complex valued Reidemeister torsion, including its phase. …
Bayesian sparsification improves complex-valued neural networks by 50-100x with minimal performance loss.
Study complex-valued VAEs for radar OOD detection.
The construction of synthetic complex-valued signals from real-valued observations is an important step in many time series analysis techniques. The most widely used approach is based on the Hilbert transform, which maps the real-valued signal into its quadrature component. In this paper, we define a probabilistic gene…
We introduce the convolutional spectral kernel (CSK), a novel family of non-stationary, nonparametric covariance kernels for Gaussian process (GP) models, derived from the convolution between two imaginary radial basis functions. We present a principled framework to interpret CSK, as well as other deep probabilistic mo…
Complex-valued neural networks are not a new concept, however, the use of real-valued models has often been favoured over complex-valued models due to difficulties in training and performance. When comparing real-valued versus complex-valued neural networks, existing literature often ignores the number of parameters, r…
We study the curvature of a manifold on which there can be defined a complex-valued submersive harmonic morphism with either, totally geodesic fibers or that is holomorphic with respect to a complex structure which is compatible with the second fundamental form. We also give a necessary curvature condition for the exis…
Usually, complex-valued RKHS are presented as an straightforward application of the real-valued case. In this paper we prove that this procedure yields a limited solution for regression. We show that another kernel, here denoted as pseudo kernel, is needed to learn any function in complex-valued fields. Accordingly, we…
Unified framework for complex-valued eigenfunctions on Riemannian symmetric spaces.
This paper describes a novel energy-based probabilistic distribution that represents complex-valued data and explains how to apply it to direct feature extraction from complex-valued spectra. The proposed model, the complex-valued restricted Boltzmann machine (CRBM), is designed to deal with complex-valued visible unit…
C-SURE improves complex-valued deep learning models by shrinking estimates, outperforming MLE and SurReal.
In this paper we study deep learning-based music source separation, and explore using an alternative loss to the standard spectrogram pixel-level L2 loss for model training. Our main contribution is in demonstrating that adding a high-level feature loss term, extracted from the spectrograms using a VGG net, can improve…
Over the last decade, both the neural network and kernel adaptive filter have successfully been used for nonlinear signal processing. However, they suffer from high computational cost caused by their complex/growing network structures. In this paper, we propose two random Euler filters for complex-valued nonlinear filt…
This paper proves a generalization bound for complex-valued neural networks scaling with spectral complexity.
Optimal transport as a loss for machine learning optimization problems has recently gained a lot of attention. Building upon recent advances in computational optimal transport, we develop an optimal transport non-negative matrix factorization (NMF) algorithm for supervised speech blind source separation (BSS). Optimal …
This research explores complex-valued neural networks and their implementation.
We propose a novel adaptive kernel based regression method for complex-valued signals: the generalized complex-valued kernel least-mean-square (gCKLMS). We borrow from the new results on widely linear reproducing kernel Hilbert space (WL-RKHS) for nonlinear regression and complex-valued signals, recently proposed by th…
New submersion proves complex-valued harmonic map existence.
Paper presents a deep learning framework for classifying respiratory anomalies and lung diseases from sound recordings.
CVNN outperforms RVNN on non-circular data.
Improving speech system performance in noisy environments remains a challenging task, and speech enhancement (SE) is one of the effective techniques to solve the problem. Motivated by the promising results of generative adversarial networks (GANs) in a variety of image processing tasks, we explore the potential of cond…
We extend the complex-valued analytic torsion, introduced by Burghelea and Haller on closed manifolds, to compact Riemannian bordisms. We do so by considering a flat complex vector bundle over a compact Riemannian manifold, endowed with a fiberwise nondegenerate symmetric bilinear form. The Riemmanian metric and the bi…
CVNNs improve performance in tasks with complex-valued inputs.
In this note, we report the back propagation formula for complex valued singular value decompositions (SVD). This formula is an important ingredient for a complete automatic differentiation(AD) infrastructure in terms of complex numbers, and it is also the key to understand and utilize AD in tensor networks.
SpecGrad improves neural vocoder sound quality by adapting diffusion noise to log-mel spectrogram.
Complex valued analytic torsion and dynamical zeta function studied on locally symmetric spaces.
We describe the relationship between complex-valued harmonic morphisms from Minkowski 4-space} and the shear-free ray congruences of mathematical physics. Then we show how a horizontally conformal submersion on a domain of Euclidean 3-space gives the boundary values at infinity of a complex-valued harmonic morphism on …
X-DC improves speech separation by making DNNs more interpretable.