A new RBM model handles both linear and log-amplitude spectrograms.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Convolutional neural networks (CNN) are widely used for speech emotion recognition (SER). In such cases, the short time fourier transform (STFT) spectrogram is the most popular choice for representing speech, which is fed as input to the CNN. However, the uncertainty principles of the short-time Fourier transform preve…
Improving speech system performance in noisy environments remains a challenging task, and speech enhancement (SE) is one of the effective techniques to solve the problem. Motivated by the promising results of generative adversarial networks (GANs) in a variety of image processing tasks, we explore the potential of cond…
A typical audio signal processing pipeline includes multiple disjoint analysis stages, including calculation of a time-frequency representation followed by spectrogram-based feature analysis. We show how time-frequency analysis and nonnegative matrix factorisation can be jointly formulated as a spectral mixture Gaussia…
iSTFTNet speeds up mel-spectrogram vocoders without sacrificing quality.
This paper aims to address two issues existing in the current speech enhancement methods: 1) the difficulty of phase estimations; 2) a single objective function cannot consider multiple metrics simultaneously. To solve the first problem, we propose a novel convolutional neural network (CNN) model for complex spectrogra…
This paper shows the susceptibility of spectrogram-based audio classifiers to adversarial attacks and the transferability of such attacks to audio waveforms. Some commonly used adversarial attacks to images have been applied to Mel-frequency and short-time Fourier transform spectrograms, and such perturbed spectrograms…
Sound source separation has attracted attention from Music Information Retrieval(MIR) researchers, since it is related to many MIR tasks such as automatic lyric transcription, singer identification, and voice conversion. In this paper, we propose an intuitive spectrogram-based model for source separation by adapting U-…
CycleGAN-VC3 improves CycleGAN-VCs for mel-spectrogram conversion.
This paper proposes a multichannel source separation technique called the multichannel variational autoencoder (MVAE) method, which uses a conditional VAE (CVAE) to model and estimate the power spectrograms of the sources in a mixture. By training the CVAE using the spectrograms of training examples with source-class l…
For training the sequence-to-sequence voice conversion model, we need to handle an issue of insufficient data about the number of speech pairs which consist of the same utterance. This study experimentally investigated the effects of Mel-spectrogram augmentation on training the sequence-to-sequence voice conversion (VC…
In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker. We achieve this by training two separate neural networks: (1) A speaker recognition network that produces speaker-discriminative embeddings; (2) …
In this paper, we address the problem of reconstructing a time-domain signal (or a phase spectrogram) solely from a magnitude spectrogram. Since magnitude spectrograms do not contain phase information, we must restore or infer phase information to reconstruct a time-domain signal. One widely used approach for dealing w…
CAP-BM learns complex-valued data's amplitude and phase distributions.
We introduce the convolutional spectral kernel (CSK), a novel family of non-stationary, nonparametric covariance kernels for Gaussian process (GP) models, derived from the convolution between two imaginary radial basis functions. We present a principled framework to interpret CSK, as well as other deep probabilistic mo…
In this paper we study deep learning-based music source separation, and explore using an alternative loss to the standard spectrogram pixel-level L2 loss for model training. Our main contribution is in demonstrating that adding a high-level feature loss term, extracted from the spectrograms using a VGG net, can improve…
Optimal transport as a loss for machine learning optimization problems has recently gained a lot of attention. Building upon recent advances in computational optimal transport, we develop an optimal transport non-negative matrix factorization (NMF) algorithm for supervised speech blind source separation (BSS). Optimal …
Quantum computer method for pricing rainbow options efficiently.
This paper shows how learning the phase-amplitude coupling improves bio-signal classification.
New insights connect strong coupling SYM amplitudes to hyperkähler geometry.
Paper presents a deep learning framework for classifying respiratory anomalies and lung diseases from sound recordings.
A fast method for estimating radar amplitude density parameters.
Singing Voice Separation (SVS) tries to separate singing voice from a given mixed musical signal. Recently, many U-Net-based models have been proposed for the SVS task, but there were no existing works that evaluate and compare various types of intermediate blocks that can be used in the U-Net architecture. In this pap…
Improved formulation of spinfoam quantum gravity with cosmological constant, ensuring all amplitudes are finite and providing semiclassical asymptotics.
SpecGrad improves neural vocoder sound quality by adapting diffusion noise to log-mel spectrogram.
Atiyah classes of DG manifolds of positive amplitude are invariant under weak equivalences.
The paper provides exact multivariate amplitude distributions for non-stationary Gaussian or algebraic fluctuations.
We define a topological quantum membrane theory on a seven dimensional manifold of holonomy. We describe in detail the path integral evaluation for membrane geometries given by circle bundles over Riemann surfaces. We show that when the target space is quantum amplitudes of non-local observables …
Paper computes Atiyah class for DG manifolds of amplitude +1.
X-DC improves speech separation by making DNNs more interpretable.
We demonstrate the equivalence of all loop closed topological string amplitudes on toric local Calabi-Yau threefolds with computations of certain knot invariants for Chern-Simons theory. We use this equivalence to compute the topological string amplitudes in certain cases to very high degree and to all genera. In parti…
PolarBM models complex-valued audio signals in polar coordinates, improving over conventional methods.
MaskCycleGAN-VC improves voice conversion without parallel data.
We introduce a fully coherent spin network amplitude whose expansion generates all SU(2) spin networks associated with a given graph. We then give an explicit evaluation of this amplitude for an arbitrary graph. We show how this coherent amplitude can be obtained from the specialization of a generating functional obtai…
CLCNet improves noise reduction in hearing aids with deep learning.
We study topological open string amplitudes on orientifolds without fixed planes. We determine the contributions of the untwisted and twisted sectors as well as the BPS structure of the amplitudes. We illustrate our general results in various examples involving D-branes in toric orientifolds. We perform the computation…
AaSP improves audio self-supervised learning by addressing aliasing issues.
Convolutional neural network (CNN) architectures have originated and revolutionized machine learning for images. In order to take advantage of CNNs in predictive modeling with audio data, standard FFT-based signal processing methods are often applied to convert the raw audio waveforms into an image-like representations…
The paper proves a category of dg manifolds with finite positive amplitude.
Algorithm finds frequencies, amplitudes, and phases of sinusoids in noisy data.
End-to-end neural network based approaches to audio modelling are generally outperformed by models trained on high-level data representations. In this paper we present preliminary work that shows the feasibility of training the first layers of a deep convolutional neural network (CNN) model to learn the commonly-used l…
We decompose the exchange rates returns of 41 currencies (incl. gold) into their sign and amplitude components. Then we group together all exchange rates with a common base currency, construct Minimal Spanning Trees for each group independently, and analyze properties of these trees. We show that both the sign and the …
iSTFTNet2 improves iSTFTNet's speed and lightness with 1D-2D CNN.
We show how the amplitude of holonomies on a vector bundle can be controlled by the integral of the curvature of the connection on a surface enclosed by the curve.
In this paper we propose a novel environmental sound classification approach incorporating unsupervised feature learning from codebook via spherical -Means++ algorithm and a new architecture for high-level data augmentation. The audio signal is transformed into a 2D representation using a discrete wavelet transform …
The paper analyzes the amplitude of functions on the sphere, improving FDA methods.
Transformers predict scattering amplitudes in theoretical physics.
The transition amplitudes between coherent states on a coherent state manifold are expressed in terms of the embedding of the coherent state manifold into a projective Hilbert space. Consequences for the dimension of projective Hilbert space and a simple geometric interpretation of Calabi's diastasis follows.