This paper proposes an end-to-end approach for single-channel speaker-independent multi-speaker speech separation, where time-frequency (T-F) masking, the short-time Fourier transform (STFT), and its inverse are represented as layers within a deep network. Previous approaches, rather than computing a loss on the recons…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper introduces a differentiable STFT for continuous window length optimization.
Recently, we proposed short-time Fourier transform (STFT)-based loss functions for training a neural speech waveform model. In this paper, we generalize the above framework and propose a training scheme for such models based on spectral amplitude and phase losses obtained by either STFT or continuous wavelet transform …
Proposes a differentiable STFT for more efficient optimization of hop length.
Many neural speech enhancement and source separation systems operate in the time-frequency domain. Such models often benefit from making their Short-Time Fourier Transform (STFT) front-ends trainable. In current literature, these are implemented as large Discrete Fourier Transform matrices; which are prohibitively inef…
iSTFTNet speeds up mel-spectrogram vocoders without sacrificing quality.
This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly. Not only amplitude spectra but also phase spectra obtained from generated speech waveforms are used to c…
Recent deep learning approaches have achieved impressive performance on speech enhancement and separation tasks. However, these approaches have not been investigated for separating mixtures of arbitrary sounds of different types, a task we refer to as universal sound separation, and it is unknown how performance on spe…
This study proposes a trainable adaptive window switching (AWS) method and apply it to a deep-neural-network (DNN) for speech enhancement in the modified discrete cosine transform domain. Time-frequency (T-F) mask processing in the short-time Fourier transform (STFT)-domain is a typical speech enhancement method. To re…
Research into automated systems for detecting and classifying marine mammals in acoustic recordings is expanding internationally due to the necessity to analyze large collections of data for conservation purposes. In this work, we present a Convolutional Neural Network that is capable of classifying the vocalizations o…
Improved speech enhancement with MNTFA using time-frequency attention.
Diffusion models enhance speech without supervision.
Hybrid model improves geopolitical conflict forecasting.
In this paper, we use several techniques with conventional vocal feature extraction (MFCC, STFT), along with deep-learning approaches such as CNN, and also context-level analysis, by providing the textual data, and combining different approaches for improved emotion-level classification. We explore models that have not…
Supervised learning based on a deep neural network recently has achieved substantial improvement on speech enhancement. Denoising networks learn mapping from noisy speech to clean one directly, or to a spectrum mask which is the ratio between clean and noisy spectra. In either case, the network is optimized by minimizi…
Convolutional neural networks (CNN) are widely used for speech emotion recognition (SER). In such cases, the short time fourier transform (STFT) spectrogram is the most popular choice for representing speech, which is fed as input to the CNN. However, the uncertainty principles of the short-time Fourier transform preve…
Speech separation refers to extracting each individual speech source in a given mixed signal. Recent advancements in speech separation and ongoing research in this area, have made these approaches as promising techniques for pre-processing of naturalistic audio streams. After incorporating deep learning techniques into…
Optimal transport as a loss for machine learning optimization problems has recently gained a lot of attention. Building upon recent advances in computational optimal transport, we develop an optimal transport non-negative matrix factorization (NMF) algorithm for supervised speech blind source separation (BSS). Optimal …
The inverse-free extreme learning machine (ELM) algorithm proposed in [4] was based on an inverse-free algorithm to compute the regularized pseudo-inverse, which was deduced from an inverse-free recursive algorithm to update the inverse of a Hermitian matrix. Before that recursive algorithm was applied in [4], its impr…
Epilepsy affects nearly 1% of the global population, of which two thirds can be treated by anti-epileptic drugs and a much lower percentage by surgery. Diagnostic procedures for epilepsy and monitoring are highly specialized and labour-intensive. The accuracy of the diagnosis is also complicated by overlapping medical …
Proof of convergence for multi-objective optimization using inverse reinforcement learning.
Two new inverse-free ELM algorithms for incremental and decremental learning are proposed.
The notion of a generalized harmonic inverse mean curvature surface in the Euclidean four-space is introduced. A backward Bäcklund transform of a generalized harmonic inverse mean curvature surface is defined. A Darboux transform of a generalized harmonic inverse mean curvature surface is constructed by a backward Bäck…
New method for estimating parameters in inverse problems using double robustness.
New findings on mesh group-planes validate Signature-inverse Theorem under specific conditions.
Seismic inversion method uses GAN to improve efficiency and accuracy.
Develops an inverse particle filter for cognitive systems.
The paper studies Möbius inversion on surfaces in Minkowski 3-space.
Inversive distance circle packing metric was introduced by P Bowers and K Stephenson \cite{BS} as a generalization of Thurston's circle packing metric \cite{T1}. They conjectured that the inversive distance circle packings are rigid. For nonnegative inversive distance, Guo \cite{Guo} proved the infinitesimal rigidity a…
The aim of the paper is to investigate the relation between inverse limit of branched manifolds and codimension zero laminations. We give necessary and sufficient conditions for such an inverse limit to be a lamination. We also show that codimension zero laminations are inverse limits of branched manifolds. The inverse…
CNN outperforms other methods in gravity inversion.
The inversion formula for conservative multifractal measures was unveiled mathematically a decade ago, which is however not well tested in real complex systems. In this Letter, we propose to verify the inversion formula using high-frequency turbulent financial data. We construct conservative volatility measure based on…
Paper proposes efficient image inversion and editing using rectified stochastic differential equations.
Global inverse function theorem proved easily using Riemannian geometry.
New filters improve radar target inference in complex scenarios.
Study uses machine learning to solve photoacoustic tomography's inverse problem.
The paper explains how microlocal analysis solves geometric inverse problems.
MCGDiff uses SGM to guide SMC for solving ill-posed linear inverse problems.
Study constructs solutions for evolving hypersurfaces using inverse spacetime mean curvature.
Deep learning methods improve subsurface flow modeling efficiency.
The aim of this paper is to show how the homotopy type of compact metric spaces can be reconstructed by the inverse limit of an inverse sequence of finite approximations of the corresponding space. This recovering allows us to define inverse persistence as a new kind of persistence process.
Machine learning models solve inverse eigenvalue problems for symmetric potentials and refractive indices.
This paper proposes the recursive and square-root BLS algorithms to improve the original BLS for new added inputs, which utilize the inverse and inverse Cholesky factor of the Hermitian matrix in the ridge inverse, respectively, to update the ridge solution. The recursive BLS updates the inverse by the matrix inversion…
Inverts operator on hyperbolic surfaces, constructing invariant distributions.
Abstract: Generalizes Milnor-Schwarz lemma to inverse monoids.
Proves rigidity of circle packings in the plane, generalizing previous work.
We give a counterexample of Bowers-Stephenson's conjecture in the spherical case: spherical inversive distance circle packings are not determined by their inversive distances.
EnKG solves inverse problems without derivatives, using diffusion models.