Study shows reverberant phase is not essential for weakly-supervised dereverberation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper, we present a reverberation removal approach for speaker verification, utilizing dual-label deep neural networks (DNNs). The networks perform feature mapping between the spectral features of reverberant and clean speech. Long short term memory recurrent neural networks (LSTMs) are trained to map corrupted…
This paper proposes RAS, a novel unsupervised loss function for speech separation.
This article addresses the modeling of reverberant recording environments in the context of under-determined convolutive blind source separation. We model the contribution of each source to all mixture channels in the time-frequency domain as a zero-mean Gaussian random variable whose covariance encodes the spatial cha…
Study neural networks learning from noisy examples via reverberation.
Unified deep learning framework improves SV in noisy, reverberant, and long non-speech segments.
This paper proposes a deep speech enhancement method which exploits the high potential of residual connections in a wide neural network architecture, a topology known as Wide Residual Network. This is supported on single dimensional convolutions computed alongside the time domain, which is a powerful approach to proces…
We consider the problem of simultaneous reduction of acoustic echo, reverberation and noise. In real scenarios, these distortion sources may occur simultaneously and reducing them implies combining the corresponding distortion-specific filters. As these filters interact with each other, they must be jointly optimized. …
We propose a spatial diffuseness feature for deep neural network (DNN)-based automatic speech recognition to improve recognition accuracy in reverberant and noisy environments. The feature is computed in real-time from multiple microphone signals without requiring knowledge or estimation of the direction of arrival, an…
Quaternion neural networks improve distant speech recognition.
We propose a direction of arrival (DOA) estimation method that combines sound-intensity vector (IV)-based DOA estimation and DNN-based denoising and dereverberation. Since the accuracy of IV-based DOA estimation degrades due to environmental noise and reverberation, two DNNs are used to remove such effects from the obs…
In the last years, increasing efforts have been put into the development of effective stress tests to quantify the resilience of financial institutions. Here we propose a stress test methodology for central counterparties based on a network characterization of clearing members, whose links correspond to direct credits …
We propose a method to generate audio adversarial examples that can attack a state-of-the-art speech recognition model in the physical world. Previous work assumes that generated adversarial examples are directly fed to the recognition model, and is not able to perform such a physical attack because of reverberation an…
In this paper we will try to assess the multifractality displayed by the high-frequency returns of Madrid's Stock Exchange IBEX35 index. A Multifractal Detrended Fluctuation Analysis shows that this index has a wide singularity spectrum which is most likely caused by its long memory. Our findings also show that this lo…
The study examines how social biases are reinforced in machine learning models used for credit scoring.
We introduce Wavesplit, an end-to-end source separation system. From a single mixture, the model infers a representation for each source and then estimates each source signal given the inferred representations. The model is trained to jointly perform both tasks from the raw waveform. Wavesplit infers a set of source re…
Financial contagion from liquidity shocks has being recently ascribed as a prominent driver of systemic risk in interbank lending markets. Building on standard compartment models used in epidemics, in this work we develop an EDB (Exposed-Distressed-Bankrupted) model for the dynamics of liquidity shocks reverberation be…
We present an unsupervised training approach for a neural network-based mask estimator in an acoustic beamforming application. The network is trained to maximize a likelihood criterion derived from a spatial mixture model of the observations. It is trained from scratch without requiring any parallel data consisting of …
This paper develops an active sensing method to estimate the relative weight (or trust) agents place on their neighbors' information in a social network. The model used for the regression is based on the steady state equation in the linear DeGroot model under the influence of stubborn agents, i.e., agents whose opinion…
Credit and liquidity risks represent main channels of financial contagion for interbank lending markets. On one hand, banks face potential losses whenever their counterparties are under distress and thus unable to fulfill their obligations. On the other hand, solvency constraints may force banks to recover lost funding…
Speaker recognition systems based on deep speaker embeddings have achieved significant performance in controlled conditions according to the results obtained for early NIST SRE (Speaker Recognition Evaluation) datasets. From the practical point of view, taking into account the increased interest in virtual assistants (…
This work introduces sequential neural beamforming, which alternates between neural network based spectral separation and beamforming based spatial separation. Our neural networks for separation use an advanced convolutional architecture trained with a novel stabilized signal-to-noise ratio loss function. For beamformi…
RoyalFlush system improves multi-speaker ASR in M2MeT challenge.
Develops a framework to assess systemic risk in the economy using bank-firm network data.
Despite significant efforts over the last few years to build a robust automatic speech recognition (ASR) system for different acoustic settings, the performance of the current state-of-the-art technologies significantly degrades in noisy reverberant environments. Convolutional Neural Networks (CNNs) have been successfu…
Neural codec for high-fidelity audio compression.
Paper presents a new method for better financial market forecasting.
Speech separation refers to extracting each individual speech source in a given mixed signal. Recent advancements in speech separation and ongoing research in this area, have made these approaches as promising techniques for pre-processing of naturalistic audio streams. After incorporating deep learning techniques into…
Hybrid diarization framework handles overlapped speech and long recordings.
X-DC improves speech separation by making DNNs more interpretable.