Bayesian method suppresses low-frequency pulses in audio recordings.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
GACELA fills long gaps in musical audio with a GAN and context conditioning.
Neural network-based vocoders have recently demonstrated the powerful ability to synthesize high-quality speech. These models usually generate samples by conditioning on spectral features, such as Mel-spectrogram and fundamental frequency, which is crucial to speech synthesis. However, the feature extraction procession…
We present a probabilistic modeling and inference framework for discriminative analysis dictionary learning under a weak supervision setting. Dictionary learning approaches have been widely used for tasks such as low-level signal denoising and restoration as well as high-level classification tasks, which can be applied…
In this paper, we address the problem of reconstructing a time-domain signal (or a phase spectrogram) solely from a magnitude spectrogram. Since magnitude spectrograms do not contain phase information, we must restore or infer phase information to reconstruct a time-domain signal. One widely used approach for dealing w…
Study subjective perception of low light restored images and develop an unsupervised QA model.
New image restoration method using localized patches and external databases.
Improved image restoration using frequency-guided sampling.
RGI improves robustness of GAN-inversion for image restoration and anomaly detection.
Deep audio prior uses neural networks to solve audio problems without data.
Bayesian filtering approach identifies nonlinear restoring forces in dynamic systems.
Unified framework for image restoration using equivariant denoisers.
This paper proposes a zero-shot learning approach for audio classification based on the textual information about class labels without any audio samples from target classes. We propose an audio classification system built on the bilinear model, which takes audio feature embeddings and semantic class label embeddings as…
Proposes a new model for image restoration combining deep learning and total variation.
Generative models improve image restoration from unknown transformations.
Pyramid Attention Networks improve image restoration by leveraging self-similarities across scales.
Recently, there has been great interest in the field of audio style transfer, where a stylized audio is generated by imposing the style of a reference audio on the content of a target audio. We improve on the current approaches which use neural networks to extract the content and the style of the audio signal and propo…
The problem of image restoration in cryo-EM entails correcting for the effects of the Contrast Transfer Function (CTF) and noise. Popular methods for image restoration include `phase flipping', which corrects only for the Fourier phases but not amplitudes, and Wiener filtering, which requires the spectral signal to noi…
Spatial audio is an essential medium to audiences for 3D visual and auditory experience. However, the recording devices and techniques are expensive or inaccessible to the general public. In this work, we propose a self-supervised audio spatialization network that can generate spatial audio given the corresponding vide…
Method learns audio embeddings with contextualized tags.
Study improves radio show segmentation using audio embeddings.
Online audio advertising is a particular form of advertising used abundantly in online music streaming services. In these platforms, which tend to host tens of thousands of unique audio advertisements (ads), providing high quality ads ensures a better user experience and results in longer user engagement. Therefore, th…
The paper develops algorithms to restore monotonicity in non-monotone functions.
Suppose a genus two handlebody is removed from a 3-manifold M and then a single meridian of the handlebody is restored. The result is a knot or link complement in M and it is natural to ask whether geometric properties of the link complement say something about the meridian that was restored. Here we consider what the …
Deep learning methods are becoming widely used for restoration of defects associated with fluorescence microscopy imaging. One of the major challenges in application of such methods is the availability of training data. In this work, we propose a unified method for reconstruction of multi-defect fluorescence microscopy…
New method restores source features for SFDA without source data.
In this paper, we propose a new framework to remove parts of the systematic errors affecting popular restoration algorithms, with a special focus for image processing tasks. Generalizing ideas that emerged for regularization, we develop an approach re-fitting the results of standard methods towards the input d…
MicAugment transfers audio style from few seconds of input to match target conditions.
Study improves animal audio classification using data augmentation.
Detects audio adversarial examples using anomalous pattern detection.
Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities, and the recent success of convolutional neural networks (CNNs) in SE, we propose an audio-visual de…
Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities, and the recent success of convolutional neural networks (CNNs) in SE, we propose an audio-visual de…
Unified model for audio control and style transfer.
Audio fingerprinting, also named as audio hashing, has been well-known as a powerful technique to perform audio identification and synchronization. It basically involves two major steps: fingerprint (voice pattern) design and matching search. While the first step concerns the derivation of a robust and compact audio si…
Transformer model estimates keywords for better audio captioning.
LVTINO improves high-definition video restoration with consistent temporal details.
We present a time-dependent Langevin description of dynamics of stock prices. Based on a simple sliding-window algorithm, the fluctuation of stock prices is discussed in the view of a time-dependent linear restoring force which is the linear approximation of the drift parameter in Langevin equation estimated from the f…
This paper describes Task 2 of the DCASE 2018 Challenge, titled "General-purpose audio tagging of Freesound content with AudioSet labels". This task was hosted on the Kaggle platform as "Freesound General-Purpose Audio Tagging Challenge". The goal of the task is to build an audio tagging system that can recognize the c…
AV-CPL uses continuous pseudo-labels for AVSR combining labeled and unlabeled data.
Time-frequency (TF) representations provide powerful and intuitive features for the analysis of time series such as audio. But still, generative modeling of audio in the TF domain is a subtle matter. Consequently, neural audio synthesis widely relies on directly modeling the waveform and previous attempts at unconditio…
HEAR benchmark evaluates audio representations for diverse tasks.
A new transformer model corrects diacritics and typos in multiple languages.
Proposes COALA method for learning audio representations aligned with tags.
This paper shows the susceptibility of spectrogram-based audio classifiers to adversarial attacks and the transferability of such attacks to audio waveforms. Some commonly used adversarial attacks to images have been applied to Mel-frequency and short-time Fourier transform spectrograms, and such perturbed spectrograms…
A new method detects hallucinations in medical image restoration using Fourier Ring Correlation.
Mic2Mic reduces microphone variability for speech systems.
As the technology is advancing, audio recognition in machine learning is improved as well. Research in audio recognition has traditionally focused on speech. Living creatures (especially the small ones) are part of the whole ecosystem, monitoring as well as maintaining them are important tasks. Species such as animals …
FSD50K provides an open dataset of over 51k audio clips for sound event recognition.