Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

35810 · Dec 201919922001200920172026
48 results for amplitude spectrogram

A typical audio signal processing pipeline includes multiple disjoint analysis stages, including calculation of a time-frequency representation followed by spectrogram-based feature analysis. We show how time-frequency analysis and nonnegative matrix factorisation can be jointly formulated as a spectral mixture Gaussia…

2019-01-31abs ↗pdf ↗

iSTFTNet speeds up mel-spectrogram vocoders without sacrificing quality.

problem Efficiently converting mel-spectrograms to speech with minimal computation.
method Replaces convolutional layers with iSTFT after frequency dimension reduction.
result Significant reduction in computational cost with comparable quality.

CycleGAN-VC3 improves CycleGAN-VCs for mel-spectrogram conversion.

problem Ambiguity in CycleGAN-VC/VC2 effectiveness for mel-spectrogram conversion.
method Proposes CycleGAN-VC3 with time-frequency adaptive normalization (TFAN).
result CycleGAN-VC3 outperforms or matches CycleGAN-VC2 for mel-spectrogram conversion.

This paper proposes a multichannel source separation technique called the multichannel variational autoencoder (MVAE) method, which uses a conditional VAE (CVAE) to model and estimate the power spectrograms of the sources in a mixture. By training the CVAE using the spectrograms of training examples with source-class l…

2018-08-02abs ↗pdf ↗

For training the sequence-to-sequence voice conversion model, we need to handle an issue of insufficient data about the number of speech pairs which consist of the same utterance. This study experimentally investigated the effects of Mel-spectrogram augmentation on training the sequence-to-sequence voice conversion (VC…

2020-01-06abs ↗pdf ↗

We introduce the convolutional spectral kernel (CSK), a novel family of non-stationary, nonparametric covariance kernels for Gaussian process (GP) models, derived from the convolution between two imaginary radial basis functions. We present a principled framework to interpret CSK, as well as other deep probabilistic mo…

2019-05-23abs ↗pdf ↗

In this paper we study deep learning-based music source separation, and explore using an alternative loss to the standard spectrogram pixel-level L2 loss for model training. Our main contribution is in demonstrating that adding a high-level feature loss term, extracted from the spectrograms using a VGG net, can improve…

2019-01-15abs ↗pdf ↗

This paper shows how learning the phase-amplitude coupling improves bio-signal classification.

problem Discarding phase component in bio-signal feature extraction leads to poor generalization.
method Introducing a novel self-supervised learning task called Phase-Swap to detect phase-amplitude coupling.
result Neural networks trained on Phase-Swap task generalize better across subjects and recording sessions.

Paper presents a deep learning framework for classifying respiratory anomalies and lung diseases from sound recordings.

problem Classifying respiratory anomalies and lung diseases from respiratory sound recordings.
method The framework uses front-end feature extraction to transform sound into spectrograms, and a deep learning network to classify these features.
result The proposed deep learning system outperforms current state-of-the-art methods on the ICBHI benchmark dataset.

A fast method for estimating radar amplitude density parameters.

problem Accurate estimation of amplitude density function parameters in radar applications.
method Projecting amplitude data onto horizontal and vertical axes, then using MLE for α\alpha-stale distribution parameters.
result The average of computed MLEs based on two projections is a fast and accurate estimator for amplitude distribution parameters.

Improved formulation of spinfoam quantum gravity with cosmological constant, ensuring all amplitudes are finite and providing semiclassical asymptotics.

problem Ensuring the finiteness of spinfoam amplitudes and providing semiclassical asymptotics for quantum gravity.
method Using state-integral model of PSL(2, C\mathbb{C}) Chern-Simons theory and implementing simplicity constraint.
result All spinfoam amplitudes are finite and provide semiclassical asymptotics with oscillatory terms related to the Regge action.

SpecGrad improves neural vocoder sound quality by adapting diffusion noise to log-mel spectrogram.

problem Improving neural vocoder sound quality, especially in high-frequency bands.
method Adapting the diffusion noise distribution to the conditioning log-mel spectrogram through time-varying filtering.
result SpecGrad generates higher-fidelity speech waveform than conventional DDPM-based neural vocoders.

Atiyah classes of DG manifolds of positive amplitude are invariant under weak equivalences.

problem Defining and studying Hochschild cohomology of DG manifolds of positive amplitude.
method Using poly-differential operators and derived intersection, proving invariance under weak equivalences.
result Hochschild cohomology of DG manifolds of positive amplitude is invariant under weak equivalences.

The paper provides exact multivariate amplitude distributions for non-stationary Gaussian or algebraic fluctuations.

problem Capturing the statistical properties of fluctuating correlations in non-stationary systems.
method Developed a random matrix model to average multivariate amplitude distributions from short time scales to large time scales.
result Explicit multivariate distributions for non-stationary correlation systems are provided, capturing the degree of non-stationarity.

We define a topological quantum membrane theory on a seven dimensional manifold of G2G_2 holonomy. We describe in detail the path integral evaluation for membrane geometries given by circle bundles over Riemann surfaces. We show that when the target space is CY3×S1CY_3\times S^1 quantum amplitudes of non-local observables …

2006-11-29abs ↗pdf ↗

X-DC improves speech separation by making DNNs more interpretable.

problem Black-box nature of DNNs in speech separation tasks.
method Introduces X-DC, a DNN architecture that interprets as spectrogram template fitting followed by Wiener filtering.
result X-DC achieves comparable speech separation performance to DC but with enhanced interpretability.

We demonstrate the equivalence of all loop closed topological string amplitudes on toric local Calabi-Yau threefolds with computations of certain knot invariants for Chern-Simons theory. We use this equivalence to compute the topological string amplitudes in certain cases to very high degree and to all genera. In parti…

2002-06-18abs ↗pdf ↗

PolarBM models complex-valued audio signals in polar coordinates, improving over conventional methods.

problem Discarding structural information in complex-valued problems simplifies models but loses important amplitude-phase relationships.
method Proposes PolarBM, a novel Boltzmann machine for complex-valued variables in polar coordinates, and LogPolarBM for logarithmic amplitude.
result PolarBM and LogPolarBM achieve superior modeling accuracy compared to conventional models, including deep neural networks.

MaskCycleGAN-VC improves voice conversion without parallel data.

problem Limited ability to convert mel-spectrogram data without parallel data.
method Integrates a novel auxiliary task called filling in frames (FIF) to learn time-frequency structures.
result MaskCycleGAN-VC outperforms existing methods with similar model size.

We introduce a fully coherent spin network amplitude whose expansion generates all SU(2) spin networks associated with a given graph. We then give an explicit evaluation of this amplitude for an arbitrary graph. We show how this coherent amplitude can be obtained from the specialization of a generating functional obtai…

2012-01-17abs ↗pdf ↗

CLCNet improves noise reduction in hearing aids with deep learning.

problem Noise reduction in hearing aids is challenging due to real-time and frequency resolution constraints.
method Proposes CLCNet, a deep learning framework based on complex linear coding.
result CLCNet outperforms traditional methods in noisy environments.

We study topological open string amplitudes on orientifolds without fixed planes. We determine the contributions of the untwisted and twisted sectors as well as the BPS structure of the amplitudes. We illustrate our general results in various examples involving D-branes in toric orientifolds. We perform the computation…

2004-11-24abs ↗pdf ↗

AaSP improves audio self-supervised learning by addressing aliasing issues.

problem Alias issues in audio spectrogram transformers.
method AaSP combines aliasing-aware patch representation, teacher-student masked modeling, cross-attention predictor, and contrastive regularization.
result AaSP learns more stable representations that integrate high-frequency cues.

Convolutional neural network (CNN) architectures have originated and revolutionized machine learning for images. In order to take advantage of CNNs in predictive modeling with audio data, standard FFT-based signal processing methods are often applied to convert the raw audio waveforms into an image-like representations…

2019-03-21abs ↗pdf ↗

The paper proves a category of dg manifolds with finite positive amplitude.

problem Understanding the structure of dg manifolds with finite positive amplitude.
method Using path spaces and homotopy transfer theorem for curved L[1]L_\infty[1]-algebras.
result Proves that dg manifolds of finite positive amplitude form a category of fibrant objects.

Algorithm finds frequencies, amplitudes, and phases of sinusoids in noisy data.

problem Finding frequencies, amplitudes, and phases of sinusoids in noisy data.
method Maximum likelihood approach to estimate tone parameters from contaminated observations. Successively estimates frequencies and jointly optimizes amplitudes and phases.
result Near-linear computational complexity (O(N)) for estimating MM number of sinusoidal sources.

End-to-end neural network based approaches to audio modelling are generally outperformed by models trained on high-level data representations. In this paper we present preliminary work that shows the feasibility of training the first layers of a deep convolutional neural network (CNN) model to learn the commonly-used l…

2017-12-01abs ↗pdf ↗

We decompose the exchange rates returns of 41 currencies (incl. gold) into their sign and amplitude components. Then we group together all exchange rates with a common base currency, construct Minimal Spanning Trees for each group independently, and analyze properties of these trees. We show that both the sign and the …

2009-11-16abs ↗pdf ↗

The paper analyzes the amplitude of functions on the sphere, improving FDA methods.

problem Analyzing trajectories on non-linear manifolds with time variability.
method Developed tools for temporal alignment, geodesic computation, and mean calculation on S2\mathbb{S}^2.
result Efficient and accurate tools for analyzing manifold-valued functions on S2\mathbb{S}^2.

Transformers predict scattering amplitudes in theoretical physics.

problem Computing exact coefficients of scattering amplitudes in N = 4 SYM theory.
method Applied Transformers to predict integer coefficients of scattering amplitudes.
result Transformers achieve high (> 98%) accuracy on predicting scattering amplitudes.

The transition amplitudes between coherent states on a coherent state manifold are expressed in terms of the embedding of the coherent state manifold into a projective Hilbert space. Consequences for the dimension of projective Hilbert space and a simple geometric interpretation of Calabi's diastasis follows.

1997-07-31abs ↗pdf ↗