Paper enhances speech by estimating RI spectrograms and optimizing multiple metrics.
problem Difficulty in phase estimation and lack of multi-metric optimization in speech enhancement.
method Proposes a CNN model for RI spectrogram estimation and multi-metrics learning.
result Unified objective function improves speech enhancement metrics.
RI-based variable ranking and selection outperforms lasso in high-dimensional datasets.
problem Challenges in variable selection and model creation with correlated predictors.
method RI measures for feature ranking and selection, including CRI.Z.
result RI-based methods outperform lasso in high-dimensional datasets, especially with correlated predictors.
A machine-learning method speeds up RIS design by predicting reflection coefficients.
problem Extensive full-wave EM simulations are time-consuming for RIS design.
method Combining MLP and dual-port network to develop a fast model.
result The proposed method significantly reduces the time for RIS design.
Machine learning speeds up RIS design for efficient RF components.
problem Designing reconfigurable intelligent surfaces (RIS) for efficient RF components is time-consuming and resource-intensive.
method Machine/deep learning techniques are used to reduce the computational cost and time of RIS inverse design.
result Machine learning techniques significantly reduce the time and computational cost of RIS design.
Study improves communication efficiency in RIS-assisted downlink communication.
problem Improving performance of RIS-aided downlink communication over heterogeneous designs.
method Distributed learning with distributionally robust optimization.
result Our algorithm achieves 50% fewer communication rounds for similar worst-case performance.
Adversarial perturbations and RIS interaction vectors improve covert communication.
problem Covert communication in the presence of RISs.
method Designing RIS interaction vectors to balance receiver and eavesdropper detection, adding adversarial perturbations to signals.
result Adversarial perturbations and RIS interaction vectors can be jointly designed to boost covert communications.
Investigates RI strategies for life insurers with LRD mortality rates.
problem Effect of long-range dependent mortality rates on RI strategies.
method Volterra mortality model, compound Poisson process, open-loop equilibrium mean-variance criterion.
result Explicit equilibrium RI controls derived and uniqueness studied.
New functions derived from arrow diagrams for spherical curves, invariant under certain deformations.
problem Defining and analyzing integer-valued functions on spherical curves.
method Introducing new functions and relators to study spherical curves and their isotopy classes.
result Functions derived from arrow diagrams are invariant under specific deformations.
iSTFTNet speeds up mel-spectrogram vocoders without sacrificing quality.
problem Efficiently converting mel-spectrograms to speech with minimal computation.
method Replaces convolutional layers with iSTFT after frequency dimension reduction.
result Significant reduction in computational cost with comparable quality.
Deep learning model predicts tropical cyclone intensification using satellite images.
problem Accurately predicting rapid intensification of tropical cyclones.
method Attention-based deep learning model using satellite images.
result Deep learning models outperform traditional methods in RI prediction.
Adversarial attacks on spectrograms can fool audio classifiers trained on waveforms.
problem Susceptibility of audio classifiers to adversarial attacks on spectrograms.
method Applying adversarial attacks to spectrograms and reconstructing audio waveforms.
result Perturbed spectrograms can fool 2D CNNs and 1D CNNs trained on audio waveforms.
Spectrogram-Channels U-Net separates sounds by treating each channel as a source's spectrogram.
problem Sound source separation in music information retrieval.
method Adapting U-Net to treat each channel of the output as a source's spectrogram, balancing volumes between sources.
result State-of-the-art performance on singing voice and multi-instrument separation.
Study counts sub-chord diagrams to classify spherical curves.
problem Classifying spherical curves using chord diagrams.
method Counting sub-chord diagrams under specific moves.
result New invariant classifies prime reduced spherical curves.
Federated edge learning improves with CSIT-free model aggregation using RIS.
problem Lack of CSIT in federated edge learning systems.
method Use RIS to align channel coefficients for model aggregation without CSIT, optimize RIS and receiver jointly.
result Achieves similar learning accuracy as CSIT-based methods without CSIT.
The study classifies rational 1-forms on the Riemann sphere with simple poles.
problem Classifying rational 1-forms on the Riemann sphere with specified pole conditions.
method Recognized three equivalent atlases, proved submanifold properties, and used PSL(2,C) action.
result Quotients of isochronous 1-forms admit stratified orbit types.
Improved speech emotion recognition using pitch-synchronous single frequency filtering spectrogram.
problem Uncertainty principles in STFT spectrogram limit time and frequency resolutions.
method Modified SFF spectrogram by averaging amplitudes between GCI locations, named pitch-synchronous SFF spectrogram.
result Improved SER accuracy (63.95% to 70.4%) on IEMOCAP dataset.
Improves off-policy RL stability with RIS.
problem Stability issues in off-policy RL due to distributional mismatch.
method Relative Importance Sampling (RIS) for off-policy actor-critic.
result RIS stabilizes RL learning by reducing variance.
Study improves voice conversion model with Mel-spectrogram augmentation.
problem Insufficient speech pairs data for training sequence-to-sequence voice conversion models.
method Experimented with Mel-spectrogram augmentation using SpecAugment policies and proposed new augmentation policies.
result Time axis warping policies showed better performance in training the voice conversion model.
Generative adversarial network improves signal reconstruction from magnitude spectrograms.
problem Reconstructing a time-domain signal from a magnitude spectrogram.
method Deep neural network and generative adversarial network approach.
result Our method reconstructs signals faster with higher quality than the Griffin-Lim method.
CycleGAN-VC3 improves CycleGAN-VCs for mel-spectrogram conversion.
problem Ambiguity in CycleGAN-VC/VC2 effectiveness for mel-spectrogram conversion.
method Proposes CycleGAN-VC3 with time-frequency adaptive normalization (TFAN).
result CycleGAN-VC3 outperforms or matches CycleGAN-VC2 for mel-spectrogram conversion.
Improved music source separation using spectrogram feature loss.
problem Music source separation quality improvement.
method Added a high-level feature loss term from spectrograms using a VGG net to a deep learning model.
result Improvement in separation quality of drums and vocals from songs.
VoiceFilter separates target speaker from multi-speaker signals.
problem Speech recognition in multi-speaker environments.
method Speaker recognition network and spectrogram masking network trained together.
result Significant reduction in speech recognition WER on multi-speaker signals.
Paper proposes MVAE for semi-blind source separation using CVAE.
problem Semi-blind source separation in multichannel mixtures.
method Multichannel variational autoencoder (MVAE) with conditional VAE (CVAE).
result MVAE outperforms baseline method in separation performance.
New spherical curve deformations solve a conjecture.
problem Solving the Östlund Conjecture for spherical curves.
method Introducing a new type of deformation (β) and proving equivalence under specific deformations.
result Equivalence of spherical curves under specific deformations.
Improved U-Nets with various intermediate blocks enhance singing voice separation.
problem Improving singing voice separation accuracy using U-Net architectures.
method Implemented and compared U-Nets with different intermediate spectrogram transformation blocks.
result A specific block type achieves state-of-the-art SDR by 0.9 dB.
Trading invariance hypothesis is revised with high correlation to trading costs.
problem Revisiting trading invariance hypothesis in metaorders.
method Empirical analysis of a large dataset of metaorders, investigating the quantity I and its correlation with trading costs. result Trading invariance hypothesis is revised; I is not invariant but highly correlated with trading costs. The performance of the Self-Organizing Map (SOM) algorithm is dependent on the initial weights of the map. The different initialization methods can broadly be classified into random and data analysis based initialization approach. In this paper, the performance of random initialization (RI) approach is compared to that…
End-to-end models perform better with learned log-scaled mel-spectrogram features.
problem End-to-end neural network models struggle with performance compared to models using high-level data representations.
method Trained first layers of a CNN model on log-scaled mel-spectrogram transformation and then used these learned features to initialize an end-to-end CNN classifier.
result Convergence and performance on ESC-50 dataset are similar to a model trained on pre-processed log-scaled mel-spectrogram features.
Novel CSK kernel improves GP model generalization for non-stationary patterns.
problem Improving generalization of Gaussian process models for non-stationary data.
method Introduced convolutional spectral kernel (CSK) derived from convolution of imaginary radial basis functions, using Fourier transform for interpretation.
result CSK improves GP model generalization on spatiotemporal datasets.
A new RBM model handles both linear and log-amplitude spectrograms.
problem Handling amplitude spectra with existing models.
method Proposed gamma-Bernoulli RBM that uses gamma distribution.
result The model can naturally handle positive numbers and log-amplitude spectrograms.
Optimal transport improves speech BSS by better aligning spectrogram frequencies.
problem Speech BSS with improved frequency alignment.
method Developed optimal transport NMF for supervised speech BSS.
result Optimal transport NMF leads to better perceptual results than Euclidean NMF.
Arnold introduced invariants J+, J− and St for generic planar curves. It is known that both J+/2+St and J−/2+St are invariants for generic spherical curves. Applying these invariants to underlying curves of knot diagrams, we can obtain lower bounds for the number of Reidemeister moves for uknotting.…
Measure contraction properties MCP(K,N) are synthetic Ricci curvature lower bounds for metric measure spaces which do not necessarily have smooth structures. It is known that if a Riemannian manifold has dimension N, then MCP(K,N) is equivalent to Ricci curvature bounded below by K. On the other hand, it was ob…
Conditional GANs enhance speech in noisy conditions.
problem Improving speech system performance in noisy environments.
method Conditional Generative Adversarial Networks (cGANs) trained on spectrograms.
result cGAN method outperforms classical SE algorithms and is comparable to deep neural networks.
Deep model generates high-quality speech from spectrograms.
problem Speech reconstruction from spectrograms.
method Deep generative model with Gaussian and von Mises distributions for magnitude and phase, variational autoencoder framework.
result Generated speech has high perceptual quality and intelligibility.
New algorithm for signal estimation in noisy matrix models.
problem Signal estimation in rectangular spiked matrix models with rotationally invariant noise.
method Orthogonal Approximate Message Passing (OAMP) algorithm for signal estimation.
result Optimal OAMP algorithm minimizes mean-squared error and achieves Bayes-optimal performance.
WaveGlow generates high-quality speech from spectrograms.
problem Speech synthesis quality and efficiency.
method WaveGlow combines Glow and WaveNet insights, using a single network and cost function for efficient, high-quality audio synthesis.
result WaveGlow produces audio samples at over 500 kHz, matching WaveNet quality.
Paper presents a deep learning framework for classifying respiratory anomalies and lung diseases from sound recordings.
problem Classifying respiratory anomalies and lung diseases from respiratory sound recordings.
method The framework uses front-end feature extraction to transform sound into spectrograms, and a deep learning network to classify these features.
result The proposed deep learning system outperforms current state-of-the-art methods on the ICBHI benchmark dataset.
Study compares new audio representation methods for limited data music retrieval.
problem Improving machine learning for audio data with limited training data.
method Investigated mel-spectrogram and Mel scattering representations, and augmented target loss function.
result All proposed methods outperform standard mel-spectrogram when using limited data.
Graph neural networks improve music genre classification on audio datasets.
problem Difficulty in applying deep learning on spectrograms due to lack of quality data and augmentation.
method Combination of CNN and Graph Neural Networks (GNN) with Siamese Neural Networks.
result Achieved state-of-the-art results on GTZAN and AudioSet datasets.
SpecGrad improves neural vocoder sound quality by adapting diffusion noise to log-mel spectrogram.
problem Improving neural vocoder sound quality, especially in high-frequency bands.
method Adapting the diffusion noise distribution to the conditioning log-mel spectrogram through time-varying filtering.
result SpecGrad generates higher-fidelity speech waveform than conventional DDPM-based neural vocoders.
X-DC improves speech separation by making DNNs more interpretable.
problem Black-box nature of DNNs in speech separation tasks.
method Introduces X-DC, a DNN architecture that interprets as spectrogram template fitting followed by Wiener filtering.
result X-DC achieves comparable speech separation performance to DC but with enhanced interpretability.
We introduce a bicomplex which computes the triple cohomology of Lie--Rinehart algebras. We prove that the triple cohomology is isomorphic to the Rinehart cohomology \cite{Ri} provided the Lie--Rinehart algebra is projective over the corresponding commutative algebra. As an application we construct a canonical class in…
MaskCycleGAN-VC improves voice conversion without parallel data.
problem Limited ability to convert mel-spectrogram data without parallel data.
method Integrates a novel auxiliary task called filling in frames (FIF) to learn time-frequency structures.
result MaskCycleGAN-VC outperforms existing methods with similar model size.
CLCNet improves noise reduction in hearing aids with deep learning.
problem Noise reduction in hearing aids is challenging due to real-time and frequency resolution constraints.
method Proposes CLCNet, a deep learning framework based on complex linear coding.
result CLCNet outperforms traditional methods in noisy environments.
Deep learning detects atrial fibrillation with high accuracy.
problem Detecting atrial fibrillation in ECG signals.
method Extracted deep features from spectrograms using convolutional networks.
result Convolutional network achieved 93.16% classification accuracy.
AaSP improves audio self-supervised learning by addressing aliasing issues.
problem Alias issues in audio spectrogram transformers.
method AaSP combines aliasing-aware patch representation, teacher-student masked modeling, cross-attention predictor, and contrastive regularization.
result AaSP learns more stable representations that integrate high-frequency cues.
iSTFTNet2 improves iSTFTNet's speed and lightness with 1D-2D CNN.
problem Efficiently synthesizing high-fidelity speech.
method Improved iSTFTNet using 1D-2D CNNs for temporal and spectrogram structures.
result iSTFTNet2 is faster and more lightweight with comparable speech quality.