Paper proposes ICCN to learn correlations between text, audio, and video for multimodal sentiment analysis.
problem Improving multimodal sentiment analysis by learning hidden correlations between text and audio/video features.
method Interaction Canonical Correlation Network (ICCN) using deep canonical correlation analysis (DCCA).
result Empirical results confirm the effectiveness of ICCN in capturing useful information from all three views.
Bayesian Gaussian process model for music audio analysis.
problem Joint estimation of music parameters from audio signals.
method Bayesian Gaussian processes incorporating prior information.
result Improved pitch estimation and missing segment inference.
End-to-end probabilistic inference improves audio signal processing.
problem Efficiently processing large audio signals with varying characteristics.
method Formulated a spectral mixture Gaussian process model with nonstationary priors, enabling infinite-horizon Gaussian process regression.
result The method outperforms standard techniques in processing audio signals with hundreds of thousands of data points.
Challenge aims to recognize music genres from audio.
problem Recognizing music genres from audio recordings.
method Open data challenge with submissions evaluated.
result Results presented from the challenge.
Generative model improves audio synthesis from TF features.
problem Challenges in generating high-quality audio from TF features.
method Used a GAN to generate invertible TF features from short-time Fourier transforms.
result Generated TF features from GAN outperformed direct waveform generation.
Deep learning predicts mental disorders from audio and text samples.
problem Predicting mental disorders from speech samples.
method Multimodal deep learning structure using various pre-trained models for audio and text embeddings, transfer learning, and auxiliary corpora.
result Acceptable accuracy in predicting mental disorders through multimodal analysis.
Efficient sparse GP model improves audio source separation.
problem Sparse Gaussian Process (GP) inference is computationally expensive for long audio frames.
method Used GP regression, spectral mixture kernels, and variational sparse GPs.
result Proposed method outperforms LD-PSDTF, KL-NMF, and IS-NMF.
Researchers create audio attacks to test speech analysis systems.
problem Adversarial attacks on speech analysis models.
method End-to-end scheme generating waveform perturbations.
result Deep neural networks show significant performance drop from adversarial attacks.
AclNet improves audio classification with high accuracy and reduced complexity.
problem Efficiently classifying audio data with high accuracy.
method Proposed AclNet architecture with data augmentation and regularization.
result Achieved state-of-the-art performance on ESC-50 corpus with 85.65% accuracy.
This paper analyzes sound event detection in synthetic office audio, comparing different systems.
problem Comparing sound event detection systems in synthetic office audio.
method Analysis of systems submitted to DCASE 2016 task, using synthetic office sounds.
result Statistical analysis of results, highlighting system performance under controlled conditions.
Two-stream model recognizes affect from audio and video.
problem Human affect recognition in real-world settings.
method Two-stream aural-visual analysis model with separate audio and visual processing.
result Model achieves promising results on Aff-Wild2 database.
A model classifies music genres from MP3 files using metric learning and feature extraction.
problem Classifying music genres from MP3 files efficiently and accurately.
method Metric learning and feature extraction using MFCC and PCA.
result Promising results in classification accuracy compared to baseline algorithms.
Automates deep model selection for audio source separation.
problem Selecting the right deep learning model for audio source separation.
method Confidence measure based on model output quality estimation.
result Confidence-based ensemble approach significantly outperforms random and oracle selections.
Paper presents an audiovisual model to recognize sounds from weakly labeled video data.
problem Sound recognition from weakly labeled video data.
method Audiovisual fusion model with attention mechanism.
result The model achieves a mean Average Precision (mAP) of 46.16 on AudioSet, outperforming state-of-the-art models.
This paper improves sentiment classification by combining text, audio, and video data using DCCA.
problem Improving sentiment classification accuracy using multi-modal data.
method Deep Canonical Correlation Analysis (DCCA) for combining text, audio, and video embeddings.
result One-Step DCCA outperforms current state-of-the-art in multi-modal embedding learning.
CrossFilter tackles noisy labels in audio tagging.
problem Noisy labels in large audio datasets.
method CrossFilter framework using multiple representations and multi-task learning.
result Improves audio tagging performance on FSDKaggle2018 and FSDKaggle2019 datasets.
Study uses neural networks to classify musical instruments from audio samples.
problem Classifying musical instruments from audio samples.
method Trained an Artificial Neural Network on audio samples transformed to frequency domain. Analyzed features in time and frequency domains.
result Accuracy decreased when using only attack features or initial 100 Hz.
Proposes a self-supervised method for generating spatial audio from monaural audio and video.
problem Generating spatial audio from monaural audio and video recordings is challenging and expensive.
method Uses a self-supervised network with an auxiliary classifier to classify video channels and generate spatial audio.
result The proposed method effectively generates spatial audio from monaural audio and video.
Graph neural networks improve music genre classification on audio datasets.
problem Difficulty in applying deep learning on spectrograms due to lack of quality data and augmentation.
method Combination of CNN and Graph Neural Networks (GNN) with Siamese Neural Networks.
result Achieved state-of-the-art results on GTZAN and AudioSet datasets.
Deep audio prior uses neural networks to solve audio problems without data.
problem Challenging audio problems like source separation, editing, and synthesis.
method Randomly-initialized neural network with carefully designed audio prior.
result Superior audio results on Universal-150 benchmark dataset.
The paper analyzes how speech enhancement and recognition can be improved in noisy environments.
problem Improving speech recognition in multi-talker scenarios with limited resources.
method Developed and trained two LSTM-based models for speech enhancement and phone recognition, then studied their joint optimization.
result Joint optimization of speech enhancement and recognition leads to a significant reduction in Phone Error Rate (PER).
Zero-shot audio classification using class label embeddings.
problem Classifying audio without labeled data.
method Bilinear model with audio feature embeddings and class label embeddings.
result Achieved accuracy up to 39.7% for natural audio categories.
A new autoencoder speeds up audio style transfer.
problem Efficiently generating stylized audio from content audio.
method Autoencoder architecture for real-time audio style transfer.
result Single forward pass for stylized audio generation.
Paper extends principal component pursuit to hypercomplex numbers for improved audio data analysis.
problem Improving robust principal component analysis for audio data.
method Extends principal component pursuit to polar n-complex and n-bicomplex numbers, deriving proximity operators for ℓ1- and trace-norm regularizers. result Our approach outperforms tensor robust principal component analysis on audio data.
Machine learning detects frog calls in audio recordings with high accuracy.
problem Estimating frog presence in audio recordings from tropical environments.
method Gaussian mixture model classification of frog calls with 20 cepstral features.
result The algorithm achieved an average weighted error rate of 0.9% for nine species classification.
This paper explores how different audio signal representations affect topological signatures and their predictive power.
problem The impact of different signal representations on topological signatures and their predictive power.
method The study compares three different signal representations (embedding, spectrogram, and spectrogram zeroes) and evaluates their topological signatures for speaker gender, vowel type, and individual prediction.
result Topological signatures from spectrogram zeroes offer the best improvement for gender prediction, and different representations are complementary.
The paper proposes a method to predict audio ad quality using acoustic features.
problem Improving user experience in online music streaming services by ensuring high quality audio advertisements.
method The paper proposes predicting audio ad quality using acoustic features and a proxy metric called Long Click Rate (LCR). A deep learning model is also introduced.
result The proposed deep learning model outperforms other models trained on hand-crafted features for audio ad quality prediction.
Method learns audio embeddings with contextualized tags.
problem Align audio and tags for cross-modal tasks.
method Audio autoencoder, word embeddings, multi-head self-attention, contrastive loss.
result Multi-head self-attention improves audio representations.
Study improves radio show segmentation using audio embeddings.
problem Automated segmentation of radio shows.
method Created audio embeddings from multi-class classification tasks on different datasets, evaluated performance against text-only baseline.
result Audio embeddings from non-speech sound event classification significantly outperformed text-only baseline by 32.3% in F1-measure.
Bayesian model reconstructs time and frequency data robustly.
problem Missing observations and noise in time/frequency data.
method Probabilistic model, Bayesian update, joint reconstruction.
result Effective joint time/frequency reconstruction with missing data.
Task describes tagging audio clips from 41 categories.
problem Tagging audio clips from diverse categories.
method Baseline system using AudioSet labels.
result Demonstrates feasibility of general-purpose audio tagging.
Unified probabilistic models improve audio signal processing efficiency and interpretability.
problem High computational cost and difficulty in interpreting probabilistic models in time-frequency analysis.
method Equivalence to Spectral Mixture Gaussian processes, state space representation, Kalman smoothing, efficient parameter learning.
result Unified models make it easier to interpret and modify model assumptions.
Deep learning model outperforms traditional methods in music mood prediction.
problem Predicting the emotional state of music from audio and lyrics.
method Implemented deep learning model alongside traditional feature engineering methods and compared their performance.
result Deep learning model outperforms traditional methods in arousal detection.
Meta-learning improves few-shot acoustic event detection.
problem Detecting new audio events with limited labeled data.
method Formulated few-shot AED problem; explored supervised and meta-learning approaches.
result Meta-learning achieves superior performance in few-shot AED.
Wearable tech detects table tennis shots with high accuracy.
problem Lack of shot detection in table tennis using wearables.
method Fusion of IMU and audio sensor data for real-time shot detection.
result 95.6% accuracy in shot detection.
AVDCNN combines audio and visual data for better speech enhancement.
problem Improving speech quality by reducing noise in audio signals.
method Proposes an AVDCNN model that integrates audio and visual streams into a unified deep CNN network for end-to-end training.
result AVDCNN outperforms audio-only and conventional SE methods in enhancing speech quality.
AVDCNN combines audio and visual data for better speech enhancement.
problem Improving speech quality in noisy environments.
method Multimodal deep CNNs model that integrates audio and visual streams.
result AVDCNN outperforms audio-only and conventional SE methods.
Acoustic sensors identify vehicles using spectral embedding.
problem Vehicle recognition from roadside audio sensors.
method Extract frequency signatures, apply spectral embedding for dimensionality reduction.
result K-nearest neighbors achieve accurate vehicle identification after dimensionality reduction.
A new VAD method uses respiration patterns from video to detect speech.
problem Improving VAD performance in noisy audio recordings.
method Extract respiration patterns from video, use neural models to detect speech.
result Efficacy demonstrated through experiments on real acoustic environments.
Adversarial attacks on spectrograms can fool audio classifiers trained on waveforms.
problem Susceptibility of audio classifiers to adversarial attacks on spectrograms.
method Applying adversarial attacks to spectrograms and reconstructing audio waveforms.
result Perturbed spectrograms can fool 2D CNNs and 1D CNNs trained on audio waveforms.
Study improves animal audio classification using data augmentation.
problem Improving automated animal audio classification accuracy.
method Exploits different data augmentation techniques for training CNNs.
result Best recognition rates on animal audio classification datasets.
MicAugment transfers audio style from few seconds of input to match target conditions.
problem Audio model robustness to diverse acquisition conditions.
method Identifies and applies transformations learned from target audio to input audio.
result MicAugment significantly improves model robustness in downstream tasks.
Detects audio adversarial examples using anomalous pattern detection.
problem Identifies adversarial audio attacks in deep neural networks.
method Applies anomalous pattern detection in activation space of audio models.
result Can detect adversarial examples with up to 0.98 AUC, no degradation on benign samples.
Wavelet Attribution Method (WAM) improves feature attribution for deep models.
problem Inability of pixel-based heatmaps to capture data structure and variability in feature attribution.
method Wavelet domain for feature attribution, leveraging spatial and scale-localized properties of wavelet coefficients.
result WAM provides quantitatively superior explanations across audio, image, and volume modalities.
Unified model for audio control and style transfer.
problem Explicit control and style transfer in music generation.
method Diffusion autoencoders for semantic feature extraction, disentanglement using adversarial criterion.
result Model generates audio matching timbre targets with specified structure.
Transformer model estimates keywords for better audio captioning.
problem Indeterminacy in word selection for audio events/scenes.
method Transformer-based model with keyword estimation.
result Achieved state-of-the-art performance in AAC.
Audio fingerprinting, also named as audio hashing, has been well-known as a powerful technique to perform audio identification and synchronization. It basically involves two major steps: fingerprint (voice pattern) design and matching search. While the first step concerns the derivation of a robust and compact audio si…
GANSynth uses GANs to efficiently synthesize high-fidelity audio.
problem Efficient and high-fidelity audio synthesis is challenging.
method Model log magnitudes and instantaneous frequencies with GANs.
result GANSynth outperforms WaveNet on automated and human evaluation metrics.