New method uses overcomplete frames for better acoustic scene analysis.
problem Improving acoustic scene analysis in real-world applications.
method Risk minimization-based overcomplete frame thresholding.
result Validated on bird activity detection task using wavelets.
New task AQA tackles acoustic reasoning from sound scenes.
problem Promote research in acoustic reasoning.
method Generate acoustic scenes from elementary sounds and formulate questions.
result Preliminary results with models FiLM and MAC show promise.
Enhances sound texture in CNN for better acoustic scene classification.
problem Limited understanding of how CNNs perceive audio scenes.
method Used Class Activation Mapping (CAM) to analyze log-Mel features and proposed edge enhancement using DoG and Sobel operators.
result Edge-enhanced log-Mel features improve CNN performance in acoustic scene classification.
This paper introduces a model of environmental acoustic scenes which adopts a morphological approach by ab-stracting temporal structures of acoustic scenes. To demonstrate its potential, this model is employed to evaluate the performance of a large set of acoustic events detection systems. This model allows us to expli…
Paper proposes a voting method to improve acoustic scene classification.
problem Improving acoustic scene classification accuracy.
method Punishment voting algorithm based on super categories construction.
result Punishment voting significantly improves classification performance.
A multi-head attention network improves ASC by recognizing overlapping sound patterns.
problem Challenging ASC due to overlapping sound patterns and complex event mixtures.
method Proposes a multi-head attention network to model complex temporal input structures.
result Achieved competitive performance on DCASE 2018 Task 5 dataset.
Improved acoustic scene classification with factorized CNN.
problem Acoustic scene classification in varying environments.
method Large-margin factorized CNN with triplet loss.
result Improved performance and better generalization on unseen data.
Review of acoustic scene classification methods in a competition.
problem Categorizing audio sequences into classes based on spectral content.
method Competition involving students and external participants, ablation study, neural network baseline comparison.
result Improved classification over neural network baseline.
A simple fusion of deep and shallow learning improves acoustic scene classification.
problem Improving acoustic scene classification accuracy.
method Combining a deep learning approach and a feature engineering approach using a late fusion strategy.
result The fused system achieves 72.8% classification accuracy, outperforming individual methods.
A new neural network learns from acoustic scenes by suppressing irrelevant patterns.
problem Acoustic scenes are rich and redundant, making classification challenging.
method Spatio-temporal attention pooling layer coupled with a convolutional recurrent neural network.
result The method outperforms a strong convolutional neural network baseline and sets new state-of-the-art performance.
This paper analyzes sound event detection in synthetic office audio, comparing different systems.
problem Comparing sound event detection systems in synthetic office audio.
method Analysis of systems submitted to DCASE 2016 task, using synthetic office sounds.
result Statistical analysis of results, highlighting system performance under controlled conditions.
CNNs improve generalization to unseen audio devices with increased width, not depth.
problem CNNs are sensitive to specific audio recording devices in acoustic scene classification.
method Investigated the relationship between over-parameterization and generalization in CNNs for audio classification.
result Increasing width improves generalization to unseen devices without increasing the number of parameters.
Transformer model estimates keywords for better audio captioning.
problem Indeterminacy in word selection for audio events/scenes.
method Transformer-based model with keyword estimation.
result Achieved state-of-the-art performance in AAC.
Study improves CNNs for audio scene classification by restricting receptive fields and adding frequency awareness.
problem Improving CNNs for robust acoustic scene classification.
method Investigated different receptive field configurations for various CNN architectures and introduced Frequency Aware CNNs.
result Several well-performing submissions to DCASE 2019 Challenge were achieved.
CLEAR dataset for acoustic reasoning tasks.
problem Acoustic reasoning and question answering.
method Data generation from elementary sounds, functional programs for question composition.
result Validation of current state-of-the-art visual Q&A models on AQA task.
A lightweight network and NAS method improve ASC tasks.
problem Heavy computational burden in acoustic scene classification.
method Inspired by MobileNetV2, unidirectional convolutions; dynamic NAS with evolutionary algorithm.
result 90.3% F1-score on DCASE2018 task 5, 25% fewer FLOPs.
Improved deep CNNs for ASC by optimizing receptive field size.
problem Deep CNNs perform poorly in ASC compared to simpler models.
method Analyzed and adapted the receptive field of ResNet and DenseNet.
result State-of-the-art performance achieved with optimized receptive field.
Enhances ASC using time- and frequency-liked CNNs and bilinear pooling.
problem Improving acoustic scene classification accuracy.
method Harmonic and percussive source separation, two-stream CNN architecture, bilinear pooling.
result Improved accuracy on DCASE 2019 sub task 1a dataset.
End-to-end DA method for domain-invariant CNNs using parallel audio recordings.
problem Distribution mismatches between training and application data in machine listening.
method Enforcing equal hidden layer representations for domain-parallel samples.
result Learn domain-invariant classifiers without requiring classification labels.
Proposes deep learning method for GCI detection from pathological speech.
problem Detecting glottal closure instants (GCI) in pathological acoustic speech.
method Convolutional neural network with fused deep acoustic speech and linear prediction residual features.
result Significantly better than state-of-the-art methods in GCI detection.
Paper proposes cost-sensitive detection for environmental acoustic sensing.
problem Infeasibility of manual analysis for large-scale acoustic data.
method Cost-sensitive classification with variational autoencoders in Neyman-Pearson framework.
result Improved control over false positive and false negative rates.
System tackles indeterminacies in automated audio captioning.
problem Word selection and sentence length indeterminacies in automated audio captioning.
method Solves caption generation and sub-indeterminacy problems through multi-task learning to estimate keywords and sentence length.
result Model achieved 20.7 SPIDEr score, significantly outperforming baseline.
Improved performance in classifying domestic activities.
problem Classifying domestic activities effectively.
method Ensemble learning system based on CNN and LSTM.
result Significant improvement in F1-score (92.19% vs baseline 84.49%).
Meta-learning improves few-shot acoustic event detection.
problem Detecting new audio events with limited labeled data.
method Formulated few-shot AED problem; explored supervised and meta-learning approaches.
result Meta-learning achieves superior performance in few-shot AED.
This paper improves speech recognition by using raw waveform signals in multi-span CNN acoustic models.
problem Improving speech recognition accuracy using raw waveform signals.
method Proposes a novel multi-span structure for acoustic modelling based on raw waveform signals with multiple CNN input layers.
result Multi-span acoustic models yield a lower word error rate (WER) than traditional FBANK feature-based models.
SCALOR learns scalable object representations for crowded scenes.
problem Scalability in scenes with many objects.
method Spatially-parallel attention and proposal-rejection mechanisms.
result SCALOR can handle up to a hundred objects in crowded scenes.
Improved ASR for English-isiZulu code-switched speech with semi-supervised training.
problem Improving ASR for code-switched speech between English and isiZulu.
method Semi-supervised training using automatic transcription of multilingual speech data.
result Semi-supervised training achieved significant WER reduction in ASR performance.
Paper tackles invariance of demodulation in shallow water acoustic communications.
problem Frequency-selective signal distortion (Doppler effect) in shallow water environments.
method Developed ML-based demodulation methods using DBN-NN and DBN-CNN.
result Demonstrated invariance of the proposed method to Doppler effect with 2dB error margin.
Paper predicts EEG features from acoustic features using RNN and GAN.
problem Predicting EEG features from acoustic features.
method Recurrent Neural Network (RNN) and Generative Adversarial Network (GAN).
result Lower RMSE and normalized RMSE values compared to generating acoustic features from EEG features.
Paper aims to find joint representation between vocal tract geometry and speech sound acoustics.
problem Finding a joint latent representation between articulatory and acoustic domains for vowel sounds.
method Invertible neural network models, convolutional autoencoder, normalizing flows, semi-supervised learning.
result Satisfactory performance in articulatory-to-acoustic and acoustic-to-articulatory mapping.
A neural scene representation framework enforcing 3D transformations.
problem Learning 3D scene representations from images without 3D supervision.
method Introducing a loss enforcing equivariance of the scene representation with 3D transformations.
result Real-time neural rendering with comparable results to models requiring minutes for inference.
Study shows integrating acoustic features in financial forecasting models can degrade performance.
problem Predicting stock market volatility from corporate earnings calls using speech features.
method Empirical investigation of acoustic feature extraction in teleconference environments using a two-stream late-fusion architecture.
result Integrating acoustic features via late fusion significantly degraded performance, reducing recall to 47.08%.
Bayesian SHMM discovers acoustic units from unlabeled speech.
problem Discovering language-specific acoustic units from unlabeled speech.
method Bayesian Subspace Hidden Markov Model (SHMM) trained on labeled data to find new acoustic units on target language.
result Significantly outperforms previous HMM-based systems and compares favorably with Variational Auto Encoder-HMM.
Combines multiple data types to predict emotions in images.
problem Predicting emotions in images using various data types.
method Combines facial features, scene extraction, audio tonality, human pose, text-based tagging, and CNN predictions.
result Improves accuracy in emotion prediction compared to baseline methods.
Polarimetric images enhance object detection in adverse weather conditions.
problem Object detection in road scenes is challenging in adverse weather conditions.
method Combining polarimetric imaging and deep learning.
result Polarimetry improves object detection by 20% to 50% compared to conventional RGB images.
This paper compares new speech synthesis methods and finds Wavenet vocoders and AR models perform best.
problem Improving speech synthesis quality using advanced machine learning techniques.
method Large-scale crowdsourced evaluation of vocoding and acoustic modeling techniques.
result Wavenet vocoders and AR models outperform conventional methods in speech synthesis quality.
New method estimates animal density using acoustic data, accounting for unknown call identities.
problem Estimating animal density or call density from acoustic data with unknown call identities.
method Monte Carlo Expectation-Maximization (MCEM) method to resolve unknown call identities.
result Estimates are within 15% of expert-constructed estimates and incorporate uncertainty about call identities.
GENESIS generates and samples 3D scenes by capturing object interactions.
problem Lack of models that explicitly capture object interactions in scene generation.
method Object-centric latent variables, spatial GMM, amortized inference, autoregressive prior.
result First object-centric generative model of 3D visual scenes.
We present a supervised neural network model for polyphonic piano music transcription. The architecture of the proposed model is analogous to speech recognition systems and comprises an acoustic model and a music language model. The acoustic model is a neural network used for estimating the probabilities of pitches in …
Efficient model for foggy scene understanding in vehicles.
problem Challenging scene understanding and segmentation under foggy conditions.
method Domain adaptation and illumination-invariant image transformation.
result Outperforms state-of-the-art models in foggy scene understanding.
New framework segments 3D scenes using neural algorithms and sub-Riemannian geometry.
problem Effective scene segmentation in 3D vision.
method Neurogeometric sub-Riemannian model, harmonic analysis, neural-based stereo correspondence.
result Sub-Riemannian metric is central to effective scene segmentation.
ConvNet classifies whale vocalizations and ambient noise in acoustic recordings.
problem Automated detection and classification of marine mammal vocalizations in acoustic recordings.
method Convolutional Neural Network with a novel acoustic representation.
result Classifier accurately detects and classifies whale vocalizations and ambient noise.
Scene text magnifier enhances readability for visually impaired.
problem Helps visually impaired read natural scene text.
method Four CNN-based networks: character erasing, extraction, magnify, synthesis.
result Effective text magnification without background alteration.
Improved multi-speaker TTS using GANs and waveform loss.
problem Training acoustic models for neural vocoders in multi-speaker TTS systems.
method Proposed frameworks incorporating Wasserstein GAN with gradient penalty (WGAN-GP) and discretized mixture logistic loss (DML) into acoustic models trained with WaveNet.
result Acoustic models trained with WGAN-GP and DML loss achieve highest subjective evaluation scores in multi-speaker TTS.
Improved acoustic word embeddings using shared decoder in multi-view encoders.
problem Learning discriminative acoustic word embeddings from text labels.
method Combining Siamese multi-view encoders with a shared decoder network to maximize the relationship between acoustic and text embeddings.
result 11.1% relative improvement in average precision on acoustic word discrimination task with WSJ dataset.
Self-supervised method detects replay spoofing using acoustic configurations.
problem Challenges in collecting large-scale datasets for replay spoofing detection.
method Self-supervised pretraining of acoustic configurations using existing datasets.
result The method outperforms baseline by 30% on ASVspoof 2019 physical access dataset.
Method converts facial expressions and voice of a source speaker into a target speaker.
problem Separate conversion of facial and acoustic features leads to unnatural results.
method Uses three neural networks: conversion, waveform generation, and image reconstruction.
result Significantly higher naturalness achieved when converting both features together.
New model separates objects in scenes, enabling novel arrangements and depth.
problem Lack of modular, compositional scene modeling in generative models.
method Ensemble of generative models (experts) compete for explaining different parts of a scene.
result Model generates scenes with novel object arrangement and depth ordering.