CNN-RNNs detect bird sounds with high accuracy.
problem Automated detection of bird sounds in varied environments.
method Convolutional Recurrent Neural Networks (CNN-RNNs) for feature extraction and dependency capture.
result 88.5% AUC score on unseen data.
Biodiversity monitoring using audio recordings is achievable at a truly global scale via large-scale deployment of inexpensive, unattended recording stations or by large-scale crowdsourcing using recording and species recognition on mobile devices. The ability, however, to reliably identify vocalising animal species is…
Study improves animal audio classification using data augmentation.
problem Improving automated animal audio classification accuracy.
method Exploits different data augmentation techniques for training CNNs.
result Best recognition rates on animal audio classification datasets.
Machine learning identifies species by voice in remote areas.
problem Continuous monitoring of endangered species in remote areas.
method Training machine learning models on audio data to recognize species.
result Machine learning can accurately classify and recognize various species sounds.
New method uses overcomplete frames for better acoustic scene analysis.
problem Improving acoustic scene analysis in real-world applications.
method Risk minimization-based overcomplete frame thresholding.
result Validated on bird activity detection task using wavelets.
System accurately identifies birds in real-world settings.
problem Identifying birds in diverse, realistic environments.
method Trained kNN and SVM classifiers on crowd-sourced audio data.
result Both classifiers perform similarly, with kNN offering flexibility.
Detects audio adversarial examples using anomalous pattern detection.
problem Identifies adversarial audio attacks in deep neural networks.
method Applies anomalous pattern detection in activation space of audio models.
result Can detect adversarial examples with up to 0.98 AUC, no degradation on benign samples.
Paper presents a provably correct algorithm for CNMF under separable conditions.
problem Convolutive nonnegative matrix factorization (CNMF) under separable assumptions.
method Algorithm exploiting NMF model and existing separable NMF algorithms.
result Guaranteed solution in low noise settings, runs in polynomial time.
Unified detection of isolated and overlapping audio events using CNN-RNN.
problem Detecting both isolated and overlapping audio events simultaneously.
method Multi-label multi-task framework based on CNN-RNN, with sequential losses.
result Good generalization on isolated and overlapping audio event detection datasets.
This paper analyzes sound event detection in synthetic office audio, comparing different systems.
problem Comparing sound event detection systems in synthetic office audio.
method Analysis of systems submitted to DCASE 2016 task, using synthetic office sounds.
result Statistical analysis of results, highlighting system performance under controlled conditions.
New method detects conflicts in police body-worn audio.
problem Poor metrics for conflict detection in police interactions.
method Adaptive noise removal, non-speech filtering, and new phrase-based measures.
result Demonstrated effectiveness on LAPD body-worn audio data.
This research aims to develop robust audio spoofing detection methods that work across various spoofing techniques.
problem Detecting audio spoofing attacks in speaker verification systems.
method Examined traditional and machine learned audio features for robust spoofing detection.
result Fused models based on both known and machine learned features achieve comparable performance with an EER of 12.
A method for semi-supervised sound event detection using teacher-student learning.
problem Weakly-labeled data in sound event detection.
method Guided Learning with a teacher model for audio tagging and a student model for boundary detection.
result The method improves boundary detection performance using unlabeled data.
End-to-end ASR error detection using audio-transcript entailment.
problem Detecting transcription errors in ASR systems to prevent error propagation.
method Proposes a novel end-to-end approach using audio-transcript entailment, with acoustic and linguistic encoders.
result Achieves CER of 26.2% on all transcription errors and 23% on medical errors specifically, improving by 12% and 15.4% respectively over a strong baseline.
CNNs perform well on large audio classification datasets.
problem Classifying soundtracks of videos with 30,871 labels.
method Used various CNN architectures (DNN, AlexNet, VGG, Inception, ResNet) and embeddings for audio classification.
result A model using embeddings from CNNs outperforms raw features on the Audio Set AED task.
Method trains deep neural networks on weakly labeled audio data efficiently.
problem Limited training data and lack of temporal labels for audio event detection.
method Multi-instance learning with a new loss function for stacked CNN-RNN.
result Improved performance on low-resource audio datasets.
SpeechYOLO detects and locates speech objects in audio signals.
problem Detecting and localizing speech in audio signals.
method Inspired by YOLO, SpeechYOLO uses a convolutional neural network with a least-mean-squares loss function.
result SpeechYOLO performs well on keyword spotting tasks, including read and spontaneous speech.
Review of deep learning techniques for audio signal processing.
problem Improving audio signal processing using deep learning.
method Analysis of various deep learning models and techniques.
result Advancements in speech, music, and environmental sound processing.
Meta-learning improves few-shot acoustic event detection.
problem Detecting new audio events with limited labeled data.
method Formulated few-shot AED problem; explored supervised and meta-learning approaches.
result Meta-learning achieves superior performance in few-shot AED.
A new VAD method uses respiration patterns from video to detect speech.
problem Improving VAD performance in noisy audio recordings.
method Extract respiration patterns from video, use neural models to detect speech.
result Efficacy demonstrated through experiments on real acoustic environments.
Improved cover song detection with neural networks.
problem Identifying cover songs from original recordings.
method Siamese Convolutional Neural Networks trained on cover song audio clips.
result Mean precision@1 of 65% over mini-batches, significantly outperforming random guessing.
Wearable tech detects table tennis shots with high accuracy.
problem Lack of shot detection in table tennis using wearables.
method Fusion of IMU and audio sensor data for real-time shot detection.
result 95.6% accuracy in shot detection.
A new system detects audio replay attacks with high accuracy.
problem Detecting and preventing audio replay attacks in speaker verification systems.
method Proposes Attentive Filtering Network combining attention-based filtering and ResNet classifier.
result Achieves EER of 8.99% on ASVspoof 2017 Version 2.0 dataset.
Classifiers and beamforming algorithms improved audio surveillance detection accuracy.
problem Detecting surveillance sound events with high accuracy and efficiency.
method Evaluated seven classifiers and two beamforming algorithms; used data augmentation and tested with varying SNR levels.
result SVM and Delay-and-Sum (DaS) combination achieved the highest accuracy (86.0%), but had high computational cost.
System tackles indeterminacies in automated audio captioning.
problem Word selection and sentence length indeterminacies in automated audio captioning.
method Solves caption generation and sub-indeterminacy problems through multi-task learning to estimate keywords and sentence length.
result Model achieved 20.7 SPIDEr score, significantly outperforming baseline.
Transformer model estimates keywords for better audio captioning.
problem Indeterminacy in word selection for audio events/scenes.
method Transformer-based model with keyword estimation.
result Achieved state-of-the-art performance in AAC.
Deep learning model outperforms traditional methods in music mood prediction.
problem Predicting the emotional state of music from audio and lyrics.
method Implemented deep learning model alongside traditional feature engineering methods and compared their performance.
result Deep learning model outperforms traditional methods in arousal detection.
The study improves pitch detection in polyphonic music by learning harmonic priors.
problem Challenges in transcribing polyphonic music due to overlapping harmonics.
method Introduced Gaussian process priors and used variational Bayes for inference.
result Learning priors that fit the frequency content of sound events improves pitch detection.
Defense against ASR attacks using dropout uncertainty.
problem Adversarial attacks on ASR systems.
method Dropout uncertainty in neural networks.
result High detection accuracy across various ASR systems and datasets.
Improved cover detection in music datasets with novel triplet loss.
problem Challenging task of automatically detecting covers in audio datasets.
method Convolutional neural network mapping melodic features to embeddings, training to minimize cover distance and maximize non-cover distance.
result New prototypical triplet loss improves accuracy for large datasets and live songs.
SVM algorithm extracts digits from audio CAPTCHAs.
problem Recognizing audio CAPTCHAs from computer programs.
method Used RastaPLP features and SVM algorithm.
result Successfully extracted digits from audio CAPTCHAs.
Improved cover detection using dominant melody embeddings.
problem Challenging cover detection in large audio databases.
method Neural network architecture for track embeddings, focusing on dominant melody.
result Improved accuracy on small and large datasets, scalable to thousands of tracks.
GACELA fills long gaps in musical audio with a GAN and context conditioning.
problem Restoring long gaps in musical audio with varying complexity and duration.
method Generative adversarial network (GAN) with five parallel discriminators and context conditioning.
result Reduced artifacts in inpaintings from unacceptable to mildly disturbing.
New model detects crying in real-world settings with improved accuracy.
problem Generalization of cry detection models to real-world environments.
method Evaluated machine learning approaches on a novel dataset of real-world infant crying.
result Improved F1 score of 0.613 for crying event recognition in real-world settings.
Improved anger detection in speech using transfer learning from SoundNet.
problem Detecting anger in speech with limited emotion datasets.
method Transfer learning from SoundNet, a multimodal audio classifier trained on video data.
result Improved performance and generalization on various speech emotion datasets.
Study on mother-infant affect communication using audio recordings.
problem Lack of accurate emotional speech databases for real-life settings.
method Used RAVDESS database and trained a Convolutional Neural Nets model.
result Dominant emotions in mother-infant speech were angry and sad.
Study shows current metrics for audio adversarial examples are unreliable for human perception.
problem The reliability of metrics for evaluating audio adversarial examples.
method Analytical framework and human evaluation experiment.
result Current metrics for audio adversarial examples are not reliable for human perception.
Bird sound data collected with unattended microphones for automatic surveys, or mobile devices for citizen science, typically contain multiple simultaneously vocalizing birds of different species. However, few works have considered the multi-label structure in birdsong. We propose to use an ensemble of classifier chain…
Paper improves sound event detection using semi-supervised learning.
problem Weakly labeled sound event detection in polyphonic audio clips.
method Combines tri-training and adversarial learning for semi-supervised learning.
result Significant performance improvement over baseline model.
Crowd-sourced mosquito audio dataset for malaria research.
problem Understanding mosquito locations for malaria reduction.
method Release of a large mosquito audio dataset with labels from contributors.
result Demonstrated the feasibility of training a CNN on mosquito audio data.
AVEC 2019 challenges AI in detecting depression and cross-cultural emotions.
problem Detecting depression and cross-cultural emotions from audiovisual data.
method Comparison of machine learning methods under standardized conditions.
result Baseline system performance on state-of-mind, depression, and cross-cultural tasks.
Researchers create audio attacks to test speech analysis systems.
problem Adversarial attacks on speech analysis models.
method End-to-end scheme generating waveform perturbations.
result Deep neural networks show significant performance drop from adversarial attacks.
First place in ABC 2018: Classify bird gender from GPS trajectories.
problem Predicting the gender of shearwaters from GPS navigation data.
method Ensemble of Gradient Boosting Classifiers (CatBoost, LightGBM, XGBoost) with feature engineering.
result Ranked first among 74 teams in the Animal Behavior Challenge.
Enhances safety of 3D object detection neural networks.
problem Ensuring robustness and safety of 3D object detection systems.
method Symbolic error propagation, specialized loss function, safety-aware non-max-inclusion algorithm.
result Improved safety and robustness of 3D object detection neural networks.
Machine learning detects and diagnoses coughs for respiratory infections.
problem Detecting and diagnosing coughs for respiratory infections.
method Convolutional Neural Networks (CNNs) to detect cough events and diagnose three illnesses.
result Accuracy over 89% for detection and diagnosis.
PointPillars improves object detection speed and accuracy in point clouds.
problem Encoding point clouds for efficient object detection.
method PointPillars uses PointNets to learn pillar representations of point clouds, combined with a lean downstream network.
result PointPillars outperforms previous encoders in both speed and accuracy.
DoPa detects various physical adversarial attacks on CNNs.
problem Vulnerability of CNNs to physical adversarial attacks.
method Interprets CNN's vulnerability, adds self-verification stage.
result Achieves 90% success rate for image attacks and 92% for audio attacks.
Early-bird tickets can be identified early in training, reducing costs.
problem Costly deep network training.
method Low-cost training schemes (early stopping, low-precision) and mask distance.
result Efficient training methods using EB tickets achieve up to 4.7x energy savings.