Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

6.3%12.5%18.8%25.0% · Oct 199319922001200920182026
48 results for Acoustic Scene Analysis

Enhances sound texture in CNN for better acoustic scene classification.

problem Limited understanding of how CNNs perceive audio scenes.
method Used Class Activation Mapping (CAM) to analyze log-Mel features and proposed edge enhancement using DoG and Sobel operators.
result Edge-enhanced log-Mel features improve CNN performance in acoustic scene classification.

Paper proposes a voting method to improve acoustic scene classification.

problem Improving acoustic scene classification accuracy.
method Punishment voting algorithm based on super categories construction.
result Punishment voting significantly improves classification performance.

A multi-head attention network improves ASC by recognizing overlapping sound patterns.

problem Challenging ASC due to overlapping sound patterns and complex event mixtures.
method Proposes a multi-head attention network to model complex temporal input structures.
result Achieved competitive performance on DCASE 2018 Task 5 dataset.

A simple fusion of deep and shallow learning improves acoustic scene classification.

problem Improving acoustic scene classification accuracy.
method Combining a deep learning approach and a feature engineering approach using a late fusion strategy.
result The fused system achieves 72.8% classification accuracy, outperforming individual methods.

A new neural network learns from acoustic scenes by suppressing irrelevant patterns.

problem Acoustic scenes are rich and redundant, making classification challenging.
method Spatio-temporal attention pooling layer coupled with a convolutional recurrent neural network.
result The method outperforms a strong convolutional neural network baseline and sets new state-of-the-art performance.

This paper analyzes sound event detection in synthetic office audio, comparing different systems.

problem Comparing sound event detection systems in synthetic office audio.
method Analysis of systems submitted to DCASE 2016 task, using synthetic office sounds.
result Statistical analysis of results, highlighting system performance under controlled conditions.

CNNs improve generalization to unseen audio devices with increased width, not depth.

problem CNNs are sensitive to specific audio recording devices in acoustic scene classification.
method Investigated the relationship between over-parameterization and generalization in CNNs for audio classification.
result Increasing width improves generalization to unseen devices without increasing the number of parameters.

Study improves CNNs for audio scene classification by restricting receptive fields and adding frequency awareness.

problem Improving CNNs for robust acoustic scene classification.
method Investigated different receptive field configurations for various CNN architectures and introduced Frequency Aware CNNs.
result Several well-performing submissions to DCASE 2019 Challenge were achieved.

Enhances ASC using time- and frequency-liked CNNs and bilinear pooling.

problem Improving acoustic scene classification accuracy.
method Harmonic and percussive source separation, two-stream CNN architecture, bilinear pooling.
result Improved accuracy on DCASE 2019 sub task 1a dataset.

End-to-end DA method for domain-invariant CNNs using parallel audio recordings.

problem Distribution mismatches between training and application data in machine listening.
method Enforcing equal hidden layer representations for domain-parallel samples.
result Learn domain-invariant classifiers without requiring classification labels.

Proposes deep learning method for GCI detection from pathological speech.

problem Detecting glottal closure instants (GCI) in pathological acoustic speech.
method Convolutional neural network with fused deep acoustic speech and linear prediction residual features.
result Significantly better than state-of-the-art methods in GCI detection.

Paper proposes cost-sensitive detection for environmental acoustic sensing.

problem Infeasibility of manual analysis for large-scale acoustic data.
method Cost-sensitive classification with variational autoencoders in Neyman-Pearson framework.
result Improved control over false positive and false negative rates.

System tackles indeterminacies in automated audio captioning.

problem Word selection and sentence length indeterminacies in automated audio captioning.
method Solves caption generation and sub-indeterminacy problems through multi-task learning to estimate keywords and sentence length.
result Model achieved 20.7 SPIDEr score, significantly outperforming baseline.

This paper improves speech recognition by using raw waveform signals in multi-span CNN acoustic models.

problem Improving speech recognition accuracy using raw waveform signals.
method Proposes a novel multi-span structure for acoustic modelling based on raw waveform signals with multiple CNN input layers.
result Multi-span acoustic models yield a lower word error rate (WER) than traditional FBANK feature-based models.

Improved ASR for English-isiZulu code-switched speech with semi-supervised training.

problem Improving ASR for code-switched speech between English and isiZulu.
method Semi-supervised training using automatic transcription of multilingual speech data.
result Semi-supervised training achieved significant WER reduction in ASR performance.

Paper tackles invariance of demodulation in shallow water acoustic communications.

problem Frequency-selective signal distortion (Doppler effect) in shallow water environments.
method Developed ML-based demodulation methods using DBN-NN and DBN-CNN.
result Demonstrated invariance of the proposed method to Doppler effect with 2dB error margin.

Paper aims to find joint representation between vocal tract geometry and speech sound acoustics.

problem Finding a joint latent representation between articulatory and acoustic domains for vowel sounds.
method Invertible neural network models, convolutional autoencoder, normalizing flows, semi-supervised learning.
result Satisfactory performance in articulatory-to-acoustic and acoustic-to-articulatory mapping.

A neural scene representation framework enforcing 3D transformations.

problem Learning 3D scene representations from images without 3D supervision.
method Introducing a loss enforcing equivariance of the scene representation with 3D transformations.
result Real-time neural rendering with comparable results to models requiring minutes for inference.

Study shows integrating acoustic features in financial forecasting models can degrade performance.

problem Predicting stock market volatility from corporate earnings calls using speech features.
method Empirical investigation of acoustic feature extraction in teleconference environments using a two-stream late-fusion architecture.
result Integrating acoustic features via late fusion significantly degraded performance, reducing recall to 47.08%.

Bayesian SHMM discovers acoustic units from unlabeled speech.

problem Discovering language-specific acoustic units from unlabeled speech.
method Bayesian Subspace Hidden Markov Model (SHMM) trained on labeled data to find new acoustic units on target language.
result Significantly outperforms previous HMM-based systems and compares favorably with Variational Auto Encoder-HMM.

Polarimetric images enhance object detection in adverse weather conditions.

problem Object detection in road scenes is challenging in adverse weather conditions.
method Combining polarimetric imaging and deep learning.
result Polarimetry improves object detection by 20% to 50% compared to conventional RGB images.

This paper compares new speech synthesis methods and finds Wavenet vocoders and AR models perform best.

problem Improving speech synthesis quality using advanced machine learning techniques.
method Large-scale crowdsourced evaluation of vocoding and acoustic modeling techniques.
result Wavenet vocoders and AR models outperform conventional methods in speech synthesis quality.

New method estimates animal density using acoustic data, accounting for unknown call identities.

problem Estimating animal density or call density from acoustic data with unknown call identities.
method Monte Carlo Expectation-Maximization (MCEM) method to resolve unknown call identities.
result Estimates are within 15% of expert-constructed estimates and incorporate uncertainty about call identities.

GENESIS generates and samples 3D scenes by capturing object interactions.

problem Lack of models that explicitly capture object interactions in scene generation.
method Object-centric latent variables, spatial GMM, amortized inference, autoregressive prior.
result First object-centric generative model of 3D visual scenes.

We present a supervised neural network model for polyphonic piano music transcription. The architecture of the proposed model is analogous to speech recognition systems and comprises an acoustic model and a music language model. The acoustic model is a neural network used for estimating the probabilities of pitches in …

2015-08-07abs ↗pdf ↗

Efficient model for foggy scene understanding in vehicles.

problem Challenging scene understanding and segmentation under foggy conditions.
method Domain adaptation and illumination-invariant image transformation.
result Outperforms state-of-the-art models in foggy scene understanding.

ConvNet classifies whale vocalizations and ambient noise in acoustic recordings.

problem Automated detection and classification of marine mammal vocalizations in acoustic recordings.
method Convolutional Neural Network with a novel acoustic representation.
result Classifier accurately detects and classifies whale vocalizations and ambient noise.

Improved multi-speaker TTS using GANs and waveform loss.

problem Training acoustic models for neural vocoders in multi-speaker TTS systems.
method Proposed frameworks incorporating Wasserstein GAN with gradient penalty (WGAN-GP) and discretized mixture logistic loss (DML) into acoustic models trained with WaveNet.
result Acoustic models trained with WGAN-GP and DML loss achieve highest subjective evaluation scores in multi-speaker TTS.

Improved acoustic word embeddings using shared decoder in multi-view encoders.

problem Learning discriminative acoustic word embeddings from text labels.
method Combining Siamese multi-view encoders with a shared decoder network to maximize the relationship between acoustic and text embeddings.
result 11.1% relative improvement in average precision on acoustic word discrimination task with WSJ dataset.

Self-supervised method detects replay spoofing using acoustic configurations.

problem Challenges in collecting large-scale datasets for replay spoofing detection.
method Self-supervised pretraining of acoustic configurations using existing datasets.
result The method outperforms baseline by 30% on ASVspoof 2019 physical access dataset.

Method converts facial expressions and voice of a source speaker into a target speaker.

problem Separate conversion of facial and acoustic features leads to unnatural results.
method Uses three neural networks: conversion, waveform generation, and image reconstruction.
result Significantly higher naturalness achieved when converting both features together.