Automatic emotion recognition (AER) is a challenging task due to the abstract concept and multiple expressions of emotion. Although there is no consensus on a definition, human emotional states usually can be apperceived by auditory and visual systems. Inspired by this cognitive process in human beings, it's natural to…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A fusion approach combines audio and video features for emotion recognition.
Paper presents an audiovisual model to recognize sounds from weakly labeled video data.
We propose a tri-modal architecture to predict Big Five personality trait scores from video clips with different channels for audio, text, and video data. For each channel, stacked Convolutional Neural Networks are employed. The channels are fused both on decision-level and by concatenating their respective fully conne…
LumièreNet creates lecture videos from audio narration.
Inspired by brain's modality fusion, this paper detects active speakers from audio and video.
Spatial audio is an essential medium to audiences for 3D visual and auditory experience. However, the recording devices and techniques are expensive or inaccessible to the general public. In this work, we propose a self-supervised audio spatialization network that can generate spatial audio given the corresponding vide…
Paper proposes new principles and framework for AVC learning from user-generated videos.
Paper proposes self-supervised method for accurate speaker diarization.
Paper proposes ICCN to learn correlations between text, audio, and video for multimodal sentiment analysis.
A new VAD method uses respiration patterns from video to detect speech.
Lipper synthesizes speech from silent videos, improving over single-view methods.
Polynomial fusion layer improves speech-driven facial animation.
Generates music with video emotion using deep neural networks.
This paper describes audEERING's submissions as well as additional evaluations for the One-Minute-Gradual (OMG) emotion recognition challenge. We provide the results for audio and video processing on subject (in)dependent evaluations. On the provided Development set, we achieved 0.343 Concordance Correlation Coefficien…
Convolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio. We use various CNN architectures to classify the soundtracks of a dataset of 70M training videos (5.24 million hours) with 30,871 video-level labels. We examine fully connected Deep Neural Networks (DNNs)…
Large-scale datasets have played a significant role in progress of neural network and deep learning areas. YouTube-8M is such a benchmark dataset for general multi-label video classification. It was created from over 7 million YouTube videos (450,000 hours of video) and includes video labels from a vocabulary of 4716 c…
Novel fusion of autoencoders predicts sleepiness from speech.
AI misidentifies facial expressions in videos, often misinterpreting happiness as sadness.
FSD50K provides an open dataset of over 51k audio clips for sound event recognition.
Paper improves video categorization using temporal coherence.
We consider the task of multimodal music mood prediction based on the audio signal and the lyrics of a track. We reproduce the implementation of traditional feature engineering based approaches and propose a new model based on deep learning. We compare the performance of both approaches on a database containing 18,000 …
Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy that goes beyond simple feature concatenation and learns to automatically align…
Two-stream model recognizes affect from audio and video.
Wearables like smartwatches which are embedded with sensors and powerful processors, provide a strong platform for development of analytics solutions in sports domain. To analyze players' games, while motion sensor based shot detection has been extensively studied in sports like Tennis, Golf, Baseball; Table Tennis and…
In the past, Acoustic Scene Classification systems have been based on hand crafting audio features that are input to a classifier. Nowadays, the common trend is to adopt data driven techniques, e.g., deep learning, where audio representations are learned from data. In this paper, we propose a system that consists of a …
Study improves animal audio classification using data augmentation.
AV-CPL uses continuous pseudo-labels for AVSR combining labeled and unlabeled data.
New method fuses audio and magnetic data to identify underlying subspaces.
This paper improves sentiment classification by combining text, audio, and video data using DCCA.
ESE-FN improves elderly activity recognition accuracy.
In this paper we address the problem of tracking multiple speakers via the fusion of visual and auditory information. We propose to exploit the complementary nature of these two modalities in order to accurately estimate smooth trajectories of the tracked persons, to deal with the partial or total absence of one of the…
Improved animated faces using audiovisual and modality dropout.
Deep learning predicts mental disorders from audio and text samples.
AECF improves multimodal inference robustness and calibration.
Curiosity enhanced by audio-visual associations improves learning efficiency.
Improves video search by balancing text and visual modalities.
This short paper describes our solution to the 2018 IEEE World Congress on Computational Intelligence One-Minute Gradual-Emotional Behavior Challenge, whose goal was to estimate continuous arousal and valence values from short videos. We designed four base regression models using visual and audio features, and then use…
Improved audio event recognition using audiovisual transformers.
Enhances speech emotion recognition by integrating visual data with attention mechanisms.
Generative AI tasks analyzed for text, images, audio, video, code, and molecules.
Human annotations serve an important role in computational models where the target constructs under study are hidden, such as dimensions of affect. This is especially relevant in machine learning, where subjective labels derived from related observable signals (e.g., audio, video, text) are needed to support model trai…
AV-ASR system improves speech recognition with visual context.
This paper addresses the problem of localizing audio sources using binaural measurements. We propose a supervised formulation that simultaneously localizes multiple sources at different locations. The approach is intrinsically efficient because, contrary to prior work, it relies neither on source separation, nor on mon…
Speech-driven visual speech synthesis involves mapping features extracted from acoustic speech to the corresponding lip animation controls for a face model. This mapping can take many forms, but a powerful approach is to use deep neural networks (DNNs). However, a limitation is the lack of synchronized audio, video, an…
Dance Dance Revolution (DDR) is a popular rhythm-based video game. Players perform steps on a dance platform in synchronization with music as directed by on-screen step charts. While many step charts are available in standardized packs, players may grow tired of existing charts, or wish to dance to a song for which no …
MEx dataset benchmarks HAR and multi-modal fusion for exercise quality.
With the increasing popularity of video sharing websites such as YouTube and Facebook, multimodal sentiment analysis has received increasing attention from the scientific community. Contrary to previous works in multimodal sentiment analysis which focus on holistic information in speech segments such as bag of words re…