Paper proposes new principles and framework for AVC learning from user-generated videos.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
AV-CPL uses continuous pseudo-labels for AVSR combining labeled and unlabeled data.
Curiosity enhanced by audio-visual associations improves learning efficiency.
In this paper, we analyzed how audio-visual speech enhancement can help to perform the ASR task in a cocktail party scenario. Therefore we considered two simple end-to-end LSTM-based models that perform single-channel audio-visual speech enhancement and phone recognition respectively. Then, we studied how the two model…
Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities, and the recent success of convolutional neural networks (CNNs) in SE, we propose an audio-visual de…
Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities, and the recent success of convolutional neural networks (CNNs) in SE, we propose an audio-visual de…
In this paper we address the problem of tracking multiple speakers via the fusion of visual and auditory information. We propose to exploit the complementary nature of these two modalities in order to accurately estimate smooth trajectories of the tracked persons, to deal with the partial or total absence of one of the…
Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy that goes beyond simple feature concatenation and learns to automatically align…
Data clustering has received a lot of attention and numerous methods, algorithms and software packages are available. Among these techniques, parametric finite-mixture models play a central role due to their interesting mathematical properties and to the existence of maximum-likelihood estimators based on expectation-m…
Speech-driven visual speech synthesis involves mapping features extracted from acoustic speech to the corresponding lip animation controls for a face model. This mapping can take many forms, but a powerful approach is to use deep neural networks (DNNs). However, a limitation is the lack of synchronized audio, video, an…
Traditional multi-view learning approaches suffer in the presence of view disagreement,i.e., when samples in each view do not belong to the same class due to view corruption, occlusion or other noise processes. In this paper we present a multi-view learning approach that uses a conditional entropy criterion to detect v…
We address the problems of multi-domain and single-domain regression based on distinct and unpaired labeled training sets for each of the domains and a large unlabeled training set from all domains. We formulate these problems as a Bayesian estimation with partial knowledge of statistical relations. We propose a worst-…
We propose a tri-modal architecture to predict Big Five personality trait scores from video clips with different channels for audio, text, and video data. For each channel, stacked Convolutional Neural Networks are employed. The channels are fused both on decision-level and by concatenating their respective fully conne…
Two-stream model recognizes affect from audio and video.
Paper proposes self-supervised method for accurate speaker diarization.
It has been suggested in developmental psychology literature that the communication of affect between mothers and their infants correlates with the socioemotional and cognitive development of infants. In this study, we obtained day-long audio recordings of 10 mother-infant pairs in order to study their affect communica…
The Audio/Visual Emotion Challenge and Workshop (AVEC 2019) "State-of-Mind, Detecting Depression with AI, and Cross-cultural Affect Recognition" is the ninth competition event aimed at the comparison of multimedia processing and machine learning methods for automatic audiovisual health and emotion analysis, with all pa…
AV-ASR system improves speech recognition with visual context.
Large-scale datasets have played a significant role in progress of neural network and deep learning areas. YouTube-8M is such a benchmark dataset for general multi-label video classification. It was created from over 7 million YouTube videos (450,000 hours of video) and includes video labels from a vocabulary of 4716 c…
Automatic prediction of emotion promises to revolutionise human-computer interaction. Recent trends involve fusion of multiple data modalities - audio, visual, and physiological - to classify emotional state. However, in practice, collection of physiological data `in the wild' is currently limited to heartbeat time ser…
This paper addresses the problem of localizing audio sources using binaural measurements. We propose a supervised formulation that simultaneously localizes multiple sources at different locations. The approach is intrinsically efficient because, contrary to prior work, it relies neither on source separation, nor on mon…
SEMI uses multisensory incongruity to self-supervise exploration in reinforcement learning.
Inspired by brain's modality fusion, this paper detects active speakers from audio and video.
NeoMLP improves neural fields by adding self-attention for better downstream tasks.
Base of fibered correspondence is arbitrary correspondence. Fibered correspondence is interesting when we consider relationship between different bundles. However composition of fibered correspondences may not always be defined. Reduced fibered correspondence is defined only between fibers over the same point of base. …
Generalizes uniformization to algebraic correspondences.
Study of holomorphic correspondences combining entire maps and Fuchsian groups.
Introduces non-abelian Hodge correspondence linking algebraic structures to geometry.
New geometric proof for rational tangles links-quivers correspondence.
The paper defines a category of Lagrangian correspondences in super Hilbert spaces and constructs a functorial field theory.
With respect to the Dolbeault complex over the flat manifold $\C^n$, an explicit description of the inverse correspondence of the twistor correspondence is given.
Constructs correspondences on hyperelliptic surfaces combining orbifold groups and Blaschke products.
The calculus correspondence has been known to exist between generic pedal evolutions and generic wave front evolutions. In this paper, we first extend the known results on the calculus correspondence to evolutions with multi-parameters, and then give applications of calculus correspondence. Moreover, we discuss the pos…
Study Kobayashi-Hitchin correspondence for special sheaves on Kähler manifolds.
The paper extends Riemann-Hilbert correspondence to foliations.
Constructs Lagrangian correspondences for Higgs bundles and holomorphic connections.
ROBOT framework solves regression without correspondence for large data and complex models.
The main result of this paper is a discrete Lawson correspondence between discrete CMC surfaces in R^3 and discrete minimal surfaces in S^3. This is a correspondence between two discrete isothermic surfaces. We show that this correspondence is an isometry in the following sense: it preserves the metric coefficients int…
In this paper, we define the corresponding submanifolds to left-invariant Riemannian metrics on Lie groups, and study the following question: does a distinguished left-invariant Riemannian metric on a Lie group correspond to a distinguished submanifold? As a result, we prove that the solvsolitons on three-dimensional s…
Overview of dynamics in algebraic correspondences and their connections.
Taxicab correspondence analysis visualizes sparse text data sets.
Proves Giroux Correspondence in 3D using Heegaard splittings.
We demonstrate an isomorphism between the homology of the strand algebra of bordered Floer homology, and the category algebra of the contact category introduced by Honda. This isomorphism provides a direct correspondence between various notions of Floer homology and arc diagrams, on the one hand, and contact geometry a…
In this paper, we want to construct a one-to-one correspondence from the set of diffeomorphism classes of spin -twisted homology $\mc P^3$ to the set of isotopy classes of the embedding from to , which is a generalization of the Montgomery-Yang correspondence. Furthermore, we will apply this generalized c…
We show Mckay correspondence of Betti numbers of Chen-Ruan coho- mology for omnioriented quasitoric orbifolds. In previous articles with M. Poddar [8], [9], we proved the correspondence for four dimension and six dimensions. Here we deal with the general case.
Paper extends Kobayashi-Hitchin correspondence to non-Kähler manifolds.
New correspondence links fluxless to fluxy flag manifolds via T-duality.
We generalize Lagrangian Floer cohomology to sequences of Lagrangian correspondences. For sequences related by the geometric composition of Lagrangian correspondences we establish an isomorphism of the Floer cohomologies. We give applications to calculations of Floer cohomology, displaceability of Lagrangian correspond…