Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities, and the recent success of convolutional neural networks (CNNs) in SE, we propose an audio-visual de…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities, and the recent success of convolutional neural networks (CNNs) in SE, we propose an audio-visual de…
AV-CPL uses continuous pseudo-labels for AVSR combining labeled and unlabeled data.
Curiosity enhanced by audio-visual associations improves learning efficiency.
In this paper, we analyzed how audio-visual speech enhancement can help to perform the ASR task in a cocktail party scenario. Therefore we considered two simple end-to-end LSTM-based models that perform single-channel audio-visual speech enhancement and phone recognition respectively. Then, we studied how the two model…
Automatic speech recognition can potentially benefit from the lip motion patterns, complementing acoustic speech to improve the overall recognition performance, particularly in noise. In this paper we propose an audio-visual fusion strategy that goes beyond simple feature concatenation and learns to automatically align…
Paper proposes new principles and framework for AVC learning from user-generated videos.
In this paper we address the problem of tracking multiple speakers via the fusion of visual and auditory information. We propose to exploit the complementary nature of these two modalities in order to accurately estimate smooth trajectories of the tracked persons, to deal with the partial or total absence of one of the…
Data clustering has received a lot of attention and numerous methods, algorithms and software packages are available. Among these techniques, parametric finite-mixture models play a central role due to their interesting mathematical properties and to the existence of maximum-likelihood estimators based on expectation-m…
Inspired by brain's modality fusion, this paper detects active speakers from audio and video.
Traditional multi-view learning approaches suffer in the presence of view disagreement,i.e., when samples in each view do not belong to the same class due to view corruption, occlusion or other noise processes. In this paper we present a multi-view learning approach that uses a conditional entropy criterion to detect v…
We address the problems of multi-domain and single-domain regression based on distinct and unpaired labeled training sets for each of the domains and a large unlabeled training set from all domains. We formulate these problems as a Bayesian estimation with partial knowledge of statistical relations. We propose a worst-…
We propose a tri-modal architecture to predict Big Five personality trait scores from video clips with different channels for audio, text, and video data. For each channel, stacked Convolutional Neural Networks are employed. The channels are fused both on decision-level and by concatenating their respective fully conne…
Two-stream model recognizes affect from audio and video.
Paper proposes self-supervised method for accurate speaker diarization.
It has been suggested in developmental psychology literature that the communication of affect between mothers and their infants correlates with the socioemotional and cognitive development of infants. In this study, we obtained day-long audio recordings of 10 mother-infant pairs in order to study their affect communica…
The Audio/Visual Emotion Challenge and Workshop (AVEC 2019) "State-of-Mind, Detecting Depression with AI, and Cross-cultural Affect Recognition" is the ninth competition event aimed at the comparison of multimedia processing and machine learning methods for automatic audiovisual health and emotion analysis, with all pa…
AV-ASR system improves speech recognition with visual context.
Large-scale datasets have played a significant role in progress of neural network and deep learning areas. YouTube-8M is such a benchmark dataset for general multi-label video classification. It was created from over 7 million YouTube videos (450,000 hours of video) and includes video labels from a vocabulary of 4716 c…
Automatic prediction of emotion promises to revolutionise human-computer interaction. Recent trends involve fusion of multiple data modalities - audio, visual, and physiological - to classify emotional state. However, in practice, collection of physiological data `in the wild' is currently limited to heartbeat time ser…
Speech-driven visual speech synthesis involves mapping features extracted from acoustic speech to the corresponding lip animation controls for a face model. This mapping can take many forms, but a powerful approach is to use deep neural networks (DNNs). However, a limitation is the lack of synchronized audio, video, an…
This paper addresses the problem of localizing audio sources using binaural measurements. We propose a supervised formulation that simultaneously localizes multiple sources at different locations. The approach is intrinsically efficient because, contrary to prior work, it relies neither on source separation, nor on mon…
SEMI uses multisensory incongruity to self-supervise exploration in reinforcement learning.
NeoMLP improves neural fields by adding self-attention for better downstream tasks.
The study examines stability of Hamiltonian Poisson integrators on both integrable and non-integrable systems.
Geometrically interprets integrability of geodesic flow using web theory.
Proof shows volume equals integral points for certain manifolds.
Integrates rough geometric forms on manifolds.
The paper defines and analyzes set-valued stochastic integrals for Lévy processes.
New integration theory on topological spaces, including fractals.
We discuss a recurrent geometrical method, due to Élie Cartan and von Weber ([1],[11]) enabling us to determine, step by step, the maximal integral manifolds of a not necessarily integrable nor regular Pfaffian system. The dimensions of such integral manifolds can, of course, vary from point to point but more so can va…
We define a non-absolutely convergent integration on integral currents of dimension 1 in Euclidean space. This integral is closely related to the Henstock-Kurzweil and Pfeffer Integrals. Using it, we prove a generalized Fundamental Theorem of Calculus on these currents. A detailed presentation of Henstock-Kurzweil Inte…
The paper defines and proves the existence of decompositions of integral varifolds.
Counterexample shows Ito integrand needn't be locally square integrable.
Study integrable geodesic flows on 2-surfaces with high-degree polynomial first integrals.
The article constructs stochastic integration in Riemannian manifolds.
We introduce renormalized integrals which generalize conventional measure theoretic integrals. One approximates the integration domain by measure spaces and defines the integral as the limit of integrals over the approximating spaces. This concept is implicitly present in many mathematical contexts such as Cauchy's pri…
TQ separates sampling and integration for high-dimensional integrals.
Method finds differential equations for integrable billiard tables.
Constructs Lie groupoid integrating elliptic tangent bundles and Poisson structures.
Investigates integrable systems with linear periodic integral for e(3) Lie algebra.
We use neural networks as control variates with geometric integration techniques.
Integral foliated simplicial volume is a version of simplicial volume combining the rigidity of integral coefficients with the flexibility of measure spaces. In this article, using the language of measure equivalence of groups we prove a proportionality principle for integral foliated simplicial volume for aspherical m…
This paper is an exposition of heuristics related to Witten's functional integral, relating it to Vassiliev invariants and to the Kontsevich integrals that can be used to produce Vassiliev invariants of knots and links.In particular, we give a simplified version of the appearance of the Kontsevich integrals in the pert…
Surveying integrability of Lie algebroids and structures.
This paper provides an existence-and-uniqueness theorem characterizing the stochastic integral with respect to a Wiener process. The integral is represented as a mapping from the space of measurable and adapted pathwise locally integrable processes to the space of continuous adapted processes. It is characterized in te…
We construct an infinite-dimensional symplectic 2-groupoid as the integration of an exact Courant algebroid. We show that every integrable Dirac structure integrates to a "Lagrangian" sub-2-groupoid of this symplectic 2-groupoid. As a corollary, we recover a result of Bursztyn-Crainic-Weinstein-Zhu that every integrabl…
Integrable LCK manifolds characterized as Kähler Lie algebras.