Improved speech recognition with audio-visual fusion.
problem Enhance speech recognition accuracy in noisy conditions.
method Proposes an attention-based audio-visual fusion strategy to align and learn from acoustic and lip motion data.
result Significant improvements in recognition accuracy (7-30%) on TCD-TIMIT dataset.
The paper tracks multiple speakers using audio and visual data.
problem Tracking multiple speakers with audio and visual data.
method Generative model with variational inference for latent variables.
result The proposed method outperforms baseline methods in real-world scenarios.
We propose a tri-modal architecture to predict Big Five personality trait scores from video clips with different channels for audio, text, and video data. For each channel, stacked Convolutional Neural Networks are employed. The channels are fused both on decision-level and by concatenating their respective fully conne…
Inspired by brain's modality fusion, this paper detects active speakers from audio and video.
problem Detecting active speakers in noisy environments.
method Inspired by brain's superior colliculus, combines audio and visual data through specialized neural networks and a novel fusion layer.
result Achieved results greatly surpassing initial expectations, confirming the effectiveness of the proposed method.
AV-CPL uses continuous pseudo-labels for AVSR combining labeled and unlabeled data.
problem Improving AVSR performance with labeled and unlabeled data.
method Semi-supervised method using continuous pseudo-labels generated by the same AVSR model.
result Significant improvements in VSR performance on LRS3 dataset.
The paper analyzes how speech enhancement and recognition can be improved in noisy environments.
problem Improving speech recognition in multi-talker scenarios with limited resources.
method Developed and trained two LSTM-based models for speech enhancement and phone recognition, then studied their joint optimization.
result Joint optimization of speech enhancement and recognition leads to a significant reduction in Phone Error Rate (PER).
Curiosity enhanced by audio-visual associations improves learning efficiency.
problem Challenges in reinforcement learning, especially predicting the future.
method Exploits multiple modalities (audio and vision) to predict novel associations.
result Improves exploration and learning efficiency in various environments.
Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities, and the recent success of convolutional neural networks (CNNs) in SE, we propose an audio-visual de…
Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities, and the recent success of convolutional neural networks (CNNs) in SE, we propose an audio-visual de…
Paper proposes new principles and framework for AVC learning from user-generated videos.
problem Challenges in learning audio-visual correspondence from short-term user-generated videos.
method Introduced new principles and a framework to facilitate AVC learning from videos' themes.
result Proposed approach outperformed baseline by 23.15% on KWAI-AD-AudVis corpus.
Bayesian deep learning predicts emotion from heartbeat data.
problem Uncertainty in emotion prediction from physiological data.
method End-to-end deep learning model with Bayesian uncertainty estimation.
result Peak classification accuracy of 90% on benchmark datasets.
This paper addresses the problem of localizing audio sources using binaural measurements. We propose a supervised formulation that simultaneously localizes multiple sources at different locations. The approach is intrinsically efficient because, contrary to prior work, it relies neither on source separation, nor on mon…
Data clustering has received a lot of attention and numerous methods, algorithms and software packages are available. Among these techniques, parametric finite-mixture models play a central role due to their interesting mathematical properties and to the existence of maximum-likelihood estimators based on expectation-m…
AVEC 2019 challenges AI in detecting depression and cross-cultural emotions.
problem Detecting depression and cross-cultural emotions from audiovisual data.
method Comparison of machine learning methods under standardized conditions.
result Baseline system performance on state-of-mind, depression, and cross-cultural tasks.
Improved robustness in multi-modal sensor fusion with deep learning.
problem Inconsistency in fusion weights leading to poor performance under sensor failures.
method Proposes deep multi-modal sensor fusion architectures with fusion weight regularization and target learning.
result Proposed architectures outperform existing deep learning methods under sensor failures.
Proposes Fusion Recurrent Neural Network for sequence data.
problem Improving sequence learning for practical applications.
method Fusion module and Transport module for sequence data.
result Fusion RNN performs comparably to state-of-the-art RNNs.
Meta Fusion integrates various multimodal data fusion strategies into a unified framework.
problem Improving predictive power of machine learning methods across diverse applications.
method Meta Fusion constructs a cohort of models based on latent representations across modalities, sharing soft information to boost performance.
result Meta Fusion consistently outperforms conventional fusion strategies in simulation and real-world applications.
We propose a generalized class of multimodal fusion operators for the task of visual question answering (VQA). We identify generalizations of existing multimodal fusion operators based on the Hadamard product, and show that specific non-trivial instantiations of this generalized fusion operator exhibit superior perform…
New insights into knot fusion numbers via cabling.
problem Understanding fusion numbers of ribbon knots and their behavior under cabling.
method Utilizing knot Floer homology and cabling formulas to analyze fusion numbers.
result The fusion number and strong homotopy fusion number of (p,1)-cable knots are preserved.
Study improves robustness of deep fusion models against single source noise.
problem Ensuring robustness of deep fusion models against noise added to a single input source.
method Proposed two approaches: a carefully designed loss function and a convolutional fusion layer.
result Deep fusion models become robust against noise applied to a single source, preserving performance on clean data.
A new memory-based fusion layer improves multi-modal deep learning performance.
problem Improving performance of multi-modal deep learning by addressing long-term dependencies.
method Introducing a Memory based Attentive Fusion (MBAF) layer that incorporates both current and long-term dependencies.
result The MBAF layer enhances fusion and improves performance across different modalities and networks.
Optimized deep learning architectures improve sensor fusion performance.
problem Sensor fusion in autonomous systems.
method Proposed two optimized architectures: coarser-grained and two-stage gated.
result Significant performance improvements and robustness in noisy conditions.
Innovative 2-categories create 4-manifold invariants.
problem Constructing invariants for 4-manifolds.
method Semisimple 2-categories, fusion 2-categories, and state-sum construction.
result Construct a state-sum invariant for 4-manifolds.
This paper improves model fusion by training-time neuron alignment, reducing barriers in multi-model fusion.
problem Diverse neuron permutations across different settings hinder model fusion performances.
method Training-time neuron alignment using fixed neuron anchors to reduce training-time permutations.
result Training-time neuron alignment improves fusion of pretrained models and federated learning performances.
We classify all fusion categories for a given set of fusion rules with three simple object types. If a conjecture of Ostrik is true, our classification completes the classification of fusion categories with three simple object types. To facilitate the discussion we describe a convenient, concrete and useful variation o…
Novel fusion network combines polarization and radiomics features for liver cancer classification.
problem Challenges in histopathological diagnosis of HCC and ICC.
method Two-tier fusion approach: feature-level and classification-level.
result Significantly enhances classification accuracy, even at reduced resolutions.
Study fusion methods for financial image views to improve robustness against attacks.
problem Improving robustness of financial image views for next-day direction prediction.
method Same-source multi-view learning with early fusion and late fusion, using OHLCV and technical-indicator views, and evaluating pixel-space L-infinity attacks.
result Early fusion can suffer negative transfer under noisy settings, while late fusion is more reliable once labels stabilize.
New proof and formula linking fusion trees to quantum knot invariants.
problem Quantum knot invariants encoding in non-semisimple TQC.
method Connection between fusion trees and Lawrence representations, using graphical calculus.
result Explicit encoding of quantum knot invariants via fusion trees.
Paper proposes a new method for Bayesian linear regression using spike-and-slab priors.
problem Identifying predictors with similar relationships in linear regression models.
method Hierarchical Bayesian models with spike-and-slab priors and a Gibbs sampler.
result The proposed method outperforms previous methods in simulations and real data analysis.
Paper proposes RMFN for multimodal language analysis.
problem Modeling interactions between language, visual, and acoustic modalities.
method Recurrent Multistage Fusion Network (RMFN) decomposes fusion into stages focusing on subsets of multimodal signals.
result RMFN achieves state-of-the-art performance across multimodal sentiment analysis, emotion recognition, and speaker traits recognition datasets.
Gradient descent constructs tight fusion frames.
problem Constructing tight fusion frames from prescribed subspaces.
method Gradient descent and symplectic geometry.
result Gradient descent can be used to construct tight fusion frames.
A lot of attention has been devoted to multimedia indexing over the past few years. In the literature, we often consider two kinds of fusion schemes: The early fusion and the late fusion. In this paper we focus on late classifier fusion, where one combines the scores of each modality at the decision level. To tackle th…
Study Coxeter groups over fusion rings and their geometric realisations.
problem Understanding Coxeter groups and their embeddings.
method Investigate faithful realisations and Vinberg systems.
result Induce embeddings of hyperplane complements.
Constructs a fusion product on spinor bundle over loop space.
problem Describes a fusion product on the spinor bundle over loop space.
method Uses Connes fusion of von Neumann bimodules and string structures.
result Establishes a novel relation between string structures, loop fusion, and Connes fusion of Fock spaces.
Bayesian fusion improves radar target recognition for UAVs.
problem Improving radar target recognition for UAVs using multistatic radar configurations.
method Proposes a fully Bayesian RATR framework using Optimal Bayesian Fusion (OBF) to aggregate classification probability vectors from multiple radars.
result Empirical results show that the OBF method significantly enhances classification accuracy compared to other fusion methods and single radar configurations.
Adversarial approach enhances sensor fusion for robust target detection.
problem Improving target detection and classification using multi-modal sensor fusion.
method Generative network learns latent space from various sensor modalities, then detects damaged sensors and safeguards performance.
result Automatic robustness against noisy/damaged sensors achieved.
Constructs a lift of fusion for spinors on the circle using Tomita-Takesaki theory.
problem Fusion of operators on spinors on the circle.
method Uses Tomita-Takesaki theory for Clifford-von Neumann algebras.
result Lifts fusion to a central extension of implementers on a Fock space.
Fusion of transformer networks using optimal transport for improved performance.
problem Improving performance of transformer-based models through fusion.
method Exploiting optimal transport for soft alignment of transformer components.
result Consistently outperforms vanilla fusion and individual parent models.
Paper proposes a method to preserve multimodal sentiment analysis fidelity.
problem Lack of fidelity in multimodal fusion for sentiment analysis.
method Variational autoencoder-based approach for modality fusion.
result Empirically shows superior performance over state-of-the-art methods.
The paper analyzes how shared priors affect Bayesian data fusion performance.
problem Effect of shared priors on Bayesian data fusion performance.
method Theoretical analysis using two divergences common in Bayesian inference.
result Theoretical analysis and experimental validation of performance behavior.
Traditional multi-view learning approaches suffer in the presence of view disagreement,i.e., when samples in each view do not belong to the same class due to view corruption, occlusion or other noise processes. In this paper we present a multi-view learning approach that uses a conditional entropy criterion to detect v…
Partial fusion combines neural networks to balance accuracy and efficiency.
problem Balancing accuracy and computational cost in neural networks.
method Extending weight aggregation methods based on neuron-level similarity, using partial optimal transport to match similar neurons.
result Achieves a flexible tradeoff between computational cost and performance.
Formula for Alexander polynomials of simple-ribbon knots derived.
problem Determining Alexander polynomials for simple-ribbon knots.
method Introduced simple-ribbon fusions and used them to derive a formula for Alexander polynomials.
result Formula for Alexander polynomials of simple-ribbon knots derived.
This paper reviews deep learning for multi-modality medical image segmentation.
problem Improving segmentation accuracy in medical images using multiple modalities.
method Overview of deep learning and multi-modal medical image segmentation, analysis of different network architectures and fusion strategies.
result Later fusion of modalities can lead to more accurate segmentation results.
We present a baseline approach for cross-modal knowledge fusion. Different basic fusion methods are evaluated on existing embedding approaches to show the potential of joining knowledge about certain concepts across modalities in a fused concept representation.
New 4-manifold invariant defined from trisection diagrams.
problem Defining a new 4-manifold invariant from trisection diagrams.
method Algebraic data from bimodule categories and spherical fusion categories, described diagrammatically.
result Includes Hopf algebraic invariants and modular fusion category invariants.
Spinor bundle constructed on loop space for string manifolds.
problem Constructing spinor bundles for string manifolds.
method Using ideas from Stolz and Teichner, a spinor bundle and fusion product are defined for a string manifold.
result Existence of spinor bundle with fusion product on a manifold X if and only if X admits a string structure.
MNIST-NET10 fusion improves MNIST classification to 0.1% error rate.
problem Improving MNIST classification accuracy.
method Complex heterogeneous fusion architecture using degree of certainty aggregation.
result MNIST-NET10 achieves 0.1% error rate with 10 misclassifications.