Paper proposes dual recurrent attention units for VQA models.
problem Comprehending visual and textual data for accurate question answering.
method Introduces and evaluates recurrent attention mechanisms in VQA models.
result Dual Recurrent Attention Units (RAUs) improve VQA performance.
Efficient CNN for VQA achieves similar performance to standard models.
problem Computational intensity of standard VQA models.
method Proposes a sparsely activated CNN architecture.
result Sparsely activated CNN achieves comparable performance.
New methods optimize training VQAs without barren plateaus, improving efficiency and applicability.
problem Barren plateaus in training variational quantum algorithms.
method Derive adaptive learning rates and use Gaussian kernels to optimize movement in parameter space.
result Optimized training methods outperform other routines and can train VQAs free of barren plateaus.
A framework isolates VQA reasoning from perception for better model evaluation.
problem Improper separation of visual perception and reasoning in VQA models.
method Introducing a framework and a top-down calibration technique to decouple reasoning from perception.
result Improved evaluation of VQA models by separating reasoning from perception.
AdvReg improves VQA models but introduces instability and bias issues.
problem VQA models over-rely on linguistic biases, ignoring visual context.
method Adversarial regularization to encourage bias-free question representations.
result AdvReg yields side-effects like unstable gradients and reduced performance on in-domain examples.
Improved VQA accuracy with generalized fusion operators.
problem Enhancing multimodal fusion for better VQA performance.
method Generalized Hadamard-Product fusion operators with Nonlinearity Ensembling, Feature Gating, and post-fusion layers.
result 1.1% absolute improvement on VQA 2.0 test-dev set.
New approach connects quantum phases to VQA trainability, enabling better scaling.
problem Scalability issues in VQAs, especially barren plateaus.
method Analog VQA ansätze composed of quenches of a disordered Ising chain, tuning disorder strength.
result Thermalized and MBL phases reach maximal expressivity at large M, but barren plateaus emerge at smaller M in the thermalized phase. We conduct large-scale studies on `human attention' in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images. We design and test multiple game-inspired novel attention-annotation interfaces that require the subject to sharpen regions of a blurred image to answer a qu…
VQAs use classical optimization to train quantum circuits, promising quantum advantage.
problem High computational cost of quantum simulations and solving large-scale problems.
method Variational Quantum Algorithms (VQAs) use classical optimizers to train parametrized quantum circuits.
result VQAs are a promising strategy for obtaining quantum advantage.
Quantum machine learning faces challenges similar to variational quantum algorithms in training.
problem Challenges in training quantum machine learning models.
method Bridge between variational quantum algorithms and quantum machine learning, applying gradient scaling results.
result Gradient scaling results for variational quantum algorithms can also be applied to quantum machine learning models, revealing new trainability issues.
Improved text-to-image alignment using iterative VQA feedback.
problem Misalignment between text prompts and generated images, especially for complex inputs.
method Decompose complex prompts into assertions, evaluate each using VQA, combine scores iteratively.
result Significantly higher correlation with human ratings compared to CLIP, BLIP scores.
Develops an analytic theory for quantum imaginary time evolution.
problem Lack of a first-principle understanding of quantum imaginary time evolution.
method Interprets QITE as a form of VQA trained with QNGD and connects it to the geometric geodesic distance in the quantum Fisher information metric.
result QITE converges faster than vanilla gradient descent-based VQAs, though the advantage is suppressed by Hilbert space dimensionality.
An important goal of computer vision is to build systems that learn visual representations over time that can be applied to many tasks. In this paper, we investigate a vision-language embedding as a core representation and show that it leads to better cross-task transfer than standard multi-task learning. In particular…
Quantum algorithm improves portfolio construction accuracy.
problem Efficiently constructing portfolios with real-world constraints.
method Sampling-based CVaR Variational Quantum Algorithm (VQA) combined with local-search post-processing.
result Achieved a relative solution error of 0.49% on IBM Heron processors.
A new approach to quantum machine learning circuits reduces training difficulties.
problem Challenges in training deep quantum circuits due to flat training landscapes.
method Variable structure approach (VAns) to build ansatzes, applying rules for gate growth and removal.
result VAns successfully mitigates trainability and noise-related issues, improving performance in various applications.
Improves VQAs by balancing classical and quantum training resources.
problem Challenges in trainability and resource costs of VQAs on quantum hardware.
method Adopting HELIA Ansatz and combining classical and quantum methods for gradient estimation and training.
result Achieves higher accuracy and success rates in VQE and improved test accuracy in quantum phase classification.
We propose a technique for making Convolutional Neural Network (CNN)-based models more transparent by visualizing input regions that are 'important' for predictions -- or visual explanations. Our approach, called Gradient-weighted Class Activation Mapping (Grad-CAM), uses class-specific gradient information to localize…
New models explain image-based questions better with fewer examples.
problem Improving understanding and efficiency in visual question answering.
method Probabilistic neural-symbolic models with interpretable latent programs.
result Models generate more understandable programs with fewer examples and allow probing reasoning.
Partial-input models fail to detect dataset artifacts, even when they perform poorly.
problem The effectiveness of partial-input models in detecting dataset artifacts is questionable.
method Design artificial datasets and identify trivial patterns in the SNLI dataset.
result Partial-input models can solve examples previously considered hard, indicating potential dataset artifacts.
This study improves quantum classifiers by optimizing data preprocessing.
problem Quantum Machine Learning advantages are not yet clearly demonstrated.
method Used Linear Discriminant Analysis (LDA) for data preprocessing.
result Variational Quantum Algorithm (VQA) outperforms classical classifiers.
Study evaluates bias mitigation methods in deep learning, finds they often exploit hidden biases.
problem Deep learning systems learn biases, affecting performance on minority groups.
method Improved evaluation protocol, new dataset, robustness across different tuning distributions.
result Bias mitigation methods often exploit hidden biases, are not robust to multiple forms of bias, and are sensitive to tuning set choice.