A framework isolates VQA reasoning from perception for better model evaluation.
problem Improper separation of visual perception and reasoning in VQA models.
method Introducing a framework and a top-down calibration technique to decouple reasoning from perception.
result Improved evaluation of VQA models by separating reasoning from perception.
Achieving artificial visual reasoning - the ability to answer image-related questions which require a multi-step, high-level process - is an important step towards artificial general intelligence. This multi-modal task requires learning a question-dependent, structured reasoning process over images from language. Stand…
ARNe model excels in abstract visual reasoning tasks.
problem Abstract visual reasoning using attention mechanisms.
method Hybrid network architecture combining self-attention and relational reasoning.
result ARNe model surpasses WReN model by 11.28 ppt on PGM datasets.
Few-shot visual reasoning model learns analogical relationships from small data.
problem Training deep models on few samples for visual reasoning tasks.
method Meta-analogical contrastive learning to enforce structural similarity between training and test samples.
result Method outperforms state-of-the-art on RAVEN dataset with scarce training data.
New task AQA tackles acoustic reasoning from sound scenes.
problem Promote research in acoustic reasoning.
method Generate acoustic scenes from elementary sounds and formulate questions.
result Preliminary results with models FiLM and MAC show promise.
Improved few-shot visual reasoning with image preprocessing.
problem Few-shot classifiers struggle with abstract visual reasoning tasks.
method Spectral feature removal to emphasize unique image parts.
result Combining spectral preprocessing with Relational Networks improves accuracy nearly 40%.
MXGNet tackles visual reasoning tasks using graph neural networks.
problem Abstract reasoning, especially in the visual domain, is challenging for AI.
method Combines object-level representations, graph neural networks, and multiplex graphs.
result Achieves state-of-the-art accuracy on Euler Diagram Syllogisms and outperforms state-of-the-art models on RPM datasets.
Constellation learns group-level visual relationships for abstract reasoning.
problem Learning configurational properties of entire groups of objects.
method Introduces Constellation, a network that learns relational abstractions over static visual scenes.
result Offers a basis for abstract relational reasoning and sensory imagination.
Disentangled representations improve abstract visual reasoning tasks.
problem The usefulness of disentangled representations for abstract visual reasoning.
method A large-scale study with 360 state-of-the-art unsupervised disentanglement models and 3600 abstract reasoning models.
result Disentangled representations lead to better down-stream performance in abstract reasoning tasks.
We introduce a general-purpose conditioning method for neural networks called FiLM: Feature-wise Linear Modulation. FiLM layers influence neural network computation via a simple, feature-wise affine transformation based on conditioning information. We show that FiLM layers are highly effective for visual reasoning - an…
Neural model predicts object states and physical parameters from visual observations.
problem Computational models struggle with physical reasoning and adapting to new environments.
method Visual prior predicts particle-based system from visual observations; inference module refines estimates subject to dynamics constraints.
result Model can infer physical properties within a few observations and adapt to unseen scenarios.
More than 50 years ago Bongard introduced 100 visual concept learning problems as a testbed for intelligent vision systems. These problems are now known as Bongard problems. Although they are well known in the cognitive science and AI communities only moderate progress has been made towards building systems that can so…
Agent learns causal relationships from visual data to perform tasks.
problem Performing tasks in novel environments with latent causal structures.
method Learning-based approach to induce causal graphs from visual observations, using attention mechanisms.
result Effective generalization to new tasks with unseen causal structures.
A CBR approach helps fraud analysts trust machine learning predictions.
problem Understanding the trustworthiness of machine learning predictions for fraud analysts.
method Case-based reasoning (CBR) approach to visualize similar previous instances and their local post-hoc explanations.
result Empirically, the visualization of similar previous instances is useful and easy to use for fraud analysts.
A new method learns object hierarchies from images to reason about physical interactions.
problem Learning about the interactions of complex objects and their dynamics.
method Unsupervised learning of object hierarchies from raw visual images.
result Improves over a strong baseline at modeling synthetic and real-world videos.
Program synthesis struggles with complex spatial relationships in image classification.
problem Challenges in solving Synthetic Visual Reasoning Test problems.
method Quantitative reanalysis of human and machine performance, improved program synthesis classifier, categorization of SVRT problems.
result Program synthesis is constrained by spatial relationships in images, not just shape specification.
Agent Trading Arena trains LLMs in real-time financial markets to improve numerical reasoning.
problem Limited real-world training for LLMs in financial markets.
method Virtual zero-sum stock market with competitive multi-agent trading.
result LLMs perform better with chart-based visualizations and a reflection module.
Visual analytics systems combine machine learning or other analytic techniques with interactive data visualization to promote sensemaking and analytical reasoning. It is through such techniques that people can make sense of large, complex data. While progress has been made, the tactful combination of machine learning a…
A novel method for visual question answering using scene graphs and reinforcement learning.
problem Answering free-form questions about images with deep linguistic and visual understanding.
method Context-driven, sequential reasoning based on scene graphs and reinforcement learning.
result Our method almost reaches human performance on the GQA dataset.
New models explain image-based questions better with fewer examples.
problem Improving understanding and efficiency in visual question answering.
method Probabilistic neural-symbolic models with interpretable latent programs.
result Models generate more understandable programs with fewer examples and allow probing reasoning.
Visualizes deep neural networks for speech recognition using learned topographic filter maps.
problem Unintuitive internal structure of deep neural networks complicates activation visualization.
method Trains a convolutional speech recognition model with filters arranged in a 2D grid, highlighting similar filters.
result Topographic filter maps visualize artificial neuron activations more intuitively.
This work uses visualizations to make generalization of neural networks more intuitive.
problem Understanding the reasons behind neural networks' ability to generalize to unseen data.
method Visualization methods to explain the geometry of loss landscapes and the role of dimensionality in optimization.
result Visualization helps in understanding how optimizers settle into minima that generalize well.
DPFRL uses particle filters for decision making with complex visual observations.
problem Decision making with partial complex visual observations.
method Discriminative Particle Filter Reinforcement Learning (DPFRL) with a differentiable particle filter in the neural network policy.
result DPFRL outperforms state-of-the-art POMDP RL models in complex visual observation tasks.
MAOP learns object dynamics from raw visual data.
problem Efficient learning of dynamics from raw visual data for multiple objects.
method Three-level learning architecture with spatial-temporal relational reasoning.
result Significantly outperforms previous methods in sample efficiency and generalization.
It is commonly believed that increasing the interpretability of a machine learning model may decrease its predictive power. However, inspecting input-output relationships of those models using visual analytics, while treating them as black-box, can help to understand the reasoning behind outcomes without sacrificing pr…
SATNet solves the Symbol Grounding Problem, enabling self-supervised learning.
problem Mapping visual inputs to symbolic variables without explicit supervision.
method Self-supervised pre-training pipeline and proofreading method.
result SATNet achieves full accuracy with no label leakage, surpassing state-of-the-art.
A neural network learns relational representations from raw data.
problem Learning reusable representations from raw pixel data.
method Explicitly relational neural network architecture trained on visual relational tasks.
result The architecture outperforms baselines on unseen tasks.
The paper offers a checklist for comparing human and machine visual perception.
problem Comparing human and machine visual perception.
method Designing, conducting, and interpreting experiments to investigate mechanisms.
result Feedback mechanisms may not be necessary for visual reasoning tasks.
MAC Net is a compositional attention network designed for Visual Question Answering. We propose a modified MAC net architecture for Natural Language Question Answering. Question Answering typically requires Language Understanding and multi-step Reasoning. MAC net's unique architecture - the separation between memory an…
Explains curves and surfaces in differential geometry.
problem Understanding smooth curves and surfaces in differential geometry.
method Problem-centered, elementary, visual approach focusing on essential techniques.
result Provides a solid foundation for further study in differential geometry.
The paper proposes a framework to reason about object dynamics for faster reinforcement learning.
problem Current reinforcement learning approaches lack prior knowledge about the environment, limiting efficiency.
method Integrates object dynamics and behavior into reinforcement learning to improve efficiency.
result Demonstrates the need for reasoning about object behavior and dynamics, leading to faster learning.
Robots learn object dynamics from visuals using GNNs and relational biases.
problem Challenging for robots to reason like humans about physical interactions.
method Graph Neural Networks (GNNs) with relational inductive bias.
result Auto-Predictor outperforms GN-based models and auto-encoder baseline.
We focus on two supervised visual reasoning tasks whose labels encode a semantic relational rule between two or more objects in an image: the MNIST Parity task and the colorized Pentomino task. The objects in the images undergo random translation, scaling, rotation and coloring transformations. Thus these tasks involve…
PyFi uses adversarial agents to train VLMs on financial image understanding.
problem Training VLMs to understand complex financial questions.
method PyFi-600K dataset and adversarial MCTS mechanism.
result Fine-tuned VLMs improve by 19.52% and 8.06% on financial question accuracy.
These notes introduce key techniques in differential geometry for curves and surfaces.
problem Understanding the basics of differential geometry for curve and surface analysis.
method Problem-centered, elementary, visual approach to teaching essential techniques.
result Provides a solid foundation for further study in differential geometry.
Paper introduces scalable neural architecture for solving NP-hard problems.
problem Solving NP-hard reasoning problems from natural inputs.
method Scalable neural architecture and loss function for discrete Graphical Models.
result Empirically shows efficient learning of NP-hard problems.
VCML learns concepts and metaconcepts from images and questions.
problem Learning concepts and metaconcepts from visual data.
method Bidirectional connection between visual concepts and metaconcepts.
result VCML can generalize from limited data and noisy inputs.
DAFT models attention as a dynamical system to make neural networks more interpretable.
problem Uninterpretable features learned by neural networks without human priors.
method DAFT models attention as a continuous dynamical system using neural ODEs.
result DAFT reduces the number of reasoning steps while maintaining similar performance.
Paper closes neural-symbolic learning loop with grammar model and back-search algorithm.
problem Slow convergence in neural-symbolic learning due to error propagation issues.
method Introduces grammar model as symbolic prior and back-search algorithm for efficient error propagation.
result Significantly outperforms RL methods in performance, converging speed, and data efficiency.
Neural network training relies on our ability to find "good" minimizers of highly non-convex loss functions. It is well-known that certain network architecture designs (e.g., skip connections) produce loss functions that train easier, and well-chosen training parameters (batch size, learning rate, optimizer) produce mi…
Neural networks struggle with reasoning tasks that require specialized structures.
problem Understanding why and when neural network structures generalize better.
method Developing a framework to characterize algorithmic alignment with reasoning tasks.
result Neural networks align better with dynamic programming (DP) for certain reasoning tasks.
Two new algorithms select matrix rows and columns to preserve distances.
problem Preserving distances in large matrix visualizations.
method Selects rows and columns to preserve distances.
result Preserves distances as closely as possible.
CausAdv detects adversarial examples using causal reasoning.
problem Vulnerability of CNNs to adversarial perturbations.
method Causal framework based on counterfactual reasoning.
result Adversarial examples exhibit different CI distributions compared to clean samples.
This paper proposes a web-based visual graph analytics platform for interactive graph mining, visualization, and real-time exploration of networks. GraphVis is fast, intuitive, and flexible, combining interactive visualizations with analytic techniques to reveal important patterns and insights for sense making, reasoni…
Recent breakthroughs in computer vision and natural language processing have spurred interest in challenging multi-modal tasks such as visual question-answering and visual dialogue. For such tasks, one successful approach is to condition image-based convolutional network computation on language via Feature-wise Linear …
New RL framework learns task completion without prior knowledge.
problem Learning task completion without linguistic or perceptual knowledge.
method Sequentially imagining visual goals and choosing actions.
result Framework outperforms flat and hierarchical architectures.
As Computer Vision moves from a passive analysis of pixels to active analysis of semantics, the breadth of information algorithms need to reason over has expanded significantly. One of the key challenges in this vein is the ability to identify the information required to make a decision, and select an action that will …
The Giroux correspondence and the notion of a near force-free magnetic field are used to topologically characterize near force-free magnetic fields which describe a variety of physical processes, including plasma equilibrium. As a byproduct, the topological characterization of force-free magnetic fields associated with…