LORL learns object-centric representations from vision and language.
problem Learning disentangled, object-centric scene representations from vision and language.
method LORL integrates unsupervised object discovery and segmentation with language input to learn object-centric concepts.
result LORL improves unsupervised object discovery methods and aids downstream tasks.
Simple object representations improve model-free RL performance.
problem Current reinforcement learning agents lack object recognition.
method Used simple, feature-engineered object representations with the Rainbow model.
result Object representations significantly boost performance on Atari games.
New method interprets object representations from human behavior.
problem Understanding how mental object representations relate to human behavior.
method Sparse, non-negative representations of objects estimated from behavioral judgments.
result Representations predict latent object similarity and are interpretable.
SCALOR learns scalable object representations for crowded scenes.
problem Scalability in scenes with many objects.
method Spatially-parallel attention and proposal-rejection mechanisms.
result SCALOR can handle up to a hundred objects in crowded scenes.
GENESIS-V2 infers unordered object representations without iterative refinement.
problem Unsupervised learning of unordered object representations for complex images.
method Stochastic stick-breaking process for clustering pixel embeddings.
result GENESIS-V2 outperforms recent baselines in unsupervised image segmentation and scene generation.
ROOTS learns to represent and render 3D scenes with object-centric models.
problem Learning to represent and render 3D scenes with object-centric compositionality.
method Probabilistic generative model for learning object representations and scene rendering from partial observations.
result The model can infer 3D object representations and render scenes from arbitrary viewpoints.
Slot Attention extracts object-centric representations from images.
problem Learning distributed representations that don't capture natural scene composition.
method Slot Attention module interfaces with CNN outputs to produce task-dependent abstract slots.
result Slot Attention enables generalization to unseen compositions.
Model learns disentangled object location and appearance representations.
problem Learning disentangled representations of object location and appearance.
method Probabilistic generative model with amortized variational inference.
result Fully disentangled object location and appearance representations.
Object-based factorizations provide a useful level of abstraction for interacting with the world. Building explicit object representations, however, often requires supervisory signals that are difficult to obtain in practice. We present a paradigm for learning object-centric representations for physical scene understan…
KINet learns object interactions without supervision for robotic pushing.
problem Lack of supervised data for object-centric forward prediction.
method End-to-end unsupervised framework using keypoint representation and contrastive estimation.
result Automatically generalizes to unseen scenarios and accurately predicts future states.
3D object recognition accuracy can be improved by learning the multi-scale spatial features from 3D spatial geometric representations of objects such as point clouds, 3D models, surfaces, and RGB-D data. Current deep learning approaches learn such features either using structured data representations (voxel grids and o…
Paper introduces a method to create robust representations against covariate shifts.
problem Distribution shift between training and testing data in machine learning.
method Introduces a variational objective with two components: discriminative representation and invariant support.
result Optimal representations ensure robustness to covariate shifts, improving performance on DomainBed.
The paper proposes a framework to reason about object dynamics for faster reinforcement learning.
problem Current reinforcement learning approaches lack prior knowledge about the environment, limiting efficiency.
method Integrates object dynamics and behavior into reinforcement learning to improve efficiency.
result Demonstrates the need for reasoning about object behavior and dynamics, leading to faster learning.
A variety of representation learning approaches have been investigated for reinforcement learning; much less attention, however, has been given to investigating the utility of sparse coding. Outside of reinforcement learning, sparse coding representations have been widely used, with non-convex objectives that result in…
A new method learns to segment and represent objects jointly without supervision.
problem Learning to segment and represent objects jointly without supervision.
method Iterative variational inference for disentangled object representations.
result System learns to inpaint occluded parts and extrapolates to unseen objects.
The paper shows how continuous attention can lead to better object representations.
problem Learning good representations from data without labeled examples.
method Integrating unsupervised learning and reinforcement learning with intrinsic motivation.
result The proposed algorithm improves object recognition in few-shot settings.
New method improves domain generalization by matching object representations.
problem Existing domain generalization methods fail to generalize to unseen domains.
method Proposes matching-based algorithms to match object representations across domains.
result MatchDG algorithm matches ground-truth object representations and improves out-of-domain accuracy.
This work learns visual representations for deformable objects using contrastive estimation.
problem Challenges in learning plannable visual representations for deformable objects.
method Jointly optimizes visual representation and dynamics models using contrastive estimation.
result Substantial improvements in performance over standard model-based learning techniques.
A new method learns object representations from motion in slot representations.
problem Unsupervised object extraction from low-level visual data.
method Contrastive learning in slot representations, focusing on moving objects and distinct entities.
result Introduced a new evaluation metric to measure diversity of slot vectors.
UNTIE learns representations of coupled categorical data.
problem Challenges in learning from unlabeled categorical data with complex couplings.
method UNTIE approach for unsupervised representation learning of heterogeneous couplings.
result UNTIE significantly improves categorical data representations on 25 diverse datasets.
Improved sample efficiency in reinforcement learning with object exchangeability.
problem Sample inefficiency in reinforcement learning, especially with complex input structures.
method Attention-based method to project inputs into an efficient representation space invariant under input ordering.
result Our representation reduces the search space by a factor of m! for m objects, improving sample efficiency.
This research explores a modified VAE model to learn disentangled representations for object recognition.
problem Learning invariant representations for object recognition from diverse appearances.
method Develops a modified Variational Autoencoder (β-VAE) to enforce disentangled representations using variational inference. result Demonstrates that the incompatibility between β-VAE's conditional independence and latent variable independence leads to non-monotonic inference performance. In this paper, we advocate for representation learning as the key to mitigating unfair prediction outcomes downstream. Motivated by a scenario where learned representations are used by third parties with unknown objectives, we propose and explore adversarial representation learning as a natural method of ensuring those…
New approach categorizes objective functions for embodied agents.
problem Understanding how objectives relate to each other and discovering new objectives.
method Introducing Action Perception Divergence (APD) to categorize objective functions.
result Introduces a spectrum of objectives from narrow to general, explaining various unsupervised objectives.
SPACE models complex scenes by decomposing objects and backgrounds.
problem Scalability and unsupervised object-oriented scene representation learning.
method Generative latent variable model combining spatial-attention and scene-mixture approaches.
result SPACE achieves factorized object representations and decomposes complex scenes.
Unified multilinear model for causal factor disentanglement.
problem Disentangling causal factors from complex data without direct manipulation.
method Hierarchical block multilinear factorization (M-mode Block SVD) and incremental approach.
result Interpretable object representation robust to occlusion and reduced training data.
Object-centric learning improves generalization and robustness in multi-object scenes.
problem Improving generalization and robustness in neural networks for scenes with multiple objects.
method Training state-of-the-art unsupervised models on multi-object datasets and evaluating segmentation metrics and downstream tasks.
result Object-centric representations are useful for downstream tasks and generally robust to most distribution shifts affecting objects, but less so for less structured shifts.
New method learns high-quality Laplacian representations for reinforcement learning.
problem Lack of accurate Laplacian representations in large or continuous state spaces.
method Reformulated spectral graph drawing objective to have eigenvectors as unique global minimizer.
result Learned Laplacian representations more faithfully approximate the ground truth.
Objects are represented in sensory systems by continuous manifolds due to sensitivity of neuronal responses to changes in physical features such as location, orientation, and intensity. What makes certain sensory representations better suited for invariant decoding of objects by downstream networks? We present a theory…
Involutive Hopf monoids yield surface invariants.
problem Involutive Hopf monoids in symmetric monoidal categories.
method Construction of invariants via (co)equalizers and images.
result Categorical generalization of quantum double models.
This paper analyzes self-supervised learning from a multi-view perspective.
problem Understanding and optimizing self-supervised learning from multi-view data.
method Information-theoretical framework to understand and design self-supervised learning objectives.
result Self-supervised representations can extract task-relevant information and discard task-irrelevant information.
DHOG improves unsupervised clustering accuracy on image benchmarks.
problem Local optima in mutual information maximisation lead to suboptimal representations.
method Deep hierarchical object grouping (DHOG) computes multiple discrete representations in a hierarchical order.
result DHOG achieves new state-of-the-art results on three main benchmarks.
New method encodes 3D object geometry into neural network weights for efficient reconstruction.
problem Efficiently representing and reconstructing 3D objects with minimal parameters.
method Mapping network that encodes object geometry into neural network weights, reconstructing objects using simple geometric spaces.
result Reconstructed objects have accuracy comparable to state-of-the-art methods with significantly fewer parameters.
PCI combines perception and control using Bayesian inference with object-based representations.
problem Separate perception and control in reinforcement learning.
method Joint Perception and Control as Inference (PCI) framework with Object-based Perception Control (OPC).
result OPC achieves good perceptual grouping quality and outperforms baselines in accumulated rewards.
C-SWMs learn structured world models from raw data.
problem Learning structured world models from raw sensory data.
method Contrastive learning with graph neural networks.
result C-SWMs outperform typical models in structured environments.
Deep latent-variable models learn representations of high-dimensional data in an unsupervised manner. A number of recent efforts have focused on learning representations that disentangle statistically independent axes of variation by introducing modifications to the standard objective function. These approaches general…
Self-supervised and supervised methods learn similar intermediate visual representations but diverge in final layers.
problem Comparing self-supervised and supervised methods for visual learning.
method Comparison of contrastive self-supervised and supervised methods on simple image data.
result Contrastive and supervised methods learn similar intermediate representations but diverge in final layers.
FairNN learns fair representations and decisions by optimizing a multi-objective loss function.
problem Fairness in machine learning models for decision-making.
method Joint feature representation and classification with multi-objective loss function.
result Joint approach outperforms separate treatment of fairness in representation learning or supervised learning.
Proposes a new objective function to learn robust deep features.
problem Learning robust deep representations of noisy or unavailable features.
method Maximizes mutual information of all subsets of features relative to supervising signal.
result Surrogate objective function encourages non-redundant and conditionally independent features.
A new method for tracking objects using diverse templates.
problem Improving visual tracking performance and robustness.
method Proposes a framework that uses additional object templates and a new diversity measure in siamese feature space.
result Achieves strong empirical results on tracking benchmarks, improving performance and robustness.
Event2vec learns object embeddings in HINs by considering both relation properties and quantities.
problem Ignoring relation properties in existing NRL methods for HINs.
method Event2vec uses events to represent relations and defines event-driven proximities to measure object relevance.
result Event2vec preserves event-driven proximities in the embedding space and outperforms state-of-the-art methods.
The paper connects hyperbolic Dehn surgery and Higgs bundles to construct model objects in representation varieties.
problem Constructing model objects in representation varieties for Higgs bundles.
method Reviewing hyperbolic Dehn surgery and bending procedures, and applying them to Higgs bundles.
result Explicit examples of model objects in representation varieties are produced.
New method finds Lie group representations without explicit groups, enabling new neural network architectures.
problem Building neural networks equivariant to arbitrary Lie groups.
method Algorithm to find Lie group representations from Lie algebra structure constants. Self-contained method for constructing Lie group-equivariant neural networks.
result First object-tracking model equivariant to the Poincaré group.
FSRL balances fairness and sufficiency in learning representations.
problem Ensuring unbiased predictions and decisions in machine learning.
method A convex combination of sufficiency and fairness objectives using distance covariance.
result FSRL achieves a superior trade-off between fairness and accuracy.
The study quantifies how many objects can be linearly classified under all views.
problem Understanding the expressivity of group-equivariant representations.
method Generalization of Cover's Function Counting Theorem to quantify separable dichotomies.
result The fraction of separable dichotomies is determined by the fixed space dimension of the group action.
When working with three-dimensional data, choice of representation is key. We explore voxel-based models, and present evidence for the viability of voxellated representations in applications including shape modeling and object classification. Our key contributions are methods for training voxel-based variational autoen…
Context-aware ZSL improves object recognition by considering object context.
problem Previous ZSL approaches ignore object context, limiting their effectiveness.
method Proposes a new approach that models the conditional likelihood of objects appearing in specific contexts.
result Contextual information significantly improves ZSL performance and is robust to class imbalance.
Learning data representations that are transferable and are fair with respect to certain protected attributes is crucial to reducing unfair decisions while preserving the utility of the data. We propose an information-theoretically motivated objective for learning maximally expressive representations subject to fairnes…