Systematic approach to generating diverse 3D scenes and images for training and benchmarking.
problem Training and benchmarking machine learning algorithms for scene understanding.
method Stochastic grammar-based pipeline for generating and rendering photorealistic 3D scenes with detailed ground truth.
result Improves performance in scene understanding tasks and provides controllable benchmarks for model diagnostics.
BPI models 2D patterns on multiple planes and 3D scene from a single image.
problem Understanding and editing images with multiple 2D planes and 3D scene from a single image.
method Box Program Induction (BPI) with neural networks and search-based algorithm.
result Holistic, structured scene representation enables 3D-aware image editing.
Generative model creates realistic scenes from pixel-wise labels.
problem Creating photo-realistic scenes from pixel-wise labels.
method Semantic bottleneck GAN model combining conditional and unconditional generation networks.
result Model outperforms state-of-the-art models in unsupervised image synthesis.
A neural scene representation framework enforcing 3D transformations.
problem Learning 3D scene representations from images without 3D supervision.
method Introducing a loss enforcing equivariance of the scene representation with 3D transformations.
result Real-time neural rendering with comparable results to models requiring minutes for inference.
ROOTS learns to represent and render 3D scenes with object-centric models.
problem Learning to represent and render 3D scenes with object-centric compositionality.
method Probabilistic generative model for learning object representations and scene rendering from partial observations.
result The model can infer 3D object representations and render scenes from arbitrary viewpoints.
Pix2Shape learns 3D scene representations from single images without supervision.
problem Learning 3D scene information from a single image without supervision.
method Pix2Shape uses an encoder, decoder, and critic network to generate 2.5D surfel-based reconstructions.
result Pix2Shape can generate complex 3D scenes from a single image, scaling with on-screen resolution.
Generative Multisensory Network learns 3D scene representations from multiple modalities.
problem Learning robust 3D scene representations from multiple sensory modalities.
method Amortized Product-of-Experts for efficient inference and cross-modal generation.
result The model can infer modality-invariant 3D scene representations efficiently from various sensory modalities.
Improved 3D scene understanding from partial point sets using multiview fusion.
problem Challenging task of 3D scene semantic understanding from partial point clouds.
method Multiview representation of 360° point clouds and fusion with original data.
result Overall increase of 31.9% and 4.3% in segmentation accuracy for partial and complete scenes.
Generates coherent 3D scenes from monocular videos without supervision.
problem Lack of 3D scene modeling in video generation models.
method Trains a model to generate 3D scenes with moving objects and a background from monocular videos.
result Trained model generates coherent 3D scenes with multiple moving objects and a background.
NeRF-VAE generates 3D scenes with geometric structure from few images.
problem Generating 3D scenes from few images with geometric consistency.
method Combines NeRF and VAE, incorporating shared geometric structure.
result NeRF-VAE can infer and render geometrically-consistent scenes from unseen environments.
GENESIS generates and samples 3D scenes by capturing object interactions.
problem Lack of models that explicitly capture object interactions in scene generation.
method Object-centric latent variables, spatial GMM, amortized inference, autoregressive prior.
result First object-centric generative model of 3D visual scenes.
ObSuRF converts a single image into a 3D model with NeRFs.
problem Creating a 3D model from a single image with object segmentation.
method Unsupervised volume segmentation using Neural Radiance Fields (NeRFs).
result ObSuRF can segment a 3D scene into objects from a single image.
Generative models learn 3D scenes without explicit maps.
problem Learning models for visual 3D localization without explicit maps.
method Generative Query Networks (GQNs) with attention mechanisms.
result GQNs can capture complex 3D scenes and perform localization.
SCENE-Net improves 3D point cloud segmentation with low resource usage and transparency.
problem Lack of resources and transparency in 3D semantic segmentation models.
method SCENE-Net uses signature shapes identified via GENEOs to achieve semantic segmentation with minimal resources.
result SCENE-Net achieves comparable IoU to state-of-the-art methods with less data and computational resources.
Topology-GS improves 3D GS for better structural and feature integrity.
problem Compromised pixel-level and feature-level integrity in 3D GS.
method Incorporates Local Persistent Voronoi Interpolation (LPVI) and PersLoss based on persistent homology.
result Topology-GS outperforms existing methods in PSNR, SSIM, and LPIPS metrics.
DreamFusion uses text-to-image diffusion models to create 3D images efficiently.
problem Lack of large-scale 3D datasets and efficient architectures for 3D synthesis.
method Adapting a 2D diffusion model to 3D synthesis using a loss based on probability density distillation.
result A 3D model can be optimized from a 2D diffusion model, allowing for text-to-3D synthesis.
MONet learns to decompose scenes into meaningful components without supervision.
problem Learning meaningful scene decompositions without labeled data.
method MONet combines a VAE and recurrent attention network to learn decompositions of 3D scenes.
result MONet can learn to represent 3D scenes into meaningful components like objects and background.
This paper learns hierarchical compositional models for image synthesis.
problem Creating interpretable and hierarchical image representations.
method Sparsifying a generator network to induce an AND-OR hierarchy.
result The method learns meaningful and interpretable hierarchical representations.
We consider the problem of learning object arrangements in a 3D scene. The key idea here is to learn how objects relate to human poses based on their affordances, ease of use and reachability. In contrast to modeling object-object relationships, modeling human-object relationships scales linearly in the number of objec…
New dataset tests mental rotation from single images, improving model understanding of 3D scenes.
problem Understanding how a scene looks from a different viewpoint using a single image.
method Created CLEVR-MRT dataset, explored neural architectures for volumetric scene representations.
result Demonstrated the effectiveness of volumetric representations in answering mental rotation questions.
Proposes a new layer for efficient 3D shape discrimination.
problem Irregular structure and redundancy in 3D point clouds hinder efficient inter-class discrimination.
method Integrates Blended Convolution and Synthesis layer that projects and synthesizes 3D point clouds, followed by 3D convolution in the unit ball.
result End-to-end architecture achieves compelling results on 3D shape recognition and retrieval.
SNP extends Neural Processes to handle temporal dependencies in sequences.
problem Handling temporal dependencies in sequences of stochastic processes.
method Integrates a temporal state-transition model into Neural Processes.
result First 4D model capable of dynamic 3D scene modeling.
Recently, multiple formulations of vision problems as probabilistic inversions of generative models based on computer graphics have been proposed. However, applications to 3D perception from natural images have focused on low-dimensional latent scenes, due to challenges in both modeling and inference. Accounting for th…
ED-NeRF efficiently edits 3D scenes using latent space NeRF and improved loss functions.
problem Slow training speeds and inadequate editing loss functions in existing NeRF editing techniques.
method Embedding real-world scenes into latent space of LDM, using a unique refinement layer and an improved loss function.
result ED-NeRF achieves faster editing speed and improved output quality compared to state-of-the-art models.
Scene text magnifier enhances readability for visually impaired.
problem Helps visually impaired read natural scene text.
method Four CNN-based networks: character erasing, extraction, magnify, synthesis.
result Effective text magnification without background alteration.
New framework segments 3D scenes using neural algorithms and sub-Riemannian geometry.
problem Effective scene segmentation in 3D vision.
method Neurogeometric sub-Riemannian model, harmonic analysis, neural-based stereo correspondence.
result Sub-Riemannian metric is central to effective scene segmentation.
Worldsheet wraps a 3D mesh sheet onto a single image to synthesize novel views.
problem Synthesizing novel views from a single image with large viewpoint changes.
method Shrink-wrapping a planar mesh sheet onto the input image, consistent with learned depth.
result Worldsheet consistently outperforms prior methods on single-image view synthesis.
StyleNeRF generates high-resolution images with 3D consistency and style control.
problem Generating high-resolution images with fine details and 3D consistency.
method Integrates NeRF into a style-based generator for efficient high-resolution image synthesis.
result Synthesizes high-resolution images at interactive rates with high 3D consistency and style control.
GIBLy adds geometric priors to 3D segmentation models, improving performance with minimal overhead.
problem Lack of explicit geometric information in 3D semantic segmentation models.
method Introduces GIBLy, a lightweight geometric inductive bias layer that integrates learnable geometric priors into existing 3D segmentation pipelines.
result Consistent performance gains across multiple benchmarks, including up to +11.5% mIoU on TS40K with PTV3.
Improves AI agents' 3D navigation by learning from failures and 3D spatial relationships.
problem Challenges in data efficiency, obstacle avoidance, and generalization in 3D visual navigation.
method Incorporates attention on 3D spatial relationships and a target skill extension module into DRL framework.
result Significantly improves navigation performance and generalization across targets and scenes.
Generative model predicts future frames efficiently and at arbitrary points.
problem Slow autoregressive video prediction models and inability to sample non-consecutive frames.
method Introduces a model that generates a latent representation from an arbitrary set of frames, enabling simultaneous and efficient sampling of future frames at arbitrary time-points.
result Substantial gains in speed and functionality without loss in fidelity, demonstrated on synthetic videos and 3D scene reconstruction datasets.
System learns user preferences to synthesize materials quickly.
problem Slow material synthesis for novice and expert users.
method Gaussian Process Regression for user preferences, neural network for real-time image predictions.
result Real-time material synthesis enables novice users to generate hundreds of models.
LION generates high-quality 3D shapes using hierarchical latent diffusion models.
problem Creating high-quality 3D shapes for digital artists.
method Hierarchical Latent Point Diffusion Model (LION) with a global shape latent and point-structured latent space.
result LION achieves state-of-the-art generation performance on ShapeNet benchmarks.
DNCF framework recovers real scenes from imperfect images robustly.
problem Recovering real scenes from imperfect images.
method Nonparametric deep network that learns physical image formation equations.
result DNCF framework robustly defends against adversarial attacks.
Generates garden paintings from text descriptions using deep learning.
problem Lack of firsthand material for traditional Chinese garden reconstruction.
method Deep learning model trained on text and paintings of Ming Dynasty gardens.
result Model generates garden paintings in Ming Dynasty style based on textual descriptions.
SPACE models complex scenes by decomposing objects and backgrounds.
problem Scalability and unsupervised object-oriented scene representation learning.
method Generative latent variable model combining spatial-attention and scene-mixture approaches.
result SPACE achieves factorized object representations and decomposes complex scenes.
Enhances neural rendering with geometry-aware attention.
problem Efficiently modeling complex 3D scenes.
method Introduces Epipolar Cross Attention (ECA) for non-local operations.
result Significant improvement in Generative Query Networks (GQN) performance.
Generative model creates user-specified textures from datasets.
problem Creating detailed textures from raw data.
method Generative adversarial networks with user control and adversarial loss.
result Model generates descriptive texture manifolds and 3D textures.
Automatically infers high dynamic range illumination from a single indoor photo.
problem Predicting accurate indoor illumination from a single image.
method End-to-end deep neural network trained in three steps: lighting classifier, scene light localization, and fine-tuning for intensity prediction.
result Significantly outperforms previous methods in recovering high-quality HDR illumination.
Proposes local coordinate frames for improving model performance in complex dynamical systems.
problem Improving model performance in complex, non-linear, and time-dependent dynamical systems.
method Introduces roto-translation invariant local coordinate frames for geometric graphs.
result The approach outperforms state-of-the-art models in various complex scenarios.
Improves synthetic data for deep model training and adaptation.
problem Evaluating and improving synthetic data for deep learning models.
method Proposes a novel learned synthesis technique using generative models for shading and rendering, and uses an ensemble of models to generate datasets.
result Improves classifier performance on real data compared to state-of-the-art methods.
Generative model learns from designs to create new shapes.
problem Designing new shapes for conceptual optimization.
method Variational autoencoder-decoder architecture for 3D shape synthesis.
result Generator maps latent space to smooth 3D surfaces for optimization.
Energy-based models can generate complex images by combining simpler concepts.
problem Generating natural images that satisfy complex logical combinations of concepts.
method Energy-based models combine probability distributions of simpler concepts to generate compositions.
result Energy-based models can generate images that satisfy conjunctions, disjunctions, and negations of concepts.
Model learns tool affordances from vision, enabling tool selection.
problem Learning tool affordances from visual input.
method Vision-based generative model with task predictor.
result Agents can select appropriate tools based on task success criteria.
Real-time scene understanding solved using Approximate Bayesian Computation.
problem Predicting human actions, object poses, and pedestrian crossings from depth images.
method Bayesian error model, neural surrogates, and adaptive discretization.
result Real-time inference on real-world problems is feasible.
3D adversarial logos can fool object detectors in real-world settings.
problem Creating robust adversarial attacks in 3D rendering views.
method Constructing 3D adversarial logos via texture mapping and differentiable rendering.
result 3D adversarial logos are more versatile and robust than traditional adversarial patches.
StackGAN++ generates high-quality images from text descriptions.
problem Generating high-quality photo-realistic images from text descriptions.
method Two-stage and multi-stage generative adversarial networks (GANs) with stacked architecture.
result StackGAN++ significantly outperforms other methods in generating photo-realistic images.
nuScenes dataset includes multimodal sensor data for autonomous vehicle training.
problem Training robust detection and tracking methods for autonomous vehicles.
method Presented the first multimodal dataset with 6 cameras, 5 radars, and 1 lidar, 360-degree field of view.
result 7x more annotations and 100x more images than KITTI dataset.