Method generates uncertainty measures for street scene segmentation.
problem Reliability and uncertainty measures in semantic segmentation of street scenes.
method Nested crops, neural network segmentation, post-processing, uncertainty heat maps.
result Significant improvements in classification and regression performance.
Real-time scene understanding solved using Approximate Bayesian Computation.
problem Predicting human actions, object poses, and pedestrian crossings from depth images.
method Bayesian error model, neural surrogates, and adaptive discretization.
result Real-time inference on real-world problems is feasible.
LangDA improves domain adaptation for semantic segmentation by learning context-aware scene descriptions.
problem Improving domain adaptation for semantic segmentation with dense prediction tasks.
method LangDA learns contextual relationships between objects via VLM-generated scene descriptions and aligns image features with text representation.
result LangDA sets new state-of-the-art across three DASS benchmarks, outperforming existing methods.
The paper proposes a model to learn street network representations directly from graphs.
problem Loss of detailed topological data in raster representations of street networks.
method Variational autoencoder with graph convolutional layers and a probabilistic fully-connected graph decoder.
result The model infers good representations directly from street networks, capturing both local structure and spatial distribution.
Gursky-Streets introduced a formal Riemannian metric on the space of conformal metrics in a fixed conformal class of a compact Riemannian four-manifold in the context of the σ2-Yamabe problem. The geodesic equation of Gursky-Streets' metric is a fully nonlinear degenerate elliptic equation and Gursky-Streets have pr…
AI helps Vancouver identify where off-street parking saves time and space.
problem On-street parking inefficiencies and associated costs in Vancouver.
method Developed AI models for on-street and off-street parking, comparing time costs.
result Many areas off-street parking saves time and space, aligning with city goals.
The design of neural network architectures is an important component for achieving state-of-the-art performance with machine learning systems across a broad array of tasks. Much work has endeavored to design and build architectures automatically through clever construction of a search space paired with simple learning …
VAE models generate urban street networks from data.
problem Quantitative description of urban patterns.
method Variational Autoencoders (VAEs) for urban street network data.
result VAEs can capture urban network metrics and generate new urban forms.
Semantic image segmentation is one the most demanding task, especially for analysis of traffic conditions for self-driving cars. Here the results of application of several deep learning architectures (PSPNet and ICNet) for semantic image segmentation of traffic stereo-pair images are presented. The images from Cityscap…
Research confirms Streets-Tian conjecture for 2-step solvmanifolds.
problem Streets-Tian conjecture for compact complex manifolds.
method Used special non-unitary frames to reveal hidden symmetries.
result Confirms conjecture for all 2-step solvmanifolds.
Solves Gursky-Streets equations for σk Yamabe problem in dimensions n≥2k.
problem Solving the σk Yamabe problem in dimensions n≥2k. method Introduced and solved the Gursky-Streets equations with uniform C1,1 estimates using concavity of the operator and Garding's theory of hyperbolic polynomials. result Established the uniqueness of the solution to the degenerate equations for the first time.
Urban2Vec combines street view imagery and POIs for better urban neighborhood embeddings.
problem Lack of comprehensive representation of urban neighborhoods using heterogeneous data.
method Unsupervised multi-modal framework using CNN for visual features and bag-of-words for POI data.
result Urban2Vec achieves better performance than baseline models and comparable to fully-supervised methods.
More than half of the world's roads lack adequate street addressing systems. Lack of addresses is even more visible in daily lives of people in developing countries. We would like to object to the assumption that having an address is a luxury, by proposing a generative address design that maps the world in accordance w…
A neural scene representation framework enforcing 3D transformations.
problem Learning 3D scene representations from images without 3D supervision.
method Introducing a loss enforcing equivariance of the scene representation with 3D transformations.
result Real-time neural rendering with comparable results to models requiring minutes for inference.
GENESIS generates and samples 3D scenes by capturing object interactions.
problem Lack of models that explicitly capture object interactions in scene generation.
method Object-centric latent variables, spatial GMM, amortized inference, autoregressive prior.
result First object-centric generative model of 3D visual scenes.
Efficient model for foggy scene understanding in vehicles.
problem Challenging scene understanding and segmentation under foggy conditions.
method Domain adaptation and illumination-invariant image transformation.
result Outperforms state-of-the-art models in foggy scene understanding.
As part of autonomous car driving systems, semantic segmentation is an essential component to obtain a full understanding of the car's environment. One difficulty, that occurs while training neural networks for this purpose, is class imbalance of training data. Consequently, a neural network trained on unbalanced data …
We continue studying a parabolic flow of almost Kähler structures introduced by Streets and Tian which naturally extends Kähler-Ricci flow onto symplectic manifolds. In the system of primarily the symplectic form, almost complex structure, Chern torsion and Chern connection, we establish new formulas for the evolutions…
Scene text magnifier enhances readability for visually impaired.
problem Helps visually impaired read natural scene text.
method Four CNN-based networks: character erasing, extraction, magnify, synthesis.
result Effective text magnification without background alteration.
In recent years, Streets and Tian introduced a series of curvature flows to study non-Kähler geometry. In this paper, we study how to construct second order curvature flows in a uniform way, under some natural assumptions which holds in Streets and Tian's works. As a result, by classifying the lower order tensors, we c…
New model separates objects in scenes, enabling novel arrangements and depth.
problem Lack of modular, compositional scene modeling in generative models.
method Ensemble of generative models (experts) compete for explaining different parts of a scene.
result Model generates scenes with novel object arrangement and depth ordering.
ROOTS learns to represent and render 3D scenes with object-centric models.
problem Learning to represent and render 3D scenes with object-centric compositionality.
method Probabilistic generative model for learning object representations and scene rendering from partial observations.
result The model can infer 3D object representations and render scenes from arbitrary viewpoints.
Despite enormous progress in object detection and classification, the problem of incorporating expected contextual relationships among object instances into modern recognition systems remains a key challenge. In this work we propose Information Pursuit, a Bayesian framework for scene parsing that combines prior models …
SPACE models complex scenes by decomposing objects and backgrounds.
problem Scalability and unsupervised object-oriented scene representation learning.
method Generative latent variable model combining spatial-attention and scene-mixture approaches.
result SPACE achieves factorized object representations and decomposes complex scenes.
NeRF-VAE generates 3D scenes with geometric structure from few images.
problem Generating 3D scenes from few images with geometric consistency.
method Combines NeRF and VAE, incorporating shared geometric structure.
result NeRF-VAE can infer and render geometrically-consistent scenes from unseen environments.
Improved acoustic scene classification with factorized CNN.
problem Acoustic scene classification in varying environments.
method Large-margin factorized CNN with triplet loss.
result Improved performance and better generalization on unseen data.
The Streets-Tian conjecture is confirmed for specific types of Hermitian manifolds.
problem The Streets-Tian conjecture on compact Hermitian manifolds.
method Elementary approach, explicit descriptions, and pathways of deformation.
result The conjecture is confirmed for special types of compact Hermitian manifolds.
A new method for estimating joint value functions in multi-scene reinforcement learning.
problem High variance in samples for policy gradient computations in multi-scene environments.
method Sparse attention mechanism over multiple value function hypotheses to approximate the true joint value function.
result Significant improvements in reward scores and enhanced navigation efficiency across OpenAI ProcGen environments.
Generative models learn from unlabeled videos via object segmentation and scene modeling.
problem Learning generative models from unlabelled videos.
method Decomposed into three subtasks: motion segmentation, background and foreground modeling, and scene sampling.
result Approach allows learning models that generalize beyond occlusions and represent scenes in a modular fashion.
Method estimates travel times on urban roads using Uber data.
problem Estimating travel times on urban roads where data is scarce.
method Graph representation, trip sampling, least-squares optimization.
result Estimates travel times on arterial roads using aggregated Uber data.
Generative Multisensory Network learns 3D scene representations from multiple modalities.
problem Learning robust 3D scene representations from multiple sensory modalities.
method Amortized Product-of-Experts for efficient inference and cross-modal generation.
result The model can infer modality-invariant 3D scene representations efficiently from various sensory modalities.
Pix2Shape learns 3D scene representations from single images without supervision.
problem Learning 3D scene information from a single image without supervision.
method Pix2Shape uses an encoder, decoder, and critic network to generate 2.5D surfel-based reconstructions.
result Pix2Shape can generate complex 3D scenes from a single image, scaling with on-screen resolution.
Paper improves reinforcement learning in multi-scene tasks.
problem Reducing sample variance in multi-scene reinforcement learning.
method Sparse dynamic value estimation using Gaussian mixture models.
result Significant improvements in reward scores and navigation efficiency.
RICH models scenes as hierarchical tree to learn and generate complex compositions.
problem Learning compositional structures between parts and objects in natural scenes.
method RICH uses a latent scene graph to organize entities into a tree structure and employs a top-down inference approach.
result RICH learns and generates complex scene hierarchies from unlabeled data.
Improves reinforcement learning agent's scene-specific value function.
problem High variance in samples for policy gradient computations in multi-scene environments.
method Proposes dynamic value estimation (DVE) for multiple MDPs, clustering value functions across scenes.
result Lower sample variance and more accurate scene-specific value function estimates.
A novel method for visual question answering using scene graphs and reinforcement learning.
problem Answering free-form questions about images with deep linguistic and visual understanding.
method Context-driven, sequential reasoning based on scene graphs and reinforcement learning.
result Our method almost reaches human performance on the GQA dataset.
A new method to decompose audio scenes using deep learning.
problem Understanding audio scenes with random microphone arrangements.
method Formulated a neural network for nonnegative tensor factorization.
result Learned sources' individual spectral dictionaries and activation patterns.
We propose a systematic learning-based approach to the generation of massive quantities of synthetic 3D scenes and arbitrary numbers of photorealistic 2D images thereof, with associated ground truth information, for the purposes of training, benchmarking, and diagnosing learning-based computer vision and robotics algor…
Study uses Viber and street polls to estimate Belarus election ratings and turnout.
problem Obtaining accurate ratings of candidates in Belarus's banned polls.
method Bayesian multilevel regression with poststratification.
result Estimated ratings and turnout contradict official election results.
Acoustic scene classification is the task of identifying the scene from which the audio signal is recorded. Convolutional neural network (CNN) models are widely adopted with proven successes in acoustic scene classification. However, there is little insight on how an audio scene is perceived in CNN, as what have been d…
Generative model creates realistic scenes from pixel-wise labels.
problem Creating photo-realistic scenes from pixel-wise labels.
method Semantic bottleneck GAN model combining conditional and unconditional generation networks.
result Model outperforms state-of-the-art models in unsupervised image synthesis.
P3I learns holistic scene representations from a single image.
problem Inferring camera poses, object locations, and global scene structures from a single image.
method Combines search-based and gradient-based algorithms.
result P3I outperforms baselines on various image manipulation tasks.
Generates coherent 3D scenes from monocular videos without supervision.
problem Lack of 3D scene modeling in video generation models.
method Trains a model to generate 3D scenes with moving objects and a background from monocular videos.
result Trained model generates coherent 3D scenes with multiple moving objects and a background.
SCALOR learns scalable object representations for crowded scenes.
problem Scalability in scenes with many objects.
method Spatially-parallel attention and proposal-rejection mechanisms.
result SCALOR can handle up to a hundred objects in crowded scenes.
This essay suggests that a proper assessment of the presently unfolding financial crisis, and its cure, requires going back at least to the late 1990s, accounting for the cumulative effect of the ITC, real-estate and financial derivative bubbles. We focus on the deep loss of trust, not only in Wall Street, but more imp…
New dataset tests mental rotation from single images, improving model understanding of 3D scenes.
problem Understanding how a scene looks from a different viewpoint using a single image.
method Created CLEVR-MRT dataset, explored neural architectures for volumetric scene representations.
result Demonstrated the effectiveness of volumetric representations in answering mental rotation questions.
The Streets-Tian conjecture is confirmed for Lie algebras with specific abelian ideals.
problem The Streets-Tian conjecture on compact complex manifolds admitting Hermitian-symplectic metrics.
method Detailed case analysis of Lie algebras with abelian ideals of codimension 2, explicit construction of Hermitian-symplectic metrics and pathways to Kähler metrics.
result The Streets-Tian conjecture is confirmed for Lie algebras containing abelian ideals of codimension 2.
Proposes Deep Scenes for interaction-aware scene understanding in reinforcement learning for autonomous driving.
problem Leveraging deep reinforcement learning for high-level decision making in autonomous driving requires handling variable-length sequences of different object types and interactions.
method Introduces Deep Scenes architecture, an extension of Deep Sets or Graph Convolutional Networks, to learn complex interaction-aware scene representations.
result Graph-Q and DeepScene-Q algorithms outperform state-of-the-art methods in evaluations with SUMO.