Cooperative perception improves 3D object detection in autonomous vehicles.
problem Limited field-of-view and occlusion in single sensor data.
method Early fusion of point clouds from multiple sensors, late fusion of independently detected bounding boxes, and hybrid combination.
result Early fusion approach outperforms late fusion by significantly higher recall (95%) compared to single-point sensing (30%).
Flexible pipeline for 3D vehicle detection from 2D images.
problem Current methods lack 3D perception of vehicles and other objects.
method Adopt any 2D detection network, fuse with 3D point cloud, develop model fitting algorithm, refine with CNN.
result 3D detection results rank second among algorithms, demonstrating competencies.
Y-GAN uses multi-camera data to estimate depth maps without expensive hardware.
problem Depth perception for autonomous systems requires accurate 3D spatial information.
method Proposes Y-GAN, a deep convolutional generative adversarial network.
result Y-GAN estimates depth maps from multi-camera stereo images without ground truth data.
Paper tackles robust spatial perception by handling outliers efficiently.
problem Robust spatial perception is challenged by incorrect data association (outliers).
method Proposes adaptive trimming algorithm to remove outliers efficiently.
result Adaptive trimming algorithm outperforms state-of-the-art methods across applications.
Global Planar Convolution boosts brain tumor segmentation by enhancing context perception.
problem Improving context perception in brain tumor segmentation networks.
method Introduced Global Planar Convolution module to enhance context aggregation in brain tumor segmentation.
result Global Planar Convolution eliminates the need for multiple representation levels in segmentation networks.
3D-PRNN generates shapes from depth images using recurrent neural networks.
problem Representing 3D shapes from limited sensor data.
method Generative Recurrent Neural Network (3D-PRNN) with Gaussian Fields.
result 3D-PRNN synthesizes plausible shapes from primitives, outperforming nearest-neighbor methods.
The paper uses facial keypoints to estimate post-surgical pain intensity.
problem Accurately assessing pain levels from self-reported ratings is challenging.
method The approach analyzes 2D and 3D facial keypoints to estimate pain intensity.
result The pain estimation model uses multiple instance learning.
Research evaluates adversarial attacks and defenses on 3D point cloud classifiers.
problem Robustness of 3D object classifiers against adversarial attacks.
method Extending 2D adversarial attacks to 3D point clouds and proposing new defenses.
result 3D point cloud classifiers are weak to adversarial attacks but more defensible.
Recently, multiple formulations of vision problems as probabilistic inversions of generative models based on computer graphics have been proposed. However, applications to 3D perception from natural images have focused on low-dimensional latent scenes, due to challenges in both modeling and inference. Accounting for th…
Improves AI agents' 3D navigation by learning from failures and 3D spatial relationships.
problem Challenges in data efficiency, obstacle avoidance, and generalization in 3D visual navigation.
method Incorporates attention on 3D spatial relationships and a target skill extension module into DRL framework.
result Significantly improves navigation performance and generalization across targets and scenes.
The paper proposes a model to learn motion perception in V1 using vector and matrix representations.
problem Motion perception in primary visual cortex (V1).
method Coupling vector representations of local contents and matrix representations of local pixel displacements.
result The model can learn Gabor-like filter pairs and infer local motions.
Humans take advantage of real world symmetries for various tasks, yet capturing their superb symmetry perception mechanism with a computational model remains elusive. Motivated by a new study demonstrating the extremely high inter-person accuracy of human perceived symmetries in the wild, we have constructed the first …
Redundancy in AI perception systems doesn't guarantee independent error occurrences.
problem Lack of direct statistical evidence of super-human automated driving performance.
method Investigated the effectiveness of redundancy in neural networks for independent error occurrences.
result Errors in neural networks for computer vision tasks are correlated, not independent.
A new model classifies surface anomalies in 3D point cloud data.
problem Accurate classification of surface anomalies in manufacturing processes.
method Deep subspace learning approach for 3D point cloud data.
result The method effectively identifies new types of anomalies.
Waymo Open Dataset provides a large, diverse, and synchronized LiDAR and camera dataset for autonomous driving research.
problem Limited diversity and scale in existing self-driving datasets hinder real-world problem alignment.
method Developed a new large-scale, high-quality, diverse dataset with synchronized LiDAR and camera data.
result The dataset is 15x more diverse than the largest existing dataset based on a proposed diversity metric.
New dataset tests mental rotation from single images, improving model understanding of 3D scenes.
problem Understanding how a scene looks from a different viewpoint using a single image.
method Created CLEVR-MRT dataset, explored neural architectures for volumetric scene representations.
result Demonstrated the effectiveness of volumetric representations in answering mental rotation questions.
PCI combines perception and control using Bayesian inference with object-based representations.
problem Separate perception and control in reinforcement learning.
method Joint Perception and Control as Inference (PCI) framework with Object-based Perception Control (OPC).
result OPC achieves good perceptual grouping quality and outperforms baselines in accumulated rewards.
Study the tradeoff between signal distortion and human perception over finite channels.
problem Characterize the distortion-perception tradeoff for finite channels with arbitrary metrics.
method Solve linear programming problems to compute the distortion-perception function and optimal reconstructions.
result DP function is piecewise linear in the perception index.
One of the key challenges of visual perception is to extract abstract models of 3D objects and object categories from visual measurements, which are affected by complex nuisance factors such as viewpoint, occlusion, motion, and deformations. Starting from the recent idea of viewpoint factorization, we propose a new app…
A framework isolates VQA reasoning from perception for better model evaluation.
problem Improper separation of visual perception and reasoning in VQA models.
method Introducing a framework and a top-down calibration technique to decouple reasoning from perception.
result Improved evaluation of VQA models by separating reasoning from perception.
End-to-end autonomous driving perception learns latent features for better performance.
problem Current autonomous driving systems are complex and require human engineering.
method Sequential latent representation learning for end-to-end perception.
result End-to-end perception model solves detection, tracking, localization, and mapping problems.
Geometric model explains music perception combining neuroscience and acoustics.
problem Rationalize and predict psycho-acoustic phenomena in music perception.
method Combining neuroscientific theories with acoustic observations, a geometric model of the space of all chords is created.
result The geometric model allows for rigorous studies of psychoacoustic quantities like roughness and harmonicity.
Safe control for vehicles using learned perception from images.
problem Controlling autonomous vehicles with partial state information from images.
method Learned perception map and safe set design for a closed loop system.
result Generalization properties of the perception-control loop are favorable.
RETR improves indoor radar perception with a novel transformer model.
problem Indoor radar perception lacks models tailored for multi-view radar settings.
method RETR extends DETR architecture with depth-prioritized feature similarity, tri-plane loss, and radar-to-camera transformation.
result RETR outperforms state-of-the-art methods by 15.38+ AP for object detection and 11.91+ IoU for instance segmentation.
Researchers study how teachers' advising relationships influence their perceptions of satisfaction and students, not policy influence.
problem Understanding the relationship between teachers' advising relationships and their perceptions of satisfaction and students.
method Proposed a novel joint model of network and item responses (JNIRM) with correlated latent variables.
result Teachers' advising relationships contribute more to satisfaction and students than to influence over educational policies.
Study how actions affect perception in embodied agents using group theory.
problem Understanding how actions influence perception in autonomous agents.
method Mathematical formalism of group theory applied to sensory commutativity of action sequences.
result Introduced Sensory Commutativity Probability (SCP) to measure action effects on perception.
Reconstruction-based learning produces uninformative features for perception tasks.
problem Misalignment between reconstruction-based learning and perception tasks.
method Investigated the impact of input space reconstruction on feature learning for perception tasks.
result Reconstruction-based learning allocates model capacity to a subspace with uninformative features for perception tasks.
The study estimates how changing words in sentences affects audience perception.
problem Estimating the causal effect of lexical choice on audience perception.
method Two classes of methods: quasi-experimental designs and classification problems.
result Algorithmic estimates align with randomized-control trials and can be transferred across domains.
Study uses deep learning to predict gender and analyze HPV vaccine perceptions on Twitter.
problem Analyzing gender differences in public perceptions on HPV vaccine using social media data.
method Convolutional neural network model trained on Twitter text for gender prediction, then applied to HPV vaccine related tweets.
result Identified gender differences in public perceptions on HPV vaccine, consistent with previous studies.
Mathematical models link perception and memory formation.
problem Linking perception and memory formation.
method Tensor decompositions and latent representations.
result Active semantic decoding process in perception.
Tensor models decode human perception and memory using SPO triples.
problem Understanding implicit and explicit perception and memory in the brain.
method Tensor models with SPO triples, dual representations, and four layers.
result Semantic memory is crucial for explicit perception and declarative memories.
Approach to develop visual perception in robots through sensorimotor interactions.
problem Developing autonomous perception in robots.
method Sensorimotor contingencies theory applied to robot exploration and learning.
result Captured sensorimotor regularities in a predictive model for visual field discovery.
Study examines perceptions and attitudes about breast cancer on Twitter.
problem Understanding public perceptions and attitudes towards breast cancer on social media.
method Identified and collected tweets, used topic modeling and sentiment analysis.
result Identified themes and quantified users' perceptions and emotions about breast cancer.
PeL separates sensory interface optimization from decision learning.
problem Optimizing sensory interfaces without task-specific information.
method Formal separation of perception and decision learning, using metrics for stability, informativeness, and geometry.
result Updates preserving invariants are orthogonal to decision gradients.
Logical scaffolds enhance AI software quality.
problem Improving AI component quality in software.
method Logical scaffolds as a method to improve AI components.
result Logical scaffolds can improve AI beyond perception systems.
PERCEPT detects changes in high-dimensional data streams using topological data analysis.
problem Detecting changes in high-dimensional data streams, especially when embedded in a low-dimensional space.
method Leverages topological data analysis to learn embedded topology as a point cloud via persistence diagrams, then applies non-parametric monitoring for detecting changes.
result Demonstrates efficient detection of online changes from high-dimensional data streams.
While perception tasks such as visual object recognition and text understanding play an important role in human intelligence, the subsequent tasks that involve inference, reasoning and planning require an even higher level of intelligence. The past few years have seen major advances in many perception tasks using deep …
Default-ERM shortcut learning persists even without additional information.
problem Default-ERM shortcut learning in perception tasks despite stable feature sufficiency.
method Studied linear perception task; developed margin control (MARG-CTRL) loss functions.
result Margin control mitigates shortcut learning on various tasks.
DGP learns speech recognition by modeling complex relationships between utterances.
problem Modeling complex relationships in speech recognition without relational data.
method Bayesian nonparametric deep learning method (DGP) that generates infinite probabilistic graphs.
result DGP successfully infers relationships among utterances without relational data during training.
Model shows neural nets can learn categorical perception.
problem Understanding how categorical perception arises in neural networks.
method Developed a neural network model to simulate learning-induced categorical perception.
result Neural nets can learn to perceive categories in a way similar to human perception.
The study improves deep learning models for safer autonomous vehicles.
problem Robustness of deep neural network models in autonomous driving.
method Analyzes and proposes solutions for deep learning model robustness.
result Enhanced deep learning models for safer autonomous vehicles.
New method allows a robot to perceive space dimensions without prior knowledge.
problem Limitation of previous methods in perceiving space dimensions with small movements.
method Non-linear dimension estimation method.
result Robots can now perceive space dimensions with larger movements.
A technique finds adversarial examples for deep neural networks using human perception.
problem Finding adversarial examples for deep neural networks without access to internal structure.
method Covariance Matrix Adaptation Evolution Strategy (CMA-ES) with perception-in-the-loop.
result CMA-ES can find adversarial examples with human feedback, showing favorable performance.
This study examines LiDAR spoofing attacks on autonomous vehicle perception.
problem Security vulnerabilities in LiDAR-based perception systems in autonomous driving.
method Formulated as an optimization problem, designed input perturbation and objective functions, combined optimization and global sampling.
result Attack success rates improved to around 75% through strategic input perturbation.
Natural speech can be easily manipulated to alter perception.
problem The susceptibility of natural speech to manipulation.
method Investigation of the McGurk effect and Yanny or Laurel illusion.
result A significant fraction of natural speech is illusionable.
This paper evaluates various representations for robotics tasks, improving performance in lifting, stacking, and pushing.
problem Improving data-efficiency in reinforcement learning for robotics with limited data.
method Systematic evaluation of common representations in three robotics tasks: lifting, stacking, and pushing.
result Some representations can perform as well as simulator states as agent inputs, challenging common intuitions.
A robot learns to classify images with limited perception using a layered reinforcement learning approach.
problem Image classification for robots with partial perception.
method Three-layer architecture using deep reinforcement learning, including meta-layer, action-layer, and classification-layer.
result The method achieves high accuracy on the MNIST dataset and provides explainability of the agent's decision-making process.
Robot learns sensorimotor relationships through exploration.
problem Autonomous acquisition of sensorimotor contingencies by robots.
method Developmental framework encoding predictive models of sensorimotor experience.
result Robot discovers the environment, objects, and visual field through internal encoding of sensorimotor contingencies.