Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

57114170227 · May 202619922001200920172026
48 results for Vision-based Control

A framework disentangles controllable objects from visual signals for improved RL.

problem Improving sample efficiency and game performance in vision-based RL.
method Action-conditioned video prediction to disentangle controllable objects.
result Improved sample efficiency and game performance in Atari games.

Weakly-supervised RL identifies meaningful tasks, improving performance in complex environments.

problem Learning to efficiently explore and distinguish between meaningful and irrelevant tasks.
method Weak supervision to automatically disentangle meaningful tasks from a large space of nonsensical tasks.
result The learned subspace of meaningful tasks leads to substantial performance gains, especially in complex environments.

Modern vision-based reinforcement learning techniques often use convolutional neural networks (CNN) as universal function approximators to choose which action to take for a given visual input. Until recently, CNNs have been treated like black-box functions, but this mindset is especially dangerous when used for control…

2018-09-14abs ↗pdf ↗

We propose the use of Bayesian networks, which provide both a mean value and an uncertainty estimate as output, to enhance the safety of learned control policies under circumstances in which a test-time input differs significantly from the training set. Our algorithm combines reinforcement learning and end-to-end imita…

2018-03-27abs ↗pdf ↗

End-to-end autonomous driving models are vulnerable to simple physical manipulations of images.

problem Vulnerability of end-to-end autonomous driving models to subtle adversarial manipulations of images.
method Developed novel end-to-end attacks using simple physical manipulations (painting black lines on the road) and used Bayesian Optimization to efficiently search for successful attacks.
result Simple physical manipulations can cause autonomous driving models to follow unintended paths, highlighting the vulnerability of these models.

This study uses reinforcement learning to mitigate imminent collisions by controlling car speed and direction.

problem Mitigating imminent collisions on roads.
method Constructed a model using camera images to predict obstacle dynamics. Trained reinforcement learning policies to control braking and steering.
result Both reinforcement learning policies outperform a baseline policy, with the injury model-based policy showing the highest performance.

VNLA uses vision and language to guide agents in finding objects in indoor environments.

problem Guiding agents in finding objects in indoor environments via language.
method Developed I3L framework for imitation learning with indirect intervention.
result Significantly improved success rate of learning agents over baselines.

We use reinforcement learning (RL) to learn dexterous in-hand manipulation policies which can perform vision-based object reorientation on a physical Shadow Dexterous Hand. The training is performed in a simulated environment in which we randomize many of the physical properties of the system like friction coefficients…

2018-08-01abs ↗pdf ↗

Hierarchical Foresight improves robot vision tasks by planning long-term goals.

problem Compounding uncertainty and scalability issues in long horizon video prediction.
method Subgoal generation and planning using hierarchical visual foresight (HVF).
result Achieves nearly 200% performance improvement in vision-based manipulation tasks.

Framework transfers limited steering angle data across multiple weather conditions.

problem Limited labeled data for diverse weather conditions in sensorimotor control.
method Teacher-student learning paradigm with image-to-image translation network.
result Framework generalizes well across multiple weather conditions using limited labels.

FedVision uses federated learning to improve object detection without transmitting data.

problem Challenges in building object detection models on large training datasets due to privacy and cost issues.
method Federated learning (FL) platform for easy integration by non-experts.
result Significant efficiency improvement and cost reduction in smart city applications.

Enhances robotic grasping efficiency with learning-adaptive imagination.

problem Improving sample efficiency and performance in robotic grasping tasks.
method Learning-adaptive imagination approach using ensemble of local dynamics models in latent space.
result Significantly improves sample efficiency and achieves near-optimal performance.

Hybrid RL combines simulated and real data for robust autonomous flight.

problem Challenges in training deep RL models for real-world robotic tasks.
method Combines real-world and simulated data to improve generalization.
result Quadrotor avoids collisions using only a monocular camera.

V-SysId identifies keypoints and 3D system from unlabeled videos.

problem Identifying keypoints and 3D system from unlabeled videos.
method Alternates between parameter estimation and extrinsic camera calibration, using motion equations as weak supervision.
result Utility of the approach demonstrated across various settings.

This paper improves robot grasping by integrating meta-control and latent-space imagination.

problem Dual-system approaches fail to consider the reliability of the learned model when making multiple-step predictions.
method A meta-controller arbitrates between model-based and model-free decisions based on local reliability, encouraging actions that improve the model and generating imagined experiences for additional training.
result Our approach learns near-optimal grasping policies in dense- and sparse-reward environments, outperforming baseline and state-of-the-art methods.

Paper establishes baselines for offline RL from visual observations.

problem Challenges in offline reinforcement learning from visual observations with continuous action spaces.
method Simple baselines and benchmarking tasks for offline RL from visual observations.
result Simple modifications to existing online RL algorithms outperform existing offline RL methods.

We present an approach for building an active agent that learns to segment its visual observations into individual objects by interacting with its environment in a completely self-supervised manner. The agent uses its current segmentation model to infer pixels that constitute objects and refines the segmentation model …

2018-06-21abs ↗pdf ↗

This work proposes a RL approach to learn versatile robotic manipulation tasks.

problem Challenging manipulation tasks in robotics and vision.
method Reinforcement learning (RL) to combine primitive skills, no intermediate rewards, few demonstrations, and efficient skill learning.
result Versatile robotic manipulation in challenging settings with temporary occlusions and dynamic scene changes.

New method converts natural language commands into reward functions for robots.

problem Creating effective reward functions for autonomous machines.
method Language-conditioned reward learning (LC-RL) using inverse reinforcement learning.
result Model learns transferable rewards from natural language commands.

This work applies deep learning to bio-sensing and video data for affective computing.

problem Lack of deep learning integration in bio-sensing for affective computing.
method Novel deep-learning-based methods applied to EEG, ECG, and video data.
result Outperforms other studies in emotion/valence/arousal/liking classification.

Paper presents a method for efficient robot adaptation using fine-tuning.

problem Continuous adaptation of robot learning systems in real-world scenarios.
method Fine-tuning previously learned policies using off-policy reinforcement learning.
result Fine-tuning leads to substantial performance gains and adaptation to new conditions.

Deep reinforcement learning, applied to vision-based problems like Atari games, maps pixels directly to actions; internally, the deep neural network bears the responsibility of both extracting useful information and making decisions based on it. By separating the image processing from decision-making, one could better …

2018-06-04abs ↗pdf ↗

We find ways to make physical signals misclassified by computer vision models.

problem Vulnerability of signal classifiers to adversarial perturbations in physical signals.
method Solving PDE-constrained optimization problems to construct imperceptible perturbations.
result Effective and physically realizable adversarial perturbations can be computed for machine learning models.

Proposes novel method for detecting novel scenarios in autonomous systems.

problem Detecting when a machine learning model makes a trustworthy prediction in dynamic, real-world situations.
method Leverages trained model's learned information and a new image similarity metric.
result Demonstrates the method's efficacy on real-world driving and indoor racing datasets.

We consider the problem of learning multi-stage vision-based tasks on a real robot from a single video of a human performing the task, while leveraging demonstration data of subtasks with other objects. This problem presents a number of major challenges. Video demonstrations without teleoperation are easy for humans to…

2018-10-25abs ↗pdf ↗

Deep transfer learning improves malware classification speed and accuracy.

problem Static malware classification accuracy and speed.
method Transfer learning from computer vision to static malware detection.
result Our method outperforms classical machine learning methods in accuracy, false positive rate, true positive rate, and F1 score.