Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

8172533 · Oct 201919922001200920182026
48 results for Camera motion

V-SysId identifies keypoints and 3D system from unlabeled videos.

problem Identifying keypoints and 3D system from unlabeled videos.
method Alternates between parameter estimation and extrinsic camera calibration, using motion equations as weak supervision.
result Utility of the approach demonstrated across various settings.

New method for separating foreground from background in noisy, moving camera video.

problem Foreground-background separation in noisy, free-moving camera video.
method Registers frames, encodes perspective as missing data, uses OptShrink for low-rank estimation, and weighted total variation for smooth foreground.
result Panoramic background component that stitches together corrupted data from overlapping frames.

The field of multiple view geometry has seen tremendous progress in reconstruction and calibration due to methods for extracting reliable point features and key developments in projective geometry. Point features, however, are not available in certain applications and result in unstructured point cloud reconstructions.…

2016-04-27abs ↗pdf ↗

Y-GAN uses multi-camera data to estimate depth maps without expensive hardware.

problem Depth perception for autonomous systems requires accurate 3D spatial information.
method Proposes Y-GAN, a deep convolutional generative adversarial network.
result Y-GAN estimates depth maps from multi-camera stereo images without ground truth data.

We present a model for the joint estimation of disparity and motion. The model is based on learning about the interrelations between images from multiple cameras, multiple frames in a video, or the combination of both. We show that learning depth and motion cues, as well as their combinations, from data is possible wit…

2013-12-12abs ↗pdf ↗

Self-supervised method estimates distances on fisheye cameras for autonomous driving.

problem Accurate Euclidean distance estimation on fisheye cameras for autonomous driving.
method Self-supervised scale-aware framework for monocular fisheye videos.
result State-of-the-art results on KITTI dataset, comparable to other methods.

Paper tackles activity recognition from body-worn video footage.

problem Classifying frames of body-worn video footage according to the wearer's activity.
method Extract motion features and semi-supervised classification.
result Method achieves comparable results to supervised and deep learning methods using less training data.

PIP-Net predicts pedestrian crossing intentions with up to 4-second lead.

problem Accurate pedestrian intention prediction for autonomous vehicles in real-world scenarios.
method Recurrent and temporal attention-based model using kinematic and spatial features.
result PIP-Net predicts pedestrian crossing intentions up to 4 seconds in advance.

New approach for camera-specific color constancy using few-shot meta-learning.

problem Domain gaps and lack of generalization across different cameras.
method Formulates color constancy as few-shot meta-learning tasks, leveraging annotated samples across different cameras.
result Significant reduction in data collection time and improved generalization to new cameras.

This paper surveys DRL for autonomous vehicle motion planning.

problem Designing intelligent motion planning for autonomous vehicles.
method Deep Reinforcement Learning (DRL) for hierarchical motion planning.
result Survey of state-of-the-art DRL solutions for autonomous vehicle motion planning.

Study tackles open-set camera model identification, improving over state-of-the-art.

problem Identifying camera models from unknown ones in open-set scenarios.
method Feature extraction algorithms and classifiers for open-set recognition, evaluating different training protocols.
result A simple open-set training protocol yields the best results, improving over state-of-the-art solutions.

Paper introduces TAP-Vid, a benchmark for tracking any point in videos.

problem Tackles the problem of tracking arbitrary physical points on surfaces over longer video clips.
method Formalizes the problem as TAP, introduces TAP-Vid benchmark, uses crowdsourced pipeline with optical flow estimates, proposes TAP-Net model.
result TAP-Net outperforms all prior methods on TAP-Vid benchmark when trained on synthetic data.

SoilingNet detects soiling on automotive cameras for better autonomous driving performance.

problem Soiling degrades the performance of automotive surround-view cameras, affecting autonomous driving.
method Created a new dataset and used a Convolutional Neural Network (CNN) architecture for soiling detection, combined with multi-task learning and data augmentation using GANs.
result Demonstrated high accuracy in soiling detection using CNN and multi-task learning.

Paper proposes an edge detection method for robot navigation using low-SNR thermal cameras.

problem Efficient edge detection for robot navigation using low-SNR thermal camera.
method Raw image denoising, Canny edge detection, CSS method, edge ranking, edge linking.
result Enhanced edge detection method effectively detects smooth edges of the surrounding environment.

Sparse subspace clustering (SSC) is an elegant approach for unsupervised segmentation if the data points of each cluster are located in linear subspaces. This model applies, for instance, in motion segmentation if some restrictions on the camera model hold. SSC requires that problems based on the l1l_1-norm are solved …

2016-09-16abs ↗pdf ↗

A method for camera calibration using heatmap regression for fisheye images.

problem Accurate and robust camera angle estimation from fisheye images in the Manhattan world.
method Heatmap regression to detect directions of labeled image coordinates, simultaneous rotation and fisheye distortion recovery.
result Our method outperforms conventional methods on large-scale datasets and with off-the-shelf cameras.

Industrial vehicle uses LiDAR and camera to detect and avoid restricted areas.

problem Avoiding collisions in industrial settings with automated vehicles.
method Combining LiDAR and camera data, using deep learning for projection, and model-predictive control.
result Reduces false positives in LiDAR detection of reflective beacons.

In a range of fields including the geosciences, molecular biology, robotics and computer vision, one encounters problems that involve random variables on manifolds. Currently, there is a lack of flexible probabilistic models on manifolds that are fast and easy to train. We define an extremely flexible class of exponent…

2015-05-17abs ↗pdf ↗

Novel algorithm for separating moving camera video into static and dynamic components.

problem Foreground-background separation in noisy, moving camera video.
method Augmented robust PCA with total variation regularization, OptShrink low-rank matrix estimator.
result Panoramic low-rank component spanning entire field of view, automatically stitching corrupted data.

Camera stickers can fool deep learning systems by manipulating the lens, achieving 49.6% misclassification rate.

problem The vulnerability of deep learning systems to physical adversarial attacks.
method Iterative procedure to update attack perturbation and threat model for physical realizability.
result Achieved 49.6% misclassification rate for targeted attacks on ImageNet classifiers.

A drone-based MOT algorithm tracks vehicles using neural network detections and TPMBM filter.

problem Tracking multiple vehicles from drone-mounted cameras.
method Neural network for object detection, TPMBM filter for trajectory estimation, von-Mises Fisher distribution for DOA.
result TPMBM filter optimally estimates vehicle trajectories.

Paper uses VAEs and GANs to estimate cryo-EM image orientation and camera parameters.

problem Estimating orientation and camera parameters from noisy cryo-EM images.
method Combines VAEs and GANs to learn latent representation, then designs estimation method.
result Geometric approach for fast cryo-EM biomolecule reconstruction.

Waymo Open Dataset provides a large, diverse, and synchronized LiDAR and camera dataset for autonomous driving research.

problem Limited diversity and scale in existing self-driving datasets hinder real-world problem alignment.
method Developed a new large-scale, high-quality, diverse dataset with synchronized LiDAR and camera data.
result The dataset is 15x more diverse than the largest existing dataset based on a proposed diversity metric.

Paper proposes a method to improve semantic segmentation for fisheye urban driving images.

problem Semantic segmentation for fisheye urban driving images is challenging due to distortion and lack of large datasets.
method A seven degrees of freedom augmentation method is proposed to transform rectilinear images into fisheye images.
result Training with seven-DoF augmentation improves model accuracy and robustness against distorted fisheye data.

Recent progress on many imaging and vision tasks has been driven by the use of deep feed-forward neural networks, which are trained by propagating gradients of a loss defined on the final output, back through the network up to the first layer that operates directly on the image. We propose back-propagating one step fur…

2016-05-23abs ↗pdf ↗

Generative Map learns interpretable neural network maps for camera localization.

problem Creating interpretable maps for neural network-based camera localization.
method Combining generative models with Kalman filters and incorporating additional sensor information.
result Generative Map predicts images closely resembling the true scene and achieves comparable localization performance.

We integrate camera pose correlations into deep models using Gaussian processes.

problem Lack of inter-frame reasoning in deep neural networks.
method Derive a principled framework combining camera pose information with deep models using a novel view kernel.
result Soft-prior knowledge aids pose-related vision tasks like novel view synthesis.

Person re-identification (re-id), an emerging problem in visual surveillance, deals with maintaining entities of individuals whilst they traverse various locations surveilled by a camera network. From a visual perspective re-id is challenging due to significant changes in visual appearance of individuals in cameras wit…

2014-06-13abs ↗pdf ↗