Scalable approach for object pose estimation across domains.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We introduce SARR for symmetric object pose estimation, improving CNN performance.
Capsule models enforce object pose relationships for robustness, explored with probabilistic generative and variational methods.
This paper presents KeypointNet, an end-to-end geometric reasoning framework to learn an optimal set of category-specific 3D keypoints, along with their detectors. Given a single image, KeypointNet extracts 3D keypoints that are optimized for a downstream task. We demonstrate this framework on 3D pose estimation by pro…
Paper introduces EnDKF for more accurate pose tracking.
Scientific imaging techniques such as optical and electron microscopy and computed tomography (CT) scanning are used to study the 3D structure of an object through 2D observations. These observations are related to the original 3D object through orthogonal integral projections. For common 3D reconstruction algorithms, …
Recognizing an object's material can inform a robot on the object's fragility or appropriate use. To estimate an object's material during manipulation, many prior works have explored the use of haptic sensing. In this paper, we explore a technique for robots to estimate the materials of objects using spectroscopy. We d…
We study the problem of forecasting volatility for the multifractal random walk model. In order to avoid the ill posed problem of estimating the correlation length T of the model, we introduce a limiting object defined in a quotient space; formally, this object is an infinite range logvolatility. For this object and th…
Geometric Capsule Autoencoders group 3D points into parts and objects.
IAGAN method improves medical image reconstruction by incorporating adaptive GAN priors.
New CNN architecture improves pediatric image segmentation by homogenizing pose and size.
We consider the problem of learning object arrangements in a 3D scene. The key idea here is to learn how objects relate to human poses based on their affordances, ease of use and reachability. In contrast to modeling object-object relationships, modeling human-object relationships scales linearly in the number of objec…
In the last decade, supervised deep learning approaches have been extensively employed in visual odometry (VO) applications, which is not feasible in environments where labelled data is not abundant. On the other hand, unsupervised deep learning approaches for localization and mapping in unknown environments from unlab…
The paper proposes a method to learn 3D object pose manifolds using GANs and elasticae.
P3I learns holistic scene representations from a single image.
Proposes a model to generate 3D-aware images from 2D images.
Paper tackles in-bed pressure-based pose estimation, improving accuracy.
Perception is fundamental to many robot application areas especially in service robotics. Our aim is to perceive and model an unprepared kitchen scenario with many objects. We start with the perception of a single target object. The modeling relies especially on fusing and merging of weak information from the sensors o…
There are many forms of feature information present in video data. Principle among them are object identity information which is largely static across multiple video frames, and object pose and style information which continuously transforms from frame to frame. Most existing models confound these two types of represen…
Novel approach to OT using kernel mean embeddings controls overfitting and achieves dimension-free sample complexity.
The paper analyzes the observability of relative pose estimation using dual quaternions.
In this work, we propose a novel node splitting method for regression trees and incorporate it into the regression forest framework. Unlike traditional binary splitting, where the splitting rule is selected from a predefined set of binary splitting rules via trial-and-error, the proposed node splitting method first fin…
Objects are composed of a set of geometrically organized parts. We introduce an unsupervised capsule autoencoder (SCAE), which explicitly uses geometric relationships between parts to reason about objects. Since these relationships do not depend on the viewpoint, our model is robust to viewpoint changes. SCAE consists …
Non-parametric approaches for analyzing network data based on exchangeable graph models (ExGM) have recently gained interest. The key object that defines an ExGM is often referred to as a graphon. This non-parametric perspective on network modeling poses challenging questions on how to make inference on the graphon und…
In this paper we propose novel Deformable Part Networks (DPNs) to learn {\em pose-invariant} representations for 2D object recognition. In contrast to the state-of-the-art pose-aware networks such as CapsNet \cite{sabour2017dynamic} and STN \cite{jaderberg2015spatial}, DPNs can be naturally {\em interpreted} as an effi…
DIVE learns video representations even with missing data.
Paper extends causal inference to non-Euclidean data like images and distributions.
Geometric variations of objects, which do not modify the object class, pose a major challenge for object recognition. These variations could be rigid as well as non-rigid transformations. In this paper, we design a framework for training deformable classifiers, where latent transformation variables are introduced, and …
For recovering 3D object poses from 2D images, a prevalent method is to pre-train an over-complete dictionary of 3D basis poses. During testing, the detected 2D pose is matched to dictionary by where , by estimating the rotation , pr…
The paper develops efficient estimators for semi-parametric binary models in distributed computing.
Augment small datasets with synthetic backgrounds to train lightweight CNNs for human pose estimation.
ConquerNet smooths quantile regression for deep learning with minimax guarantees.
Maximum likelihood estimation fails to be well-posed in Gaussian process regression.
Robotic grasping improved using evolutionary computing and deep reinforcement learning.
We consider the problem of estimating human pose and trajectory by an aerial robot with a monocular camera in near real time. We present a preliminary solution whose distinguishing feature is a dynamic classifier selection architecture. In our solution, each video frame is corrected for perspective using projective tra…
Paper optimizes estimation of quadratic functionals in nonparametric IV models.
The paper proposes a new system ID method from noisy data.
Paper examines convergence rate of PGD for BP objective in inverse problems.
Improved UAV navigation and landing using deep learning.
The assessment of Parkinson's disease (PD) poses a significant challenge as it is influenced by various factors which lead to a complex and fluctuating symptom manifestation. Thus, a frequent and objective PD assessment is highly valuable for effective health management of people with Parkinson's disease (PwP). Here, w…
Directly estimates Fisher score for likelihood maximization.
We present an unsupervised approach for learning to estimate three dimensional (3D) facial structure from a single image while also predicting 3D viewpoint transformations that match a desired pose and facial geometry. We achieve this by inferring the depth of facial keypoints of an input image in an unsupervised manne…
Study tackles inverse problems on low-dimensional manifolds, proving stability and proposing a reconstruction algorithm.
Exponentially fast SMF algorithm for multi-class classification.
New method for adaptive estimation and inference in econometric models without knowing smoothness.
Reparameterizable densities are an important way to learn probability distributions in a deep learning setting. For many distributions it is possible to create low-variance gradient estimators by utilizing a `reparameterization trick'. Due to the absence of a general reparameterization trick, much research has recently…
New analysis reveals masked self-supervised learning's effectiveness in extracting data structure.
A key step to driver safety is to observe the driver's activities with the face being a key step in this process to extracting information such as head pose, blink rate, yawns, talking to passenger which can then help derive higher level information such as distraction, drowsiness, intent, and where they are looking. I…