This paper presents KeypointNet, an end-to-end geometric reasoning framework to learn an optimal set of category-specific 3D keypoints, along with their detectors. Given a single image, KeypointNet extracts 3D keypoints that are optimized for a downstream task. We demonstrate this framework on 3D pose estimation by pro…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
V-SysId identifies keypoints and 3D system from unlabeled videos.
KINet learns object interactions without supervision for robotic pushing.
The paper uses facial keypoints to estimate post-surgical pain intensity.
We present a simple, flexible, and general framework titled Partial Registration Network (PRNet), for partial-to-partial point cloud registration. Inspired by recently-proposed learning-based methods for registration, we use deep networks to tackle non-convexity of the alignment and partial correspondence problems. Whi…
Detect facial keypoints is a critical element in face recognition. However, there is difficulty to catch keypoints on the face due to complex influences from original images, and there is no guidance to suitable algorithms. In this paper, we study different algorithms that can be applied to locate keyponits. Specifical…
This paper introduces a novel deep learning framework for image animation. Given an input image with a target object and a driving video sequence depicting a moving object, our framework generates a video in which the target object is animated according to the driving sequence. This is achieved through a deep architect…
We present an unsupervised approach for learning to estimate three dimensional (3D) facial structure from a single image while also predicting 3D viewpoint transformations that match a desired pose and facial geometry. We achieve this by inferring the depth of facial keypoints of an input image in an unsupervised manne…
End-to-end trainable graph matching using improved combinatorial solvers.
New mutual information framework improves contrastive learning for vision tasks.
New method improves human mesh recovery for obese people.
Model discovers causal relationships from video data of physical systems.
Simple methods improve regression transferability estimation.
We propose Progressive Structure-conditional Generative Adversarial Networks (PSGAN), a new framework that can generate full-body and high-resolution character images based on structural information. Recent progress in generative adversarial networks with progressive training has made it possible to generate high-resol…
We tackle here the problem of multimodal image non-rigid registration, which is of prime importance in remote sensing and medical imaging. The difficulties encountered by classical registration approaches include feature design and slow optimization by gradient descent. By analyzing these methods, we note the significa…
Meta Omnium benchmarks few-shot learning across diverse vision tasks.
Prediction and interpolation for long-range video data involves the complex task of modeling motion trajectories for each visible object, occlusions and dis-occlusions, as well as appearance changes due to viewpoint and lighting. Optical flow based techniques generalize but are suitable only for short temporal ranges. …
Quantum machine learning improves satellite image alignment.
Improved visual representation learning with conditional negative sampling.
A method for camera calibration using heatmap regression for fisheye images.
The explanation of the photoelectric effect by Einstein and Maxwell's field theory of electromagnetism have motivated De Broglie to make the hypothesis that matter exhibits both waves and particles like-properties. These representations of matter are enlightened by string theory which represents particles with stringli…
Introduces Motion Programs for better video analysis of human motion.
A new method for spotting symbols in CAD images reduces annotation costs and improves accuracy.
The paper proposes blending gradient boosted trees and neural networks for hierarchical time series forecasting.
The field of multiple view geometry has seen tremendous progress in reconstruction and calibration due to methods for extracting reliable point features and key developments in projective geometry. Point features, however, are not available in certain applications and result in unstructured point cloud reconstructions.…