Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

135270404539 · Jun 202019922001200920182026
48 results for 3D pose estimation

DepthNets learns 3D face geometry and transformations without supervision.

problem Learning 3D face geometry and transformations from a single image.
method Unsupervised learning of facial keypoints depth, using backpropable loss for 3D transformations.
result DepthNets can predict 3D transformations and re-target faces to new poses or geometries.

Method generates multiple 3D poses from 2D joint detections, addressing ambiguity and uncertainty.

problem Ambiguity and uncertainty in 3D human pose estimation from 2D joint detections.
method Generative model, compositional, anatomical constraints, removing model bias.
result Generates multiple valid 3D poses consistent with 2D joint detections.

Paper develops a differentiable approach for 3D imaging models using Fourier slice theorem.

problem Uncertainty in 3D structure modeling and pose estimation in scientific imaging.
method Differentiable probabilistic models in Fourier space with backpropagation through projection.
result Validates approach on 3D protein reconstruction and extends to probabilistic models.

This paper improves 3D pose recovery from 2D images using non-convex regularization.

problem 3D object pose recovery from 2D images.
method Proposes non-convex regularization with leaky capped ℓ1-norm (LCNR) and multi-stage optimization.
result Theoretical analysis shows estimation error decreases with optimization stages.

Proposes a model to generate 3D-aware images from 2D images.

problem Generating 3D-aware images from 2D images.
method Likelihood-based top-down model using Neural Radiance Fields and energy-based latent variables.
result Model can infer 3D object structures from 2D images and generate novel views.

We introduce SARR for symmetric object pose estimation, improving CNN performance.

problem Ambiguities in symmetric object orientations hinder deep learning pose estimation.
method Numeric rotation representation using symmetry-derived trigonometric identities.
result SARR enables standard CNNs to achieve state-of-the-art performance.

This paper explores the capabilities of convolutional neural networks to deal with a task that is easily manageable for humans: perceiving 3D pose of a human body from varying angles. However, in our approach, we are restricted to using a monocular vision system. For this purpose, we apply a convolutional neural networ…

2016-08-31abs ↗pdf ↗

New framework predicts diverse, contextually plausible 3D human motions.

problem Predicting multiple plausible future 3D poses given observed poses.
method Developed a new variational framework that conditions latent variable on past observation to encourage relevant information.
result Our approach generates motions of higher quality and preserves contextual information.

StyleNeRF generates high-resolution images with 3D consistency and style control.

problem Generating high-resolution images with fine details and 3D consistency.
method Integrates NeRF into a style-based generator for efficient high-resolution image synthesis.
result Synthesizes high-resolution images at interactive rates with high 3D consistency and style control.

We consider the problem of learning object arrangements in a 3D scene. The key idea here is to learn how objects relate to human poses based on their affordances, ease of use and reachability. In contrast to modeling object-object relationships, modeling human-object relationships scales linearly in the number of objec…

2012-06-27abs ↗pdf ↗

The field of multiple view geometry has seen tremendous progress in reconstruction and calibration due to methods for extracting reliable point features and key developments in projective geometry. Point features, however, are not available in certain applications and result in unstructured point cloud reconstructions.…

2016-04-27abs ↗pdf ↗

Tensor neural network improves human pose classification from 3D skeleton data.

problem Efficiently processing spatiotemporal data for human pose classification.
method Proposes a tensor-based neural network with three components: spatiotemporal feature construction, tensor fusion, and tensor-based neural network processing.
result Achieves state-of-the-art performance in human pose classification.

Recently, multiple formulations of vision problems as probabilistic inversions of generative models based on computer graphics have been proposed. However, applications to 3D perception from natural images have focused on low-dimensional latent scenes, due to challenges in both modeling and inference. Accounting for th…

2014-07-04abs ↗pdf ↗

New robust approach for multivariate regression with corrupted data.

problem Grossly corrupted or missing observations in multivariate linear regression.
method Explicitly considers error source and sparseness, allowing unique noise levels.
result Approach guarantees convergence to optimal solution despite non-smooth optimization.

This work generates synthetic 3D thermal facial data using 2D facial data and deep learning.

problem Creating large datasets for deep learning in computer vision.
method 3D facial modelling techniques and deep learning methodologies.
result Synthetic 3D thermal facial data created for deep learning applications.

Bayesian segmentation and uncertainty estimation improve 3D model accuracy for factory planning.

problem Generating accurate 3D models from outdated and incomplete 2D data.
method Bayesian neural network for point cloud segmentation and entropy-based uncertainty estimation.
result Bayesian segmentation network significantly improves model accuracy and object identification.

Novel GNN predicts drug-target interactions using protein-ligand 3D structures.

problem Accurate prediction of drug-target interactions for in silico drug design.
method 3D structure-embedded graph representations and distance-aware graph attention algorithm with gate augmentation.
result Our model outperforms docking and other deep learning methods in virtual screening and pose prediction.

New method embeds RNN Seq2Seq models to visualize spatiotemporal data.

problem Visualizing and interpreting spatiotemporal data in sequence prediction tasks.
method Embedding approach to visualize and interpret RNN Seq2Seq model representations.
result Embedding space projections of RNN Seq2Seq models capture spatiotemporal dynamics.

The paper learns pose variations within shape populations using constrained mixtures of factor analyzers.

problem Learning pose variations within a shape population with articulated parts and relative rotations.
method Formulated as mixtures of factor analyzers, segmentation by component posterior probabilities, and constraints on factor loading matrices for rotation matrices.
result Automatic learning of pose variations from shape populations, resulting in smooth and realistic animations.

DFKI Cabin Simulator tests visual monitoring functions in vehicles.

problem Validating novel human-vehicle interfaces and driver assistance systems.
method Driving simulator with in-cabin mock-up and camera system.
result Validation of in-cabin monitoring functions for advanced driver assistance and automated driving.

Proposes PSGAN for generating high-res anime images with structural consistency.

problem Lack of high-quality, structurally consistent full-body high-resolution anime images.
method Progressive Structure-conditional Generative Adversarial Networks (PSGAN) with progressive training.
result Demonstrates effectiveness through comparisons and diverse anime character generation.

Paper proposes Roweisposes for 3D action recognition using generalized eigenvalue problem.

problem Need for basic methods in 3D action recognition.
method Roweisposes uses Roweis discriminant analysis for generalized subspace learning.
result Roweisposes is effective for 3D action recognition.

Morph accelerates 3D CNNs for video recognition, reducing energy consumption and improving performance.

problem Efficiently accelerating 3D CNNs for video recognition is challenging due to their large memory footprint and higher dimensionality.
method Designing a flexible accelerator called Morph that adapts to different spatial and temporal tiling strategies, and codesigning a software infrastructure to control the hardware.
result Morph achieves up to 3.4x reduction in energy consumption and up to 5.1x improvement in performance/watt compared to a baseline 3D CNN accelerator.

Trans-Unet predicts brain folding patterns from 3D point-clouds using novel 3D-to-2D transformation.

problem Challenges in learning high-fidelity 3D point-cloud features, including permutation invariance and fine-grained surface reconstruction.
method Transform 3D point-clouds into a 2D grid domain, then use a U-shaped hybrid model with CNNs and self-attention mechanisms.
result Trans-Unet achieves high-resolution predictions of brain patch growth, surpassing existing methods in fidelity and accuracy.

Paper uses FMCW radar and FCN for object detection and 3D estimation.

problem Object detection and 3D estimation using FMCW radar.
method Employed deep learning (FCN) over traditional signal processing. Normalization method applied to radar signal.
result System successfully detects and estimates 3D position of objects in noisy environments.