KeypointNet learns 3D keypoints for object pose estimation without ground-truth.
problem Learning 3D keypoints for object pose estimation without manual annotations.
method End-to-end geometric reasoning framework to discover keypoints.
result End-to-end framework outperforms fully supervised baseline.
DepthNets learns 3D face geometry and transformations without supervision.
problem Learning 3D face geometry and transformations from a single image.
method Unsupervised learning of facial keypoints depth, using backpropable loss for 3D transformations.
result DepthNets can predict 3D transformations and re-target faces to new poses or geometries.
Method generates multiple 3D poses from 2D joint detections, addressing ambiguity and uncertainty.
problem Ambiguity and uncertainty in 3D human pose estimation from 2D joint detections.
method Generative model, compositional, anatomical constraints, removing model bias.
result Generates multiple valid 3D poses consistent with 2D joint detections.
Paper develops a differentiable approach for 3D imaging models using Fourier slice theorem.
problem Uncertainty in 3D structure modeling and pose estimation in scientific imaging.
method Differentiable probabilistic models in Fourier space with backpropagation through projection.
result Validates approach on 3D protein reconstruction and extends to probabilistic models.
This paper improves 3D pose recovery from 2D images using non-convex regularization.
problem 3D object pose recovery from 2D images.
method Proposes non-convex regularization with leaky capped ℓ1-norm (LCNR) and multi-stage optimization.
result Theoretical analysis shows estimation error decreases with optimization stages.
Paper predicts TUG score from gait characteristics using machine learning.
problem Assessing fall risk in the elderly.
method 3D pose estimation, gait characteristics computation, copula entropy selection, predictive models.
result Effectiveness of the proposed method demonstrated on real-world data.
Geometric Capsule Autoencoders group 3D points into parts and objects.
problem Learning object representations from 3D point clouds.
method Geometric capsules with pose and feature components, Multi-View Agreement voting mechanism.
result Learned representations enable object identification and canonical pose recovery.
Proposes a model to generate 3D-aware images from 2D images.
problem Generating 3D-aware images from 2D images.
method Likelihood-based top-down model using Neural Radiance Fields and energy-based latent variables.
result Model can infer 3D object structures from 2D images and generate novel views.
We introduce SARR for symmetric object pose estimation, improving CNN performance.
problem Ambiguities in symmetric object orientations hinder deep learning pose estimation.
method Numeric rotation representation using symmetry-derived trigonometric identities.
result SARR enables standard CNNs to achieve state-of-the-art performance.
This paper explores the capabilities of convolutional neural networks to deal with a task that is easily manageable for humans: perceiving 3D pose of a human body from varying angles. However, in our approach, we are restricted to using a monocular vision system. For this purpose, we apply a convolutional neural networ…
Generative model creates realistic dance poses from music.
problem Generating human-like dance poses from music.
method Music feature encoder, pose generator, music genre classifier integrated.
result Generative autoregressive model synthesizes dance sequences up to 5,000 frames.
Generative model disentangles 3D shapes into independent factors.
problem Learning rich representations of deformable 3D shapes.
method Supervised 3D mesh-convolutional Variational AutoEncoder with latent feature disentanglement.
result Explicit disentanglement of latent factors improves shape generation and downstream tasks.
New framework predicts diverse, contextually plausible 3D human motions.
problem Predicting multiple plausible future 3D poses given observed poses.
method Developed a new variational framework that conditions latent variable on past observation to encourage relevant information.
result Our approach generates motions of higher quality and preserves contextual information.
The paper proposes a method to learn 3D object pose manifolds using GANs and elasticae.
problem Learning image manifolds of 3D objects with limited data.
method Geom-SGAN and elasticae for geometry-preserving image interpolation.
result The method outperforms state-of-the-art GANs and VAEs in learning rotation paths.
Paper tackles unsupervised learning of 3D shapes from single images.
problem Learning 3D shapes from single images without supervision.
method Generative models, variational auto-encoders, adversarial methods.
result Model learns 3D shapes and poses from single images, showing potential for various datasets.
Scalable approach for object pose estimation across domains.
problem Object pose estimation across different datasets and models.
method Multi-path learning: shared encoder, object-specific decoders.
result Generalizes well from synthetic to real data and across various instances.
CubeNet preserves 3D shape signatures through equivariance.
problem 3D ConvNets fail to capture pose differences.
method Group Convolutional Neural Network with linear equivariance to 3D transformations.
result Achieves state-of-the-art on ModelNet10 classification.
StyleNeRF generates high-resolution images with 3D consistency and style control.
problem Generating high-resolution images with fine details and 3D consistency.
method Integrates NeRF into a style-based generator for efficient high-resolution image synthesis.
result Synthesizes high-resolution images at interactive rates with high 3D consistency and style control.
Unsupervised mesh disentanglement separates identity and pose.
problem Geometric disentanglement for 3D deformable models.
method CFAN-VAE architecture using conformal factor and normal features.
result CFAN-VAE achieves state-of-the-art performance on unsupervised geometric disentanglement.
We consider the problem of learning object arrangements in a 3D scene. The key idea here is to learn how objects relate to human poses based on their affordances, ease of use and reachability. In contrast to modeling object-object relationships, modeling human-object relationships scales linearly in the number of objec…
New method reconstructs 3D shapes from 2D images using Kendall's shape space.
problem Reconstruct 3D shapes from 2D images, especially for rare specimens.
method Kendall's shape space approach with prior information.
result More robust and plausible shapes compared to previous methods.
The field of multiple view geometry has seen tremendous progress in reconstruction and calibration due to methods for extracting reliable point features and key developments in projective geometry. Point features, however, are not available in certain applications and result in unstructured point cloud reconstructions.…
3D capsule network for point clouds handles rotations and translations.
problem Processing 3D point clouds with equivariance and invariance.
method Quaternion equivariant capsule network with dynamic routing on quaternions.
result Parsimonious capsule network disentangles geometry and pose.
Tensor neural network improves human pose classification from 3D skeleton data.
problem Efficiently processing spatiotemporal data for human pose classification.
method Proposes a tensor-based neural network with three components: spatiotemporal feature construction, tensor fusion, and tensor-based neural network processing.
result Achieves state-of-the-art performance in human pose classification.
Paper defines continuous rotation representations for neural networks.
problem Discontinuous representations of rotations in neural networks.
method Definition of continuous representations, relating to topological concepts.
result Continuous representations for 3D rotations in 5D and 6D are more suitable for neural networks.
Recently, multiple formulations of vision problems as probabilistic inversions of generative models based on computer graphics have been proposed. However, applications to 3D perception from natural images have focused on low-dimensional latent scenes, due to challenges in both modeling and inference. Accounting for th…
New robust approach for multivariate regression with corrupted data.
problem Grossly corrupted or missing observations in multivariate linear regression.
method Explicitly considers error source and sparseness, allowing unique noise levels.
result Approach guarantees convergence to optimal solution despite non-smooth optimization.
This work generates synthetic 3D thermal facial data using 2D facial data and deep learning.
problem Creating large datasets for deep learning in computer vision.
method 3D facial modelling techniques and deep learning methodologies.
result Synthetic 3D thermal facial data created for deep learning applications.
Bayesian segmentation and uncertainty estimation improve 3D model accuracy for factory planning.
problem Generating accurate 3D models from outdated and incomplete 2D data.
method Bayesian neural network for point cloud segmentation and entropy-based uncertainty estimation.
result Bayesian segmentation network significantly improves model accuracy and object identification.
LatticeNet segments 3D point clouds faster and more efficiently.
problem Challenges in applying CNNs to 3D point cloud data.
method Embeds point cloud geometry into a permutohedral lattice for fast convolutions.
result Achieves state-of-the-art performance in 3D segmentation.
Researchers extend reparameterization to Lie groups for better probability modeling.
problem Lack of reparameterization for distributions on Lie groups.
method Developed a general framework for creating reparameterizable densities on Lie groups.
result Demonstrated complex and multimodal distributions on SO(3) for pose estimation.
A^2-Net learns to estimate molecular structures from Cryo-EM data.
problem Estimating molecular structures from Cryo-EM density volumes.
method Learning-based approach using 3D detection and pose estimation.
result Achieves 91% coverage on new dataset and is hundreds of times faster.
Novel GNN predicts drug-target interactions using protein-ligand 3D structures.
problem Accurate prediction of drug-target interactions for in silico drug design.
method 3D structure-embedded graph representations and distance-aware graph attention algorithm with gate augmentation.
result Our model outperforms docking and other deep learning methods in virtual screening and pose prediction.
New method embeds RNN Seq2Seq models to visualize spatiotemporal data.
problem Visualizing and interpreting spatiotemporal data in sequence prediction tasks.
method Embedding approach to visualize and interpret RNN Seq2Seq model representations.
result Embedding space projections of RNN Seq2Seq models capture spatiotemporal dynamics.
BPI models 2D patterns on multiple planes and 3D scene from a single image.
problem Understanding and editing images with multiple 2D planes and 3D scene from a single image.
method Box Program Induction (BPI) with neural networks and search-based algorithm.
result Holistic, structured scene representation enables 3D-aware image editing.
The paper learns pose variations within shape populations using constrained mixtures of factor analyzers.
problem Learning pose variations within a shape population with articulated parts and relative rotations.
method Formulated as mixtures of factor analyzers, segmentation by component posterior probabilities, and constraints on factor loading matrices for rotation matrices.
result Automatic learning of pose variations from shape populations, resulting in smooth and realistic animations.
DFKI Cabin Simulator tests visual monitoring functions in vehicles.
problem Validating novel human-vehicle interfaces and driver assistance systems.
method Driving simulator with in-cabin mock-up and camera system.
result Validation of in-cabin monitoring functions for advanced driver assistance and automated driving.
Generative models learn 3D scenes without explicit maps.
problem Learning models for visual 3D localization without explicit maps.
method Generative Query Networks (GQNs) with attention mechanisms.
result GQNs can capture complex 3D scenes and perform localization.
Proposes PSGAN for generating high-res anime images with structural consistency.
problem Lack of high-quality, structurally consistent full-body high-resolution anime images.
method Progressive Structure-conditional Generative Adversarial Networks (PSGAN) with progressive training.
result Demonstrates effectiveness through comparisons and diverse anime character generation.
Proves Borel Conjecture for certain 3D spaces.
problem Characterizing fundamental groups of 3D Alexandrov spaces.
method Analyzes properties of Alexandrov 3-spaces.
result Proves Borel Conjecture for specific types of spaces.
Paper proposes Roweisposes for 3D action recognition using generalized eigenvalue problem.
problem Need for basic methods in 3D action recognition.
method Roweisposes uses Roweis discriminant analysis for generalized subspace learning.
result Roweisposes is effective for 3D action recognition.
Extracts airways from 3D CT data using graph refinement methods.
problem Extracting airway trees from volumetric CT data.
method Two methods: mean-field approximation and graph neural networks.
result Both methods improve airway detection accuracy.
Visualizes 3D CNNs for protein-ligand scoring.
problem Interpreting complex neural network decisions for protein-ligand scoring.
method Three visualization methods for 3D CNNs, including filters and weights.
result Visualizations aid in tuning and designing neural networks.
CNN scoring function predicts protein-ligand interactions.
problem Scoring protein-ligand interactions for drug discovery.
method Convolutional Neural Networks (CNN) for 3D protein-ligand interactions.
result CNN scoring function outperforms AutoDock Vina in ranking poses.
Morph accelerates 3D CNNs for video recognition, reducing energy consumption and improving performance.
problem Efficiently accelerating 3D CNNs for video recognition is challenging due to their large memory footprint and higher dimensionality.
method Designing a flexible accelerator called Morph that adapts to different spatial and temporal tiling strategies, and codesigning a software infrastructure to control the hardware.
result Morph achieves up to 3.4x reduction in energy consumption and up to 5.1x improvement in performance/watt compared to a baseline 3D CNN accelerator.
Improved QSM maps from MRI using deep learning.
problem Inaccurate susceptibility maps due to ill-posed dipole inversion.
method 3D GAN with increased receptive field and WGAN with gradient penalty.
result Significantly better QSM maps from single orientation phase maps.
Trans-Unet predicts brain folding patterns from 3D point-clouds using novel 3D-to-2D transformation.
problem Challenges in learning high-fidelity 3D point-cloud features, including permutation invariance and fine-grained surface reconstruction.
method Transform 3D point-clouds into a 2D grid domain, then use a U-shaped hybrid model with CNNs and self-attention mechanisms.
result Trans-Unet achieves high-resolution predictions of brain patch growth, surpassing existing methods in fidelity and accuracy.
Paper uses FMCW radar and FCN for object detection and 3D estimation.
problem Object detection and 3D estimation using FMCW radar.
method Employed deep learning (FCN) over traditional signal processing. Normalization method applied to radar signal.
result System successfully detects and estimates 3D position of objects in noisy environments.