Paper tackles novel object recognition by improving hierarchical classification.
problem Challenges in recognizing novel object classes unseen during training.
method Proposes top-down and flatten methods for hierarchical novelty detection.
result Generates a hierarchical embedding leading to improved zero-shot learning performance.
OBSER framework infers sub-environments from objects, outperforming scene-based methods.
problem Zero-shot recognition of environments from object distributions.
method Bayesian framework using metric and self-supervised learning models to estimate object distributions in latent space.
result OBSER framework reliably performs inference in open-world and photorealistic environments, outperforming scene-based methods.
DNNs generalize object recognition in novel orientations via neurons tuned to common features.
problem Understanding how DNNs generalize to objects in novel orientations.
method Training DNNs with familiar objects from multiple viewpoints and analyzing neuron responses.
result DNNs disseminate orientation-invariance from familiar objects to recognize objects in novel orientations.
The Familiarity Hypothesis explains deep open set methods' success in detecting novel objects.
problem Detecting novel objects in open set recognition problems.
method Logits-based detection of absence of familiar features.
result Familiarity-based detection fails in scenarios with both novel and familiar objects.
Enhances deep reinforcement learning with object recognition.
problem Few works consider object characteristics in deep reinforcement learning.
method Proposes a novel method to incorporate object recognition into deep reinforcement learning models.
result Shows state-of-the-art results on Atari games.
Object recognition systems usually require fully complete manually labeled training data to train the classifier. In this paper, we study the problem of object recognition where the training samples are missing during the classifier learning stage, a task also known as zero-shot learning. We propose a novel zero-shot l…
End-to-end audio recognition system improves accuracy.
problem Improving accuracy in auditory object recognition.
method Proposes an end-to-end deep neural network with an 'inception nucleus' to learn features from raw waveforms.
result Bests current state-of-the-art approaches by 10.4 percentage points on Urbansound8k dataset.
Human visual object recognition is typically rapid and seemingly effortless, as well as largely independent of viewpoint and object orientation. Until very recently, animate visual systems were the only ones capable of this remarkable computational feat. This has changed with the rise of a class of computer vision algo…
Top 8 robotic vision systems tackled lifelong object recognition challenges.
problem Lifelong learning in robotic vision for varied, dynamic environments.
method Design of a dataset with diverse conditions and rules for evaluation.
result Robotic vision systems improved over time with dynamic object appearances.
This research generates synthetic data streams for handling concept drifts and novel classes.
problem Handling concept drifts and novel classes in dynamic data streams.
method Synthetic data stream generation for both concept drifts and novel classes.
result Demonstrates the effectiveness of unsupervised drift detectors in open set recognition.
New model improves illustration classification using transfer learning.
problem Improving image classification for artistic depictions.
method Transfer learning from VGG19 pre-trained on natural images, learning new features for illustrations.
result Optimized network achieves 86.61% top-1 and 97.21% top-5 precision on illustration dataset.
ViewFool identifies adversarial viewpoints to test image recognition robustness.
problem Lack of robustness to viewpoint changes in visual recognition models.
method Neural Radiance Fields (NeRF) and entropic regularizer to find adversarial viewpoints.
result Common image classifiers are highly vulnerable to generated adversarial viewpoints.
A novel clustering method uses torque balance to group objects.
problem Grouping similar objects in various scientific fields.
method Inspired by gravitational interactions, a parameter-free clustering algorithm based on mass and distance.
result The algorithm effectively clusters objects regardless of their shape, size, or density.
Study proposes a decision tree for more accurate depression recognition in speech.
problem Subjective bias in traditional depression diagnosis methods.
method A novel speech segment fusion method based on decision tree.
result The proposed decision tree model improves depression classification performance.
New architecture for RGB-D object recognition outperforms existing methods.
problem Improving object recognition in RGB-D cameras.
method RCFusion architecture combining RGB and depth information.
result RCFusion outperforms existing methods on standard datasets.
Study neural systems' classification of perceptual manifolds.
problem Classifying neural responses to diverse sensory features.
method Statistical mechanics theory for linear classification of manifolds.
result Introduction of novel geometrical measures for manifold analysis.
A novel tracking method for dense honeybee colonies using pixel personality.
problem Tracking large numbers of densely-arranged, interacting objects in a 2D environment.
method Segmentation-based object detection followed by adaptive object recognition through visual appearance.
result Reconstructed ~46% of trajectories in 5 minutes and 71% of tracks for at least 2 minutes.
A new deep metric learning method pulls embeddings towards dense clusters to improve classification accuracy.
problem Improving classification accuracy in deep metric learning models.
method Density Aware Metric Learning (DAML) which pulls embeddings towards the densest regions of clusters for each class.
result DAML achieves faster convergence and higher generalizability compared to existing methods.
Paper introduces a novel framework for recognizing dynamic ranking structures in preference-based data.
problem Complex and noisy preference-based data often hide underlying homogeneous structures.
method Developed an approach to identify dynamic ranking groups using temporal penalties and spectral estimation. Introduced an objective function for detecting structural changes.
result Consistent recognition of ranking groups and structural changes in preference-based data.
Novel volumetric convolution for unit ball improves 3D object recognition.
problem Efficiently convolving functions in a unit ball for deep learning.
method Developed volumetric convolution using Zernike polynomials.
result Improved 3D object recognition through novel convolution.
Patch ranking improves CNN performance by focusing on object content, not location.
problem CNNs lack rotation and translation invariance, limiting model capacity.
method Patch ranking before convolution and pooling to encode invariance.
result Patch ranking module improves CNN performance on various tasks.
Improves few-shot learning for real-world recognition with novel methods.
problem Challenges in real-world recognition with heavy-tailed class distributions and cluttered scenes.
method Parameter-free improvements including better training procedures, object localization, and feature space expansion.
result Doubles accuracy of state-of-the-art models on meta-iNat while generalizing to diverse settings.
DPNs learn pose-invariant object representations.
problem Pose-invariant 2D object recognition.
method Deformable Part Networks (DPNs) as sequences of LDPM units.
result 17-layer DPN outperforms CapsNets and STNs significantly on affNIST.
Paper introduces rehearsal-free continual learning for small, non-i.i.d. batches in robotic vision.
problem Learning new objects and improving recognition in a changing robotic environment.
method Two rehearsal-free continual learning techniques (CWR* and AR1*) for small, non-i.i.d. batches.
result AR1* outperforms other techniques by more than 15% in some cases.
Framework for continual object recognition in egocentric settings.
problem Recognizing objects in a cold-start, open-world setting with limited supervision.
method Memory-based incremental framework using time and space persistence, similarity, and active learning.
result Feasibility of open-world, generic object recognition with complete user supervision.
Proposes a multi-level learning approach for 3D object recognition.
problem Improving 3D object recognition accuracy through multi-scale spatial features.
method End-to-end multi-level learning on a multi-level voxel grid.
result Comparable object recognition performance with lower memory usage.
Paper proposes redundancy-free features for zero-shot object recognition.
problem Redundant visual features degrade zero-shot object recognition.
method Project original features into a new, statistically independent space.
result RFF-GZSL achieves competitive results on benchmark datasets.
Simple object representations improve model-free RL performance.
problem Current reinforcement learning agents lack object recognition.
method Used simple, feature-engineered object representations with the Rainbow model.
result Object representations significantly boost performance on Atari games.
Context-aware ZSL improves object recognition by considering object context.
problem Previous ZSL approaches ignore object context, limiting their effectiveness.
method Proposes a new approach that models the conditional likelihood of objects appearing in specific contexts.
result Contextual information significantly improves ZSL performance and is robust to class imbalance.
Efficiently learns representations across domains and tasks with few labels.
problem Learning representations that generalize across different domains and tasks with limited labeled data.
method Combines domain adversarial loss and metric learning for representation transfer. Optimizes on both labeled and unlabeled data in the target domain.
result Significantly outperforms fine-tuning on novel classes in new domains with few labeled examples.
There has recently been significant interest in hard attention models for tasks such as object recognition, visual captioning and speech recognition. Hard attention can offer benefits over soft attention such as decreased computational cost, but training hard attention models can be difficult because of the discrete la…
Robots learn material recognition from unlabeled data using GANs.
problem Difficulties in collecting labeled training data for robots.
method Semi-supervised learning with GANs for haptic features.
result Achieves ~90% accuracy in material estimation with 92% unlabeled data.
Robots detect and recognize objects in real-time for better manipulation.
problem Real-time object detection and recognition for humanoid robots.
method Modified YOLOv3 algorithm, quantization, and re-arrangement of layers for low-compute NAO robots.
result Robots can perform real-time detection, recognition, and localization of objects.
A new model disentangles object recognition and dynamics from video data.
problem Temporal reasoning in dynamically changing video data.
method Kalman variational auto-encoder framework for unsupervised learning.
result Model disentangles object representation and dynamics, outperforming other methods.
Paper proposes IE loss for deep metric learning improving CNN performance.
problem Improving deep learning models' performance in classification tasks.
method IE loss method to force distance between samples and class centers.
result IE loss leads to great improvements on various datasets.
A new multi-layer attention mechanism improves speech keyword recognition accuracy.
problem Inaccurate attention weights in LSTM networks for speech keyword recognition.
method Introducing information from layers prior to feature extraction into attention weights calculations.
result The proposed multi-layer attention mechanism leads to more accurate attention weights and improved keyword spotting performance.
Novel approach localizes optic disc and fovea centers efficiently.
problem Localizing optic disc and fovea centers in retinal images.
method Simultaneously process optic disc and fovea, modeling their relative geometry and appearance.
result Improves localization and recognition by incorporating object-object relations.
New method improves object detection models for long-tailed datasets.
problem Classifier imbalance in long-tail object detection datasets.
method Balanced Group Softmax (BAGS) module for balanced training of classifiers.
result Significantly improves performance of object detection models.
InfoNCE objective is equivalent to ELBO in RPM, linking to self-supervised learning.
problem Improving self-supervised learning methods by connecting them to variational inference.
method Recognizing RPM and showing InfoNCE as a simplified lower bound on MI, equal to ELBO in infinite sample limit.
result The actual InfoNCE objective is equal to the ELBO (up to a constant) in the infinite sample limit.
This research explores a modified VAE model to learn disentangled representations for object recognition.
problem Learning invariant representations for object recognition from diverse appearances.
method Develops a modified Variational Autoencoder (β-VAE) to enforce disentangled representations using variational inference. result Demonstrates that the incompatibility between β-VAE's conditional independence and latent variable independence leads to non-monotonic inference performance. In this paper, we propose a novel unsupervised domain adaptation algorithm based on deep learning for visual object recognition. Specifically, we design a new model called Deep Reconstruction-Classification Network (DRCN), which jointly learns a shared encoding representation for two tasks: i) supervised classification…
Unified tensor model disentangles object appearance factors.
problem Representing hierarchical intrinsic and extrinsic causal factors of object appearance.
method Compositional hierarchical tensor factorization.
result Interpretable object representation robust to occlusion and reduced training data requirements.
Agents learn to play by predicting their own errors and exploring novel interactions.
problem Creating autonomous agents that can learn and explore in complex, unstructured environments.
method A neural network with a world-model and self-model that learns to predict and challenge its own predictions.
result The agent generates complex behaviors like ego-motion prediction and object gathering.
Generative classifiers show surprising human-like performance.
problem Comparing generative and discriminative models for object recognition.
method Built on recent advances in generative modeling to create classifiers and compared them to discriminative models.
result Generative classifiers outperform discriminative models in several key areas, including shape bias and out-of-distribution accuracy.
Proposes KMvDA for object recognition from multi-view data.
problem Recognizing objects from different views, even when views are heterogeneous.
method Introduces kernel multi-view discriminant analysis (KMvDA) and uses random Fourier features (RFF) for large-scale learning.
result KMvDA and RFF approximation improve object recognition from multi-view data.
Paper presents an unsupervised method for object recognition using pretrained CNN and associative memory.
problem Fine-tuning pretrained CNN models for new domains is time-consuming and requires labeled data.
method Uses a pretrained CNN for feature extraction and a Hopfield network associative memory bank for classification.
result Eliminates the need for backpropagation and achieves competitive performance on unseen datasets.
Designs deformable classifiers to handle geometric variations in object recognition.
problem Geometric variations of objects, including rigid and non-rigid transformations, challenge object recognition.
method Introduces latent transformation variables and computes a transformation of the object image to a reference instantiation for each class. Uses a two-step training mechanism to optimize over latent transformation variables and classifier parameters.
result Achieves state-of-the-art results on rotated MNIST and Google Earth datasets, and competitive results on MNIST and CIFAR-10.
Agent learns object segmentation through self-supervised interactions.
problem Object segmentation in visual observations.
method Self-supervised learning through interaction, robust set loss.
result Model generalizes to novel objects and backgrounds.