New architecture tracks objects in cluttered scenes without supervision.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
PiVoT improves real-time multi-object detection and tracking in clutter.
Improved RL for grasping in cluttered scenes using state representation learning.
Video prediction models combined with planning algorithms have shown promise in enabling robots to learn to perform many vision-based tasks through only self-supervision, reaching novel goals in cluttered scenes with unseen objects. However, due to the compounding uncertainty in long horizon video prediction and poor s…
MIRO learns robust latent spaces by maximizing mutual information with future information.
Object tracking is an ubiquitous problem that appears in many applications such as remote sensing, audio processing, computer vision, human-machine interfaces, human-robot interaction, etc. Although thoroughly investigated in computer vision, tracking a time-varying number of persons remains a challenging open problem.…
Hybrid approach combines transformer and Bayesian filtering for robust multiple particle tracking.
Bayesian nonparametric models improve tracking in cluttered environments.
We have developed an efficient algorithm for the maximum likelihood joint tracking and association problem in a strong clutter for GMTI data. By using an iterative procedure of the dynamic logic process "from vague-to-crisp," the new tracker overcomes combinatorial complexity of tracking in highly-cluttered scenarios a…
Robust STAP with coprime arrays reduces clutter using sparse modeling.
Traditional recognition methods typically require large, artificially-balanced training classes, while few-shot learning methods are tested on artificially small ones. In contrast to both extremes, real world recognition problems exhibit heavy-tailed class distributions, with cluttered scenes and a mix of coarse and fi…
Discrete return (DR) Laser Detection and Ranging (Ladar) systems provide a series of echoes that reflect from objects in a scene. These can be first, last or multi-echo returns. In contrast, Full-Waveform (FW)-Ladar systems measure the intensity of light reflected from objects continuously over a period of time. In a c…
The Long Short-Term Memory (LSTM) neural network based data association algorithm named as DeepDA for multi-target tracking in clutters is proposed to deal with the NP-hard combinatorial optimization problem in this paper. Different from the classical data association methods involving complex models and accurate prior…
Analytical method approximates ELBO gradient in clutter problem.
SVDD and Deep SVDD improve radar target detection in clutter.
Paper uses VAEs to detect radar targets in complex noise.
Best performing nuclear segmentation methods are based on deep learning algorithms that require a large amount of annotated data. However, collecting annotations for nuclear segmentation is a very labor-intensive and time-consuming task. Thereby, providing a tool that can facilitate and speed up this procedure is very …
Robotic grasping system learns to target objects from a single image.
Contrast enhanced ultrasound is a radiation-free imaging modality which uses encapsulated gas microbubbles for improved visualization of the vascular bed deep within the tissue. It has recently been used to enable imaging with unprecedented subwavelength spatial resolution by relying on super-resolution techniques. A t…
Novel unsupervised MIG detectors improve signal detection in cluttered environments.
CVAE detects weak complex signals in maritime radar, improving detection over classical methods.
Using low-frequency (UHF to L-band) ultra-wideband (UWB) synthetic aperture radar (SAR) technology for detecting buried and obscured targets, e.g. bomb or mine, has been successfully demonstrated recently. Despite promising recent progress, a significant open challenge is to distinguish obscured targets from other (nat…
FPI methods compute barycenters of Gaussian sets for various dissimilarity measures.
A neural scene representation framework enforcing 3D transformations.
3D filament plots visualize curves in datasets, avoiding visual clutter.
New model separates objects in scenes, enabling novel arrangements and depth.
Generative latent-variable models are emerging as promising tools in robotics and reinforcement learning. Yet, even though tasks in these domains typically involve distinct objects, most state-of-the-art generative models do not explicitly capture the compositional nature of visual scenes. Two recent exceptions, MONet …
Joint scene understanding and segmentation for automotive applications is a challenging problem in two key aspects:- (1) classifying every pixel in the entire scene and (2) performing this task under unstable weather and illumination changes (e.g. foggy weather), which results in poor outdoor scene visibility. This poo…
ROOTS learns to represent and render 3D scenes with object-centric models.
Despite enormous progress in object detection and classification, the problem of incorporating expected contextual relationships among object instances into modern recognition systems remains a key challenge. In this work we propose Information Pursuit, a Bayesian framework for scene parsing that combines prior models …
SPACE models complex scenes by decomposing objects and backgrounds.
NeRF-VAE generates 3D scenes with geometric structure from few images.
Scene text magnifier aims to magnify text in natural scene images without recognition. It could help the special groups, who have myopia or dyslexia to better understand the scene. In this paper, we design the scene text magnifier through interacted four CNN-based networks: character erasing, character extraction, char…
A new method for estimating joint value functions in multi-scene reinforcement learning.
Generative models learn from unlabeled videos via object segmentation and scene modeling.
Pix2Shape learns 3D scene representations from single images without supervision.
Paper improves reinforcement learning in multi-scene tasks.
A network removes irrelevant structures from chest radiographs for better analysis.
Improves reinforcement learning agent's scene-specific value function.
A novel method for visual question answering using scene graphs and reinforcement learning.
We propose a systematic learning-based approach to the generation of massive quantities of synthetic 3D scenes and arbitrary numbers of photorealistic 2D images thereof, with associated ground truth information, for the purposes of training, benchmarking, and diagnosing learning-based computer vision and robotics algor…
Acoustic scene classification is the task of identifying the scene from which the audio signal is recorded. Convolutional neural network (CNN) models are widely adopted with proven successes in acoustic scene classification. However, there is little insight on how an audio scene is perceived in CNN, as what have been d…
Generative model creates realistic scenes from pixel-wise labels.
P3I learns holistic scene representations from a single image.
This work considers an estimation task in compressive sensing, where the goal is to estimate an unknown signal from compressive measurements that are corrupted by additive pre-measurement noise (interference, or clutter) as well as post-measurement noise, in the specific setting where some (perhaps limited) prior knowl…
Generates coherent 3D scenes from monocular videos without supervision.
In this paper, we present an acoustic scene classification framework based on a large-margin factorized convolutional neural network (CNN). We adopt the factorized CNN to learn the patterns in the time-frequency domain by factorizing the 2D kernel into two separate 1D kernels. The factorized kernel leads to learn the m…
For embodied agents to infer representations of the underlying 3D physical world they inhabit, they should efficiently combine multisensory cues from numerous trials, e.g., by looking at and touching objects. Despite its importance, multisensory 3D scene representation learning has received less attention compared to t…