One of the greatest challenges in the design of a real-time perception system for autonomous driving vehicles and drones is the conflicting requirement of safety (high prediction accuracy) and efficiency. Traditional approaches use a single frame rate for the entire system. Motivated by the observation that the lack of…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Automatic video modification to hide faces while maintaining pose, illumination, and expression.
Objective: Ultrasound elastography is gaining traction as an accessible and useful diagnostic tool for such things as cancer detection and differentiation and thyroid disease diagnostics. Unfortunately, state of the art shear wave imaging techniques, essential to promote this goal, are limited to high-end ultrasound ha…
Paper proposes a CNN-based method for estimating intra frame bits and quality.
Sparse-RS framework efficiently attacks models with sparse perturbations.
Recent advances in video super-resolution have shown that convolutional neural networks combined with motion compensation are able to merge information from multiple low-resolution (LR) frames to generate high-quality images. Current state-of-the-art methods process a batch of LR frames to generate a single high-resolu…
CNN improves frame selection for ultrasound elastography.
New wavelet frames constructed from reproducing kernels for continuous and discrete domains.
New method for tensor classification with missing data.
Paper proposes SMFN for high-res spherical video super-resolution.
LVTINO improves high-definition video restoration with consistent temporal details.
The paper shows that random frames have full spark with high probability.
Our goal is to predict future video frames given a sequence of input frames. Despite large amounts of video data, this remains a challenging task because of the high-dimensionality of video frames. We address this challenge by proposing the Decompositional Disentangled Predictive Auto-Encoder (DDPAE), a framework that …
Deep learning predicts frame errors in CIRN using SC2 dataset.
In deep reinforcement learning (RL) tasks, an efficient exploration mechanism should be able to encourage an agent to take actions that lead to less frequent states which may yield higher accumulative future return. However, both knowing about the future and evaluating the frequentness of states are non-trivial tasks, …
The paper uses Frenet frame to unify electrical and geometric quantities.
This paper presents GRASTA (Grassmannian Robust Adaptive Subspace Tracking Algorithm), an efficient and robust online algorithm for tracking subspaces from highly incomplete information. The algorithm uses a robust -norm cost function in order to estimate and track non-stationary subspaces when the streaming data …
High-speed model accurately simulates neuromorphic devices.
VISION-XL improves HD video quality using latent image diffusion models.
Video classification is a challenging task in computer vision. Although Deep Neural Networks (DNNs) have achieved excellent performance in video classification, recent research shows adding imperceptible perturbations to clean videos can make the well-trained models output wrong labels with high confidence. In this pap…
This paper examines fairness and arbitrariness in bias mitigation methods.
The Kuperberg invariant is shown to be gauge invariant for certain framed 3-manifolds.
Deep learning has revolutionised many fields, but it is still challenging to transfer its success to small mobile robots with minimal hardware. Specifically, some work has been done to this effect in the RoboCup humanoid football domain, but results that are performant and efficient and still generally applicable outsi…
The paper addresses classification imbalance by framing it as a transfer learning problem.
Crowdsourcing provides a popular paradigm for data collection at scale. We study the problem of selecting subsets of workers from a given worker pool to maximize the accuracy under a budget constraint. One natural question is whether we should hire as many workers as the budget allows, or restrict on a small number of …
LLoCa makes any network Lorentz-equivariant, achieving high accuracy and efficiency.
Detecting the intention of drivers is an essential task in self-driving, necessary to anticipate sudden events like lane changes and stops. Turn signals and emergency flashers communicate such intentions, providing seconds of potentially critical reaction time. In this paper, we propose to detect these signals in video…
Proposes SE(3) equivariant graph neural networks with local frames for efficient geometric approximation.
Improved speech separation and enhancement using neural beamforming.
X-ray computed tomography (CT) using sparse projection views is a recent approach to reduce the radiation dose. However, due to the insufficient projection views, an analytic reconstruction approach using the filtered back projection (FBP) produces severe streaking artifacts. Recently, deep learning approaches using la…
Ball trajectory data are one of the most fundamental and useful information in the evaluation of players' performance and analysis of game strategies. Although vision-based object tracking techniques have been developed to analyze sport competition videos, it is still challenging to recognize and position a high-speed …
Confirming a conjecture, we show fundamental groups of certain abelian differentials are framed mapping class groups.
Sparse coding in learned dictionaries has been established as a successful approach for signal denoising, source separation and solving inverse problems in general. A dictionary learning method adapts an initial dictionary to a particular signal class by iteratively computing an approximate factorization of a training …
Automatic music transcription is considered to be one of the hardest problems in music information retrieval, yet recent deep learning approaches have achieved substantial improvements on transcription performance. These approaches commonly employ supervised learning models that predict various time-frequency represent…
We have recently shown that deep Long Short-Term Memory (LSTM) recurrent neural networks (RNNs) outperform feed forward deep neural networks (DNNs) as acoustic models for speech recognition. More recently, we have shown that the performance of sequence trained context dependent (CD) hidden Markov model (HMM) acoustic m…
We present a new algorithm for video coding, learned end-to-end for the low-latency mode. In this setting, our approach outperforms all existing video codecs across nearly the entire bitrate range. To our knowledge, this is the first ML-based method to do so. We evaluate our approach on standard video compression test …
Reinforcement learning is concerned with identifying reward-maximizing behaviour policies in environments that are initially unknown. State-of-the-art reinforcement learning approaches, such as deep Q-networks, are model-free and learn to act effectively across a wide range of environments such as Atari games, but requ…
MCSAE improves speaker embedding by focusing on both high- and low-level features.
Improved action recognition in live videos with hybrid FR-DL method.
Bertrand framed surfaces defined in Euclidean 3-space with applications.
Quaternionic frames' admissibility and homotopy proven.
We propose a limit order book (LOB) model with dynamics that account for both the impact of the most recent order and the shape of the LOB. We present an empirical analysis showing that the type of the last order significantly alters the submission rate of immediate future orders, even after accounting for the state of…
Introduces hyperbolic generalized framed surfaces and their properties.
Paper presents a new video generation model using diffusion probabilistic methods.
Study of generalized Bishop frames on curves in 4D space.
Noether's Theorem yields conservation laws for a Lagrangian with a variational symmetry group. The explicit formulae for the laws are well known and the symmetry group is known to act on the linear space generated by the conservation laws. The aim of this paper is to explain the mathematical structure of both the Euler…
We simplify Bayesian filtering by framing it as optimization, making it practical for high-dimensional systems.
The main drawback of the Frenet frame is that it is undefined at those points where the curvature is zero. Further- more, in the case of planar curves, the Frenet frame does not agree with the standard framing of curves in the plane. The main drawback of the Bishop frame is that the principle normal vector N is not in …