Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

3887751,1631,550 · Jun 202019922001200920182026
48 results for video learning

CB-GLNs learn video data's complex dependencies via graph representation.

problem Capturing complex dependency structures in sequential data like videos.
method Represent video data as a graph, find compositional dependencies via graph-cut and message passing.
result CB-GLNs efficiently learn video data's semantic compositional structure.

Paper proposes new principles and framework for AVC learning from user-generated videos.

problem Challenges in learning audio-visual correspondence from short-term user-generated videos.
method Introduced new principles and a framework to facilitate AVC learning from videos' themes.
result Proposed approach outperformed baseline by 23.15% on KWAI-AD-AudVis corpus.

SummaryNet automates video summarisation using deep learning.

problem Creating informative video summaries from videos.
method Two-stream convolutional network for spatial and temporal features, encoder-decoder model for salient features, sigmoid regression with LSTM for frame probability.
result SummaryNet achieves comparable or better results than state-of-the-art methods on benchmark datasets.

This paper highlights the need for explainable AI in video deep learning models.

problem Lack of explainable AI methods for video deep learning models.
method Illustrates the current state of video deep learning and highlights the need for explainability methods.
result Current explainability methods for video deep learning are insufficient and need improvement.

Generative model for high-resolution video generation.

problem Challenges in generating high-resolution videos due to memory and training stability limitations.
method Progressive growing of sliced Wasserstein GANs (SWGAN) for incremental spatiotemporal information learning.
result Generated photorealistic face videos of 256x256x32 resolution with an inception score of 14.57.

Paper introduces adversarial lossy compression for video artifacts reduction.

problem Unpleasant reconstruction artifacts in standard video coding schemes at low bit-rates.
method Adversarial lossy video compression model minimizing an adversarial distortion objective.
result Reduction of perceptual artifacts and detail reconstruction under extreme compression.

Generates, predicts, and completes human action videos with a two-stage deep framework.

problem Severe ill-posedness in video generation, prediction, and completion.
method Two-stage deep framework: 1) Generates human pose sequence from noise, 2) Converts pose sequence to video.
result Produces high-quality video generation/prediction/completion results of longer duration.

Deep networks analyze video snippets to predict outcomes, revealing a border effect that can be adjusted for better accuracy.

problem Improving the accuracy of deep networks trained on small video snippets.
method Applied the deep Taylor / LRP technique to understand and identify a border effect, tuning the step size to improve accuracy.
result The step size used to build video snippets can be adjusted to improve deep network accuracy without retraining.

New method compresses facial videos using GANs and latent space optimization.

problem Efficiently compressing facial videos at low bit rates.
method Leverages StyleGAN for latent space representation and compression, learns optimal compression through entropy model and perceptual loss.
result Significantly reduces perceptual distortion at low bit rates compared to state-of-the-art codecs.

Method learns from video demonstrations with human feedback.

problem Teaching autonomous agents using video demonstrations and human feedback.
method Constructs a mapping between standard and visual representations using a neural network.
result Effective in teaching a hopper agent to perform a backflip with minimal human feedback.

Ground-A-Video edits videos without training, preserving intended changes.

problem Complex multi-attribute video editing with omitted or wrong changes.
method Grounding-guided video-to-video translation with Cross-Frame Gated Attention.
result Zero-shot multi-attribute video editing with improved accuracy and frame consistency.

New method learns robot actions from videos without explicit labels.

problem Training robots to perform tasks from few demonstrations.
method Uses images and text for task-agnostic and general representation, synthesizes hallucinated actions, and applies dense correspondences.
result Trains robot policies solely from RGB videos, achieving diverse tasks across different robots and environments.

Deep learning method extracts medical knowledge from YouTube videos.

problem Improving healthcare information dissemination through machine learning.
method Developed a deep learning method to classify YouTube videos by medical knowledge level.
result Preliminary results show satisfactory performance in extracting medical knowledge from videos.

New method tackles anomaly detection in video surveillance using continual learning.

problem Challenges in continual learning for high-dimensional applications like video surveillance.
method Transfer learning and continual learning for online anomaly detection.
result Significantly reduces training complexity and continual learning from recent data.

Paper improves video feature learning for better downstream tasks.

problem Improving video feature learning for better performance on downstream tasks.
method Self-supervised learning approach using contrastive bidirectional transformer, extending BERT for real-valued feature vectors.
result Significantly improved performance on video classification, captioning, and segmentation tasks.

A dataset for evaluating engagement with scientific video lectures.

problem Challenges in managing learning resources due to rapid creation of video lectures.
method Introduction of VLEngagement dataset with content-based and video-specific features, and metrics related to user engagement.
result The largest and most diverse publicly available dataset for understanding context-agnostic engagement in video lectures.

System automates discovery and classification of training videos for career progression.

problem Difficulties in planning and navigating career paths due to changing job requirements and emerging sectors.
method Extracted educational videos, built a machine learning classifier, and optimized probability thresholds.
result Significant improvements in model performance by incorporating video attributes.

Generative model learns compact codes for video recovery.

problem Efficiently represent and reconstruct videos from missing data.
method Generative network trained to map compact latent codes to images, with low-rank and similarity constraints.
result Can recover true video sequences even if not in pretrained network's range.

The paper proposes a method to align surgical videos using kinematic data.

problem Difficulty in learning from comparing novice to expert surgical videos due to variability in gesture duration and execution.
method A novel technique using Dynamic Time Warping to synchronize videos of the same gesture at different speeds.
result The proposed approach allows for the alignment of surgical videos, enabling better learning for novice trainees.

DDPAE predicts video frames by decomposing and disentangling high-dimensional video data.

problem Predicting future video frames from input sequences is challenging due to high-dimensionality.
method Combines structured probabilistic models and deep networks to decompose and disentangle video components.
result DDPAE learns latent decomposition and disentanglement without supervision.

Scalable system predicts hot videos for peak VOD service.

problem Improving peak service quality of video on demand.
method Two neural networks: clustering and dispatch policy. Clustering reduces video numbers, dispatch policy ranks videos with probabilities. Networks are trained end-to-end.
result Average prediction accuracy of 17% compared to 3% baseline, for same number of dispatches.

This paper improves video summarization using a new algorithm and dataset.

problem Efficiently summarizing videos for browsing and searching.
method Improves sequential determinantal point process (SeqDPP) with a large-margin algorithm and a new probabilistic distribution.
result Significantly improved video summarization model with better user input integration and diversity.

Automated video conferencing system improves user experience with ASD and VC.

problem Improve remote video conferencing experience through automated speaker detection and virtual cinematography.
method Uses 4K wide-FOV camera, depth camera, and microphone array to extract features and train machine learning models for ASD and VC.
result System performs within 0.3 MOS of an expert cinematographer, as rated by users.

Improved video and movie description using multitask learning.

problem Lack of training data and poor generalization in video captioning.
method Multitask learning encoder-decoder framework for video sequences.
result Improved performance on multi-caption and single-caption datasets.

Proposes a new video attack method that multiplies perturbation to improve model robustness.

problem Challenges existing defense methods for video recognition models against additive adversarial attacks.
method Introduces Multiplicative Adversarial Videos (MultAV) to impose perturbation by multiplication.
result Model trained against additive attacks is less robust to MultAV.

A deep learning approach for efficient power control in wireless video transmissions.

problem Optimizing power control for real-time wireless video transmissions with quality constraints.
method Proposes a learning-based approach using a deep neural network to solve the non-convex power control problem.
result The deep neural network can quickly provide optimal power levels for given channel conditions.

We present a new model DrNET that learns disentangled image representations from video. Our approach leverages the temporal coherence of video and a novel adversarial loss to learn a representation that factorizes each frame into a stationary part and a temporally varying component. The disentangled representation can …

2017-05-31abs ↗pdf ↗