RNN-MAS predicts YouTube video popularity by integrating multiple sources of external influence.
problem Predicting popularity in asynchronous social media streams with multiple external influences.
method Recurrent Neural Network (RNN) for modeling asynchronous streams, integrating multiple sources of external influence.
result RNN-MAS outperforms state-of-the-art YouTube popularity prediction system by 17%.
Deep learning method extracts medical knowledge from YouTube videos.
problem Improving healthcare information dissemination through machine learning.
method Developed a deep learning method to classify YouTube videos by medical knowledge level.
result Preliminary results show satisfactory performance in extracting medical knowledge from videos.
Improved YouTube-8M classification accuracy using ensemble methods.
problem Improving video classification accuracy on YouTube-8M dataset.
method Ensemble methods to improve baseline predictions.
result Global prediction accuracy (GAP) increased from 77% to 80.7%.
Paper discusses YouTube-8M challenge and video understanding research.
problem Improving video understanding models for YouTube videos.
method Experimented with various models and ensemble learning techniques.
result Significant improvement in competition score through ensemble learning.
System filters inappropriate YouTube content for advertisers.
problem Inadequate detection of inappropriate content on YouTube ads.
method Proposes a system for identifying and filtering inappropriate content.
result Current countermeasures are ineffective in detecting inappropriate content.
The study explains YouTube commenters' behavior using rational inattention models.
problem Understanding and predicting YouTube commenters' behavior.
method Deep embedded clustering for user grouping, Bayesian revealed preferences for rationality testing, and behavioral economics constraints for attention span modeling.
result Most YouTube user groups optimize a Bayesian utility with rationally inattentive constraints.
Research uses machine learning to find central nodes and cliques in YouTube social networks.
problem Identifying central nodes and cliques in YouTube social networks.
method Unsupervised machine learning, Python programming, Bron-Kerbosch algorithm.
result Successfully found central nodes through clique-centrality and degree centrality.
Study finds financial YouTube channel 3PROTV predicts stock market performance and sentiment changes.
problem Determining the informational value of financial YouTube channels.
method Analyzing 3PROTV's content and its impact on stock market performance and sentiment.
result 3PROTV's content, particularly negative sentiment, predicts stock market performance and sentiment changes.
Estimates utility functions and information costs from YouTube comments.
problem Estimating rational inattention in Bayesian agents.
method Deep learning for clustering framing information, inverse reinforcement learning.
result Constructive estimates of utility and information costs.
Agent learns to play hard games by watching YouTube videos.
problem Sparse rewards in reinforcement learning environments.
method Self-supervised video mapping, YouTube video embedding, imitation reward function.
result Agent achieves human-level performance on hard games.
Proposes a deep learning model for timely and accurate recommendations.
problem Inability to provide timely recommendations and ranking issues with implicit feedback.
method Unified cross-network solution using listwise ranking for implicit data.
result Superior performance in accuracy, novelty, and diversity compared to baselines.
This study uses AI to analyze financial market coverage from YouTube videos.
problem Challenges in analyzing a large number of financial market videos.
method Used Whisper model to generate text from videos, applied natural language processing.
result Highlights dynamics of financial market coverage and identifies trending topics.
See http://www.youtube.com/watch?v=izbGXdjvK_I for a YouTube video showing part of the results in this paper.We will consider surfaces whose mean curvature at a point is a linear function of the square of the distance from that point to the vertical axis. We restrict ourselves here to surfaces which are cylinders over …
Neural M3 model adapts to diverse user behaviors over short and long timeframes.
problem Adapting to diverse user behaviors over short and long timeframes.
method Neural Multi-temporal-range Mixture Model (M3) combining short-term and long-term models with a learned gating mechanism.
result M3 consistently outperforms state-of-the-art sequential recommendation methods.
Improves YouTube's recommendation system by correcting biases in logged feedback.
problem Data biases in logged feedback from multiple behavior policies.
method Top-K off-policy correction applied to REINFORCE algorithm.
result Efficacy demonstrated through simulations and live experiments.
We present a solution to "Google Cloud and YouTube-8M Video Understanding Challenge" that ranked 5th place. The proposed model is an ensemble of three model families, two frame level and one video level. The training was performed on augmented dataset, with cross validation.
A dataset for detecting online hate speech from YouTube and Reddit comments.
problem Detecting and preventing hate speech on social media platforms.
method Created a dataset with two variants: binary and multi-label, based on YouTube and Reddit comments, using Figure-Eight crowdsourcing platform.
result Demonstrated that even a small amount of labelled data can help detect hate speech occurrences.
Proposes TDNs for learning complex video structures.
problem Complex temporal dependencies in sequential data, especially videos.
method Temporal Dependency Networks (TDNs) using graph representations and graph convolutions.
result Efficiently learns complex semantic structures of video data.
QFlow learns to prioritize video streaming to improve quality of experience.
problem Inconsistent video streaming quality due to network inefficiency.
method Develops a learning approach to dynamically allocate resources for video streaming.
result Demonstrates improved video quality for all clients at a wireless access point.
Adversarial attacks can fool copyright detection systems.
problem Vulnerability of copyright detection systems to adversarial attacks.
method Used gradient methods to create adversarial music that fooled detection systems.
result Adversarial attacks can successfully deceive industrial copyright detection tools.
New RL method learns from passive data by modeling intentions.
problem Learning from passive data like videos without rewards or actions.
method Model intentions using temporal difference learning, learning representations from raw data.
result Successfully learns features from passive data that accelerate downstream RL tasks.
The paper solves IRL for Bayesian stopping time problems.
problem Identifying optimal actions in Bayesian stopping time problems.
method Novel IRL framework using Bayesian revealed preferences.
result Identifies optimality and constructs cost function estimates.
A new method improves recommendation accuracy by learning from multiple networks and time-dependent user preferences.
problem Incomplete user profiles and dynamic user preferences degrade recommender quality.
method A cross-network time-aware recommender that learns from multiple source networks and develops current user models.
result The proposed solution achieves superior performance in accuracy, novelty, and diversity.
The Clifford group for 2 qubits is divided into 20 orbits, each with 4608 matrices.
problem Understanding the structure of the Clifford group for 2 qubits.
method Equivalence relation based on local Clifford gates and analysis of orbits.
result The Clifford group for 2 qubits is divided into 20 orbits, each with 4608 matrices.
Study builds tools to detect misinformation in online medical videos.
problem Misleading information in online medical videos.
method Manual annotation of a dataset, use of linguistic, acoustic, and user engagement features for classification models.
result Automatic models can identify misinformation with up to 74% accuracy.
Estimates LRD in sequential data, improving RNNs.
problem Quantifying LRD in sequential data for better RNNs.
method Principled estimation procedure based on LRD theory for real-valued time series.
result Estimates LRD reliably in user behavior and Wikipedia article writing.
Two-stage recommender systems show better performance when components interact rather than operate independently.
problem Two-stage recommender systems are often treated as sums of their parts, ignoring interactions between components.
method Used synthetic and real-world data to demonstrate interactions between ranker and nominators. Derived a generalization lower bound and proposed a Mixture-of-Experts approach to learn optimal item pools.
result Independent nominator training can lead to performance on par with random recommendations, highlighting the importance of interactions.
FSD50K provides an open dataset of over 51k audio clips for sound event recognition.
problem Small and domain-specific sound event recognition datasets.
method Creation of an open dataset with over 51k audio clips manually labeled using 200 classes.
result FSD50K is a new open benchmark for sound event recognition research.
VoxCeleb 2019 challenge assesses speaker recognition in uncontrolled settings.
problem Evaluate speaker recognition technology in unconstrained data.
method Public dataset, challenge, and workshop at Interspeech 2019.
result Baseline results and discussions provided.
In this paper we define a small variation of the Taylor method and a formula for the global error of this new numerical method that allows us to keep track of the round-off error and does not require previous knowledge of the exact solution. As an application we provide a rigorous proof of the construction/existence of…
In this paper we show all possible ramps where an object can move with constant speed under the effect of gravity and friction. The planar ramp are very easy to describe, just rotate a curve with velocity vector (tanh(as),sech(as)). Recall that tanh(as)^2+sech^2(as) = 1. Therefore, the solution of the planar constant s…
In this paper, we introduce a methodology that allows to model behavioral trajectories of users in online social media. First, we illustrate how to leverage the probabilistic framework provided by Hidden Markov Models (HMMs) to represent users by embedding the temporal sequences of actions they performed online. We the…
The paper establishes CLTs for Markov chains and improves sampling algorithms for heavy-tailed distributions.
problem Establishing central limit theorems for ergodic averages of Markov chains.
method Drift conditions to provide necessary and sufficient conditions for CLTs, including lower bounds on convergence rates.
result Sharp conditions and convergence rates for various MCMC algorithms on heavy-tailed targets.
New model improves sentiment analysis by fusing words from audio and video.
problem Improving sentiment analysis in noisy multimodal data.
method Gated Multimodal Embedding LSTM with Temporal Attention.
result State-of-the-art sentiment classification and regression results on CMU-MOSI dataset.
Extends Aff-Wild database for affect recognition in real-world settings.
problem Complex human emotional states in real-world settings.
method Developed deep neural architectures with attention mechanism for emotion recognition.
result Improved performance in emotion recognition using Aff-Wild2.
Deep learning approximates shortest path distances in large graphs.
problem Scaling up shortest path distance computation in large networks.
method Deep learning techniques to approximate distances using vector embeddings.
result Feedforward neural networks with embeddings can approximate distances with low distortion error.
Study links public concern in Italy to financial markets worldwide.
problem Understanding public concern's impact on financial markets during pandemics.
method Used Google Trends data from YouTube, News, and Search to measure public concern and correlate it with stock index returns.
result Public concern in Italy drives concerns in other countries and explains stock index returns of multiple nations.
Aff-Wild database expands facial expression recognition to real-world conditions.
problem Lack of spontaneous facial expression databases in real-world conditions.
method Collects spontaneous facial expressions from YouTube, annotates with valence and arousal, uses deep learning techniques.
result Developed an end-to-end DNN model achieving 0.555 CCC for valence and 0.499 CCC for arousal.
Human communication takes many forms, including speech, text and instructional videos. It typically has an underlying structure, with a starting point, ending, and certain objective steps between them. In this paper, we consider instructional videos where there are tens of millions of them on the Internet. We propose a…
Elastic-InfoGAN learns object identity in class-imbalanced data.
problem Learning disentangled representations in class-imbalanced data.
method Invariance to identity-preserving transformations to learn object identity.
result Effectiveness in disentangling object identity in imbalanced datasets.
Transformer autoencoder learns musical style from performances.
problem Learning high-level controls over symbolic music generation.
method Aggregates encodings of input data across time to obtain global style representation.
result Improves control over performance style and melody in music generation tasks.
We provide a detailed description of solutions of Curve Shortening in Rn that are invariant under some one-parameter symmetry group of the equation, paying particular attention to geometric properties of the curves, and the asymptotic properties of their ends. We find generalized helices, and a connection with curv…
LSTM LMs improve speech recognition by 8% using lattice rescoring.
problem Efficiently integrating LSTM LMs into speech recognition systems.
method Lattice rescoring algorithms using LSTM LMs.
result Reduces word error rate (WER) by 8% relative to N-gram LMs.
New graph representation learning network improves scalability and feature integration.
problem Scalability and feature integration in graph neural networks for large, dense graphs.
method Adaptive sampling of neighbours based on weighted multi-step transition probabilities.
result Comparable or better results on various graph benchmarks.
Study classifies stock price jumps as exogenous or endogenous using news data.
problem Differentiating between exogenous and endogenous price jumps.
method Synchronized news data with order book data to analyze stock price movements.
result Exogenous jumps are abrupt and follow a decaying power-law, while endogenous jumps are progressively accelerating.
Paper proposes new principles and framework for AVC learning from user-generated videos.
problem Challenges in learning audio-visual correspondence from short-term user-generated videos.
method Introduced new principles and a framework to facilitate AVC learning from videos' themes.
result Proposed approach outperformed baseline by 23.15% on KWAI-AD-AudVis corpus.
A new method matches moments exactly for large graphs, improving spectral learning.
problem Lack of exact moment matching in spectral density approximations for large graphs.
method Maximum Entropy method for spectral density approximation, with a new algorithm.
result The new method outperforms existing approaches in learning graph spectra.
Study shows online learning algorithms incentivize low-quality content, proposing new algorithms to improve quality.
problem Online learning algorithms in content recommender systems incentivize producers to create low-quality content.
method Analyzed the game between producers and content quality, designed new learning algorithms to incentivize high effort and quality.
result New algorithms incentivize producers to invest high effort and achieve high user welfare, improving content quality.