Study shows online learning algorithms incentivize low-quality content, proposing new algorithms to improve quality.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Algorithm uncovers two main patterns of online content popularity: bursty and steady.
Deep learning detects radical content on social media.
A new model considers fatigue in online content recommendation systems.
Detects radical content on Twitter using textual, psychological, and behavioral signals.
Enhances content moderation with culturally-aware models.
Study uses attention-based method to detect different types of online harassment.
One of the major hurdles preventing the full exploitation of information from online communities is the widespread concern regarding the quality and credibility of user-contributed content. Prior works in this domain operate on a static snapshot of the community, making strong assumptions about the structure of the dat…
Spotify improves content mix using contextual bandits.
Efficiently selects seed nodes to maximize content influence in unknown social networks.
Proposes a model to optimize feedback for content creators on social media.
Modeling dynamic user interests using neural matrix factorization.
New ranking algorithms improve online content delivery by learning from click data.
Stock trend prediction plays a critical role in seeking maximized profit from stock investment. However, precise trend prediction is very difficult since the highly volatile and non-stationary nature of stock market. Exploding information on Internet together with advancing development of natural language processing an…
LOLA uses LLMs to optimize content delivery, outperforming traditional methods.
Data science predicts user interest for midwifery content.
Modeling incentives for content creators on algorithm-curated platforms.
In the industry of video content providers such as VOD and IPTV, predicting the popularity of video contents in advance is critical not only from a marketing perspective but also from a network optimization perspective. By predicting whether the content will be successful or not in advance, the content file, which is l…
Applying traditional collaborative filtering to digital publishing is challenging because user data is very sparse due to the high volume of documents relative to the number of users. Content based approaches, on the other hand, is attractive because textual content is often very informative. In this paper we describe …
Bandit problem on graphs aims to recommend items with high expected ratings.
Spaced repetition is a technique for efficient memorization which uses repeated, spaced review of content to improve long-term retention. Can we find the optimal reviewing schedule to maximize the benefits of spaced repetition? In this paper, we introduce a novel, flexible representation of spaced repetition using the …
Online voting is an emerging feature in social networks, in which users can express their attitudes toward various issues and show their unique interest. Online voting imposes new challenges on recommendation, because the propagation of votings heavily depends on the structure of social networks as well as the content …
The paper tackles a bandit problem on graphs with smooth functions, aiming to recommend items with high expected ratings.
System filters inappropriate YouTube content for advertisers.
New framework predicts earnings announcements using press release content, surpassing earnings surprises.
We propose a method for building an interpretable recommender system for personalizing online content and promotions. Historical data available for the system consists of customer features, provided content (promotions), and user responses. Unlike in a standard multi-class classification setting, misclassification cost…
Massive Open Online Courses (MOOCs) bring together thousands of people from different geographies and demographic backgrounds -- but to date, little is known about how they learn or communicate. We introduce a new content-analysed MOOC dataset and use Bayesian Non-negative Matrix Factorization (BNMF) to extract communi…
New algorithm poLinUCB improves online learning in content recommendation platforms.
Online advertising is an important and huge industry. Having knowledge of the website attributes can contribute greatly to business strategies for ad-targeting, content display, inventory purchase or revenue prediction. Classical inferences on users and sites impose challenge, because the data is voluminous, sparse, hi…
Study analyzes Facebook reactions to scholarly articles.
We propose SPARFA-Trace, a new machine learning-based framework for time-varying learning and content analytics for education applications. We develop a novel message passing-based, blind, approximate Kalman filter for sparse factor analysis (SPARFA), that jointly (i) traces learner concept knowledge over time, (ii) an…
Study bandit problem on smooth graph functions for recommender systems.
The aim of this paper is to get an overview of the online buyer profile, and also some key aspects in the way the online shopping is conducted. In this project we conducted a quantitative research, consisting of a questionnaire based survey. For data processing and interpretation we used SPSS statistical software and E…
Geotagged data can be used to describe regions in the world and discover local themes. However, not all data produced within a region is necessarily specifically descriptive of that area. To surface the content that is characteristic for a region, we present the geographical hierarchy model (GHM), a probabilistic model…
Paper proposes using word embeddings to detect trolls in social media debates.
Unified framework for online LLM watermark detection using e-processes.
Paper proposes online learning for estimating AC network admittance matrix.
We introduce a novel algorithmic approach to content recommendation based on adaptive clustering of exploration-exploitation ("bandit") strategies. We provide a sharp regret analysis of this algorithm in a standard stochastic noise setting, demonstrate its scalability properties, and prove its effectiveness on a number…
Study analyzes online student behavior patterns using log data.
Quizlet is the most popular online learning tool in the United States, and is used by over 2/3 of high school students, and 1/2 of college students. With more than 95% of Quizlet users reporting improved grades as a result, the platform has become the de-facto tool used in millions of classrooms. In this paper, we expl…
Efficient sparse attention reduces self-attention complexity and improves model performance.
Understanding and predicting the popularity of online items is an important open problem in social media analysis. Considerable progress has been made recently in data-driven predictions, and in linking popularity to external promotions. However, the existing methods typically focus on a single source of external influ…
Paper develops an efficient online watermark detection for AI-generated text.
Recommendations are broadly used in marketplaces to match users with items relevant to their interests and needs. To understand user intent and tailor recommendations to their needs, we use deep learning to explore various heterogeneous data available in marketplaces. This paper focuses on the challenge of measuring re…
The social media revolution has changed the way that brands interact with consumers. Instead of spending their advertising budget on interstate billboards, more and more companies are choosing to partner with so-called Internet "influencers" --- individuals who have gained a loyal following on online platforms for the …
RALM extends exposure for long-tail contents in real-time recommender systems.
OLPA optimizes online user-centric selection with probing, achieving near-optimal regret bounds.
Noisy labeled data is more a norm than a rarity for crowd sourced contents. It is effective to distill noise and infer correct labels through aggregation results from crowd workers. To ensure the time relevance and overcome slow responses of workers, online label aggregation is increasingly requested, calling for solut…