UCB-RS uses RS to improve UCB for online advertising.
problem Improving recommendation in online advertising.
method UCB-RS, combining UCB with recommendation system.
result UCB-RS outperforms other reinforcement learning methods in RecoGym.
New algorithm reduces regret in fatigue-aware online recommendation.
problem Fatigue leads to user abandonment in online recommendation systems.
method Thompson Sampling for combinatorial bandits with fatigue model.
result Polynomial regret bound in number of items, outperforming naive approaches.
The paper analyzes regret in online recommendation systems with constraints.
problem Analyzing regret in online recommendation systems with user-item constraints.
method Theoretical analysis and algorithm design considering user-item constraints and unknown probabilities.
result Derives regret lower bounds and algorithms achieving these limits for various structural assumptions.
Optimizes recommender selection online with D-optimal design.
problem Finding the optimal recommender in online exploration-exploitation.
method Leverages D-optimal design from statistics to maximize information gain.
result Achieves maximum information gain during online exploration.
iPrescribe offers fast online offer recommendations using deep learning.
problem Online offer recommendation in real-time.
method Ensemble of deep learning and machine learning algorithms, optimized streaming technology stack, and efficient LSTM deployment.
result 90th percentile recommendation latency of 38 milliseconds.
Study shows online learning algorithms incentivize low-quality content, proposing new algorithms to improve quality.
problem Online learning algorithms in content recommender systems incentivize producers to create low-quality content.
method Analyzed the game between producers and content quality, designed new learning algorithms to incentivize high effort and quality.
result New algorithms incentivize producers to invest high effort and achieve high user welfare, improving content quality.
Recommender system improves recall of omitted foods in online dietary surveys.
problem Improving accuracy of online dietary assessment surveys through recall assistance.
method Developed a recommender algorithm to remind respondents of omitted foods based on past survey data.
result The recommender system captures more omitted foods than hand-coded prompts, but with lower precision.
Paper proposes MACDAE to infer real-time O2O contexts.
problem Difficult to infer users' real-time contexts, especially implicit ones, for O2O recommendation.
method MACDAE: a model that infers implicit contexts from user-item-explicit context interactions.
result Significant improvements in click-through rate and conversion rate in real-world traffic.
Incentive-aware recommender system for online platforms.
problem Myopic agents exploit optimal arms, not exploring alternatives.
method Model as multi-agent bandit problem, incentivizes exploration.
result Asymptotically optimal performance with ex-post fairness.
Paper tackles cold-start problems in online recommendation with few-shot learning and meta learning.
problem Cold-start problems in practical recommendations with limited interaction data.
method Combines scenario-specific learning with sequential meta-learning to create an integrated end-to-end framework.
result Significant gains over state-of-the-arts for cold-start problems in online recommendation.
We consider the problem of online collaborative filtering in the online setting, where items are recommended to the users over time. At each time step, the user (selected by the environment) consumes an item (selected by the agent) and provides a rating of the selected item. In this paper, we propose a novel algorithm …
New method uses bandit feedback to better evaluate recommender systems.
problem Traditional offline evaluation of recommender systems is inaccurate.
method Exploits bandit feedback to estimate online performance.
result Bandit feedback provides more accurate offline evaluation.
Recommenders have become widely popular in recent years because of their broader applicability in many e-commerce applications. These applications rely on recommenders for generating advertisements for various offers or providing content recommendations. However, the quality of the generated recommendations depends on …
We consider the online one-class collaborative filtering (CF) problem that consists of recommending items to users over time in an online fashion based on positive ratings only. This problem arises when users respond only occasionally to a recommendation with a positive rating, and never with a negative one. We study t…
This study improves user segmentation for online news recommendation systems.
problem Challenges in building modern recommender systems due to dynamic environments and data sparsity.
method Trend-responsive unsupervised user segmentation using multi-armed bandits.
result Significant improvements in online A/B tests compared to global-optimization algorithms.
Recommendations are broadly used in marketplaces to match users with items relevant to their interests and needs. To understand user intent and tailor recommendations to their needs, we use deep learning to explore various heterogeneous data available in marketplaces. This paper focuses on the challenge of measuring re…
LSTM improves cross-network recommendations by capturing user preference changes and irregular time intervals.
problem Offline cross-network recommender solutions fail to capture user preference changes and dynamic environments.
method Proposes a multi-layered LSTM network with attention mechanisms, higher order interactions, and time-aware gates.
result The model consistently outperforms state-of-the-art in accuracy, diversity, and novelty.
Paper introduces RTT2Vec for real-time grocery recommendations, achieving 9.4% uplift over baselines.
problem Personalized grocery recommendations to improve user experience and sales.
method RTT2Vec deep architecture for real-time recommendations, approximate inference technique.
result 9.4% uplift in prediction metrics over baseline models.
The paper shows how ignoring temporal context in recommender systems evaluation leads to false confidence, proposing a method to embed temporal context.
problem The discrepancy between offline and online recommender system performance evaluation.
method Proposes a training procedure to embed temporal context into recommender systems and validates its advantage using multi-objective optimization.
result Including temporal context in recommender systems evaluation can improve recall@20 by up to 20%.
Learning to rank is an important problem in machine learning and recommender systems. In a recommender system, a user is typically recommended a list of items. Since the user is unlikely to examine the entire recommended list, partial feedback arises naturally. At the same time, diverse recommendations are important be…
Recommender systems play a crucial role in mitigating the problem of information overload by suggesting users' personalized items or services. The vast majority of traditional recommender systems consider the recommendation procedure as a static process and make recommendations following a fixed strategy. In this paper…
Contextual bandit algorithms have become popular for online recommendation systems such as Digg, Yahoo! Buzz, and news recommendation in general. \emph{Offline} evaluation of the effectiveness of new algorithms in these applications is critical for protecting online user experiences but very challenging due to their "p…
Before A/B testing online a new version of a recommender system, it is usual to perform some offline evaluations on historical data. We focus on evaluation methods that compute an estimator of the potential uplift in revenue that could generate this new technology. It helps to iterate faster and to avoid losing money b…
In many online applications interactions between a user and a web-service are organized in a sequential way, e.g., user browsing an e-commerce website. In this setting, recommendation system acts throughout user navigation by showing items. Previous works have addressed this recommendation setup through the task of pre…
Paper tackles online ranking and diversification in recommender systems.
problem Maximizing relevance and diversity in ranked lists for online recommendation.
method CascadeHybrid approach that combines contextual bandits for relevance and topical diversity.
result CascadeHybrid outperforms baselines in real-world datasets.
Collaborative filtering, especially latent factor model, has been popularly used in personalized recommendation. Latent factor model aims to learn user and item latent factors from user-item historic behaviors. To apply it into real big data scenarios, efficiency becomes the first concern, including offline model train…
New bandit algorithms adapt to evolving user interests influenced by social circles.
problem Adapting to evolving user interests in recommendation systems.
method Online recommendation algorithms tailored for social influence, based on LinREL and Thompson Sampling.
result Our adaptations maintain asymptotic regret bounds similar to non-social cases.
A new model considers fatigue in online content recommendation systems.
problem Fatigue in users due to overexposure and boredom from similar recommendations.
method Proposed a fatigue-aware Dependent Click Model (DCM) and two learning algorithms.
result Developed algorithms with regret bounds for learning content relevance and fatigue effects.
Despite the prevalence of collaborative filtering in recommendation systems, there has been little theoretical development on why and how well it works, especially in the "online" setting, where items are recommended to users over time. We address this theoretical gap by introducing a model for online recommendation sy…
Study bandit problem on smooth graph functions for recommender systems.
problem Online learning problems involving graphs, like content-based recommendation.
method Introduced spectral bandit problem and two algorithms that scale linearly in effective dimension.
result Learned user preferences for thousands of items from just tens nodes evaluations.
A recommendation framework helps users choose healthcare interventions.
problem Choice overload in online healthcare communities.
method Multi-Armed Bandit (MAB) approach with innovative model components.
result Our recommendation design outperforms state-of-the-art systems.
Online voting is an emerging feature in social networks, in which users can express their attitudes toward various issues and show their unique interest. Online voting imposes new challenges on recommendation, because the propagation of votings heavily depends on the structure of social networks as well as the content …
Proposes a max-utility arm selection strategy for reducing cumulative regret in sequential query recommendations.
problem Reduces cumulative regret in sequential query recommendations for closed loop interactive learning settings.
method Proposes a max-utility arm selection strategy based on the maximum utility of arms.
result Improves cumulative regret substantially compared to baseline algorithms and random selection.
A model for learning customer preferences in a dynamic product launch setting.
problem Learning customer preferences in a setting with new product launches.
method Proposes a sequential multinomial logit (SMNL) model and a learning algorithm with a regret bound.
result Demonstrates the tier structure can mitigate risks associated with learning new products.
Sample-Rank simplifies MO recommendations by sampling and ranking, improving revenue with stable conversion rates.
problem Multi-objective recommendations in online food ordering systems.
method Multi-goal sampling followed by ranking, reducing MO problem to LTR model.
result Significant lift in revenue (2.64%) with stable conversion rates, no drop in last-mile traversal.
Bandit problem on graphs aims to recommend items with high expected ratings.
problem Online learning problems involving graphs, such as content-based recommendation.
method Study of a bandit problem on graphs, introducing effective dimension and proposing algorithms.
result Proposed algorithms scale linearly and sublinearly in the effective dimension, improving cumulative regret.
Recommendation systems have been integrated into the majority of large online systems to filter and rank information according to user profiles. It thus influences the way users interact with the system and, as a consequence, bias the evaluation of the performance of a recommendation algorithm computed using historical…
Collaborative filtering (CF) allows the preferences of multiple users to be pooled to make recommendations regarding unseen products. We consider in this paper the problem of online and interactive CF: given the current ratings associated with a user, what queries (new ratings) would most improve the quality of the rec…
SWAG uses graph convolutions to recommend products.
problem Product recommendation using graph-structured data.
method Graph Convolutional Network (GCN) with weighted random walks and aggregations.
result Graph-based approach improves product recommendations.
The paper tackles a bandit problem on graphs with smooth functions, aiming to recommend items with high expected ratings.
problem Online learning problems involving graphs, such as content-based recommendation.
method Introduced the notion of effective dimension and proposed two algorithms for solving the problem.
result The algorithms can learn good estimators of user preferences from just tens of nodes evaluations.
A new dataset tracks user interactions and click responses in online marketplaces.
problem Lack of exposure data in recommender systems datasets.
method Proposes a novel dataset including slates and click responses, allowing more accurate likelihood models.
result Models using exposure data show more natural likelihood, reducing bias towards previously exposed items.
Online news recommender systems aim to address the information explosion of news and make personalized recommendation for users. In general, news language is highly condensed, full of knowledge entities and common sense. However, existing methods are unaware of such external knowledge and cannot fully discover latent k…
Unified neural framework for multi-relational recommender systems.
problem Accurately capturing users' fine-grained preferences from diverse feedback types.
method Multi-Relational Memory Network (MRMN) framework that models fine-grained user-item relations and discriminates between feedback types.
result The proposed MRMN model outperforms state-of-the-art algorithms in various recommender scenarios.
Improves search performance by transferring knowledge from recommender system.
problem Cold start and feedback loop problems in search retrieval.
method Zero-Shot Heterogeneous Transfer Learning framework.
result Significant improvements in relevance and user interactions over production system.
Proposes efficient user cold start recommendation via meta parameter partition.
problem User cold start in recommendation systems.
method Divides model parameters into fixed and adaptive parts, learning them separately offline and online.
result Significant improvement in AUC (2.48% absolute improvement).
Solves online resource allocation problems with budget constraints.
problem Maximizing revenue for e-commerce platforms under budget constraints.
method Integrated online optimization and learning algorithm for non-stationary Poisson processes.
result Effective and efficient solutions for constrained resource allocation problems.
The system recommends hotels based on user preferences.
problem Overwhelm of hotel choices during trip planning.
method Used Expedia's hotel dataset to predict user preferences.
result Predicted and recommended hotels with high accuracy.
SharedMF uses secret sharing to protect privacy in distributed recommendation systems.
problem Privacy issues in multi-source data for recommendation systems.
method Federated learning and secret sharing technology.
result SharedMF achieves faster execution speed and better data adaptability compared to homomorphic encryption methods.