Study estimates 163 million online freelancers globally.
problem Estimating the number of online workers globally.
method Combining data from various online labour platforms.
result Headline estimate of 163 million registered profiles.
OpenML is an online machine learning platform where researchers can easily share data, machine learning tasks and experiments as well as organize them online to work and collaborate more efficiently. In this paper, we present an R package to interface with the OpenML platform and illustrate its usage in combination wit…
Reduces false positives in classifying rare online platforms.
problem Challenges in accurately identifying rare online platforms with ML.
method Calibrated probabilities and ensembles to reduce bias.
result Significantly reduces false positives in rare event detection.
Optimal bidding strategy for multi-platform ad auctions under budget constraints.
problem Optimizing ad placements for budget-constrained advertisers across multiple platforms.
method Developed an optimal bidding strategy for non-incentive-compatible auctions with budget constraints.
result Maximized total utility across auctions while satisfying budget constraints in expectation.
Study shows online learning algorithms incentivize low-quality content, proposing new algorithms to improve quality.
problem Online learning algorithms in content recommender systems incentivize producers to create low-quality content.
method Analyzed the game between producers and content quality, designed new learning algorithms to incentivize high effort and quality.
result New algorithms incentivize producers to invest high effort and achieve high user welfare, improving content quality.
Online trading platforms manipulate profits and losses, causing 82% of retail traders to lose money.
problem Manipulation of online trading platforms leading to financial losses for retail traders.
method Independent recording of trade details using REST API responses, comparison with broker reviews.
result 82% of retail traders lose money due to platform technical issues.
Predicts student performance in interactive online question pools using GNNs.
problem Predicting student performance in interactive online question pools with evolving knowledge.
method Proposes R^2GCN, a GNN model for heterogeneous networks to predict student performance.
result Achieves higher accuracy in student performance prediction than traditional methods.
Incentive-aware recommender system for online platforms.
problem Myopic agents exploit optimal arms, not exploring alternatives.
method Model as multi-agent bandit problem, incentivizes exploration.
result Asymptotically optimal performance with ex-post fairness.
Online advertising in E-commerce platforms provides sellers an opportunity to achieve potential audiences with different target goals. Ad serving systems (like display and search advertising systems) that assign ads to pages should satisfy objectives such as plenty of audience for branding advertisers, clicks or conver…
COBRA addresses strategic behavior in online platforms by ensuring truthful reporting without monetary incentives.
problem Ensuring truthful reporting from strategic agents in online platforms.
method Proposes COBRA, an algorithm for contextual bandits involving strategic agents that disincentivizes strategic behavior.
result COBRA achieves sub-linear regret guarantee and incentive compatibility without monetary incentives.
Dynamic assortment problem on two-sided platform with unknown parameters
problem Optimizing assortment display in an online platform with incomplete information and heterogeneous customers
method Data-driven algorithm that learns choice parameters while optimizing revenue
result Worst-case regret grows polylogarithmically over time
Taobao, as the largest online retail platform in the world, provides billions of online display advertising impressions for millions of advertisers every day. For commercial purposes, the advertisers bid for specific spots and target crowds to compete for business traffic. The platform chooses the most suitable ads to …
Motivated by the observation that overexposure to unwanted marketing activities leads to customer dissatisfaction, we consider a setting where a platform offers a sequence of messages to its users and is penalized when users abandon the platform due to marketing fatigue. We propose a novel sequential choice model to ca…
Study tackles ranking fraud in online platforms by learning robust rankings.
problem Fraudulent fake users manipulate product rankings.
method Developed algorithms for robust ranking in two informational environments.
result Our algorithms converge to optimal rankings, robust to fake users.
Study uses multiple online media to predict crude oil prices.
problem Forecasting crude oil prices using online media.
method Semantic analysis and ARIMAX models on Twitter, Google Trends, Wikipedia, and GDELT.
result Combined analysis from four platforms improves price prediction.
In this paper we consider an online recommendation setting, where a platform recommends a sequence of items to its users at every time period. The users respond by selecting one of the items recommended or abandon the platform due to fatigue from seeing less useful items. Assuming a parametric stochastic model of user …
Study proposes a machine learning method for bid shading in first-price auctions.
problem Maintaining strategy equilibrium in first-price auctions.
method Machine learning approach to model optimal bid shading.
result Demonstrates superiority and robustness of new approach across various metrics.
fintech-kMC simulates financial platforms for AI/ML model validation.
problem Validation of AI/ML models in real-world financial applications.
method Agent-based model with kinetic Monte Carlo engine.
result Generates realistic synthetic data for testing AI/ML models.
DSPN predicts advertiser satisfaction and intent for e-commerce platforms.
problem Understanding advertiser intent and satisfaction for e-commerce platforms.
method Two-stage Deep Satisfaction Prediction Network (DSPN) that models intent and satisfaction.
result DSPN outperforms state-of-the-art baselines and predicts advertiser satisfaction accurately.
New algorithms for fair item allocation with limited copies.
problem Fair division of numerous items with few copies.
method Modeling as a contextual bandit problem with sub-linear regret guarantees.
result Proposed algorithms achieve sub-linear regret in fair item allocation.
A new online learning problem, CAB, tackles matching platforms to maximize user satisfaction.
problem Maximizing matches in a matching platform can lead to dissatisfaction and churn.
method Developed CAB, an online learning problem that maximizes arm satisfaction, and analyzed algorithms like UCB and Thompson sampling.
result CAB-UCB achieves higher cumulative satisfaction than baselines in experiments.
KryptoOracle predicts cryptocurrency prices using Twitter sentiments.
problem Real-time price prediction for high-volatility cryptocurrencies.
method Spark-based architecture, sentiment analysis, online learning.
result Real-time adaptation of learning algorithms to new data.
The paper explores how to measure and optimize ad reach while maintaining user privacy.
problem Measuring ad reach while preserving user privacy in online advertising.
method Introduces k-anonymity and probabilistic discounting for frequency capping. result Privacy introduces a significant performance drop but with manageable costs.
Develops platforms to analyze social media data for human behavior and emotions.
problem Understanding human behavior and emotions from social media data.
method Self-structuring incremental machine learning, event detection, natural language processing.
result Captured salient topics and events from social media data, validated against news.
Online media provides opportunities for marketers through which they can deliver effective brand messages to a wide range of audiences. Advertising technology platforms enable advertisers to reach their target audience by delivering ad impressions to online users in real time. In order to identify the best marketing me…
A scalable system detects price anomalies in online marketplaces to improve customer experience.
problem Inaccurate prices on online marketplaces lead to poor customer experience and revenue loss.
method MoatPlus uses unsupervised statistical features and an ensemble of models to generate upper price bounds.
result Our approach improves precise anchor coverage by up to 46.6% in high-vulnerability item subsets.
An online labor platform faces an online learning problem in matching workers with jobs and using the performance on these jobs to create better future matches. This learning problem is complicated by the rise of complex tasks on these platforms, such as web development and product design, that require a team of worker…
New method minimizes experiment cost while maintaining accuracy.
problem Minimizing cost in experiments with interference or other concerns.
method Synthetically Controlled Thompson Sampling (SCTS).
result Minimizes regret and maintains inferential ability.
Develops a diamond price index for online auction platforms.
problem Tracking market trends of wholesale diamond prices.
method Modelling diamond prices to create a hedonic index.
result Provides a basis for constructing derivatives for collectables.
Dynamic promotion optimization for e-commerce platforms within financial constraints.
problem Balancing promotional costs with incremental revenue for sustainable growth.
method Knapsack Problem formulation for dynamic optimization, Retrospective Estimation, online-dynamic calibration.
result Significant increase in target outcome while staying within financial constraints.
A new model considers fatigue in online content recommendation systems.
problem Fatigue in users due to overexposure and boredom from similar recommendations.
method Proposed a fatigue-aware Dependent Click Model (DCM) and two learning algorithms.
result Developed algorithms with regret bounds for learning content relevance and fatigue effects.
Deep network optimizes ad bidding for first-price auctions.
problem Optimizing bid prices for first-price auctions in online advertising.
method Introduced a deep distribution network for optimal bidding.
result Algorithm outperforms previous methods in terms of surplus and eCPX metrics.
A new method detects fraud transactions by analyzing user behavior over time.
problem Detecting fraud transactions in online payment platforms.
method A time attention based recurrent layer framework combining static and dynamic user behaviors.
result Our method outperforms state-of-the-art methods, especially in recall at top percent.
We propose an in-depth study of lending behaviors in Kiva using a mix of quantitative and large-scale data mining techniques. Kiva is a non-profit organization that offers an online platform to connect lenders with borrowers. Their site, kiva.org, allows citizens to microlend small amounts of money to entrepreneurs (bo…
Boosts A/B test precision using auxiliary data from historical users.
problem Small sample sizes and imprecise estimates in A/B tests.
method Coupling design-based causal estimation with machine-learning models of historical user data.
result Effect estimates using auxiliary data are roughly equivalent to increasing sample size by 20%, or up to 50-80% in some cases.
CodeReef enables sharing ML models across platforms efficiently.
problem Sharing and deploying ML models across different systems efficiently.
method Developed an open platform to share ML components, automate deployment, and benchmark models.
result Demonstrated efficient deployment and benchmarking of ML models across diverse platforms.
Machine learning experiments often mislead due to unmet assumptions.
problem Machine learning experiments with pooled data may not meet necessary assumptions for unbiased causal effect estimation.
method Analysis of assumptions required for unbiased causal effect estimation in machine learning experiments.
result Practical applications of A/B-tests with machine learning models may not yield unbiased estimates of causal effect.
Novel algorithm reduces feature inclusion in online decision-making.
problem Optimizing decision-making for personalized user experiences with fairness.
method Online Batched Sequential Inclusion (OBSI) algorithm for sequential feature inclusion.
result OBSI outperforms other algorithms in terms of regret, relevance of features, and compute.
In machine learning applications for online product offerings and marketing strategies, there are often hundreds or thousands of features available to build such models. Feature selection is one essential method in such applications for multiple objectives: improving the prediction accuracy by eliminating irrelevant fe…
In this paper, we introduce a methodology that allows to model behavioral trajectories of users in online social media. First, we illustrate how to leverage the probabilistic framework provided by Hidden Markov Models (HMMs) to represent users by embedding the temporal sequences of actions they performed online. We the…
In the era of social media and networking platforms, Twitter has been doomed for abuse and harassment toward users specifically women. Monitoring the contents including sexism and sexual harassment in traditional media is easier than monitoring on the online social media platforms like Twitter, because of the large amo…
Optimizes bidding strategies for LinkedIn ads across multiple platforms.
problem Optimizing automated bidding agents for dynamic online marketplaces.
method Developed a general optimization framework for buyer's interest, agnostic to auction mechanisms.
result Automatically guarantees the optimality of budget allocation across ad units and platforms.
A dataset for detecting online hate speech from YouTube and Reddit comments.
problem Detecting and preventing hate speech on social media platforms.
method Created a dataset with two variants: binary and multi-label, based on YouTube and Reddit comments, using Figure-Eight crowdsourcing platform.
result Demonstrated that even a small amount of labelled data can help detect hate speech occurrences.
SOL is an open-source library for scalable online learning algorithms, and is particularly suitable for learning with high-dimensional data. The library provides a family of regular and sparse online learning algorithms for large-scale binary and multi-class classification tasks with high efficiency, scalability, porta…
A matching in a two-sided market often incurs an externality: a matched resource may become unavailable to the other side of the market, at least for a while. This is especially an issue in online platforms involving human experts as the expert resources are often scarce. The efficient utilization of experts in these p…
We study platforms in the sharing economy and discuss the need for incentivizing users to explore options that otherwise would not be chosen. For instance, rental platforms such as Airbnb typically rely on customer reviews to provide users with relevant information about different options. Yet, often a large fraction o…
Big data repositories from online learning platforms such as Massive Open Online Courses (MOOCs) represent an unprecedented opportunity to advance research on education at scale and impact a global population of learners. To date, such research has been hindered by poor reproducibility and a lack of replication, largel…
New approach to multi-armed bandit problem aims to maximize highest total reward.
problem Traditional multi-armed bandit problem objective of maximizing total reward is not suitable in certain applications.
method Adaptive explore-then-commit policy with confidence bounds and adaptive stopping criterion.
result Achieves asymptotic and worst-case regret bounds for the new objective.