No-regret learning with strategic experts, incentivized.
problem Online learning with strategic experts who misreport beliefs.
method Building on wagering mechanisms, we provide algorithms for no-regret and incentive compatibility in both full and partial information settings.
result Our algorithms achieve no regret and incentive compatibility for myopic experts, with comparable regret to classic no-regret algorithms and diminishing regret for forward-looking agents.
Two-stage mechanism designs reduce regret in recommender systems with stochastic covariates.
problem Designing effective recommender systems with user covariates sampled online.
method Two-stage algorithm integrating incentivized exploration with offline learning methods.
result Achieves sublinear regret while maintaining incentive compatibility.
Incentive-aware recommender system for online platforms.
problem Myopic agents exploit optimal arms, not exploring alternatives.
method Model as multi-agent bandit problem, incentivizes exploration.
result Asymptotically optimal performance with ex-post fairness.
Study assesses how much security restaking protocols need to pay for.
problem Determining the optimal security level for restaking protocols using token incentives.
method Expanding a model by Durvasula and Roughgarden to include strategic attackers and node operators, constructing an approximation algorithm for token-based incentives.
result Restaking protocols can be secure with proper incentive management, even against strategic adversaries.
COBRA addresses strategic behavior in online platforms by ensuring truthful reporting without monetary incentives.
problem Ensuring truthful reporting from strategic agents in online platforms.
method Proposes COBRA, an algorithm for contextual bandits involving strategic agents that disincentivizes strategic behavior.
result COBRA achieves sub-linear regret guarantee and incentive compatibility without monetary incentives.
A novel incentive mechanism improves fairness and participation in federated learning.
problem Low-quality clients and lack of fairness in federated learning.
method Client selection process and money transfer mechanism to ensure fairness and participation.
result The proposed incentive mechanism improves the duration and fairness of federated learning.
This paper addresses reward estimation and incentive design for agents with hidden rewards.
problem Estimating and incentivizing agents with unknown rewards in a learning setting.
method Repeated adverse selection game with a self-interested learning agent and a learning principal. Introduces an estimator for consistent reward estimation and a data-driven incentive policy.
result Finite-sample consistency of the estimator and a rigorous regret bound for the principal.
New method to understand incentives from complex models.
problem Understanding how complex models incentivize actions.
method Formulated as a Markov Decision Process (MDP) and solved using MDP tools.
result Identifies optimal actions to maximize model output.
New framework solves dynamic bilevel optimization problems in reinforcement learning.
problem Dynamic objective functions in reinforcement learning and human feedback.
method Principled penalty-based methods for bilevel reinforcement learning.
result Demonstrated effectiveness of penalty-based algorithms in simulations.
Mobile payment incentives optimized using merchant transaction networks.
problem Optimizing marketing campaigns with limited budgets.
method Graph representation learning on transaction networks.
result Effective modeling of merchant sensitivity to incentives.
We refine toxicity bounds for dynamic liquidation incentives in CP-AMM systems.
problem Ensuring stability in dynamic liquidation incentives in automated market makers.
method Derived state-dependent toxicity bounds for dynamic liquidation incentives, reconciling them with CP-AMM price dynamics.
result State-dependent bounds and liquidity-depth-only condition for dynamic liquidation incentives.
Model shows government incentives boost green bond investment.
problem Increasing green investments through government incentives.
method Optimal incentives indexed on bond prices and covariation, applied to a portfolio of bonds.
result Method outperforms current tax-incentives systems in green investments.
Study optimal incentives for cleaner energy production.
problem Accelerate transition to cleaner technologies in energy market.
method Stochastic control models for three scenarios: single firm, two firms, and two firms without incentives.
result Optimal strategies for investment and production emerge, highlighting firm interactions and incentive effects.
Modeling incentives for content creators on algorithm-curated platforms.
problem Maximizing exposure for content creators on algorithmic platforms.
method Formalized exposure game model, proving effects of algorithmic choices on equilibria, proposing tools for finding equilibria.
result Algorithmic choices significantly affect content exposure and creator behavior.
The design of personalized incentives or recommendations to improve user engagement is gaining prominence as digital platform providers continually emerge. We propose a multi-armed bandit framework for matching incentives to users, whose preferences are unknown a priori and evolving dynamically in time, in a resource c…
Exchange uses incentives to optimize limit order book dynamics.
problem Optimizing market liquidity in fragmented electronic markets.
method Modeling limit order book as SPDE and using control theory to design incentives.
result Exchange can design incentives to modify order book shape and increase liquidity.
We study platforms in the sharing economy and discuss the need for incentivizing users to explore options that otherwise would not be chosen. For instance, rental platforms such as Airbnb typically rely on customer reviews to provide users with relevant information about different options. Yet, often a large fraction o…
AI task delegation faces incentive collapse with unbounded payments as AI accuracy rises.
problem Incentive collapse in AI-assisted task delegation schemes.
method General impossibility result and sentinel-auditing payment mechanism.
result Sentinel-auditing mechanism enforces positive human effort at finite cost, independent of AI accuracy.
Study on liquidity and market efficiency in auction games with imperfect information.
problem Generating liquidity in illiquid auction markets with imperfect information.
method Characterized Nash equilibria in a two-player game with imperfect information, linking market spreads to signal strength.
result Without incentives, the market is inefficient and does not lead to trades. Quadratic fees indexed on half spread can generate liquidity.
The paper develops an economic foundation for multi-agent learning in markets.
problem Learning dynamics in markets with strategic externalities.
method A two-phase incentive mechanism that estimates and uses implementable transfers to steer long-run dynamics.
result The mechanism achieves sublinear social-welfare regret and asymptotically optimal welfare under mild rationality and exploration conditions.
Study shows online learning algorithms incentivize low-quality content, proposing new algorithms to improve quality.
problem Online learning algorithms in content recommender systems incentivize producers to create low-quality content.
method Analyzed the game between producers and content quality, designed new learning algorithms to incentivize high effort and quality.
result New algorithms incentivize producers to invest high effort and achieve high user welfare, improving content quality.
Method uses ANN to estimate incentive salience from large behavioral data.
problem Estimating incentive salience in naturalistic settings.
method Artificial Neural Networks (ANNs) for latent state approximation.
result ANNs produce better representations for predicting future behaviour.
How can we design safe reinforcement learning agents that avoid unnecessary disruptions to their environment? We show that current approaches to penalizing side effects can introduce bad incentives, e.g. to prevent any irreversible changes in the environment, including the actions of other agents. To isolate the source…
Model shows incentives in shared order book can lead to free-rider problem.
problem Incentives in shared order books can lead to free-rider problem.
method Developed a Principal-Agent model with CARA utility functions.
result Equilibrium analysis shows incentives can lead to reduced competition.
Study designs incentives for adapting multi-agent systems without knowing their learning dynamics.
problem Designing incentives for an adapting population in multi-agent systems without prior knowledge of their learning dynamics.
method Introduces a model-based non-episodic Reinforcement Learning (RL) formulation for steering Markovian agents towards desired policies, focusing on history-dependent strategies to handle model uncertainty.
result Identifies conditions for the existence of steering strategies to guide agents to desired policies and provides empirical algorithms to approximately solve the objective.
Study designs steering rewards for MFGs with unknown dynamics and model uncertainty.
problem Designing incentives for large populations of agents in MFGs with uncertain model details.
method Developed optimistic exploration algorithms for agents with no-adaptive regret behaviors.
result Sub-linear regret guarantees for cumulative gaps between agent behaviors and desired outcomes.
Study optimizes health incentives to balance efficiency and fairness.
problem Designing health incentives to balance efficiency and fairness.
method Inverse behavioral optimization framework integrating QALY-based incentives and adaptive learning.
result Modern health systems operate near an efficiency-saturated frontier, with small fairness adjustments yielding diminishing returns.
Framework trains safe agents avoiding deceptive behavior.
problem Training safe agents from unsafe incentives.
method Formal settings, causal influence analysis, maximizing non-mediated effects.
result Agents avoid manipulating delicate state for rewards.
Paper proposes incentive mechanism to encourage participation in federated learning.
problem Users are reluctant to participate in federated learning due to privacy concerns.
method Formulated as a two-stage Stackelberg game, designed an incentive mechanism to select and compensate users.
result Demonstrated effectiveness of the proposed incentive mechanism through simulations.
Study shows ethanol blends and incentives can significantly reduce transportation carbon emissions.
problem Rapid growth in electric vehicles requires complementary strategies to decarbonize transportation.
method Analysis of ethanol blending, regulatory incentives, and economic assessments.
result Ethanol blending, especially E15 and E85, can substantially reduce carbon emissions and provide economic benefits.
Paper proposes FMore to incentivize edge nodes in federated learning with MEC.
problem Incentivizing edge nodes in federated learning with MEC resources.
method Multi-dimensional procurement auction with K winners.
result FMore improves model accuracy and reduces training rounds for AI tasks.
New term ADS describes how machine learning can change user behavior.
problem Machine learning systems can unintentionally change user behavior, affecting performance.
method Introduced `unit tests` and mitigation strategy for hidden incentives in auto-induced distributional shift.
result Meta-learning and Q-learning sometimes fail unit tests but pass with mitigation strategy.
Optimal bidding strategy for multi-platform ad auctions under budget constraints.
problem Optimizing ad placements for budget-constrained advertisers across multiple platforms.
method Developed an optimal bidding strategy for non-incentive-compatible auctions with budget constraints.
result Maximized total utility across auctions while satisfying budget constraints in expectation.
New algorithm reduces regret in strategic prediction problem.
problem Designing an IC algorithm with sublinear regret for strategic experts.
method Developed a new algorithm WSU-UX and proved a worst-case regret bound.
result WSU-UX suffers a Ω ( T 2 / 3 ) Ω(T^{2/3}) Ω ( T 2/3 ) lower bound on regret. When the planning horizon is long, and the safe asset grows indefinitely, isoelastic portfolios are nearly optimal for investors who are close to isoelastic for high wealth, and not too risk averse for low wealth. We prove this result in a general arbitrage-free, frictionless, semimartingale model. As a consequence, op…
Market makers and exchanges use deep reinforcement learning to optimize fees and trading flows.
problem Optimizing fees and trading flows in a lit and dark pool market.
method Solve stochastic control problem, derive optimal contract, design deep reinforcement learning algorithms.
result Deep reinforcement learning algorithms approximate optimal controls and incentives.
Study shows visual feedback and monetary incentives reduce plugload energy consumption in commercial buildings.
problem Mitigating energy consumption in commercial buildings through occupant plugload control.
method Field experiments with visual feedback and monetary incentives in government and university buildings.
result Mean energy reduction of ~9.52% in office environments and ~21.61% in university environments with visual feedback.
A generalized gamification framework is introduced as a form of smart infrastructure with potential to improve sustainability and energy efficiency by leveraging humans-in-the-loop strategy. The proposed framework enables a Human-Centric Cyber-Physical System using an interface to allow building managers to interact wi…
dYdX updates liquidity provider incentives to enhance trading efficiency.
problem Incentivizing liquidity providers to maintain efficient market structures.
method Analyzed various metrics (makerVolume, depths, spreads) and used historical trades to update the LP Incentives Programme.
result Updated the LP Incentives Programme to encourage more active and efficient liquidity.
The study compares M6 competitors' performance to industry benchmarks and discusses incentives for investment managers.
problem Investors seek to understand the performance and skill of M6 competitors beyond the competition's metrics.
method Comparative analysis using financial metrics, factor models, and new strategies.
result Most competitors do not generate significant out-performance compared to industry benchmarks, but some show skill in recent performance.
Study allocates resources to strategic agents while balancing cost and incentives.
problem Dynamic allocation of reusable resources to strategic agents with private valuations under long-term cost constraints.
method Incentive-aware framework combining epoch-based lazy updates and randomized exploration rounds.
result Achieves i l d e O ( T ) ilde{\mathcal{O}}(\sqrt{T}) i l d e O ( T ) social welfare regret, satisfies all cost constraints, and ensures incentive alignment. Study incentive efficiency in monopoly insurance markets with hidden information.
problem Maximizing social welfare in a monopoly insurance market with hidden agent types.
method Maximizes social welfare function subject to incentive compatibility and individual rationality constraints.
result Optimal menus of contracts depend on the level of social welfare weight and agent risk attitudes.
Forward hedging reshapes incentive provision in firms.
problem How does forward hedging affect incentive provision in firms?
method We consider a CARA framework to jointly characterize optimal production, compensation, and static hedging in equilibrium.
result Delegation and external hedging are partial substitutes, and delegation can increase firm value even when the agent is more risk averse.
New algorithm for learning preferences in decentralized matching markets reduces regret to logarithmic levels.
problem Learning preferences in decentralized matching markets without direct communication.
method Introduces a new algorithm for two-sided matching markets with competition.
result The algorithm achieves logarithmic stable regret in shared preferences and quadratic regret in general preferences.
Paper tackles online learning for DR management with incentives.
problem Estimating baseline consumption in DR programs with consumer incentives.
method Online learning scheme using least-squares with perturbed reward prices.
result Achieves low regret of $\mathcal{O}\left((\log{T})^2
ight)$ compared to optimal.
New algorithms learn stable matchings from uncertain user preferences.
problem Learning stable matchings from uncertain user preferences.
method Stochastic multi-armed bandit problem, incentive-aware learning objective, primal-dual formulation.
result Near-optimal regret bounds for learning stable matchings.
Firms delay write-downs for adverse macroeconomic and industry outcomes but not for firm-specific issues.
problem Timeliness of write-downs for adverse macroeconomic and industry outcomes versus firm-specific issues.
method Comparative analysis of write-downs driven by macroeconomic and industry outcomes versus firm-specific outcomes.
result Firms delay write-downs for adverse macroeconomic and industry outcomes but not for firm-specific issues.
This research simplifies lending pools in decentralized finance for better understanding and security.
problem Complexity and lack of executable models make lending pools hard to understand and predict.
method Developed a formal model to reflect common features of lending pools and proved general properties.
result Proved correct handling of funds and described vulnerabilities and attacks.