The paper tackles adaptive policy selection to maximize social welfare, achieving optimal regret bounds.
problem Maximizing social welfare through adaptive policy selection, considering both private utility and public revenue.
method The approach involves learning response functions through experimentation, deriving lower and upper bounds for regret, and using algorithms like Exp3.
result The algorithm achieves optimal regret bounds, showing that welfare maximization is harder than multi-armed bandit problems.
This paper tackles no-regret learning for fair multi-agent social welfare optimization.
problem Maximizing social welfare in a fair manner for multiple agents.
method Developed algorithms for stochastic and adversarial multi-agent settings, proving regret bounds and tightness.
result Achieved no-regret learning for fair multi-agent social welfare optimization in various settings.
The paper develops an economic foundation for multi-agent learning in markets.
problem Learning dynamics in markets with strategic externalities.
method A two-phase incentive mechanism that estimates and uses implementable transfers to steer long-run dynamics.
result The mechanism achieves sublinear social-welfare regret and asymptotically optimal welfare under mild rationality and exploration conditions.
New framework tackles submodular welfare with multi-agent combinatorial bandits.
problem Maximizing total welfare among agents with shared constraints and submodular utilities under bandit feedback.
method Proposes an explore-then-commit strategy with randomized assignments for multi-agent combinatorial bandits.
result Achieves i l d e O ( T 2 / 3 ) ilde{\mathcal{O}}(T^{2/3}) i l d e O ( T 2/3 ) regret, first for partition-based submodular welfare problem under bandit feedback. The paper automates policy learning for nonlinear welfare criteria using machine learning and debiasing techniques.
problem Learning optimal policies from observational data with nonlinear welfare criteria.
method Modeling a nonlinear welfare criterion with a utility function, estimating propensity scores with machine learning, and using sieve approximations and cross-validation for model selection.
result The proposed policy learning method satisfies oracle inequalities, providing theoretical guarantees on performance.
Dynamic pricing improves DeFi lending efficiency by reducing regret to logarithmic levels.
problem Static pricing mechanisms in DeFi lending protocols lead to suboptimal welfare and revenue.
method Online learning model for static and dynamic pricing models in DeFi lending.
result Adaptive supply models achieve logarithmic regret, outperforming static models.
Proposes a model to incentivize exploration in web platforms with payments.
problem Learning and social welfare goals of web platforms with myopic users.
method Contextual bandit model with payments to incentivize exploration.
result Achieves sublinear regret while maximizing cumulative social welfare.
Framework for online resource allocation using social welfare functions.
problem Optimal allocation of resources over time steps in a population.
method Confidence sequence framework for SWF-based online learning and inference, valid for any monotonic, concave, and Lipschitz-continuous SWF.
result Achieves near-optimal regret of i l d e O ( n + n k T ) ilde{O}(n+\sqrt{nkT}) i l d e O ( n + nk T ) for SWF-agnostic algorithm SWF-UCB. Study recovers investor preferences from portfolio data using synthetic data and robust optimization.
problem Recovering latent investor preferences from observed portfolio allocations under uncertainty.
method Inverse portfolio optimization framework integrating robust optimization and regret-based inference.
result Accurate recovery of transaction cost parameters and partial identifiability of ESG penalties under preference misspecification and market shocks.
Optimizes treatment allocation with spillovers using experimental data.
problem Optimizing treatment allocation with network interference.
method Mixed-integer linear programming for welfare maximization.
result Strong guarantees on policy regret for optimal welfare.
Algorithm learns fair division from noisy feedback in uncertain markets.
problem Learning fair division in uncertain markets with noisy feedback.
method Wrapper algorithms using dual averaging to learn item and agent values from bandit feedback.
result Asymptotically achieves optimal Nash social welfare in linear Fisher markets.
The paper proposes a new policy for optimal treatment allocation based on quantile treatment effects.
problem Optimal treatment allocation policies that target distributional welfare, especially when individuals are heterogeneous.
method The approach involves allocating treatments based on the conditional quantile of individual treatment effects (QoTE), considering both prudent and negligent policymakers.
result The proposed minimax policies are robust to model uncertainty and can be generalized to various settings.
Develops a method to learn optimal timing of treatments from observational data.
problem Choosing the right time to start treatments in dynamic decision-making problems.
method Advantage Doubly Robust Estimator for dynamic treatment rules under sequential ignorability.
result Proves welfare regret bounds and shows promising empirical performance.
The paper tackles fair policy targeting by optimizing allocation rules to minimize unfairness.
problem Discrimination in individualized treatments of social welfare programs.
method Formulated as a mixed-integer linear program, solved using off-the-shelf algorithms, derived regret bounds and small sample guarantees.
result Designs fair and efficient treatment allocation rules within the Pareto frontier.
Optimizes long-term social welfare in recommender systems by matching users to providers.
problem Realistic recommender systems dynamics affect all agents, not just users.
method Formulated as an optimal constrained matching problem, solved using dynamical system equilibrium selection.
result Ensures maximal social welfare with diverse viable providers, improving over myopic matching.
The paper develops methods to estimate optimal treatment sequences under policy constraints.
problem Estimating the best sequence of treatments over multiple stages for individuals.
method Empirical welfare maximization approach, solving treatment assignment sequentially or simultaneously.
result Established convergence rates and upper bounds for estimation methods.
Optimizes treatment allocation using covariates for better outcomes.
problem Improving treatment allocation in multi-armed bandit problems.
method Maximizes a functional of the conditional potential outcome distribution.
result Developed expected regret lower bounds and near minimax optimal policy.
Evaluating AI investment strategies
problem Auditing a black-box algorithmic decision-maker
method Exact decomposition of cumulative regret
result Cumulative regret equals sum of per-period covariances
Wasserstein Policy Learning for Distributional Outcomes
problem Offline policy learning with distribution-valued outcomes
method Establishing statistical guarantees for policy learning framework
result Proven leading dependence on N and N-dim(Π) for finite-sample regret
The paper develops methods for welfare analysis in dynamic models.
problem Estimating welfare metrics in complex, high-dimensional models.
method Dual and doubly robust representations, Lasso and Neural Network estimators.
result Automatic debiasing of welfare metrics without needing bias correction.
New algorithm reduces unfairness in bandit problems by balancing exploration and exploitation.
problem Fairness in bandit problems where early participants can be unfairly disadvantaged.
method Introduces extsf{UCB-HARE} algorithm that balances exploration and exploitation using inverse-weighted harmonic rank schedule.
result Algorithm extsf{UCB-HARE} achieves regret matching the lower bound Ω ( σ k max ( 1 , q ) / T ) Ω(σ\sqrt{k^{\max(1,q)}/T}) Ω ( σ k m a x ( 1 , q ) / T ) for q > 1 q>1 q > 1 . Mechanism designs for unknown agent values in stochastic bandit settings.
problem Designing truthful mechanisms for maximizing social welfare in settings with unknown agent values and stochastic feedback.
method Developed a VCG-like mechanism with regret bounds for multi-round allocations, balancing agent and seller welfare.
result Achieved an $Ω(T^{rac{2}{3}})$ lower bound for the maximum of welfare, agent utilities, and mechanism utility after T T T rounds. Study shows how competition affects learning in matching markets, proving it's possible to balance stability, fairness, and regret.
problem How competition affects learning in matching markets and the impossibility of simultaneously guaranteeing stability and low optimal regret.
method Modeling a two-sided matching market with bandit learners and adding components of costs and transfers.
result It is possible to simultaneously guarantee stability, low optimal regret, fairness in the distribution of regret, and high social welfare.
BRACE addresses noncompliance in bandits, offering methods for recommendation and treatment policies.
problem Noncompliance in bandit problems complicates learning objectives and treatment effects.
method BRACE formalizes objective-choice, identifies direct-control regimes, and proposes a phase-doubling algorithm for IV inversion.
result BRACE delivers valid policy values and structural uncertainty, even under weak identification and homogeneity failure.
Unified model optimizes experiment performance and reduces duration.
problem Balancing reward maximization and experiment termination.
method Unified model that considers both within-experiment and post-experiment outcomes.
result Familiar algorithms can optimize a broad class of objectives with proper parameter adjustment.
The paper analyzes fairness and social welfare in machine learning classification.
problem The relationship between fairness and social welfare in machine learning classification.
method Welfare-based analysis of classification and fairness regimes; algorithm for linear hyperplanes.
result More strict fairness criteria can worsen welfare outcomes for disadvantaged groups.
In treatment allocation problems the individuals to be treated often arrive sequentially. We study a problem in which the policy maker is not only interested in the expected cumulative welfare but is also concerned about the uncertainty/risk of the treatment outcomes. At the outset, the total number of treatment assign…
Current methodologies in machine learning analyze the effects of various statistical parity notions of fairness primarily in light of their impacts on predictive accuracy and vendor utility loss. In this paper, we propose a new framework for interpreting the effects of fairness criteria by converting the constrained lo…
We consider the problem of a single seller repeatedly selling a single item to a single buyer (specifically, the buyer has a value drawn fresh from known distribution D D D in every round). Prior work assumes that the buyer is fully rational and will perfectly reason about how their bids today affect the seller's decisio…
The paper addresses how to complete incomplete risk markets by iteratively enhancing welfare.
problem How to complete incomplete risk markets to enhance welfare.
method Iterative mechanism to complete the market while monotonically enhancing welfare.
result Iterative completion of incomplete risk markets can enhance welfare.
New algorithm for multi-player bandits with selfish players, achieving logarithmic regret.
problem Challenges of robustness to selfish players in multi-player bandits.
method First algorithm robust to selfish players achieving logarithmic regret, with or without collision observation.
result Achieved logarithmic regret for robust algorithms to selfish players in multi-player bandits.
New job recommendation system improves job seekers' welfare through field experiments.
problem Current job recommendation systems focus on clicks and applications, not job seekers' welfare.
method Developed a job-search model with two dimensions: utility and success probability. Conducted field experiments to validate model predictions.
result Welfare-optimal job recommendation algorithms outperform existing approaches and perform close to the benchmark.
Investors suffer welfare loss despite having better information.
problem Welfare loss among investors with absolute information advantages.
method Examined financial markets with heterogenous investors and objective measures of welfare.
result Investors incur welfare loss even with better information, revealing a double loss phenomenon.
The paper addresses treatment recommendation problems by optimizing distributional characteristics.
problem Optimizing treatment recommendations based on distributional targets.
method Characterizes the problem's difficulty and proposes near-optimal policies.
result Characterizes the difficulty of the problem and proposes near-regret optimal policies.
New framework improves cost-benefit analysis of policies.
problem Limitations of MVPF in welfare analysis.
method Developed an axiomatic framework to create RPV.
result RPV provides better equity-efficiency trade-off quantification.
New method optimizes multiple objectives in A/B testing for AI and clinical trials.
problem Minimizing cumulative regret, maximizing CATE, and ensuring differential privacy in large-scale experiments.
method ConSE and DP-ConSE algorithms for sequential segmentation and elimination, achieving Pareto-optimal frontier.
result Privacy comes 'for free' in our framework, with only asymptotically negligible costs to regret and accuracy.
Although both systems analyzed are described through two theories apparently different (quantum mechanics and game theory) it is shown that both are analogous and thus exactly equivalents. The quantum analogue of the replicator dynamics is the von Neumann equation. Quantum mechanics could be used to explain more correc…
The so called "globalization" process (i.e. the inexorable integration of markets, currencies, nation-states, technologies and the intensification of consciousness of the world as a whole) has a behavior exactly equivalent to a system that is tending to a maximum entropy state. This globalization process obeys a collec…
Study bridges welfare maximization and CATE estimation in policy learning.
problem Tackles the gap between empirical welfare maximization and conditional average treatment effect estimation in policy learning.
method Shows equivalence between EWM and least squares over reparameterized policy class, proposes regularization method.
result Both approaches are interchangeable under common conditions and share theoretical guarantees.
Existence of incomplete Radner equilibrium with endogenous noise tracker.
problem Existence of incomplete Radner equilibrium in a model with endogenous noise tracker.
method Proved existence through a coupled system of ODEs, reduced to two coupled ODEs.
result Endogenous noise tracker leads to higher aggregate welfare for large stock supply.
Study uses contextual bandits to optimize charity exposure in donation solicitation.
problem Optimizing charity exposure in donation solicitation using survey responses.
method Adaptive experiment design to balance cumulative regret minimization and simple regret minimization.
result Adaptive experimentation yields better policy learning outcomes than uniform randomization.
Study allocates resources to strategic agents while balancing cost and incentives.
problem Dynamic allocation of reusable resources to strategic agents with private valuations under long-term cost constraints.
method Incentive-aware framework combining epoch-based lazy updates and randomized exploration rounds.
result Achieves i l d e O ( T ) ilde{\mathcal{O}}(\sqrt{T}) i l d e O ( T ) social welfare regret, satisfies all cost constraints, and ensures incentive alignment. New welfare-based fairness notions align with existing error rate balance and predictive parity.
problem Aligning fairness notions with welfare-based criteria.
method Discussing and establishing conditions for envy freeness and prejudice freeness.
result Envy freeness and prejudice freeness are equivalent to error rate balance and predictive parity.
New method identifies algo trading strategies as liquidity consumers or providers.
problem Determining if algo trading strategies consume or provide liquidity.
method Analyzes trade and price history to classify strategies as liquidity consumers or providers.
result Identifies net liquidity consumption or provision of algo trading strategies.
Optimal adaptive experiment for choosing best treatment with binary outcomes.
problem Choosing the best treatment from binary options in an adaptive experiment.
method Adaptive experiment with two phases: treatment allocation and choice. Neyman allocation method used.
result Neyman allocation is minimax and Bayes optimal, matching lower bounds for regret.
Adopting a zonal structure of electricity market requires specification of zones' borders. In this paper we use social welfare as the measure to assess quality of various zonal divisions. The social welfare is calculated by Market Coupling algorithm. The analyzed divisions are found by the usage of extended Locational …
Our work extends Coase's theorem to settings with uncertainty, showing how to maximize social welfare through property rights and learning.
problem Theoretical models of externality often assume perfect knowledge, limiting practical solutions.
method We extend Coase's theorem to a two-player bandit setting with uncertainty, designing a learning policy to maximize social welfare.
result We show that property rights and learning can recover Coase's theorem in settings with uncertainty.
The paper explores fairness, welfare, and equity in personalized pricing across various applications.
problem Interplay of fairness, welfare, and equity in personalized pricing based on customer features.
method Comprehensive literature review and observational metrics without underlying valuation distribution assumptions.
result Personalized pricing can expand access, improve welfare, and increase revenue or budget utilization.