The paper tackles AI advice giving by considering adherence levels and defer options.
problem Inadequate consideration of human adherence to AI recommendations.
method Sequential decision-making model that considers adherence levels and incorporates a defer option.
result Specialized learning algorithms provide better convergence and empirical performance.
Human advice improves deep learning from sparse samples.
problem Learning from sparse, noisy samples in deep models.
method Knowledge-augmented Column Networks using human advice.
result Significantly improved performance or faster convergence.
Interactive machine learning improves deep RL in Minecraft by giving action advice.
problem Training deep RL agents in high-aliasing environments like Minecraft is computationally expensive.
method Conducted experiments with two RL algorithms, Feedback Arbitration, and Newtonian Action Advice, to give action advice to human teachers.
result Action advice from human teachers can improve agent performance in high-aliasing environments.
Proposes a deep learning framework guided by human advice.
problem Learning from sparse and noisy data.
method Knowledge-augmented Column Networks, leveraging human advice.
result Improves model performance in domains with structured representations.
DPG makes reinforcement learning safer with human advice.
problem Unsafe reinforcement learning in shared environments.
method Extends Policy Gradient to incorporate human directives.
result DPG learns faster and more safely than reward-based methods.
LIMEADE improves AI advice for opaque models, enhancing accuracy and user satisfaction.
problem Lack of advice methods for opaque AI models.
method Develops a general framework to translate advice into model updates.
result Improves accuracy and user satisfaction compared to baselines.
The paper introduces a protocol for including advice in robot interaction to improve performance.
problem Improving natural language interaction between humans and robots.
method Protocol for including advice in robot interaction, evaluating on blocks world task.
result Simple advice can lead to significant performance improvements in robot interaction.
Study shows LLM-advisors match human performance in eliciting preferences but struggle with conflicting needs and trust.
problem How do LLM-advisors perform in complex financial domains where domain expertise is crucial?
method Lab-based user study with 64 participants, focusing on three challenges: preference elicitation, personalized guidance, and relationship building.
result LLM-advisors can match human performance in preference elicitation but struggle with conflicting needs and trust issues.
Advice-efficient prediction with expert advice (in analogy to label-efficient prediction) is a variant of prediction with expert advice game, where on each round of the game we are allowed to ask for advice of a limited number M M M out of N N N experts. This setting is especially interesting when asking for advice of ever…
Dynamic ensemble active learning tackles non-stationary criteria in active learning.
problem Active learning's effectiveness varies across datasets and sessions, leading to suboptimal results.
method Developed a dynamic ensemble active learner based on a non-stationary multi-armed bandit with expert advice.
result Dynamic ensemble selects the best criteria at each step, improving overall performance.
Improved learning of multivariate Gaussians with imperfect advice.
problem Learning multivariate Gaussians with inaccurate advice.
method Developed learning algorithms for multivariate Gaussians using imperfect advice.
result Achieved better sample complexity for learning multivariate Gaussians with imperfect advice.
New algorithm learns causal structure with advice, improving efficiency.
problem Learning causal structure with side information (advice).
method Adaptive search algorithm for active causal structure learning with advice.
result Intervention cost is at most O ( max { 1 , log ψ } ) O(\max\{1, \log ψ\}) O ( max { 1 , log ψ }) times the cost for verifying the true structure, matching state-of-the-art. New algorithm learns Gaussian policies from corrective human feedback, outperforming current methods.
problem Learning from corrective human feedback for complex systems.
method Gaussian Process Coach (GPC) that uses Gaussian Processes and policy uncertainty for optimal feedback selection and learning rate adaptation.
result Demonstrated superior performance in OpenAI Gym benchmarks compared to COACH.
A new method for learning to defer decisions with expert advice improves over standard methods.
problem Learning to defer decisions with expert advice in systems where expert information can be modified after selection.
method An augmented surrogate that operates on the composite expert-advice action space, providing consistency guarantees and excess-risk bounds.
result The method improves over standard Learning-to-Defer and adapts its advice acquisition behavior to the cost regime.
New algorithm uses imperfect advice to improve online bipartite matching performance.
problem Online bipartite matching with imperfect advice.
method Designing an algorithm that uses external advice to improve performance between advice-free methods and optimal ratio.
result Algorithm achieves competitive ratio interpolating between advice-free methods and optimal ratio of 1.
Investigates fast prediction rates with limited expert advice.
problem Minimizing excess generalization error with limited expert access.
method Assumes Lipschitz and strongly convex loss, designs novel algorithms.
result Achieves fast rates of O(1/T) with optimal number of expert advices.
New algorithms improve on consistency and robustness in convex function chasing with black-box advice.
problem Minimizing cost in normed vector space with black-box advice for convex function chasing.
method Two novel algorithms: INTERP and BDINTERP, exploiting convexity to achieve improved consistency and robustness.
result BDINTERP achieves near-optimal consistency-robustness trade-off for α-polyhedral cost functions.
Online symptom checkers have significant potential to improve patient care, however their reliability and accuracy remain variable. We hypothesised that an artificial intelligence (AI) powered triage and diagnostic system would compare favourably with human doctors with respect to triage and diagnostic accuracy. We per…
Study finds optimal regret bound for multi-armed bandit problem with expert advice.
problem Optimizing decision-making in a multi-armed bandit problem with expert advice.
method Proved a tight lower bound matching the upper bound of Kale (2014) for minimax expected regret.
result The minimax optimal expected regret is Θ(√(T K log (N/K))) for the problem.
A new method combines human feedback with deep learning for faster policy learning.
problem Training deep neural networks for complex decision-making problems is data-intensive and time-consuming.
method Deep COACH (D-COACH) integrates human corrective feedback with deep learning models.
result The D-COACH framework can learn policies for continuous action spaces faster than traditional deep reinforcement learning.
Conventional learning with expert advice methods assumes a learner is always receiving the outcome (e.g., class labels) of every incoming training instance at the end of each trial. In real applications, acquiring the outcome from oracle can be costly or time consuming. In this paper, we address a new problem of active…
Improved regret bounds for bandits with expert advice.
problem Optimizing decision-making in environments with expert advice.
method Proved lower and upper bounds for regret in restricted and standard feedback models.
result Proved a new upper bound of order K T ln ( N / K ) \sqrt{K T \ln(N/K)} K T ln ( N / K ) for the worst-case regret, matching a previously known lower bound. Improved algorithm reduces regret in corrupted expert advice setting.
problem Prediction with expert advice in the presence of adversarial corruption.
method Multiplicative Weights algorithm with decreasing step sizes.
result Achieves constant regret and optimal performance in various environments.
A new method for student-initiated action advice using novelty detection.
problem Exploration and sample inefficiency in RL, especially with teacher absence.
method Random Network Distillation (RND) to measure advice novelty, updates only for advised states.
result Significant performance improvement over state-of-the-art methods, especially in challenging scenarios.
Improved regret bounds for bandits with fixed expert advice using information theory.
problem Optimizing regret in bandit problems with fixed expert distributions.
method Information-theoretic analysis and KL-divergence measures.
result First regret bounds for EXP4 that can get arbitrarily close to zero under certain conditions.
XNAS optimizes neural architecture search using expert advice theory.
problem Optimizing neural architecture selection with minimal regret.
method Uses prediction with expert advice theory and dynamic architecture wiping.
result Achieves optimal worst-case regret and state-of-the-art results.
Improved bounds for online prediction with expert advice.
problem Online prediction with expert advice in finite-horizon games.
method Verification arguments from optimal control theory applied to PDEs to find sub- and supersolutions.
result Explicit bounds for any number of experts and horizon, improving upon previous results.
Generative AI reduces herd behavior in trading, but can also lead to optimal herding.
problem Impact of generative AI on financial stability and herd behavior.
method Laboratory experiments with large language models replicating human trading behavior.
result AI agents make more rational decisions than humans, reducing herd behavior but also potentially leading to optimal herding.
We provide the first algorithm for online bandit linear optimization whose regret after T rounds is of order sqrt{Td ln N} on any finite class X of N actions in d dimensions, and of order d*sqrt{T} (up to log factors) when X is infinite. These bounds are not improvable in general. The basic idea utilizes tools from con…
ACTRCE uses natural language to improve reinforcement learning performance.
problem Sparse reward in reinforcement learning.
method Extends HER framework with natural language goal representation.
result ACTRCE can solve challenging 3D navigation tasks and generalize to unseen instructions.
Improves reward bounds for prediction with expert advice using abstention.
problem Prediction with expert advice under bandit feedback with abstention.
method CBA algorithm exploiting abstention to improve reward bounds.
result Achieved significant improvement in reward bounds for general confidence-rated predictors.
Adapts model-based advice to stabilize black-box policies for nonlinear control.
problem Stabilizing machine-learned policies for nonlinear control with limited model information.
method Proposes an adaptive λ λ λ -confident policy to combine black-box and model-based advice. result Proves the stability of the adaptive λ λ λ -confident policy and its competitive ratio. WSB's investment advice significantly outperformed the S&P500 over 3 years, but not consistently.
problem Reliability of investment advice from WSB community.
method Data analysis of WSB posts and stock performance from 2019-2021.
result A WSB portfolio grew 200% over 3 years and 480% over 1 year, outperforming S&P500.
We prove non-asymptotic lower bounds on the expectation of the maximum of d d d independent Gaussian variables and the expectation of the maximum of d d d independent symmetric random walks. Both lower bounds recover the optimal leading constant in the limit. A simple application of the lower bound for random walks is an (…
Sharp bounds found on expert error in binary advice aggregation.
problem Aggregating binary advice from conditionally independent experts.
method Sharp upper and lower bounds on optimal error probability in asymmetric case.
result Sharp bounds recover and sharpen known results in symmetric case.
Optimal algorithm reduces regret in adversarial bandit problem with multiple plays.
problem Minimizing regret in adversarial bandit problem with multiple plays.
method Introducing a new expert advice algorithm for multiple-play setting, achieving minimax optimal regret bounds.
result Minimizes regret asymptotically to the best switching strategy with optimal bounds.
Generalized algorithm for translation and scale-invariant prediction.
problem Sequential prediction with expert advice, focusing on translation and scale invariance.
method Designing a generalized online algorithm using the universal prediction perspective to compete against a generic class of expert selection strategies.
result No preliminary knowledge of loss sequences is required; performance bounds are stable under arbitrary scalings and translations.
In the framework of prediction with expert advice, we consider a recently introduced kind of regret bounds: the bounds that depend on the effective instead of nominal number of experts. In contrast to the Normal- Hedge bound, which mainly depends on the effective number of experts but also weakly depends on the nominal…
Two adaptive algorithms improve tracking regret in dynamic expert advice problems.
problem Prediction with expert advice in dynamic environments.
method Developed two adaptive and efficient algorithms using online mirror descent framework.
result Achieved data-dependent tracking regret bounds for both algorithms.
New algorithms improve prediction with expert advice under local differential privacy.
problem Predicting expert advice with privacy constraints.
method Design of two new algorithms: RW-AdaBatch and RW-Meta, leveraging limited-switching behavior and random walks.
result RW-Meta outperforms classical and central DP algorithms by 1.5-3x on predicting hospital COVID patient densities.
A new framework uses deep RL to aggregate expert advice for better portfolio management.
problem Improving portfolio management through expert advice and deep reinforcement learning.
method Convolutional networks for signal aggregation and historical price data, Proximal Policy Optimization algorithm.
result Our framework can achieve 90% of the best expert's profit on average.
Survey of algorithms to correct past mistakes in prediction.
problem Improving prediction accuracy by correcting past errors.
method Defensive Forecasting as a sequential game theory approach to minimize prediction metrics.
result Simple, near-optimal algorithms for various prediction tasks.
Paper studies continuous prediction with experts' advice using differential equations.
problem Continuous prediction with experts' advice in online learning.
method Continuous-time stochastic calculus and differential equations.
result Improved guarantees for quantile regret with continuous-time algorithm.
Hedge algorithm proves optimal in stochastic expert advice problems.
problem Prediction with expert advice in stochastic setting.
method Analyzed Hedge algorithm with decreasing learning rate in online stochastic setting.
result Hedge algorithm is worst-case optimal and adaptive in stochastic setting.
The leaderboard in machine learning competitions is a tool to show the performance of various participants and to compare them. However, the leaderboard quickly becomes no longer accurate, due to hack or overfitting. This article gives two pieces of advice to prevent easy hack or overfitting. By following these advice,…
We consider an original problem that arises from the issue of security analysis of a power system and that we name optimal discovery with probabilistic expert advice. We address it with an algorithm based on the optimistic paradigm and on the Good-Turing missing mass estimator. We prove two different regret bounds on t…
The paper proposes calibration to improve algorithm performance using machine learning predictions.
problem Improving real-world performance of online algorithms with machine learning predictions.
method Calibration as a tool to bridge the gap between prediction uncertainty and algorithm design.
result Calibrated advice leads to more effective guidance in high-variance settings and significant performance improvements in real-world data.
A key challenge in online learning is that classical algorithms can be slow to adapt to changing environments. Recent studies have proposed "meta" algorithms that convert any online learning algorithm to one that is adaptive to changing environments, where the adaptivity is analyzed in a quantity called the strongly-ad…