The UCR Time Series Archive expands from 85 to 128 datasets, offering advice and insights.
problem Lack of comprehensive data sets for time series analysis.
method Periodic expansions of the archive, providing advice and novel insights.
result A significant increase in the number of datasets from 85 to 128.
Advice-efficient prediction with expert advice (in analogy to label-efficient prediction) is a variant of prediction with expert advice game, where on each round of the game we are allowed to ask for advice of a limited number M M M out of N N N experts. This setting is especially interesting when asking for advice of ever…
System solves a significant fraction of Bongard problems using visual features and pragmatic reasoning.
problem Solving Bongard problems with intelligent vision systems.
method Image processing, symbolic visual vocabulary, Bayesian inference, pragmatic reasoning.
result Good agreement between induced concepts and Bongard's solutions.
This paper is part of an ongoing investigation of "pragmatic information", defined in Weinberger (2002) as "the amount of information actually used in making a decision". Because a study of information rates led to the Noiseless and Noisy Coding Theorems, two of the most important results of Shannon's theory, we begin …
The paper tackles AI advice giving by considering adherence levels and defer options.
problem Inadequate consideration of human adherence to AI recommendations.
method Sequential decision-making model that considers adherence levels and incorporates a defer option.
result Specialized learning algorithms provide better convergence and empirical performance.
Improved learning of multivariate Gaussians with imperfect advice.
problem Learning multivariate Gaussians with inaccurate advice.
method Developed learning algorithms for multivariate Gaussians using imperfect advice.
result Achieved better sample complexity for learning multivariate Gaussians with imperfect advice.
New algorithm learns causal structure with advice, improving efficiency.
problem Learning causal structure with side information (advice).
method Adaptive search algorithm for active causal structure learning with advice.
result Intervention cost is at most O ( max { 1 , log ψ } ) O(\max\{1, \log ψ\}) O ( max { 1 , log ψ }) times the cost for verifying the true structure, matching state-of-the-art. Interactive machine learning improves deep RL in Minecraft by giving action advice.
problem Training deep RL agents in high-aliasing environments like Minecraft is computationally expensive.
method Conducted experiments with two RL algorithms, Feedback Arbitration, and Newtonian Action Advice, to give action advice to human teachers.
result Action advice from human teachers can improve agent performance in high-aliasing environments.
The paper introduces a protocol for including advice in robot interaction to improve performance.
problem Improving natural language interaction between humans and robots.
method Protocol for including advice in robot interaction, evaluating on blocks world task.
result Simple advice can lead to significant performance improvements in robot interaction.
Study finds AI-generated financial advice influences life cycle investing patterns.
problem Understanding how AI-generated financial advice impacts life cycle investing.
method Sentiment analysis of prompts from AI-generated financial advice and simulation of lifetime effects.
result AI-generated financial advice leads to life cycle investing patterns, influenced by gender and AI experience.
A new method for learning to defer decisions with expert advice improves over standard methods.
problem Learning to defer decisions with expert advice in systems where expert information can be modified after selection.
method An augmented surrogate that operates on the composite expert-advice action space, providing consistency guarantees and excess-risk bounds.
result The method improves over standard Learning-to-Defer and adapts its advice acquisition behavior to the cost regime.
Human advice improves deep learning from sparse samples.
problem Learning from sparse, noisy samples in deep models.
method Knowledge-augmented Column Networks using human advice.
result Significantly improved performance or faster convergence.
New algorithm uses imperfect advice to improve online bipartite matching performance.
problem Online bipartite matching with imperfect advice.
method Designing an algorithm that uses external advice to improve performance between advice-free methods and optimal ratio.
result Algorithm achieves competitive ratio interpolating between advice-free methods and optimal ratio of 1.
Investigates fast prediction rates with limited expert advice.
problem Minimizing excess generalization error with limited expert access.
method Assumes Lipschitz and strongly convex loss, designs novel algorithms.
result Achieves fast rates of O(1/T) with optimal number of expert advices.
New algorithms improve on consistency and robustness in convex function chasing with black-box advice.
problem Minimizing cost in normed vector space with black-box advice for convex function chasing.
method Two novel algorithms: INTERP and BDINTERP, exploiting convexity to achieve improved consistency and robustness.
result BDINTERP achieves near-optimal consistency-robustness trade-off for α-polyhedral cost functions.
Unified framework for hybrid learning and optimization via active inference.
problem Sequential decisions in black-box evaluations requiring both task improvement and uncertainty reduction.
method Pragmatic Curiosity (PraC) framework that evaluates queries by balancing information gain and pragmatic value.
result Unified approach reduces decision risk and improves coverage of critical regions without task-specific rules.
Study finds optimal regret bound for multi-armed bandit problem with expert advice.
problem Optimizing decision-making in a multi-armed bandit problem with expert advice.
method Proved a tight lower bound matching the upper bound of Kale (2014) for minimax expected regret.
result The minimax optimal expected regret is Θ(√(T K log (N/K))) for the problem.
A number of approaches to solving the well-known transfer pricing problem are known. However, few models satisfactorily resolve the core problem of allowing both the source and receiving divisions to earn a profit on transfers during a period in such a way that sub-optimal output levels are avoided. In 1969, Samuel pro…
Conventional learning with expert advice methods assumes a learner is always receiving the outcome (e.g., class labels) of every incoming training instance at the end of each trial. In real applications, acquiring the outcome from oracle can be costly or time consuming. In this paper, we address a new problem of active…
LIMEADE improves AI advice for opaque models, enhancing accuracy and user satisfaction.
problem Lack of advice methods for opaque AI models.
method Develops a general framework to translate advice into model updates.
result Improves accuracy and user satisfaction compared to baselines.
Improved regret bounds for bandits with expert advice.
problem Optimizing decision-making in environments with expert advice.
method Proved lower and upper bounds for regret in restricted and standard feedback models.
result Proved a new upper bound of order K T ln ( N / K ) \sqrt{K T \ln(N/K)} K T ln ( N / K ) for the worst-case regret, matching a previously known lower bound. Improved algorithm reduces regret in corrupted expert advice setting.
problem Prediction with expert advice in the presence of adversarial corruption.
method Multiplicative Weights algorithm with decreasing step sizes.
result Achieves constant regret and optimal performance in various environments.
A new method for student-initiated action advice using novelty detection.
problem Exploration and sample inefficiency in RL, especially with teacher absence.
method Random Network Distillation (RND) to measure advice novelty, updates only for advised states.
result Significant performance improvement over state-of-the-art methods, especially in challenging scenarios.
Proposes a deep learning framework guided by human advice.
problem Learning from sparse and noisy data.
method Knowledge-augmented Column Networks, leveraging human advice.
result Improves model performance in domains with structured representations.
Improved regret bounds for bandits with fixed expert advice using information theory.
problem Optimizing regret in bandit problems with fixed expert distributions.
method Information-theoretic analysis and KL-divergence measures.
result First regret bounds for EXP4 that can get arbitrarily close to zero under certain conditions.
XNAS optimizes neural architecture search using expert advice theory.
problem Optimizing neural architecture selection with minimal regret.
method Uses prediction with expert advice theory and dynamic architecture wiping.
result Achieves optimal worst-case regret and state-of-the-art results.
Optimizes explanations for better listener understanding.
problem Insufficient consideration of listener preferences in concept-based explanations.
method Iterative training procedure based on direct preference optimization.
result Pragmatic explanations improve both model accuracy and user understanding.
Improved bounds for online prediction with expert advice.
problem Online prediction with expert advice in finite-horizon games.
method Verification arguments from optimal control theory applied to PDEs to find sub- and supersolutions.
result Explicit bounds for any number of experts and horizon, improving upon previous results.
DPG makes reinforcement learning safer with human advice.
problem Unsafe reinforcement learning in shared environments.
method Extends Policy Gradient to incorporate human directives.
result DPG learns faster and more safely than reward-based methods.
We provide the first algorithm for online bandit linear optimization whose regret after T rounds is of order sqrt{Td ln N} on any finite class X of N actions in d dimensions, and of order d*sqrt{T} (up to log factors) when X is infinite. These bounds are not improvable in general. The basic idea utilizes tools from con…
ACTRCE uses natural language to improve reinforcement learning performance.
problem Sparse reward in reinforcement learning.
method Extends HER framework with natural language goal representation.
result ACTRCE can solve challenging 3D navigation tasks and generalize to unseen instructions.
Study shows LLM-advisors match human performance in eliciting preferences but struggle with conflicting needs and trust.
problem How do LLM-advisors perform in complex financial domains where domain expertise is crucial?
method Lab-based user study with 64 participants, focusing on three challenges: preference elicitation, personalized guidance, and relationship building.
result LLM-advisors can match human performance in preference elicitation but struggle with conflicting needs and trust issues.
Improves reward bounds for prediction with expert advice using abstention.
problem Prediction with expert advice under bandit feedback with abstention.
method CBA algorithm exploiting abstention to improve reward bounds.
result Achieved significant improvement in reward bounds for general confidence-rated predictors.
Adapts model-based advice to stabilize black-box policies for nonlinear control.
problem Stabilizing machine-learned policies for nonlinear control with limited model information.
method Proposes an adaptive λ λ λ -confident policy to combine black-box and model-based advice. result Proves the stability of the adaptive λ λ λ -confident policy and its competitive ratio. WSB's investment advice significantly outperformed the S&P500 over 3 years, but not consistently.
problem Reliability of investment advice from WSB community.
method Data analysis of WSB posts and stock performance from 2019-2021.
result A WSB portfolio grew 200% over 3 years and 480% over 1 year, outperforming S&P500.
We prove non-asymptotic lower bounds on the expectation of the maximum of d d d independent Gaussian variables and the expectation of the maximum of d d d independent symmetric random walks. Both lower bounds recover the optimal leading constant in the limit. A simple application of the lower bound for random walks is an (…
Sharp bounds found on expert error in binary advice aggregation.
problem Aggregating binary advice from conditionally independent experts.
method Sharp upper and lower bounds on optimal error probability in asymmetric case.
result Sharp bounds recover and sharpen known results in symmetric case.
New meta algorithm improves adaptability in changing environments.
problem Adapting to changing environments in online learning.
method Derives a new parameter-free algorithm for the LEA problem, inspired by coin betting.
result Strongly-adaptive regret bound is log ( T ) \sqrt{\log(T)} log ( T ) better than other algorithms. Optimal algorithm reduces regret in adversarial bandit problem with multiple plays.
problem Minimizing regret in adversarial bandit problem with multiple plays.
method Introducing a new expert advice algorithm for multiple-play setting, achieving minimax optimal regret bounds.
result Minimizes regret asymptotically to the best switching strategy with optimal bounds.
Generalized algorithm for translation and scale-invariant prediction.
problem Sequential prediction with expert advice, focusing on translation and scale invariance.
method Designing a generalized online algorithm using the universal prediction perspective to compete against a generic class of expert selection strategies.
result No preliminary knowledge of loss sequences is required; performance bounds are stable under arbitrary scalings and translations.
In the framework of prediction with expert advice, we consider a recently introduced kind of regret bounds: the bounds that depend on the effective instead of nominal number of experts. In contrast to the Normal- Hedge bound, which mainly depends on the effective number of experts but also weakly depends on the nominal…
Two adaptive algorithms improve tracking regret in dynamic expert advice problems.
problem Prediction with expert advice in dynamic environments.
method Developed two adaptive and efficient algorithms using online mirror descent framework.
result Achieved data-dependent tracking regret bounds for both algorithms.
Unsupervised machine learning helps design complex experiments more efficiently.
problem Designing experiments with many factors and constraints is challenging and costly.
method Applied a beta variational autoencoder (beta-VAE) to represent trials in a low-dimensional latent space.
result Generated pragmatic designs with fewer trials while maintaining objectives.
New algorithms improve prediction with expert advice under local differential privacy.
problem Predicting expert advice with privacy constraints.
method Design of two new algorithms: RW-AdaBatch and RW-Meta, leveraging limited-switching behavior and random walks.
result RW-Meta outperforms classical and central DP algorithms by 1.5-3x on predicting hospital COVID patient densities.
New algorithm improves online learning with expert advice and metric learning.
problem Improving online learning performance in changing environments.
method Parameter-free online learning algorithm using coin betting.
result Strongly adaptive regret bound improvement of at least sqrt(log(T)).
A new framework uses deep RL to aggregate expert advice for better portfolio management.
problem Improving portfolio management through expert advice and deep reinforcement learning.
method Convolutional networks for signal aggregation and historical price data, Proximal Policy Optimization algorithm.
result Our framework can achieve 90% of the best expert's profit on average.
Improved ML approach for quantifying stroke rehabilitation.
problem No tool to measure functional training after stroke.
method Refined machine learning algorithm and sensor configurations.
result Linear Discriminant Analysis (LDA) showed highest accuracy.
Survey of algorithms to correct past mistakes in prediction.
problem Improving prediction accuracy by correcting past errors.
method Defensive Forecasting as a sequential game theory approach to minimize prediction metrics.
result Simple, near-optimal algorithms for various prediction tasks.