Sharp bounds found on expert error in binary advice aggregation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Advice-efficient prediction with expert advice (in analogy to label-efficient prediction) is a variant of prediction with expert advice game, where on each round of the game we are allowed to ask for advice of a limited number out of experts. This setting is especially interesting when asking for advice of ever…
New algorithms reduce label collection for online prediction with expert advice.
In a binary classification problem the feature vector (predictor) is the input to a scoring function that produces a decision value (score), which is compared to a particular chosen threshold to provide a final class prediction (output). Although the normal assumption of the scoring function is important in many applic…
The paper tackles AI advice giving by considering adherence levels and defer options.
Improved learning of multivariate Gaussians with imperfect advice.
We study the multiclass online learning problem where a forecaster makes a sequence of predictions using the advice of experts. Our main contribution is to analyze the regime where the best expert makes at most mistakes and to show that when , the expected number of mistakes made by the optima…
New algorithm learns causal structure with advice, improving efficiency.
A new method for learning to defer decisions with expert advice improves over standard methods.
New algorithm uses imperfect advice to improve online bipartite matching performance.
Investigates fast prediction rates with limited expert advice.
Training deep reinforcement learning agents complex behaviors in 3D virtual environments requires significant computational resources. This is especially true in environments with high degrees of aliasing, where many states share nearly identical visual features. Minecraft is an exemplar of such an environment. We hypo…
New algorithms improve on consistency and robustness in convex function chasing with black-box advice.
Over the last few years, there has been growing interest in learning models for physically grounded language understanding tasks, such as the popular blocks world domain. These works typically view this problem as a single-step process, in which a human operator gives an instruction and an automated agent is evaluated …
Study finds optimal regret bound for multi-armed bandit problem with expert advice.
Conventional learning with expert advice methods assumes a learner is always receiving the outcome (e.g., class labels) of every incoming training instance at the end of each trial. In real applications, acquiring the outcome from oracle can be costly or time consuming. In this paper, we address a new problem of active…
LIMEADE improves AI advice for opaque models, enhancing accuracy and user satisfaction.
Improved regret bounds for bandits with expert advice.
A new method for student-initiated action advice using novelty detection.
Improved regret bounds for bandits with fixed expert advice using information theory.
Sparse reward is one of the most challenging problems in reinforcement learning (RL). Hindsight Experience Replay (HER) attempts to address this issue by converting a failed experience to a successful one by relabeling the goals. Despite its effectiveness, HER has limited applicability because it lacks a compact and un…
We provide the first algorithm for online bandit linear optimization whose regret after T rounds is of order sqrt{Td ln N} on any finite class X of N actions in d dimensions, and of order d*sqrt{T} (up to log factors) when X is infinite. These bounds are not improvable in general. The basic idea utilizes tools from con…
Recently, deep models have been successfully applied in several applications, especially with low-level representations. However, sparse, noisy samples and structured domains (with multiple objects and interactions) are some of the open challenges in most deep models. Column Networks, a deep architecture, can succinctl…
Study shows LLM-advisors match human performance in eliciting preferences but struggle with conflicting needs and trust.
Improves reward bounds for prediction with expert advice using abstention.
Adapts model-based advice to stabilize black-box policies for nonlinear control.
WSB's investment advice significantly outperformed the S&P500 over 3 years, but not consistently.
Recently, deep models have had considerable success in several tasks, especially with low-level representations. However, effective learning from sparse noisy samples is a major challenge in most deep models, especially in domains with structured representations. Inspired by the proven success of human guided machine l…
We prove non-asymptotic lower bounds on the expectation of the maximum of independent Gaussian variables and the expectation of the maximum of independent symmetric random walks. Both lower bounds recover the optimal leading constant in the limit. A simple application of the lower bound for random walks is an (…
Many currently deployed Reinforcement Learning agents work in an environment shared with humans, be them co-workers, users or clients. It is desirable that these agents adjust to people's preferences, learn faster thanks to their help, and act safely around them. We argue that most current approaches that learn from hu…
Generalized algorithm for translation and scale-invariant prediction.
In the framework of prediction with expert advice, we consider a recently introduced kind of regret bounds: the bounds that depend on the effective instead of nominal number of experts. In contrast to the Normal- Hedge bound, which mainly depends on the effective number of experts but also weakly depends on the nominal…
New algorithms improve prediction with expert advice under local differential privacy.
A new framework uses deep RL to aggregate expert advice for better portfolio management.
Survey of algorithms to correct past mistakes in prediction.
Paper studies continuous prediction with experts' advice using differential equations.
The leaderboard in machine learning competitions is a tool to show the performance of various participants and to compare them. However, the leaderboard quickly becomes no longer accurate, due to hack or overfitting. This article gives two pieces of advice to prevent easy hack or overfitting. By following these advice,…
We consider an original problem that arises from the issue of security analysis of a power system and that we name optimal discovery with probabilistic expert advice. We address it with an algorithm based on the optimistic paradigm and on the Good-Turing missing mass estimator. We prove two different regret bounds on t…
The paper proposes calibration to improve algorithm performance using machine learning predictions.
A key challenge in online learning is that classical algorithms can be slow to adapt to changing environments. Recent studies have proposed "meta" algorithms that convert any online learning algorithm to one that is adaptive to changing environments, where the adaptivity is analyzed in a quantity called the strongly-ad…
Fund2Persona creates personalized financial advisor personas from fund data, improving investment advice and manager interpretation.
Improved prediction algorithm for 'easy' sequences with reduced regret.
A simple algorithm improves model generalization in expert advice settings.
WSB community outperforms investment banks in stock picks.
New algorithm improves bandit with graph feedback by decomposing regret.
We revisit the fundamental problem of prediction with expert advice, in a setting where the environment is benign and generates losses stochastically, but the feedback observed by the learner is subject to a moderate adversarial corruption. We prove that a variant of the classical Multiplicative Weights algorithm with …
In some reinforcement learning problems an agent may be provided with a set of input policies, perhaps learned from prior experience or provided by advisors. We present a reinforcement learning with policy advice (RLPA) algorithm which leverages this input set and learns to use the best policy in the set for the reinfo…
New algorithms reduce private bandit regret to nearly non-private levels.