This paper introduces glocal explanations for expected goal models in soccer.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
GOIMDA selects inputs to maximize expected influence on a goal functional, reducing data acquisition needs.
A new EM framework for goal-conditioned RL improves performance on sparse reward tasks.
We determine the optimal amount to invest in a Black-Scholes financial market for an individual who consumes at a rate equal to a constant proportion of her wealth and who wishes to minimize the expected time that her wealth spends in drawdown during her lifetime. Drawdown occurs when wealth is less than some fixed pro…
GO-CBED optimizes experiments for specific causal queries, improving efficiency.
New approach to goal-based investing using hedging and reinforcement learning.
In Multi-Goal Reinforcement Learning, an agent learns to achieve multiple goals with a goal-conditioned policy. During learning, the agent first collects the trajectories into a replay buffer, and later these trajectories are selected randomly for replay. However, the achieved goals in the replay buffer are often biase…
A new method for experimental design focuses on predicting downstream quantities of interest.
Our goal in this paper is to propose an alternative risk measure which takes into account the fluctuations of losses and possible correlations between random variables. This new notion of risk measures, that we call Copula Conditional Tail Expectation describes the expected amount of risk that can be experienced given …
We consider a multiobjective multiarmed bandit problem with lexicographically ordered objectives. In this problem, the goal of the learner is to select arms that are lexicographic optimal as much as possible without knowing the arm reward distributions beforehand. We capture this goal by defining a multidimensional for…
Controller seeks informative system observations to predict nonlinear dynamics.
Generative neural nets learn deep policies conditioned on goals.
Proposes data-driven methods for estimating conditional expectations.
We consider a diffusion approximation to an insurance risk model where an external driver models a stochastic environment. The insurer can buy reinsurance. Moreover, investment in a financial market is possible. The financial market is also driven by the environmental process. Our goal is to maximise terminal expected …
Active inference minimizes expected free energy for optimal behavior.
A new method reduces both input and output dimensions for better goal-oriented analysis.
We reformulate data-dependent constraints to ensure they are always met with high probability.
In this paper, we introduce the Preselection Bandit problem, in which the learner preselects a subset of arms (choice alternatives) for a user, which then chooses the final arm from this subset. The learner is not aware of the user's preferences, but can learn them from observed choices. In our concrete setting, we all…
This paper calibrates Gaussian process predictive distributions for Bayesian optimization to improve sampling decisions.
We propose a planning and perception mechanism for a robot (agent), that can only observe the underlying environment partially, in order to solve an image classification problem. A three-layer architecture is suggested that consists of a meta-layer that decides the intermediate goals, an action-layer that selects local…
Paper tackles goal-directed generation of discrete structures using conditional generative models.
We consider the problem of estimating the expected value of information (the knowledge gradient) for Bayesian learning problems where the belief model is nonlinear in the parameters. Our goal is to maximize some metric, while simultaneously learning the unknown parameters of the nonlinear belief model, by guiding a seq…
In query learning, the goal is to identify an unknown object while minimizing the number of "yes" or "no" questions (queries) posed about that object. A well-studied algorithm for query learning is known as generalized binary search (GBS). We show that GBS is a greedy algorithm to optimize the expected number of querie…
Agents learn and control complex mechanical systems through shared memories.
The health outcomes of high-need patients can be substantially influenced by the degree of patient engagement in their own care. The role of care managers includes that of enrolling patients into care programs and keeping them sufficiently engaged in the program, so that patients can attain various goals. The attainmen…
A new GP interpolation method for better predictive distributions in ranges of interest.
The goal of policy gradient approaches is to find a policy in a given class of policies which maximizes the expected return. Given a differentiable model of the policy, we want to apply a gradient-ascent technique to reach a local optimum. We mainly use gradient ascent, because it is theoretically well researched. The …
A key quantity of interest in Bayesian inference are expectations of functions with respect to a posterior distribution. Markov Chain Monte Carlo is a fundamental tool to consistently compute these expectations via averaging samples drawn from an approximate posterior. However, its feasibility is being challenged in th…
Many popular reinforcement learning problems (e.g., navigation in a maze, some Atari games, mountain car) are instances of the episodic setting under its stochastic shortest path (SSP) formulation, where an agent has to achieve a goal state while minimizing the cumulative cost. Despite the popularity of this setting, t…
The paper solves a financial mathematics problem using polytopes and probability measures.
Bayesian methods improve inference for cumulative probit models on large datasets.
Experimentally, it has been observed that humans and animals often make decisions that do not maximize their expected utility, but rather choose outcomes randomly, with probability proportional to expected utility. Probability matching, as this strategy is called, is equivalent to maximum entropy reinforcement learning…
Bayesian optimization has demonstrated impressive success in finding the optimum input x* and output f* = f(x*) = max f(x) of a black-box function f. In some applications, however, the optimum output f* is known in advance and the goal is to find the corresponding optimum input x*. In this paper, we consider a new sett…
We consider an economic agent (a household or an insurance company) modelling its surplus process by a deterministic process or by a Brownian motion with drift. The goal is to maximise the expected discounted spendings/dividend payments, given that the discounting factor is given by an exponential CIR process. In the d…
The problem of multi-hypothesis testing with controlled sensing of observations is considered. The distribution of observations collected under each control is assumed to follow a single-parameter exponential family distribution. The goal is to design a policy to find the true hypothesis with minimum expected delay whi…
The goal of subspace learning is to find a -dimensional subspace of , such that the expected squared distance between instance vectors and the subspace is as small as possible. In this paper we study subspace learning in a partial information setting, in which the learner can only observe att…
This work analyzes the value of future reward information in RL.
DGC clusters data with side-information for better prediction.
The expected improvement (EI) algorithm is a popular strategy for information collection in optimization under uncertainty. The algorithm is widely known to be too greedy, but nevertheless enjoys wide use due to its simplicity and ability to handle uncertainty and noise in a coherent decision theoretic framework. To pr…
Reinforcement learning requires manual specification of a reward function to learn a task. While in principle this reward function only needs to specify the task goal, in practice reinforcement learning can be very time-consuming or even infeasible unless the reward function is shaped so as to provide a smooth gradient…
Develops hierarchical reinforcement learning value function approximators.
This paper introduces a new method for uncertainty quantification in prediction models.
As Computer Vision moves from a passive analysis of pixels to active analysis of semantics, the breadth of information algorithms need to reason over has expanded significantly. One of the key challenges in this vein is the ability to identify the information required to make a decision, and select an action that will …
One fundamental goal in any learning algorithm is to mitigate its risk for overfitting. Mathematically, this requires that the learning algorithm enjoys a small generalization risk, which is defined either in expectation or in probability. Both types of generalization are commonly used in the literature. For instance, …
We discuss two distinct approaches, for distorting risk measures of sums of dependent random variables, which preserve the property of coherence. The first, based on distorted expectations, operates on the survival function of the sum. The second, simultaneously applies the distortion on the survival function of the su…
We determine the optimal strategy for investing in a Black-Scholes market in order to maximize the probability that wealth at death meets a bequest goal , a type of goal-seeking problem, as pioneered by Dubins and Savage (1965, 1976). The individual consumes at a constant rate , so the level of wealth required fo…
We consider the applications of the Frank-Wolfe (FW) algorithm for Apprenticeship Learning (AL). In this setting, we are given a Markov Decision Process (MDP) without an explicit reward function. Instead, we observe an expert that acts according to some policy, and the goal is to find a policy whose feature expectation…
The paper learns personalized treatment rules from observational data.