AI agent learns to handle unknown unknown states in reinforcement learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new estimator for evaluating policies in unknown environments.
Develops robust MDPs for unknown disturbances with performance guarantees.
ZSPO optimizes RL from unknown link functions using human feedback.
New method optimizes portfolio weights as functions, outperforming traditional approaches.
New method for fair resource allocation in AI-aware networks with unknown utility functions.
This paper describes a new form of unsupervised learning, whose input is a set of unlabeled points that are assumed to be local maxima of an unknown value function v in an unknown subset of the vector space. Two functions are learned: (i) a set indicator c, which is a binary classifier, and (ii) a comparator function h…
New algorithms estimate function levels with near-optimal efficiency.
Study approximates unknown function levels with queries.
Optimizes risk measures given known marginal distributions of two unknown factors.
In many applications, such as economics, operations research and reinforcement learning, one often needs to estimate a multivariate regression function f subject to a convexity constraint. For example, in sequential decision processes the value of a state under optimal subsequent decisions may be known to be convex or …
We address the problem of maximizing an unknown submodular function that can only be accessed via noisy evaluations. Our work is motivated by the task of summarizing content, e.g., image collections, by leveraging users' feedback in form of clicks or ratings. For summarization tasks with the goal of maximizing coverage…
Regression analysis is a standard supervised machine learning method used to model an outcome variable in terms of a set of predictor variables. In most real-world applications we do not know the true value of the outcome variable being predicted outside the training data, i.e., the ground truth is unknown. It is hence…
We consider the estimation of two-sample integral functionals, of the type that occur naturally, for example, when the object of interest is a divergence between unknown probability densities. Our first main result is that, in wide generality, a weighted nearest neighbour estimator is efficient, in the sense of achievi…
NP-PROV separates mean and variance spaces to improve function uncertainty.
Solves inventory control with unknown demand trend using singular control.
New BO method optimizes functions efficiently even with unknown hyperparameters.
Framework for completing computational graphs using Gaussian Processes.
Interesting theoretical associations have been established by recent papers between the fields of active learning and stochastic convex optimization due to the common role of feedback in sequential querying mechanisms. In this paper, we continue this thread in two parts by exploiting these relations for the first time …
For many important problems the quantity of interest is an unknown function of the parameters, which is a random vector with known statistics. Since the dependence of the output on this random vector is unknown, the challenge is to identify its statistics, using the minimum number of function evaluations. This problem …
Applying Bayesian optimization in problems wherein the search space is unknown is challenging. To address this problem, we propose a systematic volume expansion strategy for the Bayesian optimization. We devise a strategy to guarantee that in iterative expansions of the search space, our method can find a point whose f…
We consider the valuation problem of an (insurance) company under partial information. Therefore we use the concept of maximizing discounted future dividend payments. The firm value process is described by a diffusion model with constant and observable volatility and constant but unknown drift parameter. For transformi…
This paper, to be regularly updated, lists those prime knots with the fewest possible number of crossings for which values of basic knot invariants, such as the unknotting number or the smooth 4-genus, are unknown. This list is being developed in conjunction with "KnotInfo" (www.indiana.edu/~knotinfo), a web-based tabl…
Online minimization of an unknown convex function over the interval is considered under first-order stochastic bandit feedback, which returns a random realization of the gradient of the function at each query point. Without knowing the distribution of the random gradients, a learning algorithm sequentially choo…
Paper tackles open set domain adaptation by detecting unknown classes.
GP-MRO discovers robust mixed strategies for unknown objectives.
Many problems in financial engineering involve the estimation of unknown conditional expectations across a time interval. Often Least Squares Monte Carlo techniques are used for the estimation. One method that can be combined with Least Squares Monte Carlo is the "Regress-Later" method. Unlike conventional methods wher…
Proposes a method to handle missing inputs in Bayesian optimization.
A new algorithm optimizes unknown functions with noisy data and unmatched features.
New algorithms optimize multiple tasks with shared similarities, reducing regret.
Entropy Search (ES) and Predictive Entropy Search (PES) are popular and empirically successful Bayesian Optimization techniques. Both rely on a compelling information-theoretic motivation, and maximize the information gained about the of the unknown function; yet, both are plagued by the expensive computatio…
Study one-shot strategic classification under unknown costs, improving worst-case accuracy.
We reduce boundary determination of an unknown function and its normal derivatives from the (possibly weighted and attenuated) broken ray data to the injectivity of certain geodesic ray transforms on the boundary. For determination of the values of the function itself we obtain the usual geodesic ray transform, but for…
The paper explores when and why value decomposition algorithms work in cooperative multi-agent reinforcement learning.
We consider the problem of global optimization of an unknown non-convex smooth function with zeroth-order feedback. In this setup, an algorithm is allowed to adaptively query the underlying function at different locations and receives noisy evaluations of function values at the queried points (i.e. the algorithm has ac…
A novel GPUM constructs Gaussian Processes for unknown manifolds with probabilistic metrics.
Advanced and effective collaborative filtering methods based on explicit feedback assume that unknown ratings do not follow the same model as the observed ones (\emph{not missing at random}). In this work, we build on this assumption, and introduce a novel dynamic matrix factorization framework that allows to set an ex…
Enforcing safety is a key aspect of many problems pertaining to sequential decision making under uncertainty, which require the decisions made at every step to be both informative of the optimal decision and also safe. For example, we value both efficacy and comfort in medical therapy, and efficiency and safety in robo…
Classification tasks usually assume that all possible classes are present during the training phase. This is restrictive if the algorithm is used over a long time and possibly encounters samples from unknown classes. The recently introduced extreme value machine, a classifier motivated by extreme value theory, addresse…
In this paper, we generalize Huber's criterion to multichannel sparse recovery problem of complex-valued measurements where the objective is to find good recovery of jointly sparse unknown signal vectors from the given multiple measurement vectors which are different linear combinations of the same known elementary vec…
DeepDPM clusters images without knowing the number of clusters.
Learning to make decisions from observed data in dynamic environments remains a problem of fundamental importance in a number of fields, from artificial intelligence and robotics, to medicine and finance. This paper concerns the problem of learning control policies for unknown linear dynamical systems so as to maximize…
New algorithms learn graphons in GMFGs without knowing them.
Safe Bayesian optimization method using information theory.
Proposes PE-GP-UCB for time-varying Bayesian optimisation.
New mixture models for clustering and density estimation of unknown distributions.
We study pool-based active learning with abstention feedbacks where a labeler can abstain from labeling a queried example with some unknown abstention rate. This is an important problem with many useful applications. We take a Bayesian approach to the problem and develop two new greedy algorithms that learn both the cl…
CoinDICE estimates confidence intervals for unknown behavior policies in reinforcement learning.