Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

23466891 · Jun 202019922001200920172026
48 results for Unobserved Rewards

Study of repeated games with unobserved agent rewards using MAB framework.

problem Designing policies for principals in repeated principal-agent games with unobservable agent rewards.
method Developed a policy achieving low regret (square-root regret up to a log factor) for perfect-knowledge agents.
result Constructed an estimator for agent's expected reward and designed a policy achieving low regret.

Paper tackles transfer RL under unobserved context, developing methods to reduce bias.

problem Transfer RL with unobserved contextual information leading to biased models.
method Develops causal bounds on transition and reward functions using demonstrator's data.
result Proposes Q learning and UCB-Q learning algorithms that converge to true value function without bias.

Modeling driver trajectories using inverse reinforcement learning and random utility.

problem Modeling rational driver behavior in road networks from sparse sensor data.
method Apply random utility theory to model unknown reward function, introduce extended state, and use Markov decision process.
result Maximum entropy inverse reinforcement learning is a special case of the proposed approach.

We study a variant of the stochastic multi-armed bandit (MAB) problem in which the rewards are corrupted. In this framework, motivated by privacy preservation in online recommender systems, the goal is to maximize the sum of the (unobserved) rewards, based on the observation of transformation of these rewards through a…

2017-08-16abs ↗pdf ↗

New algorithms improve efficiency in learning from personalized rewards.

problem Learning from personalized rewards in recommendation systems.
method Developed provably efficient algorithms with sublinear regret for context-dependent feedback.
result Introduced a Lipschitz reward estimator that improves generalization performance.

Off-policy evaluation for MNAR rewards in MDPs

problem Off-policy evaluation in MDPs with MNAR rewards
method Formalizing a reward-dependent propensity model and using future states as shadow variables
result Proposed an Fitted-Q-Evaluation-style estimator that propagates recovered rewards while allowing target policies to depend on past missingness indicators

This work highlights problems with off-policy estimation in recommender systems due to unobserved confounders.

problem Evaluation of recommender systems under unobserved confounders.
method Policy-based estimators and characterisation of statistical bias due to confounding.
result Naive propensity estimation under confounding leads to severely biased metric estimates.

Study tackles OPE in confounded settings, estimating policy value from proxies.

problem Difficulty in OPE due to unobserved confounders in infinite-horizon RL.
method Two-stage approach: estimating stationary distribution ratios and combining optimal balancing.
result Policy value can be identified from off-policy data with proxies and latent variable model.

Generative Flow Networks use submodular upper bounds to generate more data.

problem Generating data from unknown, complex reward functions efficiently.
method Introduce submodular upper bounds to estimate reward, use Optimism in the Face of Uncertainty principle to train GFNs.
result SUBo-GFN generates significantly more data than classical GFNs.

We tackle linear bandits with partially observable features, achieving sublinear regret.

problem Linear regret due to unobserved features in partially observable linear bandits.
method Feature augmentation with orthogonal basis vectors and a doubly robust estimator.
result Sublinear regret bound of ildeO((d+dh)T) ilde{O}(\sqrt{(d + d_h)T}).

Algorithm learns optimal coordination for strategic agents in uncertain settings.

problem Optimizing rewards for strategic agents with private types and actions.
method Combines delaying mechanism, reward angle estimation, and LinUCB algorithm.
result Near optimal regret bound of O~(T)\tilde{O}(\sqrt{T}) for learning optimal policy.

In a linear stochastic bandit model, each arm is a vector in an Euclidean space and the observed return at each time step is an unknown linear function of the chosen arm at that time step. In this paper, we investigate the problem of learning the best arm in a linear stochastic bandit model, where each arm's expected r…

2019-06-26abs ↗pdf ↗

Paper analyzes AIRL in high-dimensional spaces using random matrix theory.

problem AIRL's performance challenges in high-dimensional environments.
method Examined the rank of the matrix derived from transition matrix, applied random matrix theory.
result High-dimensional scenarios reveal transfer limitations not inherent to AIRL framework.

Proposes DRRO to mitigate over-optimization in RLHF from human feedback.

problem Over-optimization due to reward misspecification in RLHF.
method Wasserstein distributionally robust regret optimization (DRRO).
result DRRO mitigates over-optimization more effectively than existing baselines.

New method learns optimal policies in presence of unmeasured confounders.

problem Optimal policy learning with unobserved confounders.
method Causal-assisted policy learning methods using instrumental variables and negative controls.
result Policies are ildeO(n1/2) ilde{\mathscr{O}}(n^{-1/2}) quantile-optimal under mild coverage assumptions.

Imitation learning allows agents to learn complex behaviors from demonstrations. However, learning a complex vision-based task may require an impractical number of demonstrations. Meta-imitation learning is a promising approach towards enabling agents to learn a new task from one or a few demonstrations by leveraging e…

2019-06-07abs ↗pdf ↗

Bayesian bandits misspecification affects UX optimization, revealing new models.

problem Misspecification of value models in Bayesian bandits impacts UX optimization.
method Formulated UXO as a restless, sleeping bandit with unobserved confounders and optional stopping. Provided model extensions to address misspecifications.
result Common misspecifications lead to sub-optimal rewards, demonstrating overdispersion's effects on bandit performance.

New framework tackles stochastic latent subgroup heterogeneity in online decision-making.

problem Stochastic latent heterogeneity in online decision-making where individual responses vary with unobserved subgroups.
method Latent heterogeneous bandit framework using EM-greedy algorithm to learn subgroup probabilities and reward parameters.
result Achieves optimal estimation and classification guarantees, revealing a fundamental stochastic barrier in online decision-making.

KRCD detects unobserved confounders in nonlinear observational data.

problem Detecting unobserved confounders in nonlinear observational studies.
method Kernel Regression Confounder Detection (KRCD) using reproducing kernel Hilbert spaces.
result KRCD outperforms existing methods and achieves superior computational efficiency.

Valid causal inference with unobserved confounding in high-dimensional settings.

problem Estimating causal effects with unobserved confounders in high-dimensional data.
method Proposes methods to estimate causal effects with valid confidence intervals in the presence of unobserved confounders and high-dimensional nuisance models.
result Valid semiparametric inference can be obtained with unobserved confounding, and uncertainty intervals are proposed.

A new method uses randomized trials to estimate the strength of unobserved confounding.

problem Unobserved confounding compromises causal conclusions from non-randomized studies.
method Designs a statistical test to detect unobserved confounding strength and estimates a lower bound.
result Estimates an asymptotically valid lower bound on unobserved confounding strength.

DiffATD efficiently discovers targets in partially observable environments using diffusion dynamics.

problem Efficiently discovering targets in partially observable environments with limited sampling.
method DiffATD uses diffusion dynamics to maintain a belief distribution over unobserved states, balancing exploration and exploitation.
result DiffATD outperforms baselines and supervised methods in diverse domains.

New method estimates treatment effects over time with unobserved confounders.

problem Estimating treatment effects from observational data with unobserved confounders.
method Sequential Deconfounder using Gaussian process latent variable model.
result Unbiased estimates of individualized treatment responses over time.

A federated learning algorithm tackles unknown contexts in multi-arm bandits.

problem Learning optimal actions in federated multi-arm bandits with unobserved contexts.
method Elimination-based algorithm for linearly parametrized reward functions.
result Proved regret bound for linearly parametrized reward functions.

We consider reinforcement learning in parameterized Markov Decision Processes (MDPs), where the parameterization may induce correlation across transition probabilities or rewards. Consequently, observing a particular state transition might yield useful information about other, unobserved, parts of the MDP. We present a…

2014-06-29abs ↗pdf ↗

Modeling the purposeful behavior of imperfect agents from a small number of observations is a challenging task. When restricted to the single-agent decision-theoretic setting, inverse optimal control techniques assume that observed behavior is an approximately optimal solution to an unknown decision problem. These tech…

2013-08-15abs ↗pdf ↗

New method recovers predictions from unobservable source subpopulation in binary classification.

problem Challenging binary classification with unobservable subpopulation in source domain.
method Distribution matching method to estimate subpopulation proportions, rigorous derivation of prediction models.
result Our method outperforms naive benchmarks in synthetic and real-world datasets.

New method estimates policy performance under unobserved confounding.

problem Estimating policy performance when decisions depend on unobserved variables.
method Developed worst-case bounds for robust OPE under unobserved confounding.
result Efficient procedure for computing worst-case bounds, proving statistical consistency.

We propose stochastic rank-11 bandits, a class of online learning problems where at each step a learning agent chooses a pair of row and column arms, and receives the product of their values as a reward. The main challenge of the problem is that the individual values of the row and column are unobserved. We assume tha…

2016-08-10abs ↗pdf ↗

CDVAE estimates treatment effects over time by accounting for unobserved variables.

problem Estimating treatment effects over time in the presence of unobserved confounders.
method Causal Dynamic Variational Autoencoder (CDVAE) that addresses unconfoundedness and unobserved heterogeneity.
result CDVAE outperforms existing methods in estimating Conditional Average Treatment Effects (CATEs).

Proposes ρρ-GNF for sensitivity analysis of unobserved confounding.

problem Sensitivity analysis of unobserved confounding in observational studies.
method Copulas and normalizing flows to estimate average causal effect (ACE) as a function of unobserved confounding strength.
result Develops ρcurveρ_{curve} to provide bounds for ACE and identify confounding strength required to nullify ACE.

Paper adapts DML for panel data, addressing unobserved heterogeneity.

problem Estimating causal effects with panel data and unobserved heterogeneity.
method Adapting double/debiased machine learning (DML) for panel data with predictive models based on correlated random effects.
result Predictive models based on correlated random effects within DML lead to accurate coefficient estimates.

New method removes hidden confounders for unbiased treatment effect estimation.

problem Bias in treatment effect estimation due to unobserved confounders.
method Proposes a new debiased estimation approach via SVD to handle heterogeneous confounding.
result Established rate of convergence for the estimator under different noise conditions.

Proposes a new Hawkes process bandit model for disaster search and rescue.

problem Forecasting and detecting spatio-temporal events with undersampled or biased data.
method Upper confidence bound algorithm using Bayesian spatial Hawkes process estimation.
result Model outperforms state-of-the-art spatial MAB algorithms in disaster search and rescue.

The paper tackles robust domain generalization by accounting for unobserved confounders.

problem Learning robust, generalizable models from multiple datasets in the presence of unobserved confounders.
method Defines a new invariance property for causal solutions, connects it to distributionally robust optimization, and incorporates regularization to encourage partial equality of error derivatives.
result Demonstrates the empirical effectiveness of the approach on healthcare data from various modalities.

We propose a general formulation for addressing reinforcement learning (RL) problems in settings with observational data. That is, we consider the problem of learning good policies solely from historical data in which unobserved factors (confounders) affect both observed actions and rewards. Our formulation allows us t…

2018-12-26abs ↗pdf ↗

Improved method for unbiased causal discovery in presence of unobserved confounding.

problem Unbiased data synthesis for causal discovery algorithms in the presence of unobserved confounding.
method Explicit block-hierarchical ancestral sampling to address limitations of implicit parameterization.
result Our approach fully covers the space of causal models, including those generated by implicit parameterization.

New framework for estimating treatment effects in observational studies.

problem Estimating average treatment effects in the presence of unobserved confounders.
method Distributionally robust optimization, sensitivity models.
result Sharp bounds on average treatment effects under distributional assumptions.

Thompson Sampling improves decision-making in partially observed contexts.

problem Balancing exploration and exploitation in partially observed contextual bandits.
method Thompson Sampling policy for learning optimal arms from noisy linear functions of unobserved context vectors.
result Thompson Sampling achieves poly-logarithmic regret and square-root consistency of parameter estimation.

New method identifies causal effects with categorical unobserved confounders.

problem Estimating causal effects in the presence of unobserved confounders.
method Mixture learning and tensor decomposition for consistent estimation.
result Causal effects are identifiable with categorical unobserved confounders under suitable conditions.

The paper shows how to audit fairness in decisions with hidden risk factors.

problem Estimating fairness in decisions influenced by hidden, unobservable risk factors.
method Derives unbiased estimates of risk using historical data and audits existing decision-making systems.
result One can compute meaningful bounds on treatment rates for high-risk individuals, even with hidden confounders.

Develops Austen plots for assessing bias from unobserved confounding in observational studies.

problem Bias in causal estimates due to unobserved confounding.
method Formalizes confounding strength, uses Austen plots to visualize and quantify bias.
result Allows domain experts to assess the plausibility of strong confounders.