Paper improves AIRL by enhancing policy imitation and addressing reward recovery issues.
problem Inadequate policy imitation and limited transferable reward recovery in AIRL.
method Substituted built-in algorithm with SAC for policy updating and proposed PPO-AIRL + SAC hybrid framework.
result SAC improves policy imitation but hinders reward recovery; PPO-AIRL + SAC achieves satisfactory transfer effect.
A new method recovers rewards from behavior policies using classification and regression.
problem Recovering meaningful rewards from observed behavior in reinforcement learning.
method GenPQR, a modular procedure that estimates behavior policy, evaluates soft Q-function, and recovers normalized reward using classification and regression.
result GenPQR matches or improves reward recovery compared to DeepPQR, while being simpler and more modular.
The study analyzes how neural reward models learn features for policy optimization in a Gaussian single-index model.
problem Reward modeling in policy optimization and its impact on downstream value.
method Two-stage neural reward model: first learns hidden direction, then fits readout layer.
result For any feature-learning temperature above a dimension-free threshold, a constant fraction of neurons recover the hidden direction.
Reinforcement learning mimics expert behavior.
problem Learning from expert demonstrations in reinforcement learning.
method Reduction to reinforcement learning with a stationary reward.
result Expert reward can be recovered and imitation learning is bounded.
PopArt efficiently solves sparse linear bandits with tighter recovery guarantees.
problem Sparse linear bandits where rewards depend on a few covariates.
method PopArt: a simple, computationally efficient sparse linear estimation method.
result Improved regret bounds compared to state-of-the-art algorithms.
Deep RL trains a robust humanoid push-recovery policy.
problem Training robust humanoid push-recovery policies.
method Model-free Deep Reinforcement Learning.
result Policy learns robust behaviors across the entire body.
Efficient algorithms for low-rank bandits using subspace recovery.
problem Contextual bandits with low-rank reward matrices.
method Spectral methods for subspace recovery, reformulating as linear bandits.
result Nearly optimal policy evaluation and best policy identification, minimax guarantees for regret minimization.
Unified framework for estimating reward functions in competitive games.
problem Estimating unknown reward functions in competitive games.
method Unified framework with entropy regularization for reward function recovery.
result Strong theoretical guarantees and practical effectiveness demonstrated.
Adversarial RL recovers agent rewards from financial market data simulations.
problem Recovering agent rewards in volatile financial markets with unknown dynamics.
method Adversarial inverse reinforcement learning in latent space simulations.
result Adversarial RL can robustly recover agent rewards from latent space representations of real market data.
We empirically test predictability on asset price by using stock selection rules based on maximum drawdown and its consecutive recovery. In various equity markets, monthly momentum- and weekly contrarian-style portfolios constructed from these alternative selection criteria are superior not only in forecasting directio…
Paper analyzes AIRL in high-dimensional spaces using random matrix theory.
problem AIRL's performance challenges in high-dimensional environments.
method Examined the rank of the matrix derived from transition matrix, applied random matrix theory.
result High-dimensional scenarios reveal transfer limitations not inherent to AIRL framework.
New framework recovers reward and rationality parameters from game behavior.
problem Statistical ambiguity in identifying reward and rationality parameters in competitive games.
method Blind Inverse Game Theory (Blind-IGT) using entropy-regularized Quantal Response Equilibrium and Normalized Least Squares (NLS) estimator.
result Optimal convergence rate of O(N−1/2) for joint parameter recovery. The paper provides robustness guarantees for mode estimation in bandits.
problem Understanding robustness in mode estimation under adversarial data contamination.
method Simple randomization and theoretical analysis of multi-armed bandits.
result Regret guarantees for various modal bandit problems.
We consider a stochastic continuum armed bandit problem where the arms are indexed by the ℓ2 ball Bd(1+ν) of radius 1+ν in Rd. The reward functions r:Bd(1+ν)→R are considered to intrinsically depend on k≪d unknown linear parameters so that $r(\mathbf{x}) = g(\ma…
Paper proposes new γ-regret measure for non-episodic RL.
problem Measuring performance in non-episodic RL environments.
method Introduces γ-regret as a new performance measure and derives bounds. result Closed the gap between lower and upper bounds for γ-regret. New algorithm balances personalization and statistical validity in MRTs.
problem Optimizing decisions in nonstationary settings with habituation and recovery.
method ROGUE-TS Thompson Sampling with probability clipping.
result Achieves lower regret and maintains high statistical power.
We derive an arbitrage free relationship between recovery swap rates, digital default swap spreads and conventional CDS spreads, and argue that the fair forward recovery rate used in recovery swaps must contain a convexity premium over the expected recovery value.
A contraction analysis improves model-based RL's error recovery.
problem Theoretical understanding of model-based reinforcement learning.
method Contraction analysis applied to both stochastic and deterministic state transitions.
result Error reduction in cumulative reward using branched rollouts.
Method uses Seq2Seq learning to automatically generate recovery commands for ICT systems.
problem Manual decision-making for recovery commands is time-consuming and error-prone.
method Seq2Seq neural network model trained on past logs and commands.
result The model can estimate accurate recovery commands from new failures.
New methods combat data poisoning attacks in bandit algorithms using limited verification.
problem Data poisoning attacks on bandit algorithms, especially in the UCB and ETC types.
method Verification-based mechanisms to restore optimal regret with limited verifications.
result A simple modified ETC type bandit algorithm can restore optimal regret with O(logT) verifications. A new model explains U- and Swoosh-shaped stock price recovery during the COVID-19.
problem Modeling stock price recovery during the COVID-19 with V- and L-shaped recovery.
method Introducing a sentiment variable θ to quantify investor sentiment and simulate U- and Swoosh-shaped recovery. result The model explains U- and Swoosh-shaped recovery of sectoral indices with positive sentiment.
This paper improves support recovery in universal one-bit compressed sensing.
problem Support recovery in one-bit compressed sensing for sparse signals.
method Proposes approximate support recovery and superset recovery algorithms with polynomial-time complexity.
result Achieves improved support recovery with fewer measurements compared to existing methods.
This work provides a guaranteed tensor recovery method by combining low-rankness and smoothness priors.
problem Guaranteed tensor recovery with theoretical guarantees for low-rank and smoothness priors.
method Developed a new regularization term that combines low-rankness and smoothness priors, proving exact recovery guarantees.
result Rigorously proved exact recovery guarantees for tensor completion and tensor robust principal component analysis.
This paper tackles tensor recovery from noisy and multi-level quantized measurements.
problem Tensors from multi-level quantized measurements.
method Nonconvex optimization problem with alternating proximal gradient descent.
result The recovery error diminishes to zero with increasing tensor dimensions.
We discuss a general notion of "sparsity structure" and associated recoveries of a sparse signal from its linear image of reduced dimension possibly corrupted with noise. Our approach allows for unified treatment of (a) the "usual sparsity" and "usual ℓ1 recovery," (b) block-sparsity with possibly overlapping blo…
We consider the problem of signal recovery on graphs as graphs model data with complex structure as signals on a graph. Graph signal recovery implies recovery of one or multiple smooth graph signals from noisy, corrupted, or incomplete measurements. We propose a graph signal model and formulate signal recovery as a cor…
IRKSN algorithm achieves sparse recovery with wider applicability conditions.
problem Sparse recovery challenges due to NP-hard nature and restrictive conditions.
method IRKSN algorithm based on k-support norm regularizer. result Achieves sparse recovery with explicit constants and standard linear rate.
Study finds the cutoff for exact recovery in Gaussian mixture models.
problem Determining the separation of cluster centers for exact recovery in Gaussian mixture models.
method Used information theory and SDP relaxation of K-means clustering. result Sharp threshold for exact recovery of cluster labels without assuming cluster center symmetry.
In recent years research on credit risk modelling has mainly focused on default probabilities. Recovery rates are usually modelled independently, quite often they are even assumed constant. Then, however, the structural connection between recovery rates and default probabilities is lost and the tails of the loss distri…
Study optimal portfolio selection with Recovery Average Value at Risk, showing better control over liabilities.
problem Optimizing portfolios with a new risk measure under known or uncertain distributions.
method Existence results for mean-risk optimal portfolios under different distributional assumptions.
result Portfolio selection under Recovery Average Value at Risk provides better control over liabilities.
The paper improves conditions for unique recovery in homomorphic sensing of subspaces.
problem Unique recovery of points in a linear subspace from their images under linear maps.
method Tighter and simpler conditions for unique recovery in single and subspace arrangement cases, extending to noise stability.
result Conditions for unique recovery in homomorphic sensing are improved and unified.
HSNLD solves robust Hankel recovery efficiently and robustly.
problem Robust Hankel recovery of sparse outliers and missing entries.
method Hankel Structured Newton-Like Descent (HSNLD) algorithm.
result HSNLD achieves linear convergence independent of the condition number.
New method improves dictionary recovery from over-realized models.
problem Theoretical guarantees for model recovery in dictionary learning are limited.
method Search over larger over-realized models to facilitate dictionary recovery.
result Model recovery can be upper-bounded by empirical risk and generalization gap.
Unified framework for pattern recovery in penalized and thresholded estimation.
problem Pattern recovery in penalized and thresholded estimation methods.
method Defining a novel pattern notion based on subdifferentials, introducing accessibility and noiseless recovery conditions.
result Unified and extended conditions for pattern recovery in a broad class of penalized estimators.
Paper explores exact recovery of communities in weighted graphs using Gaussian and exponential distributions.
problem Exact recovery of communities in weighted graphs with Gaussian and exponential distributions.
method Introduces a new semi-metric to describe conditions for exact recovery and analyzes conditions for both complete and incomplete graphs.
result Necessary and sufficient conditions for exact recovery are asymptotically tight and applicable to both complete and incomplete graphs.
Guarantees sparse recovery for neural networks with iterative hard thresholding.
problem Recovering sparse network weights in neural networks.
method Structural properties of sparse network weights and iterative hard thresholding algorithm.
result Simple iterative hard thresholding algorithm recovers sparse network weights exactly using linear memory.
We propose and analyze a generic method for community recovery in stochastic block models and degree corrected block models. This approach can exactly recover the hidden communities with high probability when the expected node degrees are of order logn or higher. Starting from a roughly correct community partition …
New risk measure improves creditor protection in financial regulation.
problem Current solvency requirements fail to control the size of recovery on creditors' claims.
method Developed Recovery Value at Risk (Recovery VaR) to control recovery on creditors' claims.
result Recovery VaR flexibly controls recovery on creditors' claims and integrates protection needs into management incentives.
We introduce a general framework to handle structured models (sparse and block-sparse with possibly overlapping blocks). We discuss new methods for their recovery from incomplete observation, corrupted with deterministic and stochastic noise, using block-ℓ1 regularization. While the current theory provides promis…
New method improves traffic data recovery for streaming data.
problem Improve data quality in traffic data for ITS.
method Online robust tensor recovery algorithm leveraging spatio-temporal correlations and local consistency.
result Significantly improved computational efficiency and high recovery accuracy.
We consider the effect of recovery rates on a pool of credit assets. We allow the recovery rate to depend on the defaults in a general way. Using the theory of large deviations, we study the structure of losses in a pool consisting of a continuum of types. We derive the corresponding rate function and show that it has …
Paper develops TLoc framework to improve Telco outdoor position recovery.
problem High data collection cost and poor accuracy in Telco outdoor position recovery.
method Transfer learning applied to Telco outdoor position recovery.
result TLoc framework improves accuracy by 27.58% and 26.12% on 2G GSM and 4G LTE MR datasets.
Unified framework for constructing nonconvex sparse recovery methods.
problem Constructing valid nonconvex regularization functions remains open.
method Unified framework based on probability density function, using Weibull distribution.
result New nonconvex sparse recovery method based on Weibull distribution.
Improves sparse recovery with non-linear Fourier features.
problem Sparse recovery challenges with non-linear Fourier features.
method Characterizes sufficient data points for perfect recovery.
result Sufficient data points depend on kernel matrix.
We find that factors explaining bank loan recovery rates vary depending on the state of the economic cycle. Our modeling approach incorporates a two-state Markov switching mechanism as a proxy for the latent credit cycle, helping to explain differences in observed recovery rates over time. We are able to demonstrate ho…
While defaults are rare events, losses can be substantial even for credit portfolios with a large number of contracts. Therefore, not only a good evaluation of the probability of default is crucial, but also the severity of losses needs to be estimated. The recovery rate is often modeled independently with regard to th…
Study exact community recovery in noisy SBM with limited queries.
problem Community recovery in noisy stochastic block models with limited queries.
method Balanced uniform querying, two-stage adaptive strategy, sublinear queries, subsampled graph.
result Adaptive querying can improve exact recovery limits in noisy SBM.
The paper tracks patient recovery using graphs of joint movement data.
problem Tracking individual patient recovery trajectories in physical therapy.
method Bayesian learning of Random Geometric Graphs from joint movement data.
result Optimal exercise routines can be recommended based on patient recovery data.