New algorithm reduces regret by allowing free exploration in multi-armed bandits.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Early stopping method saves up to 75% computation time in policy search tasks.
Develops an anytime-valid framework for optimal policy identification from logged contextual bandit data.
The effects of saving and spending patterns on holding time distribution of money are investigated based on the ideal gas-like models. We show the steady-state distribution obeys an exponential law when the saving factor is set uniformly, and a power law when the saving factor is set diversely. The power distribution c…
Cash management is concerned with optimizing the short-term funding requirements of a company. To this end, different optimization strategies have been proposed to minimize costs using daily cash flow forecasts as the main input to the models. However, the effect of the accuracy of such forecasts on cash management pol…
Reinforcement learning is a general technique that allows an agent to learn an optimal policy and interact with an environment in sequential decision making problems. The goodness of a policy is measured by its value function starting from some initial state. The focus of this paper is to construct confidence intervals…
We consider a simple model of a closed economic system where the total money is conserved and the number of economic agents is fixed. In analogy to statistical systems in equilibrium, money and the average money per economic agent are equivalent to energy and temperature, respectively. We investigate the effect of the …
Reinforcement learning (RL) for robotics is challenging due to the difficulty in hand-engineering a dense cost function, which can lead to unintended behavior, and dynamical uncertainty, which makes exploration and constraint satisfaction challenging. We address these issues with a new model-based reinforcement learnin…
Retirees who exhaust their savings while still alive are said to experience financial ruin. These savings are typically grown during the accumulation phase then spent during the retirement decumulation phase. Extensive research into invest-and-harvest decumulation strategies has been conducted, but recommendations diff…
We consider the ideal-gas models of trading markets, where each agent is identified with a gas molecule and each trading as an elastic or money-conserving (two-body) collision. Unlike in the ideal gas, we introduce saving propensity of agents, such that each agent saves a fraction of its money and trades with t…
Paper proposes a sequential statistical test for comparing imitation learning policies with near-optimal stopping.
We have studied numerically the statistical mechanics of the dynamic phenomena, including money circulation and economic mobility, in some transfer models. The models on which our investigations were performed are the basic model proposed by A. Dragulescu and V. Yakovenko [1], the model with uniform saving rate develop…
In this paper, we study how to solve resource allocation problems in ultra-reliable and low-latency communications by unsupervised deep learning, which often yield functional optimization problems with quality-of-service (QoS) constraints. We take a joint power and bandwidth allocation problem as an example, which mini…
Stochastic variance-reduced gradient (SVRG) is an optimization method originally designed for tackling machine learning problems with a finite sum structure. SVRG was later shown to work for policy evaluation, a problem in reinforcement learning in which one aims to estimate the value function of a given policy. SVRG m…
Optimizes pension mix of PAYGO, EET, and individual savings.
Transforms any test into anytime-valid with sample savings.
We consider the ideal-gas models of trading markets, where each agent is identified with a gas molecule and each trading as an elastic or money-conserving (two-body) collision. Unlike in the ideal gas, we introduce saving propensity of agents, such that each agent saves a fraction of its money and trades with t…
Develops CLTs for Markov chain transition probabilities and policies.
This paper studies the evaluation of policies that recommend an ordered set of items (e.g., a ranking) based on some context---a common scenario in web search, ads, and recommendation. We build on techniques from combinatorial bandits to introduce a new practical estimator that uses logged data to estimate a policy's p…
Guiding the design of neural networks is of great importance to save enormous resources consumed on empirical decisions of architectural parameters. This paper constructs shallow sigmoid-type neural networks that achieve 100% accuracy in classification for datasets following a linear separability condition. The separab…
This paper proposes a paradigm shift in the valuation of long term annuities, away from classical no-arbitrage valuation towards valuation under the real world probability measure. Furthermore, we apply this valuation method to two examples of annuity products, one having annual payments linked to a mortality index and…
We propose two optimization techniques to minimize memory usage and computation while meeting system timing constraints for real-time classification in wearable systems. Our method derives a hierarchical classifier structure for Support Vector Machine (SVM) in order to reduce the amount of computations, based on the pr…
The paper optimizes LLM accuracy by stopping early based on consistent answers.
SIMPOL solves complex economic models using numerical methods.
In this paper, we assume an insure is allowed to purchase proportional reinsurance and can invest his or her wealth into the financial market where a savings account, stocks and bonds are available. Different from classical optimal investment and reinsurance problem, this paper studies the insurer's long-term investmen…
LHIEM model predicts health, income, and employment over years.
New approach uses SPG for semantic communication without a known channel model.
The paper introduces SuccessProbaMax to optimize policy success probability in online advertising.
Algorithm finds safe zones in policy Markov Decision Processes to limit trajectory escape.
Logit dynamics formula reveals self-regulation in softmax policy gradient methods.
DG separates successes and failures by gating updates with advantage and surprisal.
Wasserstein Policy Learning for Distributional Outcomes
Simulation framework assesses ROI of chronic disease adherence and policy timing.
Policy optimization is a core component of reinforcement learning (RL), and most existing RL methods directly optimize parameters of a policy based on maximizing the expected total reward, or its surrogate. Though often achieving encouraging empirical success, its underlying mathematical principle on {\em policy-distri…
Efficient dispatching rule in manufacturing industry is key to ensure product on-time delivery and minimum past-due and inventory cost. Manufacturing, especially in the developed world, is moving towards on-demand manufacturing meaning a high mix, low volume product mix. This requires efficient dispatching that can wor…
Develops a method to estimate optimal policy value in online learning.
Prize linked savings accounts provide a return in the form of randomly chosen accounts receiving large cash prizes, in lieu of a guaranteed and uniform interest rate. This model became legal for American national banks upon bipartisan passage of the American Savings Promotion Act in December 2014, and many states have …
Study minimax off-policy evaluation in multi-armed bandits with known and unknown behavior policies.
Paper derives Thiele's equation for unit-linked policies in a stochastic volatility model.
We consider the relationship between economic activity and intervention, including monetary and fiscal policy, using a universal dynamic framework. Central bank policies are designed for growth without excess inflation. However, unemployment, investment, consumption, and inflation are interlinked. Understanding dynamic…
We identify a fundamental problem in policy gradient-based methods in continuous control. As policy gradient methods require the agent's underlying probability distribution, they limit policy representation to parametric distribution classes. We show that optimizing over such sets results in local movement in the actio…
Study on pooled annuity funds and how initial savings affect income stability.
The paper tackles optimal policy learning with asymmetric counterfactual utilities in healthcare decisions.
The paper analyzes the sample complexities for policy evaluation with linear function approximation.
This work explains why online imitation learning improves faster than theory predicts.
Optimizes data labeling for causal effect estimation with missing outcomes.
We analyze the ideal gas like models of markets and review the different cases where a `savings' factor changes the nature and shape of the distribution of wealth. These models can produce similar distribution of wealth as observed across varied economies. We present a more realistic model where the saving factor can v…
Method curates cost-effective, high-quality datasets using AI models.