A new model for lifetime maximization with reneging in heterogeneous outcomes.
On-device research index
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
169,051 papers · 148 categories
Trend · papers per month
4 results for “ReNeg”
problem Maximizing lifetime in applications with reneging and heterogeneous satisfaction levels.
method Heteroscedastic linear bandits with reneging, UCB-type policy HR-UCB.
result HR-UCB achieves regret.
Study shows how learning and analytical models affect reneging and jockeying in a dual M/M/1 system.
problem How do learning and analytical models affect reneging and jockeying in a dual M/M/1 system?
method Analytical and online trained actor-critic models were used to study reneging and jockeying in a dual M/M/1 system.
result Both analytical and online trained actor-critic models yield the same asymptotic limits for reneging and jockeying, but differ in practical sizes.
A method for training autonomous vehicles using continuous human feedback to avoid sub-optimal decisions.
problem Training autonomous vehicles with limited and potentially sub-optimal human demonstrations.
method Continuous scalar feedback for each action to learn from sub-optimal demonstrations and evaluative feedback.
result The proposed method outperforms supervised learning on positive examples alone and learns from sub-optimal demonstrations.
Deep neural networks have excelled on a wide range of problems, from vision to language and game playing. Neural networks very gradually incorporate information into weights as they process data, requiring very low learning rates. If the training distribution shifts, the network is slow to adapt, and when it does adapt…