Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

1122 · Sep 201919922001200920172026
5 results for recoGym

Paper reviews methods for learning from bandit feedback in recommender systems.

problem Learning from historical data with unknown rewards.
method Counterfactual Risk Minimisation (CRM) methods using importance sampling and variance reduction.
result Comparison of different off-policy estimators' performance.

There are three quite distinct ways to train a machine learning model on recommender system logs. The first method is to model the reward prediction for each possible recommendation to the user, at the scoring time the best recommendation is found by computing an argmax over the personalized recommendations. This metho…

2019-04-24abs ↗pdf ↗