BeBold improves exploration in sparse-reward tasks by regulating visitation counts.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
RECODE uses clustering and embedding to track state visitation counts in RL.
Study geodesic paths on flat surfaces, comparing length and singularity counts.
A web browser should not be only for browsing web pages but also help users to find out their target websites and recommend similar type websites based on their behavior. Throughout this paper, we propose two methods to make a web browser more intelligent about link prediction which works during typing on address-bar a…
The paper develops personalized DAG models for web user behavior.
Bootstrap method for Markov chains in reinforcement learning.
Reinforcement learning with sparse rewards is still an open challenge. Classic methods rely on getting feedback via extrinsic rewards to train the agent, and in situations where this occurs very rarely the agent learns slowly or cannot learn at all. Similarly, if the agent receives also rewards that create suboptimal m…
Electronic medical record (EMR) data contains historical sequences of visits of patients, and each visit contains rich information, such as patient demographics, hospital utilisation and medical codes, including diagnosis, procedure and medication codes. Most existing EMR embedding methods capture visit-code associatio…
Framework identifies comorbidities for frequent ED and inpatient visits.
Curiosity-Critic improves world model training by focusing on cumulative prediction error.
Is it true that patients with similar conditions get similar diagnoses? In this paper we show NLP methods and a unique corpus of documents to validate this claim. We (1) introduce a method for representation of medical visits based on free-text descriptions recorded by doctors, (2) introduce a new method for clustering…
Langevin DQN achieves deep exploration using Gaussian noise.
Study automates detection of visitation disruptions in ICU patients.
EQO uses a simple bonus term for efficient exploration in tabular RL.
Computational models that forecast the progression of Alzheimer's disease at the patient level are extremely useful tools for identifying high risk cohorts for early intervention and treatment planning. The state-of-the-art work in this area proposes models that forecast by using latent representations extracted from t…
The paper introduces a new intrinsic reward method for exploration in reinforcement learning.
Real-Time Bidding is nowadays one of the most promising systems in the online advertising ecosystem. In the presented study, the performance of RTB campaigns is improved by optimising the parameters of the users' profiles and the publishers' websites. Most studies about optimising RTB campaigns are focused on the biddi…
Unsupervised model detects healthcare fraud from patient visit data.
The paper analyzes how SGD visits different regions of a non-convex problem's state space.
VisitHGNN predicts visit probabilities between neighborhoods and POIs using graph neural networks.
Understanding how users navigate in a network is of high interest in many applications. We consider a setting where only aggregate node-level traffic is observed and tackle the task of learning edge transition probabilities. We cast it as a preference learning problem, and we study a model where choices follow Luce's a…
New RL approach uses future state and action visitation measures for better exploration.
Paper proposes machine learning model for early Alzheimer's diagnosis.
In recent years, state-of-the-art game-playing agents often involve policies that are trained in self-playing processes where Monte Carlo tree search (MCTS) algorithms and trained policies iteratively improve each other. The strongest results have been obtained when policies are trained to mimic the search behaviour of…
In this paper, we apply neural networks into digital marketing world for the purpose of better targeting the potential customers. To do so, we model the customer online behaviours using dedicated neural network architectures. Starting from user searched keywords in a search engine to the landing page and different foll…
This article was withdrawn by the arXiv.org administrators since it plagiarizes math.GT/0011056.
Monte Carlo Tree Search (MCTS) algorithms have achieved great success on many challenging benchmarks (e.g., Computer Go). However, they generally require a large number of rollouts, making their applications costly. Furthermore, it is also extremely challenging to parallelize MCTS due to its inherent sequential nature:…
Predicting the patient's clinical outcome from the historical electronic medical records (EMR) is a fundamental research problem in medical informatics. Most deep learning-based solutions for EMR analysis concentrate on learning the clinical visit embedding and exploring the relations between visits. Although those wor…
Enhances early risk assessments for pediatric outcomes using contrastive learning.
New concept: reward hacking, where optimizing a flawed reward function can hurt performance.
We show how to learn low-dimensional representations (embeddings) of patient visits from the corresponding electronic health record (EHR) where International Classification of Diseases (ICD) diagnosis codes are removed. We expect that these embeddings will be useful for the construction of predictive statistical models…
Unified framework for distributional regret in bandits and reinforcement learning.
Modeling infection hotspots to quantify effects of contact tracing and testing.
Pediatric asthma is the most prevalent chronic childhood illness, afflicting about 6.2 million children in the United States. However, asthma could be better managed by identifying and avoiding triggers, educating about medications and proper disease management strategies. This research utilizes deep learning methodolo…
Study online RL with mismatched dynamics, achieving sublinear regret.
Counting tripods on a flat torus using lattice point counting.
Transforming sparse outcomes into dense process rewards for efficient reinforcement learning.
The paper introduces return parity for fairness in MDPs, addressing delayed and adverse effects.
In batch reinforcement learning (RL), one often constrains a learned policy to be close to the behavior (data-generating) policy, e.g., by constraining the learned action distribution to differ from the behavior policy by some maximum degree that is the same at each state. This can cause batch RL to be overly conservat…
Algorithm optimizes constrained reinforcement learning with dual variables.
Flow Matching for count data improves sample quality and efficiency.
Study how untrained policies explore in RL environments.
New theorem counts curves on orbifolds.
We present here a general framework and a specific algorithm for predicting the destination, route, or more generally a pattern, of an ongoing journey, building on the recent work of [Y. Lassoued, J. Monteil, Y. Gu, G. Russo, R. Shorten, and M. Mevissen, "Hidden Markov model for route and destination prediction," in IE…
Effective modeling of electronic health records presents many challenges as they contain large amounts of irregularity most of which are due to the varying procedures and diagnosis a patient may have. Despite the recent progress in machine learning, unsupervised learning remains largely at open, especially in the healt…
A new method, Count-MORL, improves offline reinforcement learning by using state-action frequency.
Proposes a method to reconcile count time series forecasts.
Graph neural networks struggle with counting certain substructures in graphs.