Study online RL with mismatched dynamics, achieving sublinear regret.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
What happens when the Supreme Court of the United States decides a case impacting one or more publicly-traded firms? While many have observed anecdotal evidence linking decisions or oral arguments to abnormal stock returns, few have rigorously or systematically investigated the behavior of equities around Supreme Court…
The Brazilian court system is currently the most clogged up judiciary system in the world. Thousands of lawsuit cases reach the supreme court every day. These cases need to be analyzed in order to be associated to relevant tags and allocated to the right team. Most of the cases reach the court as raster scanned documen…
Two trees in the boundary of outer space are said to be \emph{primitive-equivalent} whenever their translation length functions are equal in restriction to the set of primitive elements of . We give an explicit description of this equivalence relation, showing in particular that it is nontrivial. This question is …
New algorithm guarantees optimal convergence rate for stochastic optimization.
Robust learning method combines kernel smoothing and robust optimization.
MedGraph learns patient visit embeddings from EMRs, capturing both attributes and temporal sequences.
Framework identifies comorbidities for frequent ED and inpatient visits.
Is it true that patients with similar conditions get similar diagnoses? In this paper we show NLP methods and a unique corpus of documents to validate this claim. We (1) introduce a method for representation of medical visits based on free-text descriptions recorded by doctors, (2) introduce a new method for clustering…
Study automates detection of visitation disruptions in ICU patients.
RECODE uses clustering and embedding to track state visitation counts in RL.
Computational models that forecast the progression of Alzheimer's disease at the patient level are extremely useful tools for identifying high risk cohorts for early intervention and treatment planning. The state-of-the-art work in this area proposes models that forecast by using latent representations extracted from t…
The paper introduces a new intrinsic reward method for exploration in reinforcement learning.
Successful attempts to predict judges' votes shed light into how legal decisions are made and, ultimately, into the behavior and evolution of the judiciary. Here, we investigate to what extent it is possible to make predictions of a justice's vote based on the other justices' votes in the same case. For our predictions…
The paper analyzes how over-parameterization affects reinforcement learning performance.
BeBold improves exploration in sparse-reward tasks by regulating visitation counts.
A web browser should not be only for browsing web pages but also help users to find out their target websites and recommend similar type websites based on their behavior. Throughout this paper, we propose two methods to make a web browser more intelligent about link prediction which works during typing on address-bar a…
Unsupervised model detects healthcare fraud from patient visit data.
The paper analyzes how SGD visits different regions of a non-convex problem's state space.
VisitHGNN predicts visit probabilities between neighborhoods and POIs using graph neural networks.
Inefficient markets allow investors to consistently outperform the market. To demonstrate that inefficiencies exist in sports betting markets, we created a betting algorithm that generates above market returns for the NFL, NBA, NCAAF, NCAAB, and WNBA betting markets. To formulate our betting strategy, we collected and …
We study the problem of off-policy policy optimization in Markov decision processes, and develop a novel off-policy policy gradient method. Prior off-policy policy gradient approaches have generally ignored the mismatch between the distribution of states visited under the behavior policy used to collect data, and what …
Optimizes RTB campaigns by selecting user profiles and website configurations.
New RL approach uses future state and action visitation measures for better exploration.
New method stabilizes FQE by reweighting Bellman targets.
Surface parameterizations and registrations are important in computer graphics and imaging, where 1-1 correspondences between meshes are computed. In practice, surface maps are usually represented and stored as 3D coordinates each vertex is mapped to, which often requires lots of storage memory. This causes inconvenien…
Paper proposes machine learning model for early Alzheimer's diagnosis.
The paper develops personalized DAG models for web user behavior.
Given a positive function , we define its John-Nirenberg radius at point to be the supreme of the radius such that when , and when . We will show that for a collapsing sequence in a fixed conformal class under some curvature c…
In this paper, we apply neural networks into digital marketing world for the purpose of better targeting the potential customers. To do so, we model the customer online behaviours using dedicated neural network architectures. Starting from user searched keywords in a search engine to the landing page and different foll…
We consider the off-policy estimation problem of estimating the expected reward of a target policy using samples collected by a different behavior policy. Importance sampling (IS) has been a key technique to derive (nearly) unbiased estimators, but is known to suffer from an excessively high variance in long-horizon pr…
This article was withdrawn by the arXiv.org administrators since it plagiarizes math.GT/0011056.
New batch RL method avoids overly optimistic policies.
Enhances early risk assessments for pediatric outcomes using contrastive learning.
We show how to learn low-dimensional representations (embeddings) of patient visits from the corresponding electronic health record (EHR) where International Classification of Diseases (ICD) diagnosis codes are removed. We expect that these embeddings will be useful for the construction of predictive statistical models…
Modeling infection hotspots to quantify effects of contact tracing and testing.
Pediatric asthma is the most prevalent chronic childhood illness, afflicting about 6.2 million children in the United States. However, asthma could be better managed by identifying and avoiding triggers, educating about medications and proper disease management strategies. This research utilizes deep learning methodolo…
Improved off-policy evaluation for MDPs with weak distributional overlap.
ConCare personalizes healthcare predictions by capturing EMR features.
Unified treatment of reinforcement learning via convex duality.
Transforming sparse outcomes into dense process rewards for efficient reinforcement learning.
The paper introduces return parity for fairness in MDPs, addressing delayed and adverse effects.
Algorithm optimizes constrained reinforcement learning with dual variables.
We present the Bayesian Echo Chamber, a new Bayesian generative model for social interaction data. By modeling the evolution of people's language usage over time, this model discovers latent influence relationships between them. Unlike previous work on inferring influence, which has primarily focused on simple temporal…
Study how untrained policies explore in RL environments.
A new approach for deep exploration in sparse reward reinforcement learning.
Point-of-Interest (POI) recommender systems play a vital role in people's lives by recommending unexplored POIs to users and have drawn extensive attention from both academia and industry. Despite their value, however, they still suffer from the challenges of capturing complicated user preferences and fine-grained user…
We present here a general framework and a specific algorithm for predicting the destination, route, or more generally a pattern, of an ongoing journey, building on the recent work of [Y. Lassoued, J. Monteil, Y. Gu, G. Russo, R. Shorten, and M. Mevissen, "Hidden Markov model for route and destination prediction," in IE…