A method corrects feedback shift in predicting conversion rates with delayed feedback.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
DisCor corrects reinforcement learning issues by re-weighting collected data.
Deep Reinforcement Learning has enabled the control of increasingly complex and high-dimensional problems. However, the need of vast amounts of data before reasonable performance is attained prevents its widespread application. We employ binary corrective feedback as a general and intuitive manner to incorporate human …
Learning from human feedback is a viable alternative to control design that does not require modelling or control expertise. Particularly, learning from corrective advice garners advantages over evaluative feedback as it is a more intuitive and scalable format. The current state-of-the-art in this field, COACH, has pro…
SA-PEF improves federated learning efficiency by correcting gradient mismatches.
Industrial recommender systems deal with extremely large action spaces -- many millions of items to recommend. Moreover, they need to serve billions of users, who are unique at any point in time, making a complex user state space. Luckily, huge quantities of logged implicit feedback (e.g., user clicks, dwell time) are …
Graph-based approach repairs programs from diagnostic feedback.
Work shows hallucination detection by LLMs is impossible without expert feedback.
Develops methods to correct bias in AI feedback for more accurate alignment.
We present online boosting algorithms for multiclass classification with bandit feedback, where the learner only receives feedback about the correctness of its prediction. We propose an unbiased estimate of the loss using a randomized prediction, allowing the model to update its weak learners with limited information. …
The paper optimizes exceptions in a statistical production system using machine learning.
CAFL breaks feedback loops in recommender systems using causal inference.
xAI-GAN improves GANs by providing richer feedback, enhancing image quality.
BCCP uses bandit feedback to provide reliable predictions with limited labeled data.
Deep Reinforcement Learning (DRL) has become a powerful strategy to solve complex decision making problems based on Deep Neural Networks (DNNs). However, it is highly data demanding, so unfeasible in physical systems for most applications. In this work, we approach an alternative Interactive Machine Learning (IML) stra…
Not all types of supervision signals are created equal: Different types of feedback have different costs and effects on learning. We show how self-regulation strategies that decide when to ask for which kind of feedback from a teacher (or from oneself) can be cast as a learning-to-learn problem leading to improved cost…
Bollerslev et al. (2006) study the cross-covariances for squared returns under the Heston (1993) stochastic volatility model. In order to obtain these cross-covariances the authors use an incorrect expression for the distribution of the squared returns. Here we will obtain the correct distribution of the squared return…
New method provides fine-grained feedback on interactive student programs.
CausalRM models rewards from user feedback, overcoming noise and bias.
We present a new algorithm to generate minimal, stable, and symbolic corrections to an input that will cause a neural network with ReLU activations to change its output. We argue that such a correction is a useful way to provide feedback to a user when the network's output is different from a desired output. Our algori…
We present an approach to interactive-predictive neural machine translation that attempts to reduce human effort from three directions: Firstly, instead of requiring humans to select, correct, or delete segments, we employ the idea of learning from human reinforcements in form of judgments on the quality of partial tra…
We consider the problem of online multiclass classification with partial feedback, where an algorithm predicts a class for a new instance in each round and only receives its correctness. Although several methods have been developed for this problem, recent challenging real-world applications require further performance…
The paper tackles combinatorial pure exploration with various feedback structures and proposes efficient algorithms.
Given a binary prediction problem, which performance metric should the classifier optimize? We address this question by formalizing the problem of Metric Elicitation. The goal of metric elicitation is to discover the performance metric of a practitioner, which reflects her innate rewards (costs) for correct (incorrect)…
Combines BO with context to optimize binary feedback.
While computer and communication technologies have provided effective means to scale up many aspects of education, the submission and grading of assessments such as homework assignments and tests remains a weak link. In this paper, we study the problem of automatically grading the kinds of open response mathematical qu…
AI assistants often give convincing but incorrect responses to match user beliefs.
In autonomous vehicle (AV) control, allowing mistakes can be quite dangerous and costly in the real world. For this reason we investigate methods of training an AV without allowing the agent to explore and instead having a human explorer collect the data. Supervised learning has been explored for AV control, but it enc…
This work explores test-time scaling strategies for LLMs, improving sample efficiency and expressiveness.
Improved PAC learning algorithm for multiclass classification with bandit feedback.
New framework learns from partial feedback in multi-label tasks.
New method clusters items with bandit feedback without parametric assumptions.
Method improves volatility targeting for index construction.
A new RL method improves revenue management with delayed feedback.
Algorithm clusters items by sequentially selecting features, minimizing observations.
New algorithm improves multiclass classification regret bound.
Plug-in method improves performative prediction accuracy.
Sharp sample complexity for multiclass PAC learning with bandit feedback.
New method for LLMs to learn reasoning by optimizing latent variables.
We introduce the probably approximately correct (PAC) \emph{Battling-Bandit} problem with the Plackett-Luce (PL) subset choice model--an online learning framework where at each trial the learner chooses a subset of arms from a fixed set of arms, and subsequently observes a stochastic feedback indicating prefere…
We seek a discussion about the most suitable feedback control structure for stock trading under the consideration of proportional transaction costs. Suitability refers to robustness and performance capability. Both are tested by considering different one-step ahead prediction qualities, including the ideal case, correc…
Agents trained in simulation may make errors in the real world due to mismatches between training and execution environments. These mistakes can be dangerous and difficult to discover because the agent cannot predict them a priori. We propose using oracle feedback to learn a predictive model of these blind spots to red…
IAL uses interactive learning to improve model performance with minimal human feedback.
Study on when RLVR can learn compositional problems.
While implicit feedback (e.g., clicks, dwell times, etc.) is an abundant and attractive source of data for learning to rank, it can produce unfair ranking policies for both exogenous and endogenous reasons. Exogenous reasons typically manifest themselves as biases in the training data, which then get reflected in the l…
Various domain users are increasingly leveraging real-time social media data to gain rapid situational awareness. However, due to the high noise in the deluge of data, effectively determining semantically relevant information can be difficult, further complicated by the changing definition of relevancy by each end user…
Language models learn automotive complaints, improving defect detection.
The paper tackles fair sequential decision making with biased linear bandit feedback.