Simulation-based training (SBT) is gaining popularity as a low-cost and convenient training technique in a vast range of applications. However, for a SBT platform to be fully utilized as an effective training tool, it is essential that feedback on performance is provided automatically in real-time during training. It i…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Social media systems rely on user feedback and rating mechanisms for personalization, ranking, and content filtering. However, when users evaluate content contributed by fellow users (e.g., by liking a post or voting on a comment), these evaluations create complex social feedback effects. This paper investigates how ra…
According to the volatility feedback effect, an unexpected increase in squared volatility leads to an immediate decline in the price-dividend ratio. In this paper, we consider the properties of stock price dynamics and option valuations under the volatility feedback effect by modeling the joint dynamics of stock price,…
A microeconomic approach is proposed to derive the fluctuations of risky asset price, where the market participants are modeled as prospect trading agents. As asset price is generated by the temporary equilibrium between demand and supply, the agents' trading behaviors can affect the price process in turn, which is cal…
New algorithm allows IGL to work with action-inclusive feedback.
We introduce an autoregressive-type model of prices in financial market taking into account the self-modulation effect. We find that traders are mainly using strategies with weighted feedbacks of past prices. These feedbacks are responsible for the slow diffusion in short times, apparent trends and power law distributi…
Classifier learns to ignore unreliable feedback from end users.
User preferences for items can be inferred from either explicit feedback, such as item ratings, or implicit feedback, such as rental histories. Research in collaborative filtering has concentrated on explicit feedback, resulting in the development of accurate and scalable models. However, since explicit feedback is oft…
CausalRM models rewards from user feedback, overcoming noise and bias.
Paper tackles noisy bandit feedback for multiclass classification.
GAIF enhances online multiple testing with feedback, improving statistical power.
We present a study on reinforcement learning (RL) from human bandit feedback for sequence-to-sequence learning, exemplified by the task of bandit neural machine translation (NMT). We investigate the reliability of human bandit feedback, and analyze the influence of reliability on the learnability of a reward estimator,…
We propose a generalization of the best arm identification problem in stochastic multi-armed bandits (MAB) to the setting where every pull of an arm is associated with delayed feedback. The delay in feedback increases the effective sample complexity of standard algorithms, but can be offset if we have access to partial…
It has recently been shown that if feedback effects of decisions are ignored, then imposing fairness constraints such as demographic parity or equality of opportunity can actually exacerbate unfairness. We propose to address this challenge by modeling feedback effects as Markov decision processes (MDPs). First, we prop…
Theoretical model for iterative user discovery in recommender systems.
Active learning framework for optimizing human preferences in reinforcement learning.
This paper extends a Kyle model to include price-responsive traders, revealing new dynamics and equilibria.
We apply the potential force estimation method to artificial time series of market price produced by a deterministic dealer model. We find that dealers' feedback of linear prediction of market price based on the latest mean price changes plays the central role in the market's potential force. When markets are dominated…
Study apple tasting feedback in online binary classification, providing new insights into minimax expected mistakes.
Modeling HFT interactions reveals market instability.
LoCo-RLHF models diverse human feedback with contextual information.
REN addresses uncertainty in user feedbacks for better recommendation systems.
New algorithms improve efficiency in learning from personalized rewards.
Interactive Machine Learning is concerned with creating systems that operate in environments alongside humans to achieve a task. A typical use is to extend or amplify the capabilities of a human in cognitive or physical ways, requiring the machine to adapt to the users' intentions and preferences. Often, this takes the…
Online learning with delayed feedback has received increasing attention recently due to its several applications in distributed, web-based learning problems. In this paper we provide a systematic study of the topic, and analyze the effect of delay on the regret of online learning algorithms. Somewhat surprisingly, it t…
Novel algorithms for online learning with uncertain feedback graphs reduce regret.
Learning from human feedback is a viable alternative to control design that does not require modelling or control expertise. Particularly, learning from corrective advice garners advantages over evaluative feedback as it is a more intuitive and scalable format. The current state-of-the-art in this field, COACH, has pro…
Traditional collaborative filtering (CF) based recommender systems tend to perform poorly when the user-item interactions/ratings are highly scarce. To address this, we propose a learning framework that improves collaborative filtering with a synthetic feedback loop (CF-SFL) to simulate the user feedback. The proposed …
Recommender systems play a crucial role in mitigating the problem of information overload by suggesting users' personalized items or services. The vast majority of traditional recommender systems consider the recommendation procedure as a static process and make recommendations following a fixed strategy. In this paper…
We attempt to unveil the fine structure of volatility feedback effects in the context of general quadratic autoregressive (QARCH) models, which assume that today's volatility can be expressed as a general quadratic form of the past daily returns. The standard ARCH or GARCH framework is recovered when the quadratic kern…
Model shows how confidence feedback can lead to different crisis outcomes.
A new one-point feedback scheme improves ZO algorithms for black-box optimization.
Algorithm provides online learning guarantees against general comparators in full and bandit feedback.
Paper measures cognitive bias in positive feedback trading using diffusion process estimates.
New approach detects cyber-attacks in real-time.
DOPL learns from preference feedback to solve RMAB problems.
Paper addresses generalization error bounds for learning with censored feedback.
We introduce an autoregressive-type model with self-modulation effects for a foreign exchange rate by separating the foreign exchange rate into a moving average rate and an uncorrelated noise. From this model we indicate that traders are mainly using strategies with weighted feedbacks of the past rates in the exchange …
Enhances AI models with human feedback for noisy data.
With the growing importance of personalized recommendation, numerous recommendation models have been proposed recently. Among them, Matrix Factorization (MF) based models are the most widely used in the recommendation field due to their high performance. However, MF based models suffer from cold start problems where us…
In semantic parsing for question-answering, it is often too expensive to collect gold parses or even gold answers as supervision signals. We propose to convert model outputs into a set of human-understandable statements which allow non-expert users to act as proofreaders, providing error markings as learning signals to…
Not all types of supervision signals are created equal: Different types of feedback have different costs and effects on learning. We show how self-regulation strategies that decide when to ask for which kind of feedback from a teacher (or from oneself) can be cast as a learning-to-learn problem leading to improved cost…
IAL uses interactive learning to improve model performance with minimal human feedback.
Paper presents a novel framework for OOD learning with human feedback.
AutoStan improves Bayesian models via predictive feedback.
A PID-based feedback-control system improves multiple KPIs in RTB display advertising.
In this short note, we will show how to optimize the portfolio of a large trader whose hedging strategy affects the price of his assets.
The study aims to prevent unfair content presentation in recommender systems.