New method learns from either positive or negative feedback alone.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper measures cognitive bias in positive feedback trading using diffusion process estimates.
Recommender systems play a crucial role in mitigating the problem of information overload by suggesting users' personalized items or services. The vast majority of traditional recommender systems consider the recommendation procedure as a static process and make recommendations following a fixed strategy. In this paper…
Framework uses human feedback to safely set OOD detection thresholds, reducing false positives.
Adaptive sampler improves recommendation for implicit feedback data.
Study how communication and feedback graphs affect learning outcomes.
Anomaly detectors are often used to produce a ranked list of statistical anomalies, which are examined by human analysts in order to extract the actual anomalies of interest. Unfortunately, in realworld applications, this process can be exceedingly difficult for the analyst since a large fraction of high-ranking anomal…
Social media systems rely on user feedback and rating mechanisms for personalization, ranking, and content filtering. However, when users evaluate content contributed by fellow users (e.g., by liking a post or voting on a comment), these evaluations create complex social feedback effects. This paper investigates how ra…
We analyze a controlled price formation experiment in the laboratory that shows evidence for bubbles. We calibrate two models that demonstrate with high statistical significance that these laboratory bubbles have a tendency to grow faster than exponential due to positive feedback. We show that the positive feedback ope…
We apply the potential force estimation method to artificial time series of market price produced by a deterministic dealer model. We find that dealers' feedback of linear prediction of market price based on the latest mean price changes plays the central role in the market's potential force. When markets are dominated…
Whenever a social media user decides to share a story, she is typically pleased to receive likes, comments, shares, or, more generally, feedback from her followers. As a result, she may feel compelled to use the feedback she receives to (re-)estimate her followers' preferences and decides which stories to share next to…
New oracle uses uncertainty for active classification with noisy feedback.
The starting point of this paper is the so-called Robust Positive Expectation (RPE) Theorem, a result which appears in literature in the context of Simultaneous Long-Short stock trading. This theorem states that using a combination of two specially-constructed linear feedback trading controllers, one long and one short…
Proposes a method to train classifiers with delayed feedback using a time window.
In display advertising, predicting the conversion rate, that is, the probability that a user takes a predefined action on an advertiser's website, such as purchasing goods is fundamental in estimating the value of displaying the advertisement. However, there is a relatively long time delay between a click and its resul…
Given a set of objects, an online ranking system outputs at each time step a full ranking of the set, observes a feedback of some form and suffers a loss. We study the setting in which the (adversarial) feedback is an element in , and the loss is the position (0th, 1st, 2nd...) of the item in the outputted r…
Multimodal analysis assesses job interview performance and provides feedback.
Presentation bias is one of the key challenges when learning from implicit feedback in search engines, as it confounds the relevance signal with uninformative signals due to position in the ranking, saliency, and other presentation factors. While it was recently shown how counterfactual learning-to-rank (LTR) approache…
Recent advances in multi-modal vision and language tasks enable a new set of applications. In this paper, we consider the task of generating natural language fashion feedback on outfit images. We collect a unique dataset, which contains outfit images and corresponding positive and constructive fashion feedback. We trea…
By combining (i) the economic theory of rational expectation bubbles, (ii) behavioral finance on imitation and herding of investors and traders and (iii) the mathematical and statistical physics of bifurcations and phase transitions, the log-periodic power law (LPPL) model has been developed as a flexible tool to detec…
New framework learns from partial feedback in multi-label tasks.
New ranking algorithms improve online content delivery by learning from click data.
We extend a model of positive feedback and contagion in large mean-field systems, by introducing a common source of noise driven by Brownian motion. Although the driving dynamics are continuous, the positive feedback effect can lead to `blow-up' phenomena whereby solutions develop jump-discontinuities. Our main results…
Recommender systems widely use implicit feedback such as click data because of its general availability. Although the presence of clicks signals the users' preference to some extent, the lack of such clicks does not necessarily indicate a negative response from the users, as it is possible that the users were not expos…
New algorithms control loss and constraints in uncertain, changing environments.
Conventional collaborative filtering techniques treat a top-n recommendations problem as a task of generating a list of the most relevant items. This formulation, however, disregards an opposite - avoiding recommendations with completely irrelevant items. Due to that bias, standard algorithms, as well as commonly used …
The paper tackles efficient change point detection with limited samples.
We study an online classification problem with partial feedback in which individuals arrive one at a time from a fixed but unknown distribution, and must be classified as positive or negative. Our algorithm only observes the true label of an individual if they are given a positive classification. This setting captures …
Proposes a new IPW-based ranking metric for two-sided markets.
SetRank tackles collaborative ranking from implicit feedback using setwise Bayesian approach.
Online learning with one-sided feedback aims to maximize accuracy while ensuring fairness.
Knowledge distillation (KD) is a well-known method to reduce inference latency by compressing a cumbersome teacher model to a small student model. Despite the success of KD in the classification task, applying KD to recommender models is challenging due to the sparsity of positive feedback, the ambiguity of missing fee…
A general theory of innovation and progress in human society is outlined, based on the combat between two opposite forces (conservatism/inertia and speculative herding "bubble" behavior). We contend that human affairs are characterized by ubiquitous ``bubbles'', which involve huge risks which would not otherwise be tak…
LLM trading agents show risk feedback can improve alignment without fine-tuning.
Online algorithms stabilize in feedback loops of performative prediction.
Approach collects missing outcomes to improve fairness in classification.
Stochastic linear bandits are a natural and well-studied model for structured exploration/exploitation problems and are widely used in applications such as online marketing and recommendation. One of the main challenges faced by practitioners hoping to apply existing algorithms is that usually the feedback is randomly …
In autonomous vehicle (AV) control, allowing mistakes can be quite dangerous and costly in the real world. For this reason we investigate methods of training an AV without allowing the agent to explore and instead having a human explorer collect the data. Supervised learning has been explored for AV control, but it enc…
Multi-label network classification is a well-known task that is being used in a wide variety of web-based and non-web-based domains. It can be formalized as a multi-relational learning task for predicting nodes labels based on their relations within the network. In sparse networks, this prediction task can be very chal…
In display advertising, predicting the conversion rate (CVR), meaning the probability that a user takes a predefined action on an advertiser's website, is a fundamental task for estimating the value of displaying an advertisement to a user. There are two main challenges in CVR prediction due to delayed feedback. First,…
A new RL method improves revenue management with delayed feedback.
A new federated bandit problem with multiple adversaries, solved with a near-optimal algorithm.
Latent factor models for Recommender Systems with implicit feedback typically treat unobserved user-item interactions (i.e. missing information) as negative feedback. This is frequently done either through negative sampling (point--wise loss) or with a ranking loss function (pair-- or list--wise estimation). Since a ze…
Work shows hallucination detection by LLMs is impossible without expert feedback.
Study online control of unknown time-varying systems with negative and positive results.
Study personalizes user experience to maximize rewards with patience budget.
Financial time series exhibit a number of interesting properties that are difficult to explain with simple models. These properties include fat-tails in the distribution of price fluctuations (or returns) that are slowly removed at longer timescales, strong autocorrelations in absolute returns but zero autocorrelation …
Following a long tradition of physicists who have noticed that the Ising model provides a general background to build realistic models of social interactions, we study a model of financial price dynamics resulting from the collective aggregate decisions of agents. This model incorporates imitation, the impact of extern…