Neural ODEs control graph dynamics with low energy feedback.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
DART optimizes subset selection in non-linear bandit problems.
Master-slave architecture tackles combinatorial multi-armed bandits with diversity constraints.
A predictor that is deployed in a live production system may perturb the features it uses to make predictions. Such a feedback loop can occur, for example, when a model that predicts a certain type of behavior ends up causing the behavior it predicts, thus creating a self-fulfilling prophecy. In this paper we analyze p…
Learning weights in a spiking neural network with hidden neurons, using local, stable and online rules, to control non-linear body dynamics is an open problem. Here, we employ a supervised scheme, Feedback-based Online Local Learning Of Weights (FOLLOW), to train a network of heterogeneous spiking neurons with hidden l…
Neural algorithms optimize arm selection with human preference feedback for complex reward functions.
In this work, we have presented a simple analytical approximation scheme for generic non-linear FBSDEs. By treating the interested system as the linear decoupled FBSDE perturbed with non-linear generator and feedback terms, we have shown that it is possible to carry out a recursive approximation to an arbitrarily highe…
In this paper, we propose an efficient Monte Carlo implementation of non-linear FBSDEs as a system of interacting particles inspired by the ideas of branching diffusion method. It will be particularly useful to investigate large and complex systems, and hence it is a good complement of our previous work presenting an a…
Learning optimal feedback control laws capable of executing optimal trajectories is essential for many robotic applications. Such policies can be learned using reinforcement learning or planned using optimal control. While reinforcement learning is sample inefficient, optimal control only plans an optimal trajectory fr…
Study non-linear combinatorial bandits with polynomial rewards, finding significant differences from linear cases.
Empirical data reveals that the liquidity flow into the order book (depositions, cancellations andmarket orders) is influenced by past price changes. In particular, we show that liquidity tends todecrease with the amplitude of past volatility and price trends. Such a feedback mechanism inturn increases the volatility, …
A new algorithm for conversational recommendation systems using dueling bandits in GLMs.
Many real-world problems like Social Influence Maximization face the dilemma of choosing the best out of options at a given time instant. This setup can be modeled as a combinatorial bandit which chooses out of arms at each time, with an aim to achieve an efficient trade-off between exploration and expl…
This paper presents a new causal network learning algorithm (FSNN, Feedback System Neural Network) based on the construction and analysis of a non-linear system of Ordinary Differential Equations (ODEs). The constructed system provides insight into the mechanisms responsible for generating the past and potential future…
Combining simple elements from the literature, we define a linear model that is geared toward sparse data, in particular implicit feedback data for recommender systems. We show that its training objective has a closed-form solution, and discuss the resulting conceptual insights. Surprisingly, this simple model achieves…
Causal inference uses observations to infer the causal structure of the data generating system. We study a class of functional models that we call Time Series Models with Independent Noise (TiMINo). These models require independent residual time series, whereas traditional methods like Granger causality exploit the var…
Improved regret bound for multinomial logistic bandits with non-linearity.
In recent years, deep neural networks have yielded state-of-the-art performance on several tasks. Although some recent works have focused on combining deep learning with recommendation, we highlight three issues of existing models. First, these models cannot work on both explicit and implicit feedback, since the networ…
This paper studies the continuous time mean-variance portfolio selection problem with one kind of non-linear wealth dynamics. To deal the expectation constraint, an auxiliary stochastic control problem is firstly solved by two new generalized stochastic Riccati equations from which a candidate portfolio in feedback for…
Imitation learning is a control design paradigm that seeks to learn a control policy reproducing demonstrations from expert agents. By substituting expert demonstrations for optimal behaviours, the same paradigm leads to the design of control policies closely approximating the optimal state-feedback. This approach requ…
We propose a simple non-equilibrium model of a financial market as an open system with a possible exchange of money with an outside world and market frictions (trade impacts) incorporated into asset price dynamics via a feedback mechanism. Using a linear market impact model, this produces a non-linear two-parametric ex…
Neural networks are commonly trained to make predictions through learning algorithms. Contrastive Hebbian learning, which is a powerful rule inspired by gradient backpropagation, is based on Hebb's rule and the contrastive divergence algorithm. It operates in two phases, the forward (or free) phase, where the data are …
DFA trains deep networks by aligning weights then memorizing data.
In this paper, we study the non-linear diffusion equation associated with a particle system where the common drift depends on the rate of absorption of particles at a boundary. We provide an interpretation as a structural credit risk model with default contagion in a large interconnected banking system. Using the metho…
Stylized facts of empirical assets log-returns include the existence of (semi) heavy tailed distributions and a non-linear spectrum of Hurst exponents . Empirical data considered are daily prices of 10 large indices from 01/01/1990 to 12/31/2004. We propose a stylized model of price dynamics which is…
Model predicts insolvency risks in banks due to liquidity and credit risks.
Oracle-efficient algorithms reduce combinatorial semi-bandit regret to logarithmic time.
New insights into cascade feedback linearization of control systems.
New method learns from either positive or negative feedback alone.
User preferences for items can be inferred from either explicit feedback, such as item ratings, or implicit feedback, such as rental histories. Research in collaborative filtering has concentrated on explicit feedback, resulting in the development of accurate and scalable models. However, since explicit feedback is oft…
This paper improves image retrieval accuracy through novel relevance feedback methods.
Recommender systems recommend items more accurately by analyzing users' potential interest on different brands' items. In conjunction with users' rating similarity, the presence of users' implicit feedbacks like clicking items, viewing items specifications, watching videos etc. have been proved to be helpful for learni…
Study evaluates new models using human feedback from another model.
Develops a new model for RLHF accounting for partially observed states and intermediate feedback.
Classifier learns to ignore unreliable feedback from end users.
MOCA uses modular attention to estimate causal effects from complex data.
Paper characterizes minimax regret rates for online ranking with top-k feedback.
Study on sample complexity for pure exploration in feedback graph settings.
One-bit feedback suffices for a bandit problem's optimal strategy.
Study how communication and feedback graphs affect learning outcomes.
New algorithms tackle RKHS bandits with reduced complexity and improved performance.
New method uses correlated auxiliary feedback to reduce regret in parameterized bandits.
Test for linearizing 2-input systems with 2D feedback.
The problem of feedback equivalence for control systems is considered. An algebra of differential invariants and criteria for the feedback equivalence for regular control systems are found.
Reduces user feedback needed for accurate recommender systems.
We present a study on reinforcement learning (RL) from human bandit feedback for sequence-to-sequence learning, exemplified by the task of bandit neural machine translation (NMT). We investigate the reliability of human bandit feedback, and analyze the influence of reliability on the learnability of a reward estimator,…
New algorithm for recommending best arms with aggregated feedback.
Study of reinforcement learning with additional feedback observations.