Modyn automates continuous ML model training on growing datasets.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We consider the problem of building a state representation model for control, in a continual learning setting. As the environment changes, the aim is to efficiently compress the sensory state's information without losing past knowledge, and then use Reinforcement Learning on the resulting features for efficient policy …
This paper proposes a novel scheme for the watermarking of Deep Reinforcement Learning (DRL) policies. This scheme provides a mechanism for the integration of a unique identifier within the policy in the form of its response to a designated sequence of state transitions, while incurring minimal impact on the nominal pe…
Graph-Triggered Bandits unify rested and restless bandits with graph-defined arm interactions.
Study shows how macroprudential policies affect credit growth in Israel, especially in housing and business sectors.
This study examine the difference in the size of avalanches among industries triggered by demand shocks, which can be rephrased by control of the economy or fiscal policy, and by using the production-inventory model and observed data. We obtain the following results. (1) The size of avalanches follows power law. (2) Th…
We consider a model of contagion in financial networks recently introduced in the literature, and we characterize the effect of a few features empirically observed in real networks on the stability of the system. Notably, we consider the effect of heterogeneous degree distributions, heterogeneous balance sheet size and…
In this paper, we study the combinatorial multi-armed bandit problem (CMAB) with probabilistically triggered arms (PTAs). Under the assumption that the arm triggering probabilities (ATPs) are positive for all arms, we prove that a class of upper confidence bound (UCB) policies, named Combinatorial UCB with exploration …
We extend in a minimal way the stylized model introduced in in "Tipping Points in Macroeconomic Agent Based Models" [JEDC 50, 29-61 (2015)], with the aim of investigating the role and efficacy of monetary policy of a `Central Bank' that sets the interest rate such as to steer the economy towards a prescribed inflation …
Reinforcement learning agents that operate in diverse and complex environments can benefit from the structured decomposition of their behavior. Often, this is addressed in the context of hierarchical reinforcement learning, where the aim is to decompose a policy into lower-level primitives or options, and a higher-leve…
Study of participating policies with guaranteed minimum interest rate and surrender option.
Neural backdoor attack is emerging as a severe security threat to deep learning, while the capability of existing defense methods is limited, especially for complex backdoor triggers. In the work, we explore the space formed by the pixel values of all possible backdoor triggers. An original trigger used by an attacker …
ReSkill reconciles RL skill creation with policy optimization.
Proximal policy optimization (PPO) is one of the most successful deep reinforcement-learning methods, achieving state-of-the-art performance across a wide range of challenging tasks. However, its optimization behavior is still far from being fully understood. In this paper, we show that PPO could neither strictly restr…
Voice-triggered smart assistants often rely on detection of a trigger-phrase before they start listening for the user request. Mitigation of false triggers is an important aspect of building a privacy-centric non-intrusive smart assistant. In this paper, we address the task of false trigger mitigation (FTM) using a nov…
Study improves CTS's approximation regret for combinatorial bandits.
MISA detects Trojan triggers in neural networks at inference time.
Recent work has identified that classification models implemented as neural networks are vulnerable to data-poisoning and Trojan attacks at training time. In this work, we show that these training-time vulnerabilities extend to deep reinforcement learning (DRL) agents and can be exploited by an adversary with access to…
Common event-triggered state estimation (ETSE) algorithms save communication in networked control systems by predicting agents' behavior, and transmitting updates only when the predictions deviate significantly. The effectiveness in reducing communication thus heavily depends on the quality of the dynamics models used …
PHAZE framework uses zkML and hashing for fast, verifiable LHC trigger decisions.
New algorithm for contextual combinatorial bandits with probabilistic arm triggering.
New framework for reinforcement learning with sporadic state observations.
AMUSE uses reinforcement learning to predict optimal model updates.
Dynamic treatment strategies on networks amplify policy impact through spillovers.
Backdoor attacks make models predict a specific class near triggers, smoothing their decision function.
Paper improves CMAB regret bounds by reducing batch-size dependency.
We consider a model in which a trader aims to maximize expected risk-adjusted profit while trading a single security. In our model, each price change is a linear combination of observed factors, impact resulting from the trader's current and prior activity, and unpredictable random effects. The trader must learn coeffi…
BadGD identifies gradient descent vulnerabilities through strategic backdoor attacks.
Corporate defaults may be triggered by some major market news or events such as financial crises or collapses of major banks or financial institutions. With a view to develop a more realistic model for credit risk analysis, we introduce a new type of reduced-form intensity-based model that can incorporate the impacts o…
In this paper, we present an online reinforcement learning algorithm, called Renewal Monte Carlo (RMC), for infinite horizon Markov decision processes with a designated start state. RMC is a Monte Carlo algorithm and retains the advantages of Monte Carlo methods including low bias, simplicity, and ease of implementatio…
We describe the design of a voice trigger detection system for smart speakers. In this study, we address two major challenges. The first is that the detectors are deployed in complex acoustic environments with external noise and loud playback by the device itself. Secondly, collecting training examples for a specific k…
Transformer learns to search through reinforcement learning, mimicking DFS.
Improved speech recognition for voice assistants by analyzing speech data.
In classical Hawkes process, the baseline intensity and triggering kernel are assumed to be a constant and parametric function respectively, which limits the model flexibility. To generalize it, we present a fully Bayesian nonparametric model, namely Gaussian process modulated Hawkes process and propose an EM-variation…
High and volatile global food prices have led to food riots and played a critical role in triggering the Arab Spring revolutions in recent years. The severe drought in the US in the summer of 2012 led to a new increase in food prices. Through the fall, they remained at a threshold above which the riots and revolutions …
We study combinatorial multi-armed bandit with probabilistically triggered arms (CMAB-T) and semi-bandit feedback. We resolve a serious issue in the prior CMAB-T studies where the regret bounds contain a possibly exponentially large factor of , where is the minimum positive probability that an arm is trigg…
This paper uses advanced math to price special insurance bonds.
The problem of resource allocation of nonlinear networked control systems is investigated, where, unlike the well discussed case of triggering for stability, the objective is optimal triggering. An approximate dynamic programming approach is developed for solving problems with fixed final times initially and then it is…
In this paper we discuss the issue of computation of the bilateral credit valuation adjustment (CVA) under rating triggers, and in presence of ratings-linked margin agreements. Specifically, we consider collateralized OTC contracts, that are subject to rating triggers, between two parties -- an investor and a counterpa…
Unified reinforcement learning and stochastic processes with action-driven processes.
The commercialization of deep learning creates a compelling need for intellectual property (IP) protection. Deep neural network (DNN) watermarking has been proposed as a promising tool to help model owners prove ownership and fight piracy. A popular approach of watermarking is to train a DNN to recognize images with ce…
Federated Q-Learning achieves linear regret speedup with low communication cost.
In this paper, we propose and analyze SPARQ-SGD, which is an event-triggered and compressed algorithm for decentralized training of large-scale machine learning models. Each node can locally compute a condition (event) which triggers a communication where quantized and sparsified local model parameters are sent. In SPA…
ET-GP-UCB optimizes time-varying functions without knowing change rates.
Paper proposes FedQ-Advantage for federated Q-learning with near-optimal regret and low communication cost.
Recent droughts in the midwestern United States threaten to cause global catastrophe driven by a speculator amplified food price bubble. Here we show the effect of speculators on food prices using a validated quantitative model that accurately describes historical food prices. During the last six years, high and fluctu…
The causal effect of a treatment can vary from person to person based on their individual characteristics and predispositions. Mining for patterns of individual-level effect differences, a problem known as heterogeneous treatment effect estimation, has many important applications, from precision medicine to recommender…
We analyze the regret of combinatorial Thompson sampling (CTS) for the combinatorial multi-armed bandit with probabilistically triggered arms under the semi-bandit feedback setting. We assume that the learner has access to an exact optimization oracle but does not know the expected base arm outcomes beforehand. When th…