Stabilizes online learning by using weighted reservoir sampling.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Detecting aggressive cancer tumors using ctDNA dynamics from few blood samples.
Paper uses stats to predict treatment choice based on illness probability.
Self-training with noisy student-teacher boosts keyword spotting accuracy.
Experience replay is an important technique for addressing sample-inefficiency in deep reinforcement learning (RL), but faces difficulty in learning from binary and sparse rewards due to disproportionately few successful experiences in the replay buffer. Hindsight experience replay (HER) was recently proposed to tackle…
The diagonal effect of orders is well documented in different markets, which states that orders are more likely to be followed by orders of the same aggressiveness and implies the presence of short-term correlations in order flows. Based on the order flow data of 43 Chinese stocks, we investigate if there are long-rang…
New algorithm improves active learning in agnostic pool-based classification.
The last financial and economic crisis demonstrated the dysfunctional long-term effects of aggressive behaviour in financial markets. Yet, evolutionary game theory predicts that under the condition of strategic dependence a certain degree of aggressive behaviour remains within a given population of agents. However, as …
Improved Thompson Sampling reduces regret in contextual bandits and reinforcement learning.
In order-driven markets, limit-order book (LOB) resiliency is an important microscopic indicator of market quality when the order book is hit by a liquidity shock and plays an essential role in the design of optimal submission strategies of large orders. However, the evolutionary behavior of LOB resilience around liqui…
Price changes are induced by aggressive market orders in stock market. We introduce a bivariate marked Hawkes process to model aggressive market order arrivals at the microstructural level. The order arrival intensity is marked by an exogenous part and two endogenous processes reflecting the self-excitation and cross-e…
New online learning algorithm combines PA and TER for binary classification.
Irregular features disrupt the desired classification. In this paper, we consider aggressively modifying scales of features in the original space according to the label information to form well-separated clusters in low-dimensional space. The proposed method exploits spectral clustering to derive scaling factors that a…
A method learns common bias for multiple low-variance tasks without hyper-parameter tuning.
We propose a new Bayesian model for flexible nonlinear regression and classification using tree ensembles. The model is based on the RuleFit approach in Friedman and Popescu (2008) where rules from decision trees and linear terms are used in a L1-regularized regression. We modify RuleFit by replacing the L1-regularizat…
A new method quantifies a specific class without labeled data.
Learning the minimum/maximum mean among a finite set of distributions is a fundamental sub-task in planning, game tree search and reinforcement learning. We formalize this learning task as the problem of sequentially testing how the minimum mean among a finite set of distributions compares to a given threshold. We deve…
Kelly betting is a prescription for optimal resource allocation among a set of gambles which are typically repeated in an independent and identically distributed manner. In this setting, there is a large body of literature which includes arguments that the theory often leads to bets which are "too aggressive" with resp…
PPG separates policy and value function training phases for better reinforcement learning efficiency.
Nearest neighbor (k-NN) graphs are widely used in machine learning and data mining applications, and our aim is to better understand what they reveal about the cluster structure of the unknown underlying distribution of points. Moreover, is it possible to identify spurious structures that might arise due to sampling va…
RSO uses random weight perturbations to train deep networks without gradients.
Addressing the ongoing examination of high-frequency trading practices in financial markets, we report the results of an extensive empirical study estimating the maximum possible profitability of the most aggressive such practices, and arrive at figures that are surprisingly modest. By "aggressive" we mean any trading …
We address the problem of multi-class classification in the case where the number of classes is very large. We propose a double sampling strategy on top of a multi-class to binary reduction strategy, which transforms the original multi-class problem into a binary classification problem over pairs of examples. The aim o…
We propose a general framework to describe the impact of different events in the order book, that generalizes previous work on the impact of market orders. Two different modeling routes can be considered, which are equivalent when only market orders are taken into account. One model posits that each event type has a te…
We present a selective sampling method designed to accelerate the training of deep neural networks. To this end, we introduce a novel measurement, the minimal margin score (MMS), which measures the minimal amount of displacement an input should take until its predicted classification is switched. For multi-class linear…
Improved analysis shows Maillard sampling achieves optimal regret bounds.
A survey of existing methods for stopping active learning (AL) reveals the needs for methods that are: more widely applicable; more aggressive in saving annotations; and more stable across changing datasets. A new method for stopping AL based on stabilizing predictions is presented that addresses these needs. Furthermo…
New autoencoder learns structured representations without regularization.
CSER improves SGD efficiency by resetting errors and partial synchronization.
In this paper, we focus on quantifying model stability as a function of random seed by investigating the effects of the induced randomness on model performance and the robustness of the model in general. We specifically perform a controlled study on the effect of random seeds on the behaviour of attention, gradient-bas…
Urban traffic systems worldwide are suffering from severe traffic safety problems. Traffic safety is affected by many complex factors, and heavily related to all drivers' behaviors involved in traffic system. Drivers with aggressive driving behaviors increase the risk of traffic accidents. In order to manage the safety…
Online Passive-Aggressive (PA) learning is a class of online margin-based algorithms suitable for a wide range of real-time prediction tasks, including classification and regression. PA algorithms are formulated in terms of deterministic point-estimation problems governed by a set of user-defined hyperparameters: the a…
The kind of realized mission inflows the sensitivity to risk. Among other factors, the risk results from decision about liquid assets investment level and liquid assets financing. The higher the risk exposure, the higher the level of liquid assets. If the specific risk exposure is smaller, the more aggressive could be …
Investors' strategies in a market influenced by price impact are analyzed, showing aggressive behavior when impact exceeds a critical point.
New method preserves spectral clustering performance under aggressive sparsification and quantization.
The study analyzes how large language models form and express investor risk profiles.
Spectral dimensionality reduction methods enable linear separations of complex data with high-dimensional features in a reduced space. However, these methods do not always give the desired results due to irregularities or uncertainties of the data. Thus, we consider aggressively modifying the scales of the features to …
We consider the problem of recovering low-rank matrices from random rank-one measurements, which spans numerous applications including covariance sketching, phase retrieval, quantum state tomography, and learning shallow polynomial neural networks, among others. Our approach is to directly estimate the low-rank factor …
Similarity/Distance measures play a key role in many machine learning, pattern recognition, and data mining algorithms, which leads to the emergence of metric learning field. Many metric learning algorithms learn a global distance function from data that satisfy the constraints of the problem. However, in many real-wor…
In compressed sensing MRI (CS-MRI), k-space measurements are under-sampled to achieve accelerated scan times. CS-MRI presents two fundamental problems: (1) where to sample and (2) how to reconstruct an under-sampled scan. In this paper, we tackle both problems simultaneously for the specific case of 2D Cartesian sampli…
In this paper, we propose exact passive-aggressive (PA) online algorithms for learning to rank. The proposed algorithms can be used even when we have interval labels instead of actual labels for examples. The proposed algorithms solve a convex optimization problem at every trial. We find exact solution to those optimiz…
New theory explains contrastive learning via overlapping augmented views.
Conformal Candidate Certification advances offline MBO by certifying candidate designs with statistical guarantees.
Soft Actor-Critic (SAC) is an off-policy actor-critic deep reinforcement learning (DRL) algorithm based on maximum entropy reinforcement learning. By combining off-policy updates with an actor-critic formulation, SAC achieves state-of-the-art performance on a range of continuous-action benchmark tasks, outperforming pr…
We consider the problem of demixing a sequence of source signals from the sum of noisy bilinear measurements. It is a generalized mathematical model for blind demixing with blind deconvolution, which is prevalent across the areas of dictionary learning, image processing, and communications. However, state-of- the-art c…
ESPO optimizes LLMs for complex tasks by balancing fine-grained updates and stability.
The Mike-Farmer (MF) model was constructed empirically based on the continuous double auction mechanism in an order-driven market, which can successfully reproduce the cubic law of returns and the diffusive behavior of stock prices at the transaction level. However, the volatility (defined by absolute return) in the MF…
Many data-fitting applications require the solution of an optimization problem involving a sum of large number of functions of high dimensional parameter. Here, we consider the problem of minimizing a sum of functions over a convex constraint set where both and are lar…