A blindfolded LLM trading framework validates market signals without ticker memorization.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Benchmark evaluates LLM trading agents by masking identifiers to prevent memory leaks.
Research proposes classifiers to distinguish tweets with conflicting cashtags.
A framework for anonymized risk sharing without revealing identities or preferences.
Anonymization reduces economic signal extraction from financial texts.
Recommender systems are widely used to predict personalized preferences of goods or services using users' past activities, such as item ratings or purchase histories. If collections of such personal activities were made publicly available, they could be used to personalize a diverse range of services, including targete…
For a dataset of label-count pairs, an anonymized histogram is the multiset of counts. Anonymized histograms appear in various potentially sensitive contexts such as password-frequency lists, degree distribution in social networks, and estimation of symmetric properties of discrete distributions. Motivated by these app…
We study a variant of the stochastic -armed bandit problem, which we call "bandits with delayed, aggregated anonymous feedback". In this problem, when the player pulls an arm, a reward is generated, however it is not immediately observed. Instead, at the end of each round the player observes only the sum of a number…
Graph matching in noisy environments with Markovian errors.
The social media revolution has produced a plethora of web services to which users can easily upload and share multimedia documents. Despite the popularity and convenience of such services, the sharing of such inherently personal data, including speech data, raises obvious security and privacy concerns. In particular, …
The study examines how brokers' identity affects their trading strategies on the Toronto Stock Exchange.
Adaptive MAB algorithms handle composite, anonymous feedback without reward interval knowledge.
Text-based analysis methods allow to reveal privacy relevant author attributes such as gender, age and identify of the text's author. Such methods can compromise the privacy of an anonymous author even when the author tries to remove privacy sensitive content. In this paper, we propose an automatic method, called Adver…
Preserving the privacy of individuals by protecting their sensitive attributes is an important consideration during microdata release. However, it is equally important to preserve the quality or utility of the data for at least some targeted workloads. We propose a novel framework for privacy preservation based on the …
We explore a novel setting of the Multi-Armed Bandit (MAB) problem inspired from real world applications which we call bandits with "stochastic delayed composite anonymous feedback (SDCAF)". In SDCAF, the rewards on pulling arms are stochastic with respect to time but spread over a fixed number of time steps in the fut…
TIPRDC anonymizes data features to protect privacy while retaining useful information.
This work synthesizes realistic data from neural excitation patterns to anonymize private data.
The paper proposes a privacy-preserving algorithm for decentralized learning using public-key cryptography.
Motion sensors such as accelerometers and gyroscopes measure the instant acceleration and rotation of a device, in three dimensions. Raw data streams from motion sensors embedded in portable and wearable devices may reveal private information about users without their awareness. For example, motion data might disclose …
The paper analyzes how clustering sensitive data can improve model generalization without revealing individual information.
WAFFLe anonymizes federated learning weights to protect data privacy and fairness.
Statistical methods protecting sensitive information or the identity of the data owner have become critical to ensure privacy of individuals as well as of organizations. This paper investigates anonymization methods based on representation learning and deep neural networks, and motivated by novel information theoretica…
Anonymizing company names in financial news improves trading performance, contrary to initial expectations.
Zero-Copy Architecture Detects Cross-Company Financial Signals Instantly.
In this paper, we investigate the effect of machine learning based anonymization on anomalous subgroup preservation. In particular, we train a binary classifier to discover the most anomalous subgroup in a dataset by maximizing the bias between the group's predicted odds ratio from the model and observed odds ratio fro…
Sensitive statistics are often collected across sets of users, with repeated collection of reports done over time. For example, trends in users' private preferences or software usage may be monitored via such reports. We study the collection of such statistics in the local differential privacy (LDP) model, and describe…
Study non-oblivious adversarial bandits with delayed feedback and propose algorithms with improved regret bounds.
We analyze a proprietary dataset of trades by a single asset manager, comparing their price impact with that of the trades of the rest of the market. In the context of a linear propagator model we find no significant difference between the two, suggesting that both the magnitude and time dependence of impact are univer…
The paper explores how to measure and optimize ad reach while maintaining user privacy.
The task of representing entire graphs has seen a surge of prominent results, mainly due to learning convolutional neural networks (CNNs) on graph-structured data. While CNNs demonstrate state-of-the-art performance in graph classification task, such methods are supervised and therefore steer away from the original pro…
This study addresses an issue of co-adaptation between a feature extractor and a classifier in a neural network. A naive joint optimization of a feature extractor and a classifier often brings situations in which an excessively complex feature distribution adapted to a very specific classifier degrades the test perform…
Differentially private GANs improve image privacy without significant quality loss.
In this work, we propose an order book model with herd behavior. The proposed model is built upon two distinct approaches: a recent empirical study of the detailed order book records by Kanazawa et al. [Phys. Rev. Lett. 120, 138301] and financial herd behavior model. Combining these approaches allows us to propose a mo…
Breaks circular dependency in synthetic option pricing with a novel model.
This study analyzes dynamic connectedness in global supply chain infrastructure portfolios, identifying key risk factors and extreme events.
EB improves asset pricing by mining large strategies without lookahead bias.
CAWs learn temporal network dynamics without node identities or edge attributes.
Efficient sequential matching of supply and demand is a problem of interest in many online to offline services. For instance, Uber, Lyft, Grab for matching taxis to customers; Ubereats, Deliveroo, FoodPanda etc for matching restaurants to customers. In these online to offline service problems, individuals who are respo…
Data diversity is critical to success when training deep learning models. Medical imaging data sets are often imbalanced as pathologic findings are generally rare, which introduces significant challenges when training deep learning models. In this work, we propose a method to generate synthetic abnormal MRI images with…
Framework for AI customer support that protects privacy and reduces costs.
It is not until recently that graph neural networks (GNNs) are adopted to perform graph representation learning, among which, those based on the aggregation of features within the neighborhood of a node achieved great success. However, despite such achievements, GNNs illustrate defects in identifying some common struct…
In the IEEE Investment ranking challenge 2018, participants were asked to build a model which would identify the best performing stocks based on their returns over a forward six months window. Anonymized financial predictors and semi-annual returns were provided for a group of anonymized stocks from 1996 to 2017, which…
Interest surrounding cryptocurrencies, digital or virtual currencies that are used as a medium for financial transactions, has grown tremendously in recent years. The anonymity surrounding these currencies makes investors particularly susceptible to fraud---such as "pump and dump" scams---where the goal is to artificia…
We confirm the square-root law of market impact on Apple Inc. using a large dataset.
SMOTE-DP enhances synthetic data privacy without sacrificing utility.
New measure corrects news bias in NLP stock return forecasting.
Most users of online services have unique behavioral or usage patterns. These behavioral patterns can be exploited to identify and track users by using only the observed patterns in the behavior. We study the task of identifying users from statistics of their behavioral patterns. Specifically, we focus on the setting i…
Uncoupled regression is the problem to learn a model from unlabeled data and the set of target values while the correspondence between them is unknown. Such a situation arises in predicting anonymized targets that involve sensitive information, e.g., one's annual income. Since existing methods for uncoupled regression …