A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
We study the phenomenon of bias amplification in classifiers, wherein a machine learning model learns to predict classes with a greater disparity than the underlying ground truth. We demonstrate that bias amplification can arise via an inductive bias in gradient descent methods that results in the overestimation of the…
This study quantifies systemic importance in global banks using a continuous framework that amplifies localized shocks.
problem Analyzing financial contagion and systemic risk in global banks.
method Developed a continuous framework incorporating geographic proximity and interbank network linkages, using a master equation and Feynman-Kac representation.
result The amplification factor correctly identifies systemically important institutions and predicts crisis outcomes.
Latent-state environments with long horizons, such as those faced by recommender systems, pose significant challenges for reinforcement learning (RL). In this work, we identify and analyze several key hurdles for RL in such environments, including belief state error and small action advantage. We develop a general prin…
A fundamental result in differential privacy states that the privacy guarantees of a mechanism are preserved by any post-processing of its output. In this paper we investigate under what conditions stochastic post-processing can amplify the privacy of a mechanism. By interpreting post-processing as the application of a…
Many real world learning tasks involve complex or hard-to-specify objectives, and using an easier-to-specify proxy can lead to poor performance or misaligned behavior. One solution is to have humans provide a training signal by demonstrating or judging performance, but this approach fails if the task is too complicated…
Differential privacy comes equipped with multiple analytical tools for the design of private data analyses. One important tool is the so-called "privacy amplification by subsampling" principle, which ensures that a differentially private mechanism run on a random subsample of a population provides higher privacy guaran…
Language is increasingly being used to define rich visual recognition problems with supporting image collections sourced from the web. Structured prediction models are used in these tasks to take advantage of correlations between co-occurring labels and visual input but risk inadvertently encoding social biases found i…
We mathematically compare four competing definitions of group-level nondiscrimination: demographic parity, equalized odds, predictive parity, and calibration. Using the theoretical framework of Friedler et al., we study the properties of each definition under various worldviews, which are assumptions about how, if at a…
Most of the econometric and econophysics models have been borrowed from the statistical physics, and as a cosequence, a new interdisciplinary science called econophysics has emerged. In this paper we planned to extend the analogy between different economic processes or phenomena and processes and phenomena from differe…
Synthetic data can amplify privacy in linear regression models.
problem Understanding how synthetic data can enhance privacy in linear regression models.
method Investigated through the linear regression framework, analyzing synthetic data generated from random inputs and controlled inputs.
result Releasing a limited number of synthetic data points amplifies privacy beyond the model's inherent guarantees when inputs are random, but not when inputs are controlled by an adversary.
We find empirically a characteristic sharp peak-flat trough pattern in a large set of commodity prices. We argue that the sharp peak structure reflects an endogenous inter-market organization, and that peaks may be seen as local ``singularities'' resulting from imitation and herding. These findings impose a novel strin…
This work studies differential privacy in the context of the recently proposed shuffle model. Unlike in the local model, where the server collecting privatized data from users can track back an input to a specific user, in the shuffle model users submit their privatized inputs to a server anonymously. This setup yields…
Study shows how sentiment shocks affect equity markets, revealing asymmetries and state-dependent effects.
problem Understanding how sentiment shocks propagate through equity markets and their impact on different investor groups.
method Used four independent proxies with sign-aligned kappa-rho parameters, calibrated a structural model to link sentiment to returns.
result A one standard deviation sentiment shock has a 1.06 basis point impact, with effects amplified over 11.2 months and concentrated in retail-tilted stocks.
Given data drawn from an unknown distribution, D, to what extent is it possible to ``amplify'' this dataset and output an even larger set of samples that appear to have been drawn from D? We formalize this question as follows: an (n,m)amplification procedure takes as input n independent draws from an …
Estimating properties of discrete distributions is a fundamental problem in statistical learning. We design the first unified, linear-time, competitive, property estimator that for a wide class of properties and for all underlying distributions uses just 2n samples to achieve the performance attained by the empirical…