A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Gradient-based methods can be biased by distributional asymmetries in bivariate categorical data.
problem Gradient-based causal discovery methods can be biased by distributional asymmetries in bivariate categorical data.
method Identified and examined two distributional biases: Marginal Distribution Asymmetry and Marginal Distribution Shift Asymmetry. Employed two simple models to demonstrate and control these biases.
result Gradient-based methods can be biased by distributional asymmetries, and these biases can be controlled.
Semantic paraphrases can fool financial sentiment classifiers due to geometric shifts in model representations.
problem Semantic paraphrase attacks on financial sentiment classifiers
method Developing a continuous local model of semantic paraphrase perturbations
result The worst-case local displacement of the target representation is governed by the largest generalised eigenvalue of a matrix pencil constructed from the Jacobians of the two embedding maps.
Skew Gaussian Processes improve classification performance by allowing asymmetry.
problem Limited use of Gaussian processes in applications requiring asymmetry.
method Propose Skew-Gaussian processes (SkewGPs) as a non-parametric prior over functions, extending the multivariate Unified Skew-Normal distribution to stochastic processes.
result SkewGPs provide better performance than symmetric Gaussian processes in classification tasks.
Faced with distribution shift between training and test set, we wish to detect and quantify the shift, and to correct our classifiers without test set labels. Motivated by medical diagnosis, where diseases (targets) cause symptoms (observations), we focus on label shift, where the label marginal p(y) changes but the …
Inverse statistics in economics is considered. We argue that the natural candidate for such statistics is the investment horizons distribution. This distribution of waiting times needed to achieve a predefined level of return is obtained from (often detrended) historic asset prices. Such a distribution typically goes t…
RLSbench benchmarks domain adaptation under label proportion shifts, revealing widespread failures and proposing a two-step meta-algorithm.
problem Domain adaptation under label proportion shifts is poorly understood and inconsistent across methods.
method RLSbench introduces a large-scale benchmark with 500 distribution shift pairs. It proposes a two-step meta-algorithm to improve domain adaptation methods under label proportion shifts.
result The two-step meta-algorithm improves domain adaptation methods by 2-10% accuracy points under large label proportion shifts.
Domain adaptation is an important technique to alleviate performance degradation caused by domain shift, e.g., when training and test data come from different domains. Most existing deep adaptation methods focus on reducing domain shift by matching marginal feature distributions through deep transformations on the inpu…
This paper explores how effective sample size, dimensionality, and model performance are related in covariate shift adaptation.
problem Understanding the relationship between effective sample size, dimensionality, and generalization in covariate shift adaptation.
method Building a unified theory connecting effective sample size, data dimensionality, and generalization in the context of covariate shift adaptation.
result Dimensionality reduction or feature selection can increase effective sample size, supporting the practice of reducing dimensionality before covariate shift adaptation.
Modified Jones-Faddy skew t-distribution captures asymmetry in stock returns.
problem Negative skew and positive mean in stock returns due to broken symmetry of stochastic volatility.
method Modified Jones-Faddy skew t-distribution applied to split gains and losses, using stochastic differential equations for stock returns and volatility.
result The modified distribution effectively captures the asymmetry in daily S&P500 returns, including its tails.
The percolation model of stock market speculation allows an asymmetry (in the return distribution) leading to fast downward crashes and slow upward recovery. We see more small upturns and more intermediate downturns.
Recent studies have revealed a number of striking dependence patterns in high frequency stock price dynamics characterizing probabilistic interrelation between two consequent price increments x (push) and y (response) as described by the bivariate probability distribution P(x,y) [1,2,3,4]. There are two properties, the…
Market Mill is a complex dependence pattern leading to nonlinear correlations and predictability in intraday dynamics of stock prices. The present paper puts together previous efforts to build a dynamical model reflecting the market mill asymmetries. We show that certain properties of the conditional dynamics at a sing…