Style Miner generates stable and significant style factors for time series analysis.
problem Finding significant and stable explanatory factors in high-dimensional time series data.
method Proposes a reinforcement learning method to balance explanatory power and stability constraints.
result Outperforms existing methods by a large margin and achieves a 10% gain in R-squared explanatory power.
Realized moments of higher order computed from intraday returns are introduced in recent years. The literature indicates that realized skewness is an important factor in explaining future asset returns. However, the literature mainly focuses on the whole market and on the monthly or weekly scale. In this paper, we cond…
We investigate entropy as a financial risk measure. Entropy explains the equity premium of securities and portfolios in a simpler way and, at the same time, with higher explanatory power than the beta parameter of the capital asset pricing model. For asset pricing we define the continuous entropy as an alternative meas…
Although interactive learning puts the user into the loop, the learner remains mostly a black box for the user. Understanding the reasons behind queries and predictions is important when assessing how the learner works and, in turn, trust. Consequently, we propose the novel framework of explanatory interactive learning…
Paper uses interbank contagion to predict U.S. bank defaults, finding it highly explanatory.
problem Predicting U.S. bank defaults using interbank contagion.
method Regression and neural network models were used to analyze U.S. commercial bank data.
result Interbank contagion is highly explanatory in default prediction, often outperforming established metrics.
Develops method to assess feature importance in black-box models for unconditional distribution.
problem Lack of methods to analyze feature importance in black-box models for unconditional distribution.
method Approximation method to compute feature importance curves for unconditional distribution.
result Produces sparse and faithful results, computationally efficient.
Bayesian framework explains diverse explanatory values.
problem Understanding and predicting human preferences for explanations.
method Developed a Bayesian account to integrate various explanatory values.
result Core values from psychology, statistics, and philosophy emerge from a common framework.
Deep learning extracts terrain texture covariates for geostatistical modeling.
problem Improving prediction accuracy in geostatistical modeling using terrain texture data.
method Deep learning approach to automatically derive optimal terrain texture covariates from SRTM 90m DEM.
result Deep learning-derived covariates have strong explanatory power (R-squared around 0.6) for geochemical data.
New models reduce discrimination in machine learning without sacrificing explanatory bias.
problem Discrimination and explanatory bias in fairness measures.
method Causal effect estimators using propensity score analysis.
result Theoretical and practical superiority of FairCEEs over existing models.
Interpretable machine learning uncovers ESG's explanatory power on equity returns across sectors and capitalizations.
problem Explaining equity returns beyond market factors using ESG data.
method Interpretable machine learning models, cross-validation scheme, random company-wise validation.
result Gradient boosting models explain unaccounted price returns, with ESG data outperforming basic fundamental features.
We present explicit formulas - that are also computer code - for 101 real-life quantitative trading alphas. Their average holding period approximately ranges 0.6-6.4 days. The average pair-wise correlation of these alphas is low, 15.9%. The returns are strongly correlated with volatility, but have no significant depend…
The paper models intraday power prices using fundamental drivers.
problem Lack of research on drivers for intraday price processes.
method Modelling location, shape, and scale of intraday price distribution using fundamental variables.
result Significant improvements in probabilistic forecasting performance, especially in tails.
PiNets provide faithful explanations for neural networks.
problem Lack of true explanations for neural network predictions.
method Pointwise-interpretable Networks (PiNets) that form linear models instance-wise.
result PiNets offer explanations that are meaningful, aligned, robust, and sufficient.
Method identifies causal interactions between time series using extreme eigenvalue variability.
problem Detecting causal interactions between time series.
method Largest eigenvalue of lagged correlation matrices, measuring causal interactions through variability.
result The method outperforms traditional Granger causality tests in detecting structural changes.
This paper evaluates the impact of the power extent on price in the electricity market. The competitiveness extent of the electricity market during specific times in a day is considered to achieve this. Then, the effect of competitiveness extent on the forecasting precision of the daily power price is assessed. A price…
Subset selection in multiple linear regression aims to choose a subset of candidate explanatory variables that tradeoff fitting error (explanatory power) and model complexity (number of variables selected). We build mathematical programming models for regression subset selection based on mean square and absolute errors…
Study tests if equity factors explain Bitcoin's risk and returns.
problem Explaining Bitcoin's risk and return with equity factors.
method Applied statistical methods to test Fama-French factors on Bitcoin's excess returns.
result Fama-French factors have explanatory power on Bitcoin's risk and returns.
Interpretable representations improve explainable AI by translating complex data into understandable concepts.
problem Many explainers use interpretable representations but overlook their full potential and assumptions.
method An in-depth analysis of interpretable representations for tabular, image, and text data, identifying strengths, weaknesses, and desiderata.
result Linear model quantifies interpretable concepts' influence on black-box predictions, revealing their explanatory properties and manipulability.
PROD method improves high-dimensional regression by handling strong correlations.
problem Violation of Irrepresentable Condition in LASSO for high-dimensional data.
method PROD procedure based on orthogonal decomposition of design matrix.
result PROD enhances performance of high-dimensional penalized regression.
FFRK automatically extracts features for spatial interpolation without external variables.
problem Spatial interpolation challenges, especially nonstationarity and lack of explanatory variables.
method Feature-Free Regression Kriging (FFRK) method that extracts geospatial features.
result FFRK outperforms classical methods in predicting heavy metal concentrations.
Skewness dispersion predicts future stock market returns, especially in months with monetary policy announcements.
problem Predicting future stock market returns using skewness dispersion.
method Cross-sectional analysis of firm-level realized skewness and stock market returns.
result Skewness dispersion is a significant predictor of future stock market returns, robust to various estimation methods.
In light of the power problems of statistical tests and undisciplined use of alpha-based statistics to compare models, this paper proposes a unified set of distance-based performance metrics, derived as the square root of the sum of squared alphas and squared standard errors. The Bayesian investor views model performan…
We statistically investigate the distribution of share price and the distributions of three common financial indicators using data from approximately 8,000 companies publicly listed worldwide for the period 2004-2013. We find that the distribution of share price follows Zipf's law; that is, it can be approximated by a …
Subset selection for multiple linear regression aims to construct a regression model that minimizes errors by selecting a small number of explanatory variables. Once a model is built, various statistical tests and diagnostics are conducted to validate the model and to determine whether the regression assumptions are me…
XGL uses global explanations to guide human supervision in machine learning.
problem Improving model quality through human-machine interaction.
method XGL employs global explanations to guide human selection of informative examples.
result XGL avoids overselling the model's quality and performs comparably to other strategies.
Defines explainability as reasoning under background knowledge.
problem Lack of agreed definitions in explainable AI.
method Reviews philosophical and social foundations, translates to tech realm.
result Defines explainability as logical reasoning under background knowledge.
New framework predicts earnings announcements using press release content, surpassing earnings surprises.
problem Predicting stock returns based on earnings press releases.
method Compared traditional and BERT-based embeddings of press releases, finding content as informative as earnings surprises.
result FinBERT yields highest predictive power for earnings announcement returns.
Knockoffs method selects financial factors, controlling false discoveries.
problem Controlling false discoveries in financial factor selection.
method Apply knockoff procedure to build fake factors.
result Shows versatility in fund replication and network inference.
Statistical boosting algorithms have triggered a lot of research during the last decade. They combine a powerful machine-learning approach with classical statistical modelling, offering various practical advantages like automated variable selection and implicit regularization of effect estimates. They are extremely fle…
A new baseline for Shapley values in MLPs considers model use.
problem Lack of a robust baseline for Shapley values in neural networks.
method Proposes a neutrality-based baseline for Shapley values.
result Empirically validated the proposed baseline for binary classification tasks.
Machine learning predicts CO2 emissions in power grids, reducing uncertainty.
problem Forecasting CO2 emission intensities in power grids.
method Developed a machine learning algorithm using LASSO, feature selection, and Softmax weighted average.
result Marginal emissions are independent of DK2 zone conditions, suggesting external generators.
New research shows input-gradients can be manipulated without changing model's core function, challenging their use for model interpretation.
problem Current methods for model interpretability using input-gradients are flawed due to their arbitrary manipulability.
method Investigated by reinterpreting logits as unnormalized log-densities, proposing novel approximations for score-matching.
result Improving alignment between implicit density model and data distribution enhances gradient structure and explanatory power.
Study finds it hard to establish common factor pricing in corporate bonds.
problem Difficulty in establishing common factor pricing in corporate bonds.
method Portfolio- and bond-level analyses using multifactor models.
result Common factor pricing in corporate bonds is not significantly explanatory.
There has recently been a surge of work in explanatory artificial intelligence (XAI). This research area tackles the important problem that complex machines and algorithms often cannot provide insights into their behavior and thought processes. XAI allows users and parts of the internal system to be more transparent, p…
RelatIF selects more intuitive training examples for explaining model predictions.
problem Influence functions identify outliers as explanatory examples, leading to poor explanations.
method RelatIF separates global and local influence, optimizing for local relative to global effects.
result Examples selected by RelatIF are more intuitive than those from influence functions.
Estimates self- and cross-impact concavity and decay patterns in financial markets.
problem Understanding the impact of financial transactions on market dynamics.
method Nonparametric estimation of concave multi-asset propagator models using metaorders and order flow data.
result Concave self-impact with shifted power-law decay, significant gain from cross-impact, and improved predictive accuracy.
Survey examines deep neural networks' ability to approximate functions.
problem Approximation of target functions by deep neural networks.
method Examination of feed-forward and residual architectures, focusing on optimization problems in regression and classification.
result Deep neural networks can approximate functions effectively, especially with ReLU activation functions.
We find that factors explaining bank loan recovery rates vary depending on the state of the economic cycle. Our modeling approach incorporates a two-state Markov switching mechanism as a proxy for the latent credit cycle, helping to explain differences in observed recovery rates over time. We are able to demonstrate ho…
A new PCR method using SVD with sparse regularization.
problem Lack of response variable information in traditional PCR.
method One-stage SVD approach with two loss functions and sparse regularization.
result Obtains principal component loadings with response variable information.
This paper investigates the risk-return relationship in determination of housing asset pricing. In so doing, the paper evaluates behavioral hypotheses advanced by Case and Shiller (1988, 2002, 2009) in studies of boom and post-boom housing markets. The paper specifies and tests a multi-factor housing asset pricing mode…
A new estimator learns sparse linear models with context-dependent coefficients.
problem Sparse linear models lack flexibility compared to deep neural networks for handling feature groups.
method Contextual lasso estimator using a deep neural network with lasso regularization.
result Learned models can be sparser than standard lasso without sacrificing predictive power.
Study shows cryptocurrency market impact on DeFi returns stronger than other drivers.
problem Understanding drivers of DeFi returns and their relative importance.
method Investigated four drivers: cryptocurrency market exposure, network effect, investor attention, and valuation ratio. Designed a new market index, DeFiX.
result Cryptocurrency market impact on DeFi returns is stronger than other drivers and provides superior explanatory power.
The paper predicts and explains the decay of stock anomaly performance over time.
problem Predicting and explaining the drop in risk-adjusted performance of stock anomalies.
method The authors propose ex-ante characteristics based on hypotheses of out-of-sample decay and in-sample overfitting.
result The year of publication explains 30% of the variance in Sharpe decay across factors.
New statistical factors improve portfolio risk estimation.
problem Improving estimation of portfolio risk using new statistical factors.
method Matrix factor models and statistical methods (partial F test, double selection LASSO).
result New statistical factors add explanatory power in asset pricing.
Improved disability insurance model with collective health claims.
problem Enhance disability insurance model with collective health claims.
method Expand classic semi-Markov model with collective health claims, solve many-body problem using mean-field approach.
result Mean-field approach simplifies complex model into a transparent pricing method.
Motivated by the practical challenge in monitoring the performance of a large number of algorithmic trading orders, this paper provides a methodology that leads to automatic discovery of the causes that lie behind a poor trading performance. It also gives theoretical foundations to a generic framework for real-time tra…
This paper examines the possibility of using derivative-implied risk premia to explain stock returns. The rapid development of derivative markets has led to the possibility of trading various kinds of risks, such as credit and interest rate risk, separately from each other. This paper uses credit default swaps and equi…
This text discusses several popular explanatory methods that go beyond the error measurements and plots traditionally used to assess machine learning models. Some of the explanatory methods are accepted tools of the trade while others are rigorously derived and backed by long-standing theory. The methods, decision tree…