This paper diagnoses factor-model pricing errors using a new method.
problem Measuring pricing errors in factor models with general characteristic axes.
method Developed a method to measure factor-model pricing errors as bridge-alpha curves, using a predetermined characteristic order and prefix portfolios.
result Adding a counterpart factor flips the curve's sign on every axis, but only HML and CMA overcorrect enough to be rejected.
Study macroscopic equity market properties affecting active strategies.
problem Lack of adequate models for active equity strategies.
method Empirical study using CRSP Database, focusing on market capitalizations and returns.
result Highlight stylized facts and open questions in equity markets.
Develops a new method for optimizing portfolios in stochastic markets.
problem Optimizing functionally generated portfolios in stochastic portfolio theory.
method Optimizes over a family of rank-based portfolios parameterized by an exponentially concave function.
result Proves existence and uniqueness of the optimization problem and provides stability estimates.
Study decomposes market portfolio into body and tail legs, revealing systematic differences.
problem Understanding the relationship between body and tail components in market portfolios.
method Decomposes CRSP market portfolio into body and tail legs, analyzes their recombination identity.
result Recombination identity holds for all models but not for all, indicating systematic differences.
Researchers have constantly asked whether stock returns can be predicted by some macroeconomic data. However, it is known that macroeconomic data may exhibit nonstationarity and/or heavy tails, which complicates existing testing procedures for predictability. In this paper we propose novel empirical likelihood methods …
In this article, the long-term behavior of the stock market index of the New York Stock Exchange is studied, for the period 1950 to 2013. Specifically, the CRSP Value-Weighted and CRSP Equal-Weighted index are analyzed in terms of market efficiency, using the standard ratio variance test, considering over 1600 one week…
Benchmarking deep time series models for equity portfolios
problem Selecting the best deep time series model for equity portfolios
method Using a CRSP benchmark and multi-criteria acceptability analysis
result No architecture dominates the benchmark, with TransEnc-8 having the highest rank-1 acceptability
The paper diagnoses factor models using characteristic axes and zero-curve restrictions.
problem Tackles systematic sign reversals and overcorrections in factor model pricing errors.
method Extends cap-axis integral diagnostic to general characteristic axes, measuring pricing errors as bridge-alpha curves.
result Axis-level pricing errors are nearly orthogonal to maximum-Sharpe gains, showing systematic sign reversals and overcorrections.
The paper diagnoses factor-model pricing errors using characteristic axes and bridge-alpha curves.
problem Tackles systematic sign reversals and overcorrections in factor-model pricing errors.
method Extends cap-axis integral diagnostic to characteristic axes, measures pricing errors as bridge-alpha curves, and uses a predetermined characteristic order to generate zero-curve restrictions.
result Axis-level pricing errors are nearly orthogonal to maximum-Sharpe gains, showing significant sign reversals and overcorrections.
Proposes a diagnostic method to evaluate factor models using cap-axis integrals.
problem Improving factor model evaluation in low-dimensional spaces.
method Lifts pricing errors into a bridge-alpha curve along the market-capitalization rank axis.
result The cap-axis norm is distinct from Sharpe gain and size exposure.
Proposes a diagnostic method to evaluate factor models using cap-axis integrals.
problem Improving factor model evaluation for low-dimensional models.
method Lifts pricing errors into a bridge-alpha curve along the market-capitalization rank axis.
result The cap-axis norm is distinct from Sharpe gain and size exposure.
The MSPI predicts market stress with machine learning.
problem Estimating the probability of high market stress.
method L1-regularized logistic regression on stock fragility signals.
result MSPI tracks major stress episodes and improves accuracy.
Market portfolio decomposed into body and tail legs
problem Separation of market portfolio into body and tail legs
method Dynamic value-weighted body and tail legs
result Recombination identity holds for all models
Bayesian approach confirms no return predictability for 1926-2004 data, weak evidence for 1953-2021.
problem Investigating return predictability using Bayesian methods.
method Developed a new shrinkage type prior for a model parameter in a VAR system, compared to other estimation methods.
result Bayesian approach outperforms reduced-bias estimator in terms of size and power.
Test-asset construction affects factor model performance.
problem How test assets are constructed impacts factor model performance.
method Forming characteristic-unsorted random portfolios and varying stock selection, initial weighting, holding, and rebalancing.
result Test-asset construction shifts factor model rankings materially.
The study uses equity order flow to forecast stock returns and resolves the liquidity premium puzzle.
problem The liquidity premium and its relation to investment horizons.
method Directly estimated Kyle's price-impact coefficient λ from daily equity order flow data.
result Signed order flow predicts stock returns, with volume volatility predicting lower returns.
Tests factor models by decomposing market into body and tail legs, revealing inconsistent results.
problem Inconsistency between factor models and market behavior.
method Decomposes market into body and tail legs, testing factor models at daily and monthly frequencies.
result q5 model shows inconsistent results, with negative body and positive tail alphas at all split ratios.
New method identifies algo trading strategies as liquidity consumers or providers.
problem Determining if algo trading strategies consume or provide liquidity.
method Analyzes trade and price history to classify strategies as liquidity consumers or providers.
result Identifies net liquidity consumption or provision of algo trading strategies.
New method accurately reconstructs Russell 3000 index, revealing crowded portfolios.
problem Crowding in index portfolios during reconstitution events.
method Developed a Python package for accurate index reconstruction using CRSP US Stock data.
result Annual Russell 3000 portfolios are more crowded than quarterly ones, suggesting lower transaction costs.
This paper improves risk control for financial markets by calibrating VaR forecasts using conformal methods.
problem Nonstationary and regime-dependent losses in financial markets.
method Regime-weighted conformal risk control (RWC) for VaR forecasting.
result RWC improves regime-conditional stability in some settings with modest conservativeness changes.
Research evaluates three risk models for portfolio construction during market downturns.
problem Challenges in constructing quantitative portfolios using statistical risk models.
method Three statistical risk models tested on 1,000 stocks across four periods.
result Models consistently outperform market returns in various crises.
The study extends SPT to account for real-world transaction costs, improving portfolio performance.
problem Real-world transaction costs affect portfolio performance, especially during market stress.
method Developed a continuous-time model with stochastic transaction costs and derived lower bounds for cost-adjusted wealth.
result Functionally generated portfolios can still achieve relative arbitrage after accounting for transaction costs.
New model predicts stock performance in large equity markets.
problem Predicting stock performance in large equity markets over long time horizons.
method Rank-based volatility stabilized models calibrated to empirical data.
result The model exhibits relative arbitrage and statistically fits empirical features.
Detecting changes in asset co-movements is of much importance to financial practitioners, with numerous risk management benefits arising from the timely detection of breakdowns in historical correlations. In this article, we propose a real-time indicator to detect temporary increases in asset co-movements, the Autoenco…
Develops a continuous compliance index for Islamic equity screening.
problem Binary rulebooks lead to inconsistent compliance assessment of firms.
method Integrates six leading financial and business activity standards into a single continuous index.
result Firms with the same pass/fail label can differ significantly in compliance strength.
Performance of investment managers are evaluated in comparison with benchmarks, such as financial indices. Due to the operational constraint that most professional databases do not track the change of constitution of benchmark portfolios, standard tests of performance suffer from the "look-ahead benchmark bias," when t…
Continuous Hidden Markov Models for Equity Returns
problem Generating synthetic equity returns that match real return characteristics
method Continuous Hidden Markov Models
result Recovered volatility clustering and narrowed kurtosis gap
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
Prevents sensitive data generation in diffusion models using labeled and unlabeled data.
problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.
Study reveals Data Shapley's inconsistent performance in data selection tasks.
problem Inconsistency of Data Shapley's performance in data selection across different settings.
method Hypothesis testing framework and identification of utility functions.
result Data Shapley's performance is no better than random selection without specific constraints.
PRRO generates synthetic tabular data that improves SL performance and class distribution.
problem Low SL utility of synthetic data due to class imbalance and overlooked data relationships.
method Data pruning and column reordering to optimize SL utility.
result Synthetic data generated with PRRO enhances predictive performance and class distribution.
Defines data science as a natural ecosystem with challenges and missions.
problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.
Differences in data size per class, also known as imbalanced data distribution, have become a common problem affecting data quality. Big Data scenarios pose a new challenge to traditional imbalanced classification algorithms, since they are not prepared to work with such amount of data. Split data strategies and lack o…
Synthetic data enhances analytics but requires careful volume management.
problem Accuracy of statistical methods on synthetic data vs. raw data.
method Synthetic Data Generation for Analytics framework using tabular diffusion models.
result Error rate decreases with more synthetic data but may stabilize or increase.
New test ensures quality of shared data in machine learning.
problem Ensuring quality of external data in machine learning tasks.
method Distribution-free two-sample testing procedures grounded in conformal outlier detection.
result Identifies valuable external data agents for model personalization.
Differential privacy allows quantifying privacy loss resulting from accessing sensitive personal data. Repeated accesses to underlying data incur increasing loss. Releasing data as privacy-preserving synthetic data would avoid this limitation, but would leave open the problem of designing what kind of synthetic data. W…
Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.
problem Ensuring fair predictions across sensitive attributes in synthetic data.
method Equalizing target probability distributions across sensitive attributes in synthetic data generation.
result Synthetic data provides strong fair predictions, equal across all thresholds.
DPA preserves data distribution in reduced dimensions.
problem Loss of data distribution in dimension reduction.
method DPA combines encoder and decoder to match data distribution.
result DPA successfully reconstructs data distribution.
Efficient synthetic data generation improves model performance on tabular data.
problem Improving model robustness and performance with scarce or low-quality data.
method Hardness characterization to identify high-value training points, generating synthetic data only from these points.
result Synthetic data generated from hardest points outperforms non-targeted methods on tabular datasets.
For most problems in science and engineering we can obtain data sets that describe the observed system from various perspectives and record the behavior of its individual components. Heterogeneous data sets can be collectively mined by data fusion. Fusion can focus on a specific target relation and exploit directly ass…
DAERNN models censored data using neural networks with data augmentation.
problem Handling censored data in expectile regression.
method Data augmentation based Expectile Regression Neural Networks (ERNNs).
result DAERNN outperforms existing censored ERNNs methods and achieves comparable predictive performance to fully observed data.
Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…
Data stream classification methods demonstrate promising performance on a single data stream by exploring the cohesion in the data stream. However, multiple data streams that involve several correlated data streams are common in many practical scenarios, which can be viewed as multi-task data streams. Instead of handli…
Data collection is a major bottleneck in machine learning and an active research topic in multiple communities. There are largely two reasons data collection has recently become a critical issue. First, as machine learning is becoming more widely-used, we are seeing new applications that do not necessarily have enough …
This paper quantifies uncertainty in Data Shapley using statistical inference.
problem Uncertainty in data valuation due to dynamic data distribution.
method Established relationship with U-statistics and quantified uncertainty using statistical inference.
result Confidence intervals for Data Shapley estimations are provided.
DCoM uses deep neural networks to detect semantic data types from raw column values.
problem Detecting semantic data types from dirty and unseen data.
method DCoM employs multi-input NLP-based deep neural networks trained on 686,765 data columns.
result DCoM outperforms existing methods significantly on 78 different semantic data types.
Causal inference from observational data is the goal of many data analyses in the health and social sciences. However, academic statistics has often frowned upon data analyses with a causal objective. The introduction of the term "data science" provides a historic opportunity to redefine data analysis in such a way tha…
Task-agnostic data valuation without validation requirements.
problem Valuing data without specific task assumptions.
method Estimating data diversity and relevance through queries without raw data.
result Estimates capture the diversity and relevance of seller's data for the buyer.