New method tests Granger non-causality in panel data with cross-sectional dependencies.
problem Testing Granger non-causality in panel data with cross-sectional dependencies.
method Proposes a new approach to aggregate p-values from panel members to test Granger non-causality, showing lower FDR.
result Our approach discovers true causal relations in panel data, unlike state-of-the-art methods.
Estimates mean and covariance for large, unbalanced stock returns panels.
problem Estimating mean and covariance in large, unbalanced panel data.
method Nonparametric, kernel-based joint estimator for conditional mean and covariance matrices.
result The idiosyncratic risk explains more than 75% of cross-sectional variance.
Paper uses machine learning for nowcasting corporate earnings from mixed-frequency data.
problem Predicting corporate earnings for a large cross-section of firms with different frequency data.
method Structured machine learning regressions with sparse-group LASSO regularization for panel data.
result Machine learning models outperform traditional methods in nowcasting corporate earnings.
Extracts representative scenarios from large data panels.
problem Creating representative scenarios from large data panels.
method Two novel algorithms: one identifies new scenarios, the other selects known important data points.
result Efficient algorithms for consistent scenario-based modeling and multi-dimensional numerical integration.
Micro-panel data are collected and analysed in many research and industry areas. Cluster analysis of micro-panel data is an unsupervised learning exploratory method identifying subgroup clusters in a data set which include homogeneous objects in terms of the development dynamics of monitored variables. The supply of cl…
Method estimates group structure in panel data using variance information.
problem Estimating group structure in panel data with unknown groups.
method Proposes a method to estimate unobserved groupings for panel data models using variance information.
result Superior performance compared to existing methods in simulations and empirical applications.
Paper adapts DML for panel data, addressing unobserved heterogeneity.
problem Estimating causal effects with panel data and unobserved heterogeneity.
method Adapting double/debiased machine learning (DML) for panel data with predictive models based on correlated random effects.
result Predictive models based on correlated random effects within DML lead to accurate coefficient estimates.
Method for factor analysis in short panels without assuming sphericity or Gaussianity.
problem Factor analysis in short panels without assuming sphericity or Gaussianity.
method Pseudo maximum likelihood method and asymptotically uniformly most powerful invariant test.
result Systematic risk explains a large part of cross-sectional total variance in bear markets but is not spanned by observed factors.
We propose a new approach for constructing synthetic pseudo-panel data from cross-sectional data. The pseudo panel and the preferences it intends to describe is constructed at the individual level and is not affected by aggregation bias across cohorts. This is accomplished by creating a high-dimensional probabilistic m…
QRAFTI uses multi-agent framework to improve equity factor research.
problem Replicating and developing new equity factors in large financial datasets.
method Integrates a research toolkit with MCP servers for data access and custom coding operations.
result Improves performance and explainability in multi-step empirical tasks.
The paper improves machine learning for heavy-tailed panel data.
problem Improving estimates for financial and economic data with fat tails.
method Sparse-group LASSO regularization and Fuk-Nagaev concentration inequality.
result Oracle inequalities for panel data estimators.
Paper introduces Functional Effects Models to account for individual heterogeneity in panel data.
problem Accounting for preference heterogeneity in panel data with machine learning.
method Functional Effects Models using gradient boosting decision trees and deep neural networks to learn individual-specific preference parameters.
result Functional Effects Models outperform traditional models in learning inter-individual heterogeneity and predictive performance.
New method for estimating heterogeneous treatment effects in panel data.
problem Estimating heterogeneous treatment effects in non-stationary, temporally dependent panel data.
method Proposes H1SL and H2SL, synthetic learners for panel data, based on existing non-panel data estimators.
result Established convergence rates for proposed estimators and demonstrated superior performance.
We build a simple diagnostic criterion for approximate factor structure in large cross-sectional equity datasets. Given a model for asset returns with observable factors, the criterion checks whether the error terms are weakly cross-sectionally correlated or share at least one unobservable common factor. It only requir…
Study uses exchangeable GPs for staggered-adoption policy evaluation in panel data.
problem Evaluating the impact of staggered treatments in panel data settings.
method Exchangeable multi-task Gaussian processes (GPs) with flexible kernels.
result Flexible tool for policy evaluation in panel data settings.
Paper develops a new estimator for panel data with endogenous treatments, improving causal inference.
problem Challenges in causal inference for static panel data with endogenous treatments and confounding variables.
method Develops Double Machine Learning (DML) estimator for static panel models with endogenous treatments (panel IV DML). Introduces weak-identification diagnostics.
result Panel IV DML estimator improves estimation accuracy and delivers more reliable inference under weak identification.
Estimates heterogeneous treatment effects in panel data with a new method.
problem Estimating heterogeneous treatment effects in panel data with general treatment patterns.
method Partition observations into clusters with similar treatment effects using a regression tree, then estimate average treatment effects for each cluster.
result Our method achieves superior accuracy compared to alternative approaches.
Proposes CoDEAL for estimating heterogeneous treatment effects in panel data models.
problem Estimating heterogeneous treatment effects in causal panel data models with covariate effects.
method Covariate-Adjusted Deep Causal Learning (CoDEAL) integrating neural networks and autoencoders.
result Establishes theoretical guarantees and demonstrates compelling performance in simulations and real data.
P-Trees improve investment performance by optimizing the efficient frontier.
problem Optimizing investment performance in complex financial markets.
method Introducing P-Trees, a new tree-based model for analyzing panel data.
result P-Trees significantly advance the efficient frontier and outperform existing models.
A new method for online prediction uncertainty quantification in non-exchangeable panel data.
problem Challenges in quantifying predictive uncertainty for non-exchangeable panel data.
method Online conformal prediction framework for non-exchangeable panel data, using similarity weights and adaptive miscoverage levels.
result Improves coverage on worst-covered target units through adaptive interval-width allocation.
Improved causal inference with panel data using deep learning.
problem Causal inference challenges in social science with panel data.
method Adapted N-BEATS deep neural architecture for time series forecasting.
result SyNBEATS estimator outperforms existing methods in panel data settings.
The paper develops coresets for panel data regression problems.
problem Efficiently summarize panel data regression problems.
method Introduced coreset construction for panel data regression problems using the Feldman-Langberg framework.
result Constructs coresets of size polynomial in 1/ε and number of parameters, independent of panel data size. New complete panel dataset for LMICs helps analyze innovation and development.
problem Lack of complete data for empirical analyses in LMICs.
method Predictive Mean Matching multiple imputation technique.
result Created a large dataset of 47 variables for 82 LMICs from 2005-2019.
Paper develops a new estimator for high-dimensional panel data with common shocks.
problem Cross-sectionally dependent errors driven by common shocks in high-dimensional panel data.
method Factor-augmented sparse-group LASSO estimator combining MIDAS aggregation with latent factors.
result The estimator outperforms standard LASSO for prediction and estimation in settings with cross-sectional dependence.
Causalfe estimates treatment effects in panel data with fixed effects.
problem Spurious heterogeneity in treatment effect estimates due to fixed effects in panel data.
method CFFE approach with node-level residualization during tree construction.
result Validates the estimator's performance through simulation studies.
Improved forecasting of investment dynamics across heterogeneous panels using a two-stage model.
problem Forecasting investment dynamics in heterogeneous panels with varying dynamics.
method Two-stage architecture: global pooled AR(1) for shared persistence, local models for residual dynamics.
result Significant improvement in out-of-sample R2 from 0.630 to 0.677, with a gain of 0.047. Flexible variable selection handles missing data for better biomarker panels.
problem Identifying relevant features from incomplete data sets.
method Nonparametric variable selection combined with multiple imputation.
result Improved biomarker panels with higher classification and variable selection performance.
Gradient boosting algorithm for spatial panel models improves estimation in high-dimensional settings.
problem Estimation failure in high-dimensional spatial panel models.
method Model-based gradient boosting algorithm for spatial panel models with random and fixed effects.
result Feasibility and interpretability in both low- and high-dimensional settings.
The paper tackles temporal coverage bias in financial panel data, proposing a structuring framework to correct for incomplete histories.
problem Incomplete histories of financial instruments lead to biased panel data.
method Formalizes the problem and proposes a coverage-aware structuring framework using structured metadata and an availability matrix.
result The framework reveals substantial distortions in return dynamics and volatility when naive temporal alignment is used.
We present the first framework for Gaussian-process-modulated Poisson processes when the temporal data appear in the form of panel counts. Panel count data frequently arise when experimental subjects are observed only at discrete time points and only the numbers of occurrences of the events between subsequent observati…
Simple method for estimating missing panel data entries with confidence intervals.
problem Estimating missing values in panel data with staggered adoption.
method Simple matrix algebra and singular value decomposition for estimation, with data-driven confidence intervals.
result Confidence intervals match non-asymptotic lower bounds, proving instance optimality.
In this paper we develop a data-driven smoothing technique for high-dimensional and non-linear panel data models. We allow for individual specific (non-linear) functions and estimation with econometric or machine learning methods by using weighted observations from other individuals. The weights are determined by a dat…
New tests for identifying the number of latent factors in short panels with small time dimensions.
problem Determining the number of latent factors in short panels with small time dimensions.
method Eigenvalue tests based on variance-covariance matrices of asset returns, with assumptions on spherical errors or instrumental variables for factor betas.
result Established asymptotic distributional results and proposed a novel statistical test for weak factors.
Develops DML for nonlinear panel data models with fixed effects.
problem Estimating causal effects in nonlinear panel data models with fixed effects.
method Double machine learning (DML) procedures for approximating nuisance functions.
result First-differencing yields the least constraints on fixed effects distribution.
R package xtdml uses DML for panel data models with fixed effects.
problem Estimating structural parameters in panel data models with fixed effects.
method Combines machine learning with statistical estimation for inference.
result Demonstrates improved performance in learning nuisance functions.
Proposes new methods for Markov chain choice models with panel data.
problem Dependence among transactions for the same customer in historical data.
method Expectation-maximization (EM) algorithms incorporating partial-ordering preference information.
result EM algorithms outperform traditional methods on synthetic and real datasets.
Surveying machine learning methods for economic forecasting.
problem Improving accuracy of economic forecasts using machine learning.
method Nowcasting, textual data, panel and tensor data, high-dimensional Granger causality tests, time series cross-validation, classification with economic losses.
result Recent advances in machine learning methods enhance economic forecasting accuracy.
Optimal tensor PCA for estimating factors and loadings in high-dimensional panel data.
problem Estimating factors and loadings in high-dimensional panel data with non-negligible correlations.
method Tensor Principal Component Analysis (TPCA) for estimating factors and loadings in a tensor factor model.
result Simple TPCA is optimal for strong factors and can be improved for weak factors with alternating least-squares iterations.
Adaptive PCR improves panel data analysis with uniform guarantees.
problem Adaptive data collection in panel data settings.
method Adapting PCR to online settings using martingale concentration.
result Time-uniform guarantees for adaptive PCR in panel data.
Paper develops robust methods for panel data with latent groups, improving inference under group separation violations.
problem Inference in latent group panel models under group separation violations.
method Selective conditional inference approach to derive conditional distribution of coefficients given estimated group structure.
result Valid inference under violations of group separation, superior to traditional asymptotic methods.
FOCUS method forecasts counterfactuals in panel data with time series dynamics.
problem Forecasting unobserved potential outcomes in causal inference with missing entries and latent factors.
method FOCUS extends matrix completion methods by leveraging time series dynamics of latent factors.
result FOCUS method outperforms existing benchmarks in predicting future counterfactuals.
Researchers often summarize their work in the form of posters. Posters provide a coherent and efficient way to convey core ideas from scientific papers. Generating a good scientific poster, however, is a complex and time consuming cognitive task, since such posters need to be readable, informative, and visually aesthet…
This paper presents a novel scheme, based on a unique combination of genetic algorithms (GAs) and deep learning (DL), for the automatic reconstruction of Portuguese tile panels, a challenging real-world variant of the jigsaw puzzle problem (JPP) with important national heritage implications. Specifically, we introduce …
Estimates treatment effects in panel data with general intervention patterns.
problem Estimating average treatment effects in panel data with heterogeneous treatment effects.
method Extends synthetic control framework to allow rate-optimal recovery of average treatment effects for general intervention patterns.
result First rate-optimal guarantees for general intervention patterns in estimating average treatment effects.
Panel count data describes aggregated counts of recurrent events observed at discrete time points. To understand dynamics of health behaviors, the field of quantitative behavioral research has evolved to increasingly rely upon panel count data collected via multiple self reports, for example, about frequencies of smoki…
Proposes a new estimator for weak instrumental variables in panel data models.
problem Weak instrumental variables due to ignored nonlinearities in panel data.
method Triangular simultaneous equation model with a nonlinear reduced form equation and a control function approach using Super Learner.
result The proposed SLCF estimator is consistent and asymptotically normal, achieving a parametric rate of convergence.
Unified model combines scores and rankings for grant panel review.
problem Combining scores and rankings for quality assessment in panel review.
method Mallows-Binomial model with tree-search algorithm for exact MLE.
result Model combines scores and rankings to quantify object quality and measure consensus.
We analyze three sets of income data: the US Panel Study of Income Dynamics PSID), the British Household Panel Survey (BHPS), and the German Socio-Economic Panel (GSOEP). It is shown that the empirical income distribution is consistent with a two-parameter lognormal function for the low-middle income group (97%-99% of …