New method creates synthetic pseudo-panel data to study transport preferences.
problem Aggregation bias in cross-sectional data.
method Conditional Variational Autoencoder (CVAE) framework for probabilistic model.
result Revealed dynamics of transport preferences and categorized individuals.
Method for factor analysis in short panels without assuming sphericity or Gaussianity.
problem Factor analysis in short panels without assuming sphericity or Gaussianity.
method Pseudo maximum likelihood method and asymptotically uniformly most powerful invariant test.
result Systematic risk explains a large part of cross-sectional total variance in bear markets but is not spanned by observed factors.
The study predicts surgical complications in Crohn's disease patients using machine learning.
problem Predicting surgical complications in Crohn's disease patients.
method Developed a novel algorithm using ensemble machine learning on 29 baseline covariates.
result Proposed pseudo-observation based estimators for evaluating predictive performance.
The paper improves machine learning for heavy-tailed panel data.
problem Improving estimates for financial and economic data with fat tails.
method Sparse-group LASSO regularization and Fuk-Nagaev concentration inequality.
result Oracle inequalities for panel data estimators.
New method tests Granger non-causality in panel data with cross-sectional dependencies.
problem Testing Granger non-causality in panel data with cross-sectional dependencies.
method Proposes a new approach to aggregate p-values from panel members to test Granger non-causality, showing lower FDR.
result Our approach discovers true causal relations in panel data, unlike state-of-the-art methods.
Paper develops a new estimator for panel data with endogenous treatments, improving causal inference.
problem Challenges in causal inference for static panel data with endogenous treatments and confounding variables.
method Develops Double Machine Learning (DML) estimator for static panel models with endogenous treatments (panel IV DML). Introduces weak-identification diagnostics.
result Panel IV DML estimator improves estimation accuracy and delivers more reliable inference under weak identification.
Method improves developability of B-spline surfaces for architectural design.
problem Developability of B-spline surfaces for architectural freeform surfaces.
method Algorithm to thin the Gauss image by increasing planarity of neighborhoods, allowing for paneling with specific developable types.
result Enforced developable panels for architectural design, reducing manufacturing costs.
New method for estimating heterogeneous treatment effects in panel data.
problem Estimating heterogeneous treatment effects in non-stationary, temporally dependent panel data.
method Proposes H1SL and H2SL, synthetic learners for panel data, based on existing non-panel data estimators.
result Established convergence rates for proposed estimators and demonstrated superior performance.
A novel clustering method for micro-panel data improves efficiency and accuracy.
problem Insufficient clustering methods for micro-panel data.
method Feature-based two-step approach for identifying homogeneous clusters.
result The CluMP method outperforms existing methods in clustering micro-panel data.
A new method for online prediction uncertainty quantification in non-exchangeable panel data.
problem Challenges in quantifying predictive uncertainty for non-exchangeable panel data.
method Online conformal prediction framework for non-exchangeable panel data, using similarity weights and adaptive miscoverage levels.
result Improves coverage on worst-covered target units through adaptive interval-width allocation.
Estimates mean and covariance for large, unbalanced stock returns panels.
problem Estimating mean and covariance in large, unbalanced panel data.
method Nonparametric, kernel-based joint estimator for conditional mean and covariance matrices.
result The idiosyncratic risk explains more than 75% of cross-sectional variance.
Proposes CoDEAL for estimating heterogeneous treatment effects in panel data models.
problem Estimating heterogeneous treatment effects in causal panel data models with covariate effects.
method Covariate-Adjusted Deep Causal Learning (CoDEAL) integrating neural networks and autoencoders.
result Establishes theoretical guarantees and demonstrates compelling performance in simulations and real data.
Paper uses machine learning for nowcasting corporate earnings from mixed-frequency data.
problem Predicting corporate earnings for a large cross-section of firms with different frequency data.
method Structured machine learning regressions with sparse-group LASSO regularization for panel data.
result Machine learning models outperform traditional methods in nowcasting corporate earnings.
The 2007 Midwest Geometry Conference included a panel discussion devoted to open problems and the general direction of future research in fields related to the main themes of the conference. This paper summarizes the comments made during the panel discussion.
Estimates heterogeneous treatment effects in panel data with a new method.
problem Estimating heterogeneous treatment effects in panel data with general treatment patterns.
method Partition observations into clusters with similar treatment effects using a regression tree, then estimate average treatment effects for each cluster.
result Our method achieves superior accuracy compared to alternative approaches.
Improved causal inference with panel data using deep learning.
problem Causal inference challenges in social science with panel data.
method Adapted N-BEATS deep neural architecture for time series forecasting.
result SyNBEATS estimator outperforms existing methods in panel data settings.
Improved forecasting of investment dynamics across heterogeneous panels using a two-stage model.
problem Forecasting investment dynamics in heterogeneous panels with varying dynamics.
method Two-stage architecture: global pooled AR(1) for shared persistence, local models for residual dynamics.
result Significant improvement in out-of-sample R2 from 0.630 to 0.677, with a gain of 0.047. The paper develops coresets for panel data regression problems.
problem Efficiently summarize panel data regression problems.
method Introduced coreset construction for panel data regression problems using the Feldman-Langberg framework.
result Constructs coresets of size polynomial in 1/ε and number of parameters, independent of panel data size. Paper develops a new estimator for high-dimensional panel data with common shocks.
problem Cross-sectionally dependent errors driven by common shocks in high-dimensional panel data.
method Factor-augmented sparse-group LASSO estimator combining MIDAS aggregation with latent factors.
result The estimator outperforms standard LASSO for prediction and estimation in settings with cross-sectional dependence.
Gradient boosting algorithm for spatial panel models improves estimation in high-dimensional settings.
problem Estimation failure in high-dimensional spatial panel models.
method Model-based gradient boosting algorithm for spatial panel models with random and fixed effects.
result Feasibility and interpretability in both low- and high-dimensional settings.
Flexible variable selection handles missing data for better biomarker panels.
problem Identifying relevant features from incomplete data sets.
method Nonparametric variable selection combined with multiple imputation.
result Improved biomarker panels with higher classification and variable selection performance.
The paper tackles temporal coverage bias in financial panel data, proposing a structuring framework to correct for incomplete histories.
problem Incomplete histories of financial instruments lead to biased panel data.
method Formalizes the problem and proposes a coverage-aware structuring framework using structured metadata and an availability matrix.
result The framework reveals substantial distortions in return dynamics and volatility when naive temporal alignment is used.
Causalfe estimates treatment effects in panel data with fixed effects.
problem Spurious heterogeneity in treatment effect estimates due to fixed effects in panel data.
method CFFE approach with node-level residualization during tree construction.
result Validates the estimator's performance through simulation studies.
We analyze linear factor models for asset pricing panels.
problem Characterizing cross-sectional and inter-temporal properties of returns and factors.
method Conditional means and covariances, review of Kozak and Nagel (2024) conditions.
result Low-dimensional factor portfolios can span efficient portfolios in unbalanced panels.
We present the first framework for Gaussian-process-modulated Poisson processes when the temporal data appear in the form of panel counts. Panel count data frequently arise when experimental subjects are observed only at discrete time points and only the numbers of occurrences of the events between subsequent observati…
Proposes a robust EM algorithm for analyzing incomplete panel count data.
problem Missing reports in panel count data.
method Functional EM algorithm for non-parametric counting process mean function estimation.
result Robust to misspecification of Poisson process assumption and missing completely at random.
Optimal tensor PCA for estimating factors and loadings in high-dimensional panel data.
problem Estimating factors and loadings in high-dimensional panel data with non-negligible correlations.
method Tensor Principal Component Analysis (TPCA) for estimating factors and loadings in a tensor factor model.
result Simple TPCA is optimal for strong factors and can be improved for weak factors with alternating least-squares iterations.
Method estimates group structure in panel data using variance information.
problem Estimating group structure in panel data with unknown groups.
method Proposes a method to estimate unobserved groupings for panel data models using variance information.
result Superior performance compared to existing methods in simulations and empirical applications.
Develops DML for nonlinear panel data models with fixed effects.
problem Estimating causal effects in nonlinear panel data models with fixed effects.
method Double machine learning (DML) procedures for approximating nuisance functions.
result First-differencing yields the least constraints on fixed effects distribution.
Simple method for estimating missing panel data entries with confidence intervals.
problem Estimating missing values in panel data with staggered adoption.
method Simple matrix algebra and singular value decomposition for estimation, with data-driven confidence intervals.
result Confidence intervals match non-asymptotic lower bounds, proving instance optimality.
Paper introduces Functional Effects Models to account for individual heterogeneity in panel data.
problem Accounting for preference heterogeneity in panel data with machine learning.
method Functional Effects Models using gradient boosting decision trees and deep neural networks to learn individual-specific preference parameters.
result Functional Effects Models outperform traditional models in learning inter-individual heterogeneity and predictive performance.
R package xtdml uses DML for panel data models with fixed effects.
problem Estimating structural parameters in panel data models with fixed effects.
method Combines machine learning with statistical estimation for inference.
result Demonstrates improved performance in learning nuisance functions.
Proposes new methods for Markov chain choice models with panel data.
problem Dependence among transactions for the same customer in historical data.
method Expectation-maximization (EM) algorithms incorporating partial-ordering preference information.
result EM algorithms outperform traditional methods on synthetic and real datasets.
Researchers often summarize their work in the form of posters. Posters provide a coherent and efficient way to convey core ideas from scientific papers. Generating a good scientific poster, however, is a complex and time consuming cognitive task, since such posters need to be readable, informative, and visually aesthet…
A hybrid scheme uses GAs and DL for tile panel reconstruction.
problem Reconstructing Portuguese tile panels with real-world effects.
method Enhanced GA-based puzzle solver with novel DLCM.
result 82% accuracy for tile reconstruction compared to 3.5% for best known method.
Study shows stick numbers for specific graphs and explains a protein structure.
problem Determining the minimum number of sticks for knotless embeddings of graphs.
method Analyzing specific graphs like K4 and K5, and using probability theory for K3,3.
result Minimum stick numbers for K4 and K5, and probability of a specific protein structure.
Paper develops robust methods for panel data with latent groups, improving inference under group separation violations.
problem Inference in latent group panel models under group separation violations.
method Selective conditional inference approach to derive conditional distribution of coefficients given estimated group structure.
result Valid inference under violations of group separation, superior to traditional asymptotic methods.
Adaptive PCR improves panel data analysis with uniform guarantees.
problem Adaptive data collection in panel data settings.
method Adapting PCR to online settings using martingale concentration.
result Time-uniform guarantees for adaptive PCR in panel data.
Surveying machine learning methods for economic forecasting.
problem Improving accuracy of economic forecasts using machine learning.
method Nowcasting, textual data, panel and tensor data, high-dimensional Granger causality tests, time series cross-validation, classification with economic losses.
result Recent advances in machine learning methods enhance economic forecasting accuracy.
Causal forests underestimate treatment effect heterogeneity, a correction is proposed.
problem Causal forests underestimate treatment effect heterogeneity in fixed-effects panel settings.
method Cross-fitted correction to estimate and restore the spread of conditional average treatment effects.
result The correction reduces mean-squared error by 25-42% in simulations and restores heterogeneity in a real-world panel study.
QRAFTI uses multi-agent framework to improve equity factor research.
problem Replicating and developing new equity factors in large financial datasets.
method Integrates a research toolkit with MCP servers for data access and custom coding operations.
result Improves performance and explainability in multi-step empirical tasks.
FOCUS method forecasts counterfactuals in panel data with time series dynamics.
problem Forecasting unobserved potential outcomes in causal inference with missing entries and latent factors.
method FOCUS extends matrix completion methods by leveraging time series dynamics of latent factors.
result FOCUS method outperforms existing benchmarks in predicting future counterfactuals.
Paper adapts DML for panel data, addressing unobserved heterogeneity.
problem Estimating causal effects with panel data and unobserved heterogeneity.
method Adapting double/debiased machine learning (DML) for panel data with predictive models based on correlated random effects.
result Predictive models based on correlated random effects within DML lead to accurate coefficient estimates.
Estimates treatment effects in panel data with general intervention patterns.
problem Estimating average treatment effects in panel data with heterogeneous treatment effects.
method Extends synthetic control framework to allow rate-optimal recovery of average treatment effects for general intervention patterns.
result First rate-optimal guarantees for general intervention patterns in estimating average treatment effects.
Extracts representative scenarios from large data panels.
problem Creating representative scenarios from large data panels.
method Two novel algorithms: one identifies new scenarios, the other selects known important data points.
result Efficient algorithms for consistent scenario-based modeling and multi-dimensional numerical integration.
Proposes a new estimator for weak instrumental variables in panel data models.
problem Weak instrumental variables due to ignored nonlinearities in panel data.
method Triangular simultaneous equation model with a nonlinear reduced form equation and a control function approach using Super Learner.
result The proposed SLCF estimator is consistent and asymptotically normal, achieving a parametric rate of convergence.
Develops a data-driven smoothing technique for high-dimensional, non-linear panel data.
problem Improving prediction accuracy in high-dimensional, non-linear panel data models.
method Adaptive discrete smoothing with data-driven weights based on individual function similarity.
result Significant improvement in prediction accuracy compared to traditional linear panel data estimators.
We analyze three sets of income data: the US Panel Study of Income Dynamics PSID), the British Household Panel Survey (BHPS), and the German Socio-Economic Panel (GSOEP). It is shown that the empirical income distribution is consistent with a two-parameter lognormal function for the low-middle income group (97%-99% of …