Paper improves Lasso for S&P500 index tracking with post-selection inference.
problem Index tracking for S&P500 with many applications.
method Used Lasso for dimension reduction and post-selection inference.
result Lasso method for S&P500 index tracking shows high performance.
New method corrects selection bias in post-selective inference for Group LASSO.
problem Inference after Group LASSO selection is unreliable.
method Develops a consistent, post-selective Bayesian method to adjust for selection bias.
result Corrects bias in recovering effects of selected variables.
Proposes HSIC-Lasso for selective inference in non-linear data.
problem Detecting influential features in non-linear and high-dimensional data.
method Model-free HSIC-Lasso based on truncated Gaussians and polyhedral lemma.
result Tight control of type-I error even for small sample sizes.
Develops methods to adjust prediction set coverage based on post-selection analysis.
problem Adjusting prediction set coverage after initial analysis to better fit specific needs.
method Post-selection conformal inference to adjust miscoverage levels.
result Allows for trade-off between coverage and prediction set quality.
We develop a general approach to valid inference after model selection. At the core of our framework is a result that characterizes the distribution of a post-selection estimator conditioned on the selection event. We specialize the approach to model selection by the lasso to form valid confidence intervals for the sel…
New framework for valid hypothesis testing in complex data settings.
problem Challenges in classical hypothesis testing frameworks.
method Add and subtract external noise to partition data, orthogonalize, and test hypotheses.
result Valid hypothesis tests can be conducted under minimal assumptions.
The paper discusses methods for interval estimation of coefficients in penalized regression models for insurance data.
problem Valid inference on coefficients after feature selection in GLM family for insurance data.
method Proposes methodologies for constructing confidence intervals of coefficients after feature selection in GLM family.
result Valid inference on coefficients after feature selection in GLM family for insurance data.
Paper simplifies data carving inference with a parametric distribution.
problem Valid inference after selection with data carving.
method Developed a parametric distribution for data carving inference.
result Exact inference for data carving can be computed trivially.
"Which Generative Adversarial Networks (GANs) generates the most plausible images?" has been a frequently asked question among researchers. To address this problem, we first propose an \emph{incomplete} U-statistics estimate of maximum mean discrepancy MMDinc to measure the distribution discrepancy betwee…
Study examines inference methods after variable selection in Cox models.
problem Bias and misleading inference after variable selection in Cox models.
method Simulation study of inference procedures for Lasso and adaptive Lasso in Cox models.
result Performance of inference procedures varies, with debiased Lasso showing promise.
A method to split a data point into two parts that individually cannot reconstruct the whole, but together can.
problem Splitting a single data point into two parts such that neither can reconstruct the whole but together can.
method Borrowing ideas from Bayesian inference to achieve a continuous analog of data splitting.
result A method to achieve data fission, enabling post-selection inference in finite samples.
New samplers minimize KL divergence for constrained and non-Euclidean geometries.
problem Efficient sampling from constrained and non-Euclidean distributions.
method Stein Variational Mirror Descent and Mirrored Stein Variational Gradient Descent.
result New samplers converge more rapidly and accurately than prior methods.
We propose a novel kernel based post selection inference (PSI) algorithm, which can not only handle non-linearity in data but also structured output such as multi-dimensional and multi-label outputs. Specifically, we develop a PSI algorithm for independence measures, and propose the Hilbert-Schmidt Independence Criteri…
New algorithms for sampling in constrained domains without learning rates.
problem Sampling in constrained domains with fairness constraints and post-selection inference.
method Coin betting ideas from convex optimisation and a unifying framework for constrained sampling.
result Our algorithms achieve competitive performance without hyperparameter tuning.
Measuring divergence between two distributions is essential in machine learning and statistics and has various applications including binary classification, change point detection, and two-sample test. Furthermore, in the era of big data, designing divergence measure that is interpretable and can handle high-dimensiona…
New method for estimating high-dimensional binary time series coefficients.
problem Statistical inference for high-dimensional binary time series.
method Post-selection estimator and second-order wild bootstrap algorithm.
result Good finite-sample performance of the proposed method.
PS-DME evaluates model performance and reliability after data-dependent selection.
problem Evaluating model performance and reliability when data is used for selection and evaluation.
method Post-selection distributional model evaluation (PS-DME) using e-values to control false coverage rate.
result PS-DME provides reliable comparison of model configurations across different reliability levels.
Proposes MinPEN framework for estimating relationships in multivariate models.
problem Estimating relationships between multivariate outcomes in statistical learning.
method MinPEN framework using minimum function penalty for non-convex optimization.
result Theoretical and practical validation of MinPEN framework for multivariate models.
Finding statistically significant high-order interaction features in predictive modeling is important but challenging task. The difficulty lies in the fact that, for a recent applications with high-dimensional covariates, the number of possible high-order interaction features would be extremely large. Identifying stati…
New method splits unknown covariance Gaussians into independent parts.
problem Splitting multivariate Gaussian data with unknown covariance.
method Developed a general algorithm for decomposing unknown covariance Gaussians.
result Demonstrated decomposition for single multivariate Gaussian with unknown covariance.
While statistics and machine learning offers numerous methods for ensuring generalization, these methods often fail in the presence of adaptivity---the common practice in which the choice of analysis depends on previous interactions with the same dataset. A recent line of work has introduced powerful, general purpose a…
Due to the increasing availability of high-dimensional empirical applications in many research disciplines, valid simultaneous inference becomes more and more important. For instance, high-dimensional settings might arise in economic studies due to very rich data sets with many potential covariates or in the analysis o…
Selective inference for group lasso estimators across various distributions and covariates.
problem Developing selective inference methods for group lasso estimators.
method Randomized group-regularized optimization problem with post-selection likelihood.
result Selective point estimator and Wald-type confidence regions for regression parameters.
We propose a statistical inference framework for the component-wise functional gradient descent algorithm (CFGD) under normality assumption for model errors, also known as L2-Boosting. The CFGD is one of the most versatile tools to analyze data, because it scales well to high-dimensional data sets, allows for a very…
In this paper, we provide efficient estimators and honest confidence bands for a variety of treatment effects including local average (LATE) and local quantile treatment effects (LQTE) in data-rich environments. We can handle very many control variables, endogenous receipt of treatment, heterogeneous treatment effects,…
PANDA augments data to regularize GLM estimation and inference.
problem Regularizing estimation and inference in GLMs with noisy data.
method Iteratively optimizes augmented noise data to converge to regularized model estimates.
result Established convergence and asymptotic distributions for regularized parameters.
Study on estimating causal effects with limited data and multiple environments.
problem Estimating causal effects under hidden confounding with unpaired data and sparse effects.
method Instrumental variable (IV) regression with cross-fold sample splitting and ℓ1-regularized estimation. result Proposed GMM-type estimator is consistent as the number of environments grows.
DebiNet uses over-parameterized neural networks to improve linear model performance and debiasing.
problem Improving linear model performance and debiasing in high-dimensional settings.
method Incorporates over-parameterized neural networks into semi-parametric models to estimate parameters consistently.
result DebiNet offers valid inference and accurate prediction by leveraging neural networks' universal approximation and linear model's interpretability.
Refining one's hypotheses in the light of data is a common scientific practice; however, the dependency on the data introduces selection bias and can lead to specious statistical analysis. An approach for addressing this is via conditioning on the selection procedure to account for how we have used the data to generate…
DM framework improves robustness and efficiency in latent-mixture models.
problem Efficient and robust inference in latent-mixture models.
method Divergence-minimization framework with monotonic convergence and robustness guarantees.
result DM yields consistent and asymptotically normal estimators under correct specification.
New framework for cyclic quantum causal models with graph separation property.
problem Understanding causal relationships in feedback processes and exotic scenarios.
method Introducing a robust probability rule and a novel graph-separation property, p-separation.
result Established graph-separation properties for all consistent cyclic causal models.
New method for estimating value of optimal policies in uncertain scenarios.
problem Inference for optimal policies when they are non-unique or nearly deterministic.
method Semiparametric efficiency bound, uniformly weighted estimator, NSAVE method.
result Proposes NSAVE method for robust inference in uncertain optimal policies.
We address the problem of non-parametric multiple model comparison: given l candidate models, decide whether each candidate is as good as the best one(s) or worse than it. We propose two statistical tests, each controlling a different notion of decision errors. The first test, building on the post selection inference…
Ever since the proof of asymptotic normality of maximum likelihood estimator by Cramer (1946), it has been understood that a basic technique of the Taylor series expansion suffices for asymptotics of M-estimators with smooth/differentiable loss function. Although the Taylor series expansion is a purely deterministic …
We propose an AdaPtive Noise Augmentation (PANDA) technique to regularize the estimation and construction of undirected graphical models. PANDA iteratively optimizes the objective function given the noise augmented data until convergence to achieve regularization on model parameters. The augmented noises can be designe…
A transfer learning method builds high-dimensional models using disparate datasets.
problem Building comprehensive prediction models with small sample sizes and limited features.
method Transfer learning approach using external data to build a reduced model and apply calibration equations.
result Proposes a penalized generalized method of moment framework for inference and one-step estimation.
We study the convergence of the predictive surface of regression trees and forests. To support our analysis we introduce a notion of adaptive concentration for regression trees. This approach breaks tree training into a model selection phase in which we pick the tree splits, followed by a model fitting phase where we f…
Develops a new framework for causal models on cyclic graphs, solving unique solvability issues.
problem Challenges in specifying unique probability distributions for cyclic functional causal models.
method Introduces a new probability rule and graph-separation property (p-separation) for cyclic fCMs.
result Proves p-separation is sound and complete for all consistent cyclic fCMs, recovering d-separation for DAGs.
CAT method learns causal structure of directed trees efficiently.
problem Learning causal structure from directed trees.
method Chu-Liu-Edmonds algorithm for fast and scalable structure learning.
result Consistency in asymptotic regime with vanishing identifiability gap for Gaussian errors.
Proposes a method to compute valid lower confidence bounds for multiple models selected based on their performance.
problem Model selection and evaluation in machine learning.
method Interprets model selection as a simultaneous inference problem, uses bootstrap tilting and maxT-type multiplicity correction.
result Yields valid lower confidence bounds that are at least as good as standard approaches and reliably reach nominal coverage probability.
CAP algorithm controls FCR in online selective prediction.
problem Online predictive tasks with temporal multiplicity and FCR control.
method CAP framework with adaptive pick rule and calibration set construction.
result CAP achieves exact selection-conditional coverage guarantee and FCR control.
Bayesian approach selects subsets of variables for interpretable prediction and identifies key factors in educational outcomes.
problem Challenges in subset selection for stability, regularization, and inference.
method Bayesian perspective on subset selection, deriving optimal subsets and variable importance metrics.
result Better prediction, interval estimation, and variable selection compared to competing methods.
Novel causal effect estimators and distributionally robust prediction methods.
problem Estimating causal effects and distributional robustness in statistical models.
method Developed novel estimators and proposed a general framework for distributional robustness.
result Mean squared error improvements in causal effect estimation compared to existing methods.
Estimates true Sharpe ratio of selected assets with various methods.
problem Estimating the true Sharpe ratio of a selected asset with high in-sample ratio.
method Polyhedral lemma, James Stein shrinkage, debiasing, thresholding, empirical Bayes.
result James Stein estimator performs best across various parameter values.
Solar algorithm selects variables faster and more accurately in high-dimensional data.
problem Variable selection in high-dimensional data with high accuracy and stability.
method Subsample-ordered least-angle regression (solar) and its coordinate descent generalization (solar-cd) using L0 norm solution path averaging. result Solar selects variables with high accuracy and stability, reducing redundant variable selection.
Paper proposes a new method combining random forests and Lasso selection.
problem Improving random forest performance by applying Lasso regression.
method Adaptive Lasso weighting applied to random forest predictions.
result Unified framework strictly outperforms other methods in simulations and real-world datasets.
DL/FBF improves GPSR solutions by selecting compact, generalising expressions.
problem Overfitting and structural bloat in symbolic regression with genetic programming.
method Description length (DL) and fractional Bayes factor (FBF) criteria for selecting compact, generalising expressions.
result DL/FBF post-selection improves test performance compared to AIC/BIC baseline.
Proposes DR-ME test for interpretable distributional treatment effects.
problem Detects invisible differences in treatment effects on distributional outcomes.
method Semiparametrically efficient finite-location test using kernel witnesses and orthogonal features.
result DR-ME reveals causal-discrepancy coordinates and has noncentral chi-square local power.