ARK improves knockoffs robustness to feature distribution misspecification.
problem Robustness of knockoffs inference to misspecified feature distributions.
method Coupling approximate knockoffs with model-X knockoffs to achieve FDR and FWER control.
result The approximate knockoffs procedure can control FDR and FWER asymptotically.
Extends model-x framework to handle missing data.
problem Inability to control false selections in missing data settings.
method Posterior sampled imputation, univariate imputation, joint imputation and sampling knockoffs.
result Preserves theoretical guarantees of model-x framework in missing data setting.
Unified framework for FDR control in knockoffs, validating Gaussian knockoffs.
problem Asymptotic FDR control in knockoffs with user-specified distributions.
method Unified theoretical framework, three conditions on approximate knockoff statistics, Gaussian knockoffs generator based on moments matching.
result Gaussian knockoffs generator achieves asymptotic FDR control.
GRASP tests goodness-of-fit for binary classifiers without parametric assumptions.
problem Assessing the fit of a binary classifier to the underlying conditional law of labels given features.
method Formulates a tolerance hypothesis testing problem and proposes a novel test called GRASP.
result Proposes GRASP and Model-X GRASP tests for assessing goodness-of-fit in finite sample settings.
Model-X test detects conditional independence in streaming data.
problem Detecting conditional independence in data streams with arbitrary dependency.
method Sequential testing inspired by model-X and testing by betting.
result Significantly reduces type-I error rate and enhances data efficiency.
Paper proposes a privacy-preserving knockoff inference method.
problem Ensuring privacy in model-X knockoff inference.
method Differential privacy framework for knockoff inference.
result Guaranteed FDR control with privacy protection.
Sequential Kernel-based Conditional Independence Testing via Adaptive Betting
problem Testing conditional independence
method Testing-by-betting on an adaptively optimized Kernel Conditional Independence statistic
result Significantly reduces Type I error inflation while preserving high power
Conditional modeling x \to y is a central problem in machine learning. A substantial research effort is devoted to such modeling when x is high dimensional. We consider, instead, the case of a high dimensional y, where x is either low dimensional or high dimensional. Our approach is based on selecting a small subset y_…
This paper introduces a machine for sampling approximate model-X knockoffs for arbitrary and unspecified data distributions using deep generative models. The main idea is to iteratively refine a knockoff sampling mechanism until a criterion measuring the validity of the produced knockoffs is optimized; this criterion i…
Paper improves power of conditional randomization tests.
problem Improving power of conditional randomization tests.
method Introducing a new cost function to maximize test statistic power.
result Consistently increases the number of correct discoveries.
A new efficient test addresses limitations of knockoffs for conditional independence testing.
problem Testing conditional independence under model-X assumptions.
method Leave-One-Covariate-Out Conditional Randomization Test (LOCO-CRT)
result LOCO-CRT produces valid p-values for familywise error rate control with minimal variability. Novel privatization framework for high-dimensional variable selection with differential privacy.
problem High-dimensional controlled variable selection with rigorous FDR control under differential privacy constraints.
method Gaussian Johnson-Lindenstrauss Transformation for privatizing the knockoff matrix.
result The proposed private variable selection procedure maintains statistical power even under strict privacy budgets.
The paper analyzes the power of MX CI tests and finds likelihood-based statistics most powerful.
problem Testing conditional independence under model-X assumptions.
method Conditional randomization test (CRT) and MX knockoffs.
result Likelihood-based statistics are most powerful in MX CI tests.
We present an Automatic Relevance Determination prior Bayesian Neural Network(BNN-ARD) weight l2-norm measure as a feature importance statistic for the model-x knockoff filter. We show on both simulated data and the Norwegian wind farm dataset that the proposed feature importance statistic yields statistically signific…
The Kähler-Ricci flow yields bounded diameter and Ricci curvature for minimal models.
problem Estimating the diameter and Ricci curvature of long-time solutions of the Kähler-Ricci flow.
method Analyzing the semi-ample canonical line bundle and using Perelman's estimates.
result Uniform bounds on diameter and Ricci curvature for long-time solutions.
Diamond method controls FDR for trustworthy feature interaction discovery in ML models.
problem Limited interpretability of ML models due to black box nature.
method Diamond method integrates model-X knockoffs framework to control FDR for non-additive interactions.
result Diamond method ensures accurate discovery of feature interactions with FDR control.
Efficient knockoffs for large-scale feature selection.
problem Large-scale feature selection problems.
method Gaussian model-X knockoffs with efficient methods for solving semidefinite programs.
result Efficient knockoffs can be generated with linear complexity in the dimension.
Study a relative aspherical conjecture and prove 3-manifold obstruction to positive scalar curvature.
problem Obstructing the existence of positive scalar curvature in higher dimensions.
method Introduced a relative aspherical condition and a new geometric quantity called spherical width.
result Proved results on how 3-manifolds obstruct the existence of positive scalar curvature.
Private CI tests for continuous Z with privacy constraints.
problem Testing conditional independence under differential privacy constraints.
method Developed two private CI testing procedures based on generalized covariance and conditional randomization tests.
result First private CI tests with rigorous theoretical guarantees for continuous Z.
Inferring the causal structure of a set of random variables from a finite sample of the joint distribution is an important problem in science. Recently, methods using additive noise models have been suggested to approach the case of continuous variables. In many situations, however, the variables of interest are discre…
New model reduces matrix factorization bias, yielding truly low-rank solutions.
problem Gradient descent's implicit bias in matrix factorization.
method Introducing a new factorization model with constrained factors and diagonal components.
result The new model consistently exhibits a strong implicit bias, yielding truly low-rank solutions.
The Model-X knockoff procedure has recently emerged as a powerful approach for feature selection with statistical guarantees. The advantage of knockoff is that if we have a good model of the features X, then we can identify salient features without knowing anything about how the outcome Y depends on X. An important dra…
Improved FDR control for sparse financial index tracking.
problem Maintaining FDR control in high-dimensional financial data with strong variable dependencies.
method Expanding T-Rex framework to handle overlapping groups of correlated variables with nearest neighbors penalization.
result Accurately tracks the S&P 500 index using only a small number of stocks.
We prove a uniform diameter bound for long time solutions of the normalized Kahler-Ricci flow on an n-dimensional projective manifold X with semi-ample canonical bundle under the assumption that the Ricci curvature is uniformly bounded for all time in a fixed domain containing a fibre of X over its canonical mode…
FlowSelect uses normalizing flows to control FDR in feature selection.
problem Controlled feature selection with knockoffs often fails to control false discovery rate (FDR).
method FlowSelect uses normalizing flows for accurate feature modeling and a novel MCMC-based p-value calculation to enforce knockoff properties.
result FlowSelect consistently controls FDR and demonstrates greater power compared to competing methods.
Sharp bounds on K-semistable Fano varieties for low dimensions.
problem Establishing bounds on the height of K-semistable Fano varieties.
method Analyzing canonical integral models of toric Fano varieties, using the gap hypothesis and Donaldson's modular height.
result Sharp lower bounds on the height of toric Fano varieties, with applications to Mabuchi functional and Odaka's modular height.
New algorithm achieves optimal clustering for sparse centers with high dimensions.
problem Statistical and computational limits of clustering sparse centers with high dimensions.
method Sparse clustering algorithm based on sparse PCA.
result Achieves minimax optimal misclustering rate under certain conditions.
Consistent model selection for spiked Wigner model via AIC-type criteria.
problem Estimating the number of spiked eigenvalues in the spiked Wigner model.
method AIC-type model selection criteria with parameters γ.
result Strong consistency for γ > 2 and weak consistency for γ = 2 + δ_N.
The paper analyzes ridge regression with random features for non-identically distributed data.
problem Analyzing ridge regression performance for data with heterogeneous variance profiles.
method Combining linear-plus-chaos approximation and operator-valued free probability.
result Derives asymptotic equivalents for training and test risks under non-identically distributed data.
DDLK uses deep learning to find important features in models.
problem Discovering important features in black box models like deep neural networks.
method DDLK directly minimizes KL divergence to generate knockoffs that obey the swap property.
result DDLK outperforms baselines in discovering important features while controlling false discovery rate.
Novel method separates astrophysical components from noisy data.
problem Separating and reconstructing astrophysical components from noisy data.
method Latent-space field tension for automated component separation.
result High accuracy in reconstructing astrophysical components.
Interpretability and stability are two important features that are desired in many contemporary big data applications arising in economics and finance. While the former is enjoyed to some extent by many existing forecasting approaches, the latter in the sense of controlling the fraction of wrongly discovered features w…
DRCD identifies causal direction between continuous and discrete variables using density ratio monotonicity.
problem Inferring causal direction between continuous and discrete variables from observational data.
method Density Ratio-based Causal Discovery (DRCD) method.
result DRCD identifies causal direction between continuous and discrete variables using density ratio monotonicity.
This work introduces a novel estimation method, called LOVE, of the entries and structure of a loading matrix A in a sparse latent factor model X = AZ + E, for an observable random vector X in Rp, with correlated unobservable factors Z \in RK, with K unknown, and independent noise E. Each row of A is scaled and sparse.…