Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Feb 199419922001200920182026
48 results for Statistical features

Paper proposes a statistical test for feature selection pipelines using selective inference.

problem Assessing the significance of feature selection pipelines in data analysis.
method Selective inference technique applied to feature selection pipelines composed of various algorithms.
result The proposed statistical test controls false positive feature selection probabilities.

New statistics improve feature importance detection with false discovery guarantees.

problem Identifying truly correlated features from observational data.
method Developed efficient knockoff generation from Bayesian Networks and new statistics.
result Improved power and efficiency of feature importance detection.

The paper introduces a statistical distance matrix for better feature representation and clustering.

problem Lack of detailed distance representation between feature elements.
method Extended traditional statistical distance to a matrix form (statistical distance matrix) and applied hierarchical clustering.
result The statistical distance matrix with clustering (Information Mandala) provides clearer and geometrically arranged feature representations.

Laplace kernel feature selection offers statistical guarantees for nonparametric models with few samples.

problem Statistical guarantees for kernel-based feature selection in nonconvex optimization problems.
method Sharp characterization of the gradient of the objective function for Laplace kernel feature selection.
result Model-selection consistency for Laplace kernel-based feature selection in nonparametric settings with nlogpn \sim \log p samples.

The paper develops a method to identify conditionally relevant features with statistical guarantees.

problem Identifying features that are relevant given the values of other features.
method A generalization of the knockoff procedure that controls a generalized FDR for conditional feature selection.
result The method provides a statistical guarantee for conditional feature selection.

Paper proposes Gini distance statistics for estimating feature-label dependence.

problem Identifying statistical dependence between features and categorical labels.
method Generalized Gini distance in RKHS for feature-label dependence estimation.
result Gini distance statistics converge faster and have tighter error bounds than distance covariance.

Bayesian neural network improves feature selection and prediction.

problem Improving feature selection and prediction accuracy in neural networks.
method BNN-ARD with l2-norm feature importance measure.
result Improves variable selection and predictive performance on real-world data.

New algorithm finds significant interactions between continuous features.

problem Finding statistically significant interactions between continuous features is challenging due to combinatorial explosion.
method Proposes an algorithm that derives a lower bound on p-values for each interaction, pruning non-significant interactions.
result Efficiently detects all significant interactions in synthetic and real-world datasets.

Paper introduces PTL-SI for statistical inference in TL-HDR, controlling FPR.

problem Quantifying statistical significance in TL-HDR with limited data.
method PTL-SI framework for valid pp-values in TL-HDR feature selection.
result Valid pp-values and controlled FPR in TL-HDR feature selection.

A new framework learns system design using neural features in function space.

problem Learning system design with neural feature extractors.
method Introduces feature geometry in function space, nesting technique for optimal feature approximation.
result Optimal features found from data samples using off-the-shelf architectures and optimizers.

We develop a theory of higher-order feature attribution for complex models.

problem Interpreting feature contributions in models with interactions is challenging.
method We extend Integrated Gradients (IG) to higher-order feature attributions.
result We establish natural connections to statistics and topological signal processing.

Finding statistically significant high-order interaction features in predictive modeling is important but challenging task. The difficulty lies in the fact that, for a recent applications with high-dimensional covariates, the number of possible high-order interaction features would be extremely large. Identifying stati…

2015-06-26abs ↗pdf ↗

The study explores statistical methods to interpret radiological models and identify key features.

problem Interpreting complex radiological models for clinical use.
method Exploration of statistical techniques to assess relationships between radiomic features.
result Identification of key relationships and features for improved interpretability.

Two statistical tasks are shown to have equivalent sample complexity.

problem Determining if a function depends on only a few variables and identifying those variables.
method Proved statistical equivalence of feature selection and junta testing through sample complexity analysis.
result Brute-force algorithm is sample-optimal for both tasks with optimal sample size.

Spofe bridges statistical rigor and interpretability in feature extraction from tabular data.

problem Ensuring statistical rigor and interpretability in feature extraction from complex tabular data.
method Spofe combines kernel principal components and sparse polynomial functions with a multi-objective knockoff selection procedure.
result Spofe consistently outperforms other methods in feature selection for regression and classification tasks.

Paper improves understanding of random Fourier features for kernel ridge regression.

problem Understanding statistical properties of random Fourier features for kernel ridge regression.
method Spectral matrix approximation approach to analyze random Fourier features.
result Proves statistical guarantees for kernel ridge regression using random Fourier features.

This work analyzes tree-based methods from a ranking perspective, providing insights and new statistics.

problem Understanding the effectiveness of tree-based methods in finite-sample settings, especially symbolic feature selection.
method Local ranking perspective, finite-sample analysis, oracle bounds, posterior contraction results, concordant divergence statistics.
result New insights and statistics for evaluating symbolic feature mappings.

A novel feature selection method using noise-based hypothesis testing improves feature selection accuracy.

problem Challenges in feature selection for complex, high-dimensional datasets.
method Introduces multiple random noise features and evaluates feature importance against noise feature maxima using non-parametric bootstrap-based hypothesis testing.
result Outperforms existing methods in simulated and real-world datasets.

Paper extracts features from time series to improve forecasting accuracy.

problem Forecasting time series generated by Itô-type processes with unknown coefficients.
method Statistical adjustment of mixture-type models to extract features from time series data.
result Additional statistical features enhance time series prediction accuracy.

Proposes φφ-table for statistical SHAP explanations in regression models.

problem Lack of clear directional summaries, uncertainty, and fidelity in SHAP feature importance.
method SHAP importance selection, fitting a standardized linear surrogate, reporting coefficients, uncertainty, fidelity, and stability.
result Extends SHAP into a statistical global explanation with direction, uncertainty, fidelity, and stability.

Paper tackles class-incremental time series classification with dual-stream feature extraction.

problem Class-incremental continual learning for multivariate time series data.
method Dual-stream feature extraction pipeline combining deep temporal embedding features and statistical features.
result Competitive average accuracy across multiple datasets with low forgetting rates.

A new method uses bandits to select summary statistics for Bayesian inference.

problem Dynamic selection of summary statistics for likelihood-free inference.
method Treats summary statistic selection as a multi-armed bandit problem.
result Improves efficiency and scalability of approximate Bayesian computation.

New method predicts protein features using statistical relational learning.

problem Challenges in automatic protein feature annotation due to limited homology data.
method Introduces Semantically Based Regularization to incorporate prior knowledge.
result Improved overall prediction quality with constraints.

SES finds multiple feature subsets with similar predictive power.

problem Finding a single optimal feature subset for predictive accuracy.
method SES algorithm identifies multiple statistically equivalent feature subsets.
result Multiple feature subsets can achieve close to maximal predictive accuracy.

Proposes a PSI framework for feature selection in divergences.

problem Measuring divergence between two distributions in high-dimensional data.
method Additive MMD estimator with incomplete U-statistics for feature selection.
result Successfully detects statistically significant features in synthetic and real-world data.

New method learns domain-invariant local feature patterns for unsupervised domain adaptation.

problem Performance degradation due to domain-shift in unsupervised domain adaptation.
method Jointly learns domain-invariant local feature patterns and holistic feature distributions.
result Superior performance on benchmark datasets compared to state-of-the-art methods.

Unified framework for rare feature selection and aggregation in high-dimensional statistics.

problem Challenges in high-dimensional statistics due to rare features.
method Developed a unified computational framework for norms promoting discrete structures, using orthogonal projection oracle.
result Proposed estimation procedure for automatic feature selection and aggregation with statistical bounds.

Adaptive feature normalization improves model robustness to extraneous variables.

problem Degrading model performance due to extraneous variables in deep learning.
method Adaptive feature normalization using instance normalization instead of batch normalization.
result Adaptive normalization leads to significant performance gains across different datasets and architectures.

Detects which features have shifted in data distributions.

problem Identifying which specific features have caused a distribution shift.
method Formalizes the problem as multiple conditional distribution hypothesis tests, proposes non-parametric and parametric statistical tests, and uses a test statistic based on the density model score function.
result Demonstrates methods for identifying when and where a shift occurs in multivariate time-series data.

We study the relationship between social media output and National Football League (NFL) games, using a dataset containing messages from Twitter and NFL game statistics. Specifically, we consider tweets pertaining to specific teams and games in the NFL season and use them alongside statistical game data to build predic…

2013-10-25abs ↗pdf ↗

DiffKnock improves feature selection in neural networks with complex dependencies and non-linear associations.

problem Selecting important features in neural networks with complex dependencies and non-linear associations.
method DiffKnock uses diffusion models to generate knockoffs and neural network statistics to measure feature importance.
result DiffKnock outperforms existing methods in detecting non-linear associations and preserving feature dependencies.

We study the generalization properties of ridge regression with random features in the statistical learning framework. We show for the first time that O(1/n)O(1/\sqrt{n}) learning bounds can be achieved with only O(nlogn)O(\sqrt{n}\log n) random features rather than O(n)O({n}) as suggested by previous results. Further, we prove fa…

2016-02-14abs ↗pdf ↗

The paper proposes a method to test features selected by SeqFS-DA with controlled FPR.

problem Ensuring reliability of feature selection after domain adaptation in high-dimensional regression.
method Proposes a novel method to test features selected by SeqFS-DA with controlled FPR.
result The proposed method controls FPR below a significance level αα (e.g., 0.05) and enhances statistical power.

New statistical methods improve explainability of boosting models.

problem Uncertainty quantification for boosting models is computationally intensive and hard to interpret.
method Derive methods for statistical inference using gradient boosting and Boulevard regularization.
result Achieve asymptotically normal predictions with theoretical guarantees and runtime independent of data size.

Paper classifies plant electrical signals to identify external stimuli.

problem Classifying external stimuli using plant electrical response.
method Computed 11 statistical features from plant electrical signals and used discriminant analysis.
result Raw electrical signals contain enough information for stimulus classification.

We analyze random feature maps for high-dimensional data using spectral methods.

problem Understanding the spectrum of random feature maps for high-dimensional data.
method We use concentration phenomena from random matrix theory to analyze the Gram matrix of random feature maps for Gaussian mixture models.
result Our results provide insights into the interplay between nonlinearity and data statistics.