Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,982 papers · 148 categories

Trend · papers per month

6481,2971,9452,593 · Jun 202019922001200920172026
48 results for permutations of variables

CPI overcomes limitations of permutation importance by providing accurate variable selection.

problem Misidentification of unimportant variables in complex models due to covariate correlations.
method Developed a model agnostic and computationally lean Conditional Permutation Importance (CPI) approach.
result CPI provides accurate type-I error control and more parsimonious variable selection.

Random Forest permutation importance measure is asymptotically unbiased in sparse regression models.

problem Challenges in selecting informative variables in high-dimensional regression problems.
method Theoretical guarantees and asymptotic unbiasedness of permutation importance measure under specific assumptions.
result Permutation importance measure in Random Forest is asymptotically unbiased.

New model preserves symmetry in multivariate time series, improving performance.

problem Implicit ordering in MTS models violates inherent exchangeability.
method Permutation-equivariant 2D state space model with canonical architecture.
result Eliminates sequential dependency chains and simplifies stability analysis.

A new method reduces computational costs for testing RF variable importance measures.

problem Testing variable importance measures from random forests is computationally expensive and challenging.
method Sequential permutation testing and sequential p-value estimation to reduce computational costs.
result Theoretical properties of sequential tests are confirmed, maintaining type-I error and high power.

The paper derives theoretical foundations for two common machine learning variable importance measures.

problem Understanding variable importance in machine learning problems.
method The paper derives closed-form expressions for Permute-and-Predict (PaP) and Leave-One-Covariate-Out (LOCO) methods.
result Theoretical derivations explain the behavior of PaP and LOCO under collinearity, linking them to coefficients and predictor variability.

We prove that certain dual pairs of Calabi-Yau manifolds have matching orbifold Euler characteristics.

problem Constructing mirror symmetric Calabi-Yau manifolds with specific symmetry groups.
method Generalized Berglund-Hübsch-Henningson construction to include permutations of variables.
result Reduced orbifold Euler characteristics of dual pairs coincide up to sign under cyclic permutation groups satisfying parity condition.

A new method improves feature importance and model stress-testing reliability.

problem Estimating feature contributions in machine learning models for trust and transparency.
method Replacing multiple random permutations with a single, deterministic, and optimal permutation.
result Improved bias-variance tradeoffs and accuracy in challenging scenarios.

New methods reduce extrapolation errors in feature importance.

problem Flawed feature importance methods using unrestricted permutations lead to extrapolation errors.
method Three new approaches: conditional model reliance, Knockoffs with Gaussian transformation, and restricted ALE plot designs.
result Theoretical and numerical results show our strategies reduce/eliminate extrapolation.

Prob-PIT improves speech separation by considering output-label permutations as random variables.

problem Overconfident output-label assignment in PIT leads to unreliable speech separation.
method Prob-PIT treats output-label permutations as a discrete latent random variable with a uniform prior distribution and maximizes the log-likelihood function.
result Prob-PIT significantly outperforms PIT in terms of Signal to Distortion Ratio and Signal to Interference Ratio.

A new nonparametric test measures dependence between variables using decision trees.

problem Measuring statistical dependence between two variables robustly and efficiently.
method An ensemble of decision trees discriminates between observed and permuted samples without generating the latter.
result The method effectively detects complex relationships from noisy data.

This work improves multi-modal generative models by using permutation-invariant neural networks.

problem Improving multi-modal generative models with tighter variational objectives.
method Developed more flexible aggregation schemes based on permutation-invariant neural networks.
result Our variational objective and flexible aggregation models can better approximate the true joint distribution.

In regression analysis of multivariate data, it is tacitly assumed that response and predictor variables in each observed response-predictor pair correspond to the same entity or unit. In this paper, we consider the situation of "permuted data" in which this basic correspondence has been lost. Several recent papers hav…

2017-10-16abs ↗pdf ↗

PIVID infers DAG structures from data using variational inference and permutations.

problem Estimating the structure of Bayesian networks from observational data.
method PIVID uses variational inference and continuous relaxations of discrete distributions to infer a distribution over permutations and DAGs.
result PIVID outperforms deterministic and Bayesian approaches in estimating DAG structures from data.

To model categorical response variables given their covariates, we propose a permuted and augmented stick-breaking (paSB) construction that one-to-one maps the observed categories to randomly permuted latent sticks. This new construction transforms multinomial regression into regression analysis of stick-specific binar…

2016-12-30abs ↗pdf ↗

Paper proves a Central Limit Theorem for Random Forest Permutation Importance Measure.

problem Lack of theoretical analysis of Random Forest Permutation Importance Measure (RFPIM).
method Formal proof using U-Statistics theory, deviating from conventional Random Forest model.
result Established a Central Limit Theorem for RFPIM.

VarPro selects features without model dependence, achieving balanced performance.

problem Finding a small set of features with high explanatory power.
method Rule-based variable priority approach, avoiding model-specific methods and artificial data.
result VarPro has a consistent filtering property for noise variables and achieves balanced performance.

The introduction of convolutional layers greatly advanced the performance of neural networks on image tasks due to innately capturing a way of encoding and learning translation-invariant operations, matching one of the underlying symmetries of the image domain. In comparison, there are a number of problems in which the…

2016-12-14abs ↗pdf ↗

Study examines challenges in variable importance ranking due to feature correlation.

problem Challenges in variable importance ranking under correlation.
method Simulation study and theoretical analysis of feature knockoffs and conditional predictive impact (CPI).
result Highly correlated features increase the correlation of knockoff variables, posing a limitation for CPI.

New method identifies latent causal variables from observed data, overcoming indeterminacies.

problem Identifying latent causal variables from observed data, especially when latent variables are weight-variant.
method Introduces a novel identifiability condition for latent causal models, proposing SuaVE method.
result Identifies latent causal variables up to trivial permutation and scaling, demonstrating consistency and efficacy.

A new method handles mismatched data in multivariate regression.

problem Handling mismatched data in multivariate linear regression.
method Two-stage approach: first stage estimates parameters, second stage estimates permutation.
result Permutation recovery conditions become less stringent with increasing number of responses.

In literature there are several studies on the performance of Bayesian network structure learning algorithms. The focus of these studies is almost always the heuristics the learning algorithms are based on, i.e. the maximisation algorithms (in score-based algorithms) or the techniques for learning the dependencies of e…

2011-01-27abs ↗pdf ↗

Databases in domains such as healthcare are routinely released to the public in aggregated form. Unfortunately, naive modeling with aggregated data may significantly diminish the accuracy of inferences at the individual level. This paper addresses the scenario where features are provided at the individual level, but th…

2016-05-14abs ↗pdf ↗

We investigate the problem of testing whether dd random variables, which may or may not be continuous, are jointly (or mutually) independent. Our method builds on ideas of the two variable Hilbert-Schmidt independence criterion (HSIC) but allows for an arbitrary number of variables. We embed the dd-dimensional joint …

2016-03-01abs ↗pdf ↗

This paper tackles permutation recovery in unlabeled sensing from multiple measurement vectors.

problem Permutation recovery in unlabeled sensing from multiple measurement vectors.
method The paper studies the case of multiple noisy measurement vectors (MMVs) resulting from a common permutation and proposes computational schemes for permutation recovery.
result A large stable rank of the signal significantly reduces the required signal-to-noise ratio (SNR) for permutation recovery, and the problem can be solved efficiently using ADMM.

Paper proposes a differentially private test for joint dependence among random vectors.

problem Detecting joint dependence among sensitive data while maintaining privacy.
method Differentially private permutation methodology for dHSIC test.
result Proposed test attains minimax optimal power across privacy regimes.

A new metric assesses causal graphs using node permutations to detect inconsistencies.

problem Quantifying the goodness of causal graphs and distinguishing them from random graphs.
method Constructing a baseline through node permutations and comparing inconsistencies.
result The proposed metric can distinguish between true and wrong causal graphs.

New method for regression in high-dimensional space using mixture modeling and optimal transport.

problem Regression in high-dimensional space with unordered data.
method Mixture modeling and optimal transport for permutation recovery and denoising.
result Explicit upper bounds on mean squared denoising error for Gaussian noise.

Paper proposes a chi-square test for distance correlation.

problem Testing distance correlation is computationally expensive.
method Proposes a chi-square test for distance correlation, non-parametric, fast, applicable to various metrics.
result Chi-square test exhibits similar power to permutation test and can be valid and universally consistent for testing independence.