Develops a new method for nonlinear dimension reduction using random features.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We show that the Alexander-Conway polynomial Delta is obtainable via a particular one-variable reduction of each two-variable Links-Gould invariant LG^{m,1}, where m is a positive integer. Thus there exist infinitely many two-variable generalisations of Delta. This result is not obvious since in the reduction, the repr…
Direct contextual policy search methods learn to improve policy parameters and simultaneously generalize these parameters to different context or task variables. However, learning from high-dimensional context variables, such as camera images, is still a prominent problem in many real-world tasks. A naive application o…
Variable selection and dimension reduction are two commonly adopted approaches for high-dimensional data analysis, but have traditionally been treated separately. Here we propose an integrated approach, called sparse gradient learning (SGL), for variable selection and dimension reduction via learning the gradients of t…
New method reduces high-dimensional data to key features.
A/B testing improves marketing decisions by selecting effective stratification variables.
Proposes an online method for high-dimensional streaming data.
Bayesian neural networks improve uncertainty quantification in non-linear dimensionality reduction.
Bayesian optimization reduces hyperparameters for mixed variable design problems.
CIR method preserves relation for case-control studies.
Proposes a simple solution to Gini importance bias in random forests.
A typical goal of supervised dimension reduction is to find a low-dimensional subspace of the input space such that the projected input variables preserve maximal information about the output variables. The dependence maximization approach solves the supervised dimension reduction problem through maximizing a statistic…
We show a connection between the Fourier spectrum of Boolean functions and the REINFORCE gradient estimator for binary latent variable models. We show that REINFORCE estimates (up to a factor) the degree-1 Fourier coefficients of a Boolean function. Using this connection we offer a new perspective on variance reduction…
Two novel methods estimate multiple FDR directions for binary categorical responses.
Paper improves Gumbel-Softmax estimator variance reduction.
Discrete random variables are natural components of probabilistic clustering models. A number of VAE variants with discrete latent variables have been developed. Training such methods requires marginalizing over the discrete latent variables, causing training time complexity to be linear in the number clusters. By appl…
A new method reduces variance in training discrete latent variable models.
The un-reduction procedure introduced previously in the context of Mechanics is extended to covariant Field Theory. The new covariant un-reduction procedure is applied to the problem of shape matching of images which depend on more than one independent variable (for instance, time and an additional labelling parameter)…
LMMVAE improves VAE for correlated data by separating latent variables into fixed and random parts.
Proposes flexible auto-encoders for varying data dimensions.
A novel MM algorithm optimizes DCOV for SDR and SVS.
Paper reduces turbomachinery CFD simulations by identifying key dimensions.
To address the challenge of backpropagating the gradient through categorical variables, we propose the augment-REINFORCE-swap-merge (ARSM) gradient estimator that is unbiased and has low variance. ARSM first uses variable augmentation, REINFORCE, and Rao-Blackwellization to re-express the gradient as an expectation und…
Dimensionality reduction is ubiquitous in analysis of complex dynamics. The conventional dimensionality reduction techniques, however, focus on reproducing the underlying configuration space, rather than the dynamics itself. The constructed low-dimensional space does not provide complete and accurate description of the…
Paper automates variable selection for network anomaly detection.
Proposes variance reduction for optimizing permutation models.
Dimensionality reduction is one of the key issues in the design of effective machine learning methods for automatic induction. In this work, we introduce recursive maxima hunting (RMH) for variable selection in classification problems with functional data. In this context, variable selection techniques are especially a…
New method reduces bias in estimating causal effects from discretized variables.
Decision trees with binary splits are popularly constructed using Classification and Regression Trees (CART) methodology. For binary classification and regression models, this approach recursively divides the data into two near-homogenous daughter nodes according to a split point that maximizes the reduction in sum of …
The process of un-reduction, a sort of reversal of reduction by the Lie group symmetries of a variational problem, is explored in the setting of field theories. This process is applied to the problem of curve matching in the plane, when the curves depend on more than one independent variable. This situation occurs in a…
Model reduction methods aim to describe complex dynamic phenomena using only relevant dynamical variables, decreasing computational cost, and potentially highlighting key dynamical mechanisms. In the absence of special dynamical features such as scale separation or symmetries, the time evolution of these variables typi…
New method estimates Gaussian vector functions more efficiently.
Novel technique reduces Bayesian network complexity while preserving inference accuracy.
PSMM method optimizes matrix sufficient dimension reduction.
We introduce a multiscale supervised dimension reduction method for SPatial Interaction Network (SPIN) data, which consist of a collection of spatially coordinated interactions. This type of predictor arises when the sampling unit of data is composed of a collection of primitive variables, each of them being essentiall…
A new framework for bilevel optimization tackles stochastic and global variance reduction.
Super learner uses diverse screeners to improve prediction performance.
Study on reducing dimensionality in high-dimensional regression with kernel methods and stability analysis.
This paper proposes a novel kernel approach to linear dimension reduction for supervised learning. The purpose of the dimension reduction is to find directions in the input space to explain the output as effectively as possible. The proposed method uses an estimator for the gradient of regression function, based on the…
The current study proposes a dimension reduction method, stepwise support vector machine (SVM), to reduce the dimensions of large p small n datasets. The proposed method is compared with other dimension reduction methods, namely, the Pearson product difference correlation coefficient (PCCs), recursive feature eliminati…
EBM reduces dimensionality for estimating heterogeneous CATEs.
AI-generated variables bias regression estimates; methods correct for invalid inference.
Experimental life sciences like biology or chemistry have seen in the recent decades an explosion of the data available from experiments. Laboratory instruments become more and more complex and report hundreds or thousands measurements for a single experiment and therefore the statistical methods face challenging tasks…
We consider the problem of high-dimensional classification between the two groups with unequal covariance matrices. Rather than estimating the full quadratic discriminant rule, we propose to perform simultaneous variable selection and linear dimension reduction on original data, with the subsequent application of quadr…
Policy evaluation is a crucial step in many reinforcement-learning procedures, which estimates a value function that predicts states' long-term value under a given policy. In this paper, we focus on policy evaluation with linear function approximation over a fixed dataset. We first transform the empirical policy evalua…
Inspired by the success of deep learning techniques in the physical and chemical sciences, we apply a modification of an autoencoder type deep neural network to the task of dimension reduction of molecular dynamics data. We can show that our time-lagged autoencoder reliably finds low-dimensional embeddings for high-dim…
New sampling strategy preserves relationships in multivariate scientific data.
A restricted Boltzmann machine (RBM) is a two-layer neural network with shared weights and has been extensively studied for dimensionality reduction, data representation and recommendation systems in the literature. The traditional RBM requires a probabilistic interpretation of the values on both layers and a Markov ch…