The paper constructs a complex for the Dirac operator in 4 dimensions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Variable selection and dimension reduction are two commonly adopted approaches for high-dimensional data analysis, but have traditionally been treated separately. Here we propose an integrated approach, called sparse gradient learning (SGL), for variable selection and dimension reduction via learning the gradients of t…
Develops a new method for nonlinear dimension reduction using random features.
Proposes an online method for high-dimensional streaming data.
We extend the Bayesian Information Criterion (BIC), an asymptotic approximation for the marginal likelihood, to Bayesian networks with hidden variables. This approximation can be used to select models given large samples of data. The standard BIC as well as our extension punishes the complexity of a model according to …
Paper introduces new bounds linking data compressibility to generalization error.
Many 0/1 datasets have a very large number of variables; on the other hand, they are sparse and the dependency structure of the variables is simpler than the number of variables would suggest. Defining the effective dimensionality of such a dataset is a nontrivial problem. We consider the problem of defining a robust m…
Random feature matrices' singular values concentrate near their full expectation in high dimensions.
Bayesian non-parametric model selects latent dimensions automatically.
We study a method of reducing space dimension in multi-dimensional Black-Scholes partial differential equations as well as in multi-dimensional parabolic equations. We prove that a multiplicative transformation of space variables in the Black-Scholes partial differential equation reserves the form of Black-Scholes part…
A new method detects unknown classes and adapts to extra dimensions in high-dimensional classification.
In this study, we propose an automatic learning method for variables selection based on Lasso in epidemiology context. One of the aim of this approach is to overcome the pretreatment of experts in medicine and epidemiology on collected data. These pretreatment consist in recoding some variables and to choose some inter…
Model complexity is an important factor to consider when selecting among graphical models. When all variables are observed, the complexity of a model can be measured by its standard dimension, i.e. the number of independent parameters. When hidden variables are present, however, standard dimension might no longer be ap…
This paper proposes a general adaptive procedure for budget-limited predictor design in high dimensions called two-stage Sampling, Prediction and Adaptive Regression via Correlation Screening (SPARCS). SPARCS can be applied to high dimensional prediction problems in experimental science, medicine, finance, and engineer…
Paper improves deep learning convergence rates for low-dimensional data.
CIR method preserves relation for case-control studies.
This work develops rigorous theoretical basis for the fact that deep Bayesian neural network (BNN) is an effective tool for high-dimensional variable selection with rigorous uncertainty quantification. We develop new Bayesian non-parametric theorems to show that a properly configured deep BNN (1) learns the variable im…
When can reliable inference be drawn in the "Big Data" context? This paper presents a framework for answering this fundamental question in the context of correlation mining, with implications for general large scale inference. In large scale data applications like genomics, connectomics, and eco-informatics the dataset…
A new algorithm reduces regret in high-dimensional online learning problems.
DPA preserves data distribution in reduced dimensions.
A Kronecker product model is the set of visible marginal probability distributions of an exponential family whose sufficient statistics matrix factorizes as a Kronecker product of two matrices, one for the visible variables and one for the hidden variables. We estimate the dimension of these models by the maximum rank …
A typical goal of supervised dimension reduction is to find a low-dimensional subspace of the input space such that the projected input variables preserve maximal information about the output variables. The dependence maximization approach solves the supervised dimension reduction problem through maximizing a statistic…
Bayesian optimization reduces hyperparameters for mixed variable design problems.
GIV methodology extends instrumental variable estimation for high-dimensional data.
VAE global minima can learn correct manifold dimensions, even with conditioning variables.
Bayesian network models with latent variables are widely used in statistics and machine learning. In this paper we provide a complete algebraic characterization of Bayesian network models with latent variables when the observed variables are discrete and no assumption is made about the state-space of the latent variabl…
Sparse PCA selects variables with FDR control for improved performance.
We consider non-parametric estimation and inference of conditional moment models in high dimensions. We show that even when the dimension of the conditioning variable is larger than the sample size , estimation and inference is feasible as long as the distribution of the conditioning variable has small intrinsic…
Generative models improve angular variable simulation in high dimensions.
We show that every Sasakian manifold in dimension is locally generated by a free real function of variables. This function is a Sasakian analogue of the Kähler potential for Kähler geometry. It is also shown that every locally Sasakian-Einstein manifold in dimensions is generated by a locally Kähler-…
Variable screening is a fast dimension reduction technique for assisting high dimensional feature selection. As a preselection method, it selects a moderate size subset of candidate variables for further refining via feature selection to produce the final model. The performance of variable screening depends on both com…
It is now known that an extended Gaussian process model equipped with rescaling can adapt to different smoothness levels of a function valued parameter in many nonparametric Bayesian analyses, offering a posterior convergence rate that is optimal (up to logarithmic factors) for the smoothness class the true function be…
New BO method efficiently optimizes high-dimensional functions by automatically selecting variables.
We introduce a new approach to variable selection, called Predictive Correlation Screening, for predictor design. Predictive Correlation Screening (PCS) implements false positive control on the selected variables, is well suited to small sample sizes, and is scalable to high dimensions. We establish asymptotic bounds f…
Variational Auto-Encoder (VAE) has been widely applied as a fundamental generative model in machine learning. For complex samples like imagery objects or scenes, however, VAE suffers from the dimensional dilemma between reconstruction precision that needs high-dimensional latent codes and probabilistic inference that f…
We present the Mixed Likelihood Gaussian process latent variable model (GP-LVM), capable of modeling data with attributes of different types. The standard formulation of GP-LVM assumes that each observation is drawn from a Gaussian distribution, which makes the model unsuited for data with e.g. categorical or nominal a…
Two novel methods estimate multiple FDR directions for binary categorical responses.
Smoothness of Sklyanin algebras examined in 3D and 4D cases.
The current study proposes a dimension reduction method, stepwise support vector machine (SVM), to reduce the dimensions of large p small n datasets. The proposed method is compared with other dimension reduction methods, namely, the Pearson product difference correlation coefficient (PCCs), recursive feature eliminati…
Novel privatization framework for high-dimensional variable selection with differential privacy.
A new kernel Stein test assesses fit for variable-length sequential data.
The monodromy conjecture states that every pole of the topological (or related) zeta function induces an eigenvalue of monodromy. This conjecture has already been studied a lot; however, in full generality it is proven only for zeta functions associated to a polynomial in two variables. In this article we consider zeta…
The paper provides risk bounds for learning many response functions using linear regression.
The main contribution of our paper is to give a partial classification of the quasi-exactly solvable Lie algebras of first order differential operators in three variables, and to show how this can be applied to the construction of new quasi-exactly solvable Schrödinger operators in three dimensions.
Improved CRT for sparse logistic regression in high dimensions.
tvGP-VAE models tensor-valued latent variables with Gaussian processes for better data structure representation.
The paper introduces a method for forecasting corporate sales growth using multiple reference variables.
Proposes VAE-KRnet for density estimation and variational Bayes.