The article proposes modified Gower's coefficients for handling mixed type variables in nearest neighbor methods.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study proves boundedness of operators in variable exponent Morrey spaces.
Bayesian test assesses dependence between mixed data types.
Improves Gower's similarity for mixed-type variables with automatic weighting.
A new method clusters mixed-type data efficiently.
New method for mixed data types in graphical models.
The notion of type of a differential 2-form in four variables is introduced and for 2-forms of type < 4, local normal models are given. If the type of a 2-form is 4, then the equivalence under diffeomorphisms of is reduced to the equivalence of a symplectic linear frame functorially attached to . As the equi…
A novel graph spectral method for mixed categorical and numerical data.
Derives derivatives and geometric framework for functions with non-independent variables.
Energy trees handle complex data structures with multiple variable types.
Clustering is an essential technique for discovering patterns in data. The steady increase in amount and complexity of data over the years led to improvements and development of new clustering algorithms. However, algorithms that can cluster data with mixed variable types (continuous and categorical) remain limited, de…
Two Bayesian optimization methods tackle dynamic design spaces with mixed variables.
Paper extends FOFC algorithm to work with mixed data types.
The paper investigates causal relationships in heart failure prediction using machine learning.
Finite mixture model is an important branch of clustering methods and can be applied on data sets with mixed types of variables. However, challenges exist in its applications. First, it typically relies on the EM algorithm which could be sensitive to the choice of initial values. Second, biomarkers subject to limits of…
The aim of this paper is to introduce a risk measure that extends the Gini-type measures of risk and variability, the Extended Gini Shortfall, by taking risk aversion into consideration. Our risk measure is coherent and catches variability, an important concept for risk management. The analysis is made under the Choque…
A new method clusters mixed-type data tables effectively.
New model clusters cells and individuals, revealing genetic influences on cell types.
New assumptions help identify causal relationships in data.
We introduce a Bernstein-type inequality which serves to uniformly control quadratic forms of gaussian variables. The latter can for example be used to derive sharp model selection criteria for linear estimation in linear regression and linear inverse problems via penalization, and we do not exclude that its scope of a…
New bounds on continuous random variables' right-tail probabilities.
A hybrid model for Bayesian optimization handles mixed variables using MCTS for categorical and GP for continuous.
Modern data acquisition based on high-throughput technology is often facing the problem of missing data. Algorithms commonly used in the analysis of such large-scale data often depend on a complete set. Missing value imputation offers a solution to this problem. However, the majority of available imputation methods are…
Study on how intraclass variability affects Temporal Ensembling accuracy.
CPI overcomes limitations of permutation importance by providing accurate variable selection.
LambdaNet infers TypeScript types using graph neural networks.
A new GP framework for discovering unknown functions and hypergraph structure.
Generalized Precision Matrix for scalable estimation of nonparametric Markov networks.
In this note we present a generative model of natural images consisting of a deep hierarchy of layers of latent random variables, each of which follows a new type of distribution that we call rectified Gaussian. These rectified Gaussian units allow spike-and-slab type sparsity, while retaining the differentiability nec…
In regression settings where explanatory variables have very low correlations and there are relatively few effects, each of large magnitude, we expect the Lasso to find the important variables with few errors, if any. This paper shows that in a regime of linear sparsity---meaning that the fraction of variables with a n…
We classify the dispersive Poisson brackets with one dependent variable and two independent variables, with leading order of hydrodynamic type, up to Miura transformations. We show that, in contrast to the case of a single independent variable for which a well known triviality result exists, the Miura equivalence class…
Statistical boosting algorithms have triggered a lot of research during the last decade. They combine a powerful machine-learning approach with classical statistical modelling, offering various practical advantages like automated variable selection and implicit regularization of effect estimates. They are extremely fle…
Derives derivatives of risk measures for various types of portfolio losses.
HNet detects significant associations in mixed data types efficiently.
We improve Gaussian copula models for imputing mixed data types with precise approximations.
Harmoniums model multiple time-to-event variables in survival analysis.
Functions of several octonion variables are investigated and integral representation theorems for them are proved. With the help of them solutions of the -equations are studied. More generally functions of several Cayley-Dickson variables are considered. Integral formulas of the Martinelli-Bochner,…
Proves new concentration inequalities for sub-gaussian and sub-exponential variables.
We describe a method to reduce partial differential equations of Monge-Ampère type in 4 variables to complex partial differential equations in 2 variables. To illustrate this method, we construct explicit holomorphic solutions of the special lagrangian equation, the real Monge-Ampère equations and the Plebanski equatio…
Improves Bayesian optimization efficiency for mixed variable spaces.
We prove semi-empirical concentration inequalities for random variables which are given as possibly nonlinear functions of independent random variables. These inequalities describe concentration of random variable in terms of the data/distribution-dependent Efron-Stein (ES) estimate of its variance and they do not requ…
We present the Mixed Likelihood Gaussian process latent variable model (GP-LVM), capable of modeling data with attributes of different types. The standard formulation of GP-LVM assumes that each observation is drawn from a Gaussian distribution, which makes the model unsuited for data with e.g. categorical or nominal a…
Quantitatively assessing relationships between latent variables and observed variables is important for understanding and developing generative models and representation learning. In this paper, we propose latent-observed dissimilarity (LOD) to evaluate the dissimilarity between the probabilistic characteristics of lat…
Deep learning had been used in program analysis for the prediction of hidden software defects using software defect datasets, security vulnerabilities using generative adversarial networks as well as identifying syntax errors by learning a trained neural machine translation on program codes. However, all these approach…
This paper explores what causal structures can be distinguished by observational and interventional probing schemes.
New method uses CDMs to improve CI testing without distributional assumptions.
One popular approach for nonstructural economic and financial forecasting is to include a large number of economic and financial variables, which has been shown to lead to significant improvements for forecasting, for example, by the dynamic factor models. A challenging issue is to determine which variables and (their)…
We introduce Thurstonian Boltzmann Machines (TBM), a unified architecture that can naturally incorporate a wide range of data inputs at the same time. Our motivation rests in the Thurstonian view that many discrete data types can be considered as being generated from a subset of underlying latent continuous variables, …