Sparse reduced-rank regression selects variables and ranks via manifold optimization.
problem Traditional rank selection fails when true rank is high.
method Sparse regularization and manifold optimization for rank and variable selection.
result Accurate estimation of coefficient parameter with high true rank.
R package varrank ranks variables based on mutual information for multivariate data analysis.
problem Selecting and ranking variables for multivariate datasets.
method Minimum redundancy maximum relevance (mRMRe) model based on information theory.
result Flexible implementation for discrete and continuous data.
Develops methods to estimate high rank tensors from noisy data.
problem Estimating high rank tensors from noisy observations.
method Generative latent variable tensor model, polynomial-time spectral algorithm.
result Achieves computationally optimal rate for signal tensor estimation.
Paper solves low-rank Boolean matrix approximation using integer programming.
problem Finding low-rank approximations to Boolean matrices.
method Integer programming formulation with polynomial variables and constraints.
result First computationally tractable integer programming approach.
The paper introduces an inference algorithm for graded Bayesian networks.
problem Inference in graded Bayesian networks.
method Tropicalization of the marginal distribution of observed variables, rank-by-rank evaluation of hidden variables.
result Established an inference algorithm for graded Bayesian networks.
Paper develops compact formulations for optimization problems with rank-one convex functions and indicator variables.
problem Optimization problems involving rank-one convex functions with support constraints.
method Perspective reformulation techniques to exploit conic structure and establish convex hull results.
result Systematic perspective formulations for convex hull descriptions of sets with nonlinear separable or non-separable objective functions and combinatorial constraints.
Sequential regression procedures can include spurious variables early, even in sparse settings.
problem Sequential regression procedures can select spurious variables early in rankings.
method Analysis of three sequential procedures: forward stepwise, lasso, and least angle regression.
result The first spurious variable is selected earlier as coefficients become denser.
Proposes a method to learn sparse and low-rank interactions in Ising models with latent variables.
problem Learning sparse interactions in Ising models with latent variables.
method Sparse + low-rank decomposition of Ising model parameters using convex regularized likelihood problem.
result Consistency properties in high-dimensional settings with growing number of variables and samples.
RI-based variable ranking and selection outperforms lasso in high-dimensional datasets.
problem Challenges in variable selection and model creation with correlated predictors.
method RI measures for feature ranking and selection, including CRI.Z.
result RI-based methods outperform lasso in high-dimensional datasets, especially with correlated predictors.
Matrices of (approximate) low rank are pervasive in data science, appearing in recommender systems, movie preferences, topic models, medical records, and genomics. While there is a vast literature on how to exploit low rank structure in these datasets, there is less attention on explaining why the low rank structure ap…
Novel Fréchet regression method handles errors-in-variables with low-rank covariates.
problem Regression with noisy and limited covariate data.
method Combines global Fréchet regression and principal component regression for low-rank structure.
result Improved efficiency and accuracy in high-dimensional and noisy data settings.
We develop latent variable models for Bayesian learning based low-rank matrix completion and reconstruction from linear measurements. For under-determined systems, the developed methods are shown to reconstruct low-rank matrices when neither the rank nor the noise power is known a-priori. We derive relations between th…
Study on diagonal and separating coordinates for symmetric spaces of rank 1.
problem Existence and nonexistence of diagonal and separating coordinates for symmetric spaces of rank 1.
method Generalization of results by Gauduchon and Moroianu, 2020, and analysis of constant sectional curvature and orthogonal separation of variables.
result Diagonal coordinates exist if and only if the symmetric space has constant sectional curvature.
The most popular approach for analyzing survival data is the Cox regression model. The Cox model may, however, be misspecified, and its proportionality assumption may not always be fulfilled. An alternative approach for survival prediction is random forests for survival outcomes. The standard split criterion for random…
RaFM improves FM performance with variable-rank embeddings.
problem Learning pairwise interactions with varying feature frequencies.
method Introduces RaFM with variable-rank embeddings for better performance.
result RaFM achieves better performance on real-world datasets.
New methods rank variables for Gaussian processes better than automatic relevance determination.
problem Variable selection for Gaussian process models using inverse length-scale parameters has limitations.
method Two novel methods rank variables based on their predictive relevance using posterior predictive distribution predictions.
result Improved variable selection compared to automatic relevance determination in terms of variability and predictive performance.
We formulate and solve a tensor model using a latent-variable approach.
problem Parameter inference for Poisson canonical polyadic tensor models.
method Latent-variable formulation, Expectation-Maximization algorithms, Fisher information matrices.
result Derivation of Fisher information for PCP models, insights into model well-posedness.
New insights into identifying mixtures of product distributions using Hadamard extensions.
problem Identifying mixtures of product distributions on binary variables.
method Analysis of Hadamard extensions of matrix products.
result Conditions for full column rank of Hadamard extensions.
Estimating the strength of dependency between two variables is fundamental for exploratory analysis and many other applications in data mining. For example: non-linear dependencies between two continuous variables can be explored with the Maximal Information Coefficient (MIC); and categorical variables that are depende…
RaSE screens variables via random subspaces, identifying joint effects.
problem Missing joint effects of predictors in ultra-high dimensional data.
method Random Subspace Ensemble (RaSE) framework combining subspace evaluation criteria.
result RaSE identifies signals with no marginal effect or high-order interactions.
Distributed algorithm finds global solutions for low-rank matrices.
problem Finding global solutions for low-rank matrices in distributed systems.
method Distributed Gradient Descent (DGD+) with LOCAL variables.
result DGD+LOCAL converges to global minimizer with exact consensus.
FLAMBE tackles RL in low rank MDPs by learning features.
problem Dealing with the curse of dimensionality in RL.
method Develops FLAMBE, a method that engages in exploration and representation learning for RL in low rank transition models.
result FLAMBE efficiently learns features for RL in low rank transition models.
Bayesian method combines expert and user rankings using copulas.
problem Combining expert and user rankings for accurate predictions.
method Bayesian inference with copula modeling latent variables.
result Predictive distribution of user rankings can be approximated accurately.
SOLVAR efficiently analyzes cryo-EM data's structural variability.
problem Analyzing continuous heterogeneity in cryo-EM data.
method Low-rank assumption on covariance matrix for tractable estimation.
result Accurately captures dominant components of structural variability.
We generalise surface cluster algebras to the case of infinite surfaces where the surface contains finitely many accumulation points of boundary marked points. To connect different triangulations of an infinite surface, we consider infinite mutation sequences. We show transitivity of infinite mutation sequences on tria…
New method learns graphical models with latent variables for extreme events.
problem Learning graphical models with latent variables for multivariate extremes.
method Tractable convex program exttt{eglatent} for Hüsler-Reiss models.
result Consistently recovers conditional graph and latent variables.
The paper tackles fair ranking in ranked data by addressing causal discrimination.
problem Fairness in predictive models for ranked data.
method Mapping rank positions to continuous scores, building causal graphs, and using path-specific effects.
result Effective algorithms for discovering and removing discrimination from ranked datasets.
Solar algorithm selects variables faster and more accurately in high-dimensional data.
problem Variable selection in high-dimensional data with high accuracy and stability.
method Subsample-ordered least-angle regression (solar) and its coordinate descent generalization (solar-cd) using L0 norm solution path averaging. result Solar selects variables with high accuracy and stability, reducing redundant variable selection.
We study the problem of learning latent variables in Gaussian graphical models. Existing methods for this problem assume that the precision matrix of the observed variables is the superposition of a sparse and a low-rank component. In this paper, we focus on the estimation of the low-rank component, which encodes the e…
New algorithm learns low-rank matrices with linear number of samples.
problem Learning low-rank matrices efficiently in latent-variable applications.
method Proposed algorithm that uses linear number of samples in high dimension.
result Learning kimesk, rank-r, matrices requires $Ω(rac{kr}{ε^2})$ samples. The paper improves multi-task learning by selecting variables and grouping tasks.
problem Improving generalization performance in multi-task learning.
method Factorizes a coefficient matrix into two matrices with sparsity for variable selection and overlapping group structure among tasks. Minimized using alternating optimization methods.
result Validated the effectiveness of the method on both synthetic and real-world datasets.
SON-NMF estimates nonnegative rank on-the-fly for NMF.
problem Estimating the nonnegative rank of data in NMF.
method Sum-of-norms (SON) regularization to reduce rank, combined with a first-order BCD algorithm.
result SON-NMF can automatically estimate the rank from data without prior knowledge.
New method improves image and signal processing with nonconvex rank surrogates and dual momentum.
problem Optimizing nonconvex rank minimization problems in image processing.
method Proposes a novel nonconvex rank surrogate, uses ADMM with dual momentum trick.
result Effective in image and signal processing applications, outperforming state-of-the-art methods.
This paper studies simultaneous feature selection and extraction in supervised and unsupervised learning. We propose and investigate selective reduced rank regression for constructing optimal explanatory factors from a parsimonious subset of input features. The proposed estimators enjoy sharp oracle inequalities, and w…
A new method for decomposing non-negative tensors using energy-based modeling.
problem Challenges in traditional tensor decomposition methods, especially global optimization and rank selection.
method Energy-based modeling of tensors, considering interactions between modes for global optimization.
result Demonstrates effectiveness in tensor completion and approximation, revealing a relationship between many-body and low-rank approximations.
iSplit LBI predicts individualized partial rankings from ties, outperforming state-of-the-art methods.
problem Predicting partial rankings from pairwise comparisons with ties, considering individual preferences.
method Variable splitting-based algorithm (iSplit LBI) that generates a sequence of estimations with a regularization path, decomposing parameters into abnormal signals, personalized signals, and random noise.
result iSplit LBI significantly outperforms state-of-the-art alternatives in predicting individualized partial rankings.
New method selects features via tensor decomposition and submodular optimization.
problem Feature selection for high-dimensional data.
method Low-rank tensor model, submodular optimization, greedy algorithm.
result Proposed method outperforms state-of-the-art feature selection.
In the era of big data, reducing data dimensionality is critical in many areas of science. Widely used Principal Component Analysis (PCA) addresses this problem by computing a low dimensional data embedding that maximally explain variance of the data. However, PCA has two major weaknesses. Firstly, it only considers li…
We classify invariant Lagrangians of the form L(gij,gij,k,gij,kl,DI,DI,j) depending at most quadratically on the variables gij,k,gij,kl and DI,DI,j, where g is a Lorentz metric and D is a tensor field of arbitrary rank on a smooth manifold. As a corollary, we prove a conjecture of Bray'…
Develops a new fairness learning approach for multi-task regression models.
problem Fairness in multi-task regression models with biased datasets.
method Uses rank-based non-parametric independence test (Mann Whitney U statistic) and reformulates as non-convex optimization problem.
result Outperforms state-of-the-art methods on fairness metrics.
We study the estimation of the latent variable Gaussian graphical model (LVGGM), where the precision matrix is the superposition of a sparse matrix and a low-rank matrix. In order to speed up the estimation of the sparse plus low-rank components, we propose a sparsity constrained maximum likelihood estimator based on m…
This article reviews ranking problems and their solutions.
problem Ranking problems in statistical learning.
method Systematic review of ranking problems and optimization techniques.
result Unified notation for optimization problems and identification of strengths and limitations of algorithms.
In this paper, we propose a low-rank approximation method based on discrete least-squares for the approximation of a multivariate function from random, noisy-free observations. Sparsity inducing regularization techniques are used within classical algorithms for low-rank approximation in order to exploit the possible sp…
New method identifies causal direction with latent confounders.
problem Identifying causal direction in presence of multiple latent variables.
method Use of joint higher-order cumulant matrix properties.
result Causal asymmetry can be seen from rank deficiency properties of cumulant matrices.
Paper applies ANOVA decomposition for interpretable data approximation.
problem High-dimensional data interpretation and dimensionality reduction.
method ANOVA decomposition and Grouped Transformations for interpretability.
result Ability to rank variable interactions and unimportant variables.
Study examines challenges in variable importance ranking due to feature correlation.
problem Challenges in variable importance ranking under correlation.
method Simulation study and theoretical analysis of feature knockoffs and conditional predictive impact (CPI).
result Highly correlated features increase the correlation of knockoff variables, posing a limitation for CPI.
In this present paper, we study geometric structures of rank two prolongations of implicit second-order partial differential equations (PDEs) for two independent and one dependent variables and characterize the type of these PDEs by the topology of fibers of the rank two prolongations. Moreover, by using properties of …
Enhances interpretability of linear latent spaces through automated clustering and ranking.
problem Severe interpretability issues in latent directions of PCA, ICA, CCA, and FA.
method LS-PIE framework automates clustering and ranking of latent vectors.
result Enhanced interpretability of latent vectors through LR, LS, LC, and LCON.