Prevalidated ridge regression simplifies logistic regression for high-dimensional data.
problem Efficient probabilistic classification in high-dimensional data with logistic regression.
method Developed a prevalidated ridge regression model that matches logistic regression's performance but is more computationally efficient.
result Prevalidated ridge regression achieves similar classification error and log-loss to logistic regression for high-dimensional data.
Proposes a model selection procedure for high-dimensional binary classification using sparse logistic regression.
problem High-dimensional binary classification with sparse logistic regression.
method Penalized maximum likelihood with complexity penalty on model size, Slope estimator for logistic regression.
result Proposed complexity penalty is rate-optimal in the minimax sense.
Improved VB algorithm for high-dimensional logistic regression with theoretical guarantees.
problem Sparse high-dimensional logistic regression model selection.
method Spike and slab variational Bayes approximation.
result Optimal convergence rates in ℓ2 and prediction loss for sparse truths. The paper shows a phase transition for the existence of MLE in high-dimensional logistic regression.
problem The existence of the maximum likelihood estimate in high-dimensional logistic regression models.
method Established a phase transition boundary curve hextMLE parameterized by scalars measuring the magnitude of regression coefficients. result The existence of the MLE in high-dimensional logistic regression models undergoes a sharp phase transition.
Regularization improves logistic regression performance in high-dimensional settings.
problem Improving logistic regression in scenarios with many parameters and observations.
method Introducing a convex regularizer to the negative log-likelihood function to encourage desired structures.
result Explicit expressions for various performance metrics of regularized logistic regression are derived.
Efficiently selects important variables in high-dimensional logistic regression.
problem Variable selection in high-dimensional logistic regression with binary responses.
method Developed a variational empirical Bayes approach for efficient model space marginal distribution.
result The variational approximation inherits strong selection consistency from the posterior distribution.
New method calibrates logistic regression tuning parameters.
problem Difficulty in calibrating tuning parameters in logistic regression.
method Simple tests along the tuning parameter path.
result Optimal guarantees for feature selection.
SLOE speeds up logistic regression in high dimensions with accurate signal strength estimation.
problem Poor performance of logistic regression in high-dimensional settings.
method SLOE reparameterizes the signal strength for faster and more accurate estimation.
result SLOE provides a fast and accurate method for dimensionality correction in logistic regression.
The l1-regularized logistic regression (or sparse logistic regression) is a widely used method for simultaneous classification and feature selection. Although many recent efforts have been devoted to its efficient implementation, its application to high dimensional data still poses significant challenges. In this paper…
Optimizes link prediction using matrix logistic regression.
problem Predicting links in networks with limited data.
method Formulated as matrix logistic regression, analyzed in high dimensions, combinatorial estimator with penalized maximum likelihood.
result Achieves minimax rate for Frobenius-norm risk, cannot be computed efficiently.
Paper generalizes Gaussian universality and CGMT to dependent data, impacting data augmentation in high-dimensional logistic regression.
problem Limitation of Gaussian universality and CGMT in handling dependent data.
method Generalizes Gaussian universality and CGMT to dependent data (block dependence, m-dependence, mixing). Establishes a novel CGMT framework.
result Gaussian universality holds for high-dimensional logistic regression under various types of dependence.
Sparse multinomial logistic regression for multiclass classification with feature selection.
problem High-dimensional multiclass classification with a focus on sparse models.
method Penalized maximum likelihood with complexity penalty, feature selection using group Lasso and Slope classifiers.
result Achievement of minimax order in both small and large number of classes regimes.
FILTER model uses fusion penalized logistic threshold regression for high-dimensional data with unknown cut points.
problem Modeling high-dimensional data with unknown cut points and binary responses.
method Fusion penalized logistic threshold regression (FILTER) model with fused lasso penalty for variable selection.
result Established non-asymptotic error bounds for coefficient estimation and model selection consistency.
Study high-dimensional logistic regression with missing data, providing exact error characterizations.
problem High-dimensional logistic regression with missing or corrupted covariates.
method Exact characterizations of prediction and estimation errors under independence and moment conditions.
result Characterizations are universal and hold for various imputation strategies.
Proposes a method to detect and explain outliers using localized logistic regression.
problem Detecting and explaining outliers in high-dimensional data.
method Localized logistic regression for density ratio estimation.
result The method successfully detects important features for outliers and outperforms existing algorithms.
The paper proves asymptotic normality for multinomial logistic regression on null covariates.
problem Classical asymptotic normality results fail in high-dimensional multinomial logistic models.
method Developed asymptotic normality and chi-square results for multinomial logistic MLE on null covariates.
result Validated new methodology to test feature significance in high-dimensional classification problems.
Annotation errors can significantly hurt classifier performance, yet datasets are only growing noisier with the increased use of Amazon Mechanical Turk and techniques like distant supervision that automatically generate labels. In this paper, we present a robust extension of logistic regression that incorporates the po…
We propose a new algorithm called PLUTO for building logistic regression trees to binary response data. PLUTO can capture the nonlinear and interaction patterns in messy data by recursively partitioning the sample space. It fits a simple or a multiple linear logistic regression model in each partition. PLUTO employs th…
Early stopping improves logistic regression's calibration and consistency in high dimensions.
problem Improving the statistical performance of gradient descent in overparameterized logistic regression.
method Investigates the effects of early stopping on gradient descent in logistic regression.
result Early-stopped gradient descent is well-calibrated and statistically consistent, while asymptotic gradient descent is not.
The paper forecasts corporate distress using a novel MIDAS logistic regression method.
problem Forecasting corporate distress with right-censored data, high-dimensional predictors, and mixed-frequency data.
method The paper introduces a novel high-dimensional censored MIDAS logistic regression method that handles censoring through inverse probability weighting and employs a sparse-group penalty for mixed-frequency predictors.
result The method achieves accurate estimation and superior performance in predicting financial distress of Chinese-listed firms.
The likelihood ratio test in high-dimensional logistic regression is not a chi-square, but a rescaled one.
problem The incorrect chi-square approximation in high-dimensional logistic regression.
method Proving the rescaled chi-square distribution and solving nonlinear equations.
result The likelihood ratio test is a rescaled chi-square, not a standard chi-square.
Faster and more accurate image classification via label embeddings.
problem Efficiently training multi-label, large-scale image classification models.
method Embedding labels onto a dense sphere and treating classification as cosine proximity regression.
result 7% higher mean average precision compared to logistic regression.
We develop a first order expansion for convex penalized estimators in high-dimensional regression.
problem High-dimensional regression problems with random designs.
method Construct a first order expansion η of the penalized estimator β^. result The risk of β^ is asymptotically the same as the risk of η. New weighted Lasso estimates improve logistic regression performance with measurement error.
problem Improper Lasso estimates in sparse logistic regression with equal penalties.
method Proposed weighted Lasso estimates using McDiarmid inequality for non-asymptotic oracle inequalities.
result Finite sample behavior illustrated by non-asymptotic oracle inequalities for estimation and prediction errors.
Improved logistic regression for multi-omics data improves prediction and variable selection.
problem Predicting binary class labels from multi-omics datasets with varying characteristics.
method Two-step penalized logistic regression with separate variable selection for each data layer.
result Our approach selects more relevant predictors and achieves comparable prediction performance.
The paper identifies universal features for high-dimensional data inference.
problem Identifying universal low-dimensional features from high-dimensional data for inference tasks.
method Introduces natural notions of universality and shows a local equivalence among them, using information geometry.
result Reveals the complementary roles of various data analysis techniques.
New method distinguishes predictive distribution estimators in high-dimensional inputs.
problem Difficulty in evaluating predictive distributions for high-dimensional inputs.
method Introduces dyadic sampling to focus on predictive distributions associated with pairs of inputs.
result Demonstrates efficient distinction of predictive distribution estimators in high-dimensional examples.
High-dimensional feature selection arises in many areas of modern science. For example, in genomic research we want to find the genes that can be used to separate tissues of different classes (e.g. cancer and normal) from tens of thousands of genes that are active (expressed) in certain tissue cells. To this end, we wi…
New AMP algorithm detects change points in high-dimensional GLMs.
problem Detecting change points in high-dimensional GLMs.
method Approximate Message Passing (AMP) algorithm for estimating signals and change points.
result Characterization of AMP algorithm's performance in high-dimensional limit.
A new method speeds up learning sparse binary networks.
problem Learning sparse binary pairwise Markov networks efficiently.
method Formulated as sparse multiple logistic regression, uses coordinate descent with strong screening rules.
result Substantial speedup with no loss of accuracy, more stable on unbalanced data.
Bayesian method tackles variable selection in high-dimensional data.
problem Challenges in Bayesian variable selection with large P.
method Efficient MCMC scheme with sublinear cost per iteration, extended to generalized linear models.
result Demonstrated effectiveness on cancer and maize genomic data.
Improved CRT for sparse logistic regression in high dimensions.
problem Accurate inference in high-dimensional sparse logistic regression.
method Variable-distillation and decorrelation steps in CRT-logit.
result CRT-logit provides a more powerful solution with theoretical guarantees.
New method uses nuclear and ℓ1 penalties for matrix regression, improving brain disorder detection.
problem Modeling high-dimensional matrix predictors with binary responses.
method Convex optimization with ADMM for low-rank and sparse structures.
result Effective in identifying brain disorder-related connectivity patterns.
Trans-GCR uses GCR model for node classification, providing theoretical guarantees and superior performance.
problem Challenges in obtaining node classification labels in real-world scenarios.
method Graph Convolutional Multinomial Logistic Regression (GCR) model and transfer learning method based on GCR.
result Trans-GCR provides superior empirical performance and theoretical guarantees.
Network Lasso classifies partially labeled data with high-dimensional features.
problem Classifying data points with limited labeled data and high-dimensional features.
method Logistic Network Lasso using total variation regularization and primal-dual splitting.
result Accurate classification achieved from limited labeled data via network structure.
Study shows how classifiers can approach Bayes error in high-dimensional settings.
problem Generalization error in high-dimensional perceptrons.
method Proved a formula for generalization error using convex optimization and observed that logistic and hinge regression can approach Bayes error closely.
result Logistic and hinge regression can approach Bayes-optimal generalization error closely in high-dimensional settings.
Efficiently solves inverse classification problems for logistic and softmax models.
problem Finding instances that change classifier predictions.
method Closed-form solution for logistic regression, iterative optimization for softmax.
result Fast, exact solutions for high-dimensional instances and many classes.
New algorithm learns sparse GLMs for binary outcomes efficiently.
problem Sparse modeling of binary outcomes in high-dimensional data.
method Iterative hard thresholding algorithm (BIHT) for sparse GLMs.
result BIHT achieves statistical optimality for logistic regression.
Logistic regression connects to perceptron learning via gradient ascent.
problem No specific problem stated; focuses on connection between algorithms.
method Gradient ascent for logistic regression compared to perceptron learning.
result Gradient ascent for logistic regression is a soft variant of perceptron learning.
High-dimensional regression models struggle with resampling methods.
problem Estimating uncertainty in high-dimensional supervised regression tasks.
method Investigation of bootstrap, subsampling, and jackknife methods in high-dimensional generalized linear models.
result Resampling methods exhibit double-descent behavior and are inconsistent in high dimensions.
Picasso is a new library for sparse learning problems in R and Python.
problem Sparse learning problems in high-dimensional data analysis.
method Unified framework of pathwise coordinate optimization with efficient active set selection strategies.
result picasso can efficiently handle large-scale problems.
Characterizes uncertainty in high-dimensional linear classification models.
problem Assessing uncertainty in high-dimensional linear classification models.
method Approximate message passing algorithm for posterior marginals, closed-form formula for joint statistics.
result Closed-form formula for joint statistics between logistic classifier, Bayesian uncertainty, and ground-truth probit uncertainty.
New method estimates density ratio for well-separated distributions using multi-class logistic regression.
problem Challenges in estimating density ratio for well-separated distributions.
method Uses multi-class logistic regression with auxiliary densities to estimate log(p/q).
result Demonstrates superior performance on density ratio estimation, mutual information, and representation learning tasks.
Paper finds a lower bound for estimating low-rank matrices in logistic regression.
problem Estimating low-rank coefficient matrices in logistic regression.
method Derives a minimax lower bound on the risk.
result The bound depends on matrix dimensions, rank, and sample size.
Hybrid machine learning improves gallstone risk prediction.
problem Complex gallstone disease risk factors and interactions.
method Adaptive LASSO for variable selection, BART for interactions, differential equations for interpretation.
result Enhanced prediction accuracy and actionable insights.
Developed efficient distributed logistic regression for large datasets.
problem Communication inefficiency and sparsity issues in distributed training of large-scale models.
method Iterative local optimization of a surrogate likelihood to improve initial solutions, handling sparsity and diverging updates.
result Learned a communication-efficient distributed logistic regression model for millions of features.
Novel bounds for logistic regression coreset construction and feature selection.
problem Efficiently summarize and reduce logistic regression inputs.
method Feature space sketching for logistic regression.
result Tight bounds for coreset construction and feature selection.
Unified framework for sparse logistic regression with nonconvex regularization.
problem Sparse logistic regression with nonconvex regularization.
method Unified framework, line search criteria for nonconvex terms.
result Effective classification and feature selection at lower computational cost.