A distributed bootstrap method for high-dimensional data reduces communication rounds efficiently.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper studies schemes to de-bias the Lasso in a linear model where the goal is to construct confidence intervals for in a direction , where has iid rows. We show that previously analyzed propositions to de-bias the Lasso require a modification in order to enjoy efficiency in a f…
The paper develops inference methods for high-dimensional multi-task regression with row-sparse coefficients.
Paper compares Bayesian and de-biased estimators for low-rank matrix completion.
Novel characterization of augmented balancing weights combining outcome and weighting models.
We propose a robust inferential procedure for assessing uncertainties of parameter estimation in high-dimensional linear models, where the dimension can grow exponentially fast with the sample size . Our method combines the de-biasing technique with the composite quantile function to construct an estimator that …
The Whittle likelihood is a widely used and computationally efficient pseudo-likelihood. However, it is known to produce biased parameter estimates for large classes of models. We propose a method for de-biasing Whittle estimates for second-order stationary stochastic processes. The de-biased Whittle likelihood can be …
We study high-dimensional Gaussian mixture classification using statistical physics methods.
New method for estimating treatment effects without complex propensity models.
Many machine learning algorithms are trained and evaluated by splitting data from a single source into training and test sets. While such focus on in-distribution learning scenarios has led to interesting advancement, it has not been able to tell if models are relying on dataset biases as shortcuts for successful predi…
Performing statistical inference in high-dimension is an outstanding challenge. A major source of difficulty is the absence of precise information on the distribution of high-dimensional estimators. Here, we consider linear regression in the high-dimensional regime . In this context, we would like to perform in…
Although a majority of the theoretical literature in high-dimensional statistics has focused on settings which involve fully-observed data, settings with missing values and corruptions are common in practice. We consider the problems of estimation and of constructing component-wise confidence intervals in a sparse high…
Most modern supervised statistical/machine learning (ML) methods are explicitly designed to solve prediction problems very well. Achieving this goal does not imply that these methods automatically deliver good estimators of causal parameters. Examples of such parameters include individual regression coefficients, avera…
Noisy matrix completion aims at estimating a low-rank matrix given only partial and corrupted entries. Despite substantial progress in designing efficient estimation algorithms, it remains largely unclear how to assess the uncertainty of the obtained estimates and how to perform statistical inference on the unknown mat…
Paper addresses eigenvector perturbation in small eigen-gap scenarios.
Proposes FARM model combining latent factor and sparse regression.
Estimates social network structure from random walk subgraphs.
A new method corrects bias in high-dimensional ridge regression.
Visually predicting the stability of block towers is a popular task in the domain of intuitive physics. While previous work focusses on prediction accuracy, a one-dimensional performance measure, we provide a broader analysis of the learned physical understanding of the final model and how the learning process can be g…
PLS-Lasso integrates dimension reduction into regression for financial index tracking.
New method constructs confidence bands for ODE models with unknown regulatory effects.
The paper examines Adaptive Lasso and Transfer Lasso, highlighting their differences and proposing a new method.
DFR reduces the computational cost of sparse-group lasso and adaptive sparse-group lasso.
Anonymizing company names in financial news improves trading performance, contrary to initial expectations.
A fast method for Lasso and Logistic Lasso problems.
We leverage recent advances in high-dimensional statistics to derive new L2 estimation upper bounds for Lasso and Group Lasso in high-dimensions. For Lasso, our bounds scale as --- is the size of the design matrix and the dimension of the ground truth ---and match t…
We introduce an application of the group lasso to design of experiments. Note that we are NOT trying to explain experimental design for the group lasso. Conversely, we explain how we can use the idea of the group lasso in experimental design, showing that the problem of constructing an optimal design matrix can be tran…
New method for tuning Graphical Lasso hyperparameters.
This paper examines AI and ML bias and fairness issues.
The Bayesian Lasso is constructed in the linear regression framework and applies the Gibbs sampling to estimate the regression parameters. This paper develops a new sparse learning model, named the Bayesian Lasso Sparse (BLS) model, that takes the hierarchical model formulation of the Bayesian Lasso. The main differenc…
LLM-Lasso uses LLMs to improve feature selection in Lasso regression.
The solution path of the 1D fused lasso for an -dimensional input is piecewise linear with segments (Hoefling et al. 2010 and Tibshirani et al 2011). However, existing proofs of this bound do not hold for the weighted fused lasso. At the same time, results for the generalized lasso, of which the wei…
Proposes MM-DUST for efficient generalized lasso solution paths.
Bayesian approach improves network lasso for multi-task learning.
We propose a method for finding alternate features missing in the Lasso optimal solution. In ordinary Lasso problem, one global optimum is obtained and the resulting features are interpreted as task-relevant features. However, this can overlook possibly relevant features not selected by the Lasso. With the proposed met…
This review summarizes five Lasso optimization algorithms.
In high dimensional settings, sparse structures are crucial for efficiency, both in term of memory, computation and performance. It is customary to consider penalty to enforce sparsity in such scenarios. Sparsity enforcing methods, the Lasso being a canonical example, are popular candidates to address high dim…
We propose an improved LASSO estimation technique based on Stein-rule. We shrink classical LASSO estimator using preliminary test, shrinkage, and positive-rule shrinkage principle. Simulation results have been carried out for various configurations of correlation coefficients (), size of the parameter vector (), …
Paper examines LASSO for high-dimensional predictive regression, improving its performance in forecasting unemployment.
Exponential Lasso improves Lasso's robustness to outliers and heavy-tailed noise.
Exclusive Group Lasso improves feature selection in correlated biological data.
A new method speeds up overlapping group lasso computations.
The "least absolute shrinkage and selection operator" (Lasso) method has been adapted recently for networkstructured datasets. In particular, this network Lasso method allows to learn graph signals from a small number of noisy signal samples by using the total variation of a graph signal for regularization. While effic…
We compare alternative computing strategies for solving the constrained lasso problem. As its name suggests, the constrained lasso extends the widely-used lasso to handle linear constraints, which allow the user to incorporate prior information into the model. In addition to quadratic programming, we employ the alterna…
Paper improves Lasso for S&P500 index tracking with post-selection inference.
The sparse group lasso optimization problem is solved using a coordinate gradient descent algorithm. The algorithm is applicable to a broad class of convex loss functions. Convergence of the algorithm is established, and the algorithm is used to investigate the performance of the multinomial sparse group lasso classifi…
Proposes a new Lasso method with performance constraints.
We study the property of the Fused Lasso Signal Approximator (FLSA) for estimating a blocky signal sequence with additive noise. We transform the FLSA to an ordinary Lasso problem. By studying the property of the design matrix in the transformed Lasso problem, we find that the irrepresentable condition might not hold, …