Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

3673109145 · Jun 202019922001200920182026
48 results for sparse GLM

Paper proposes robust and sparse GLM regression using stochastic optimization.

problem Sparse GLM's lack robustness against outliers in high-dimensional data.
method Robust and sparse linear regression based on γγ-divergence with stochastic optimization.
result The proposed method outperforms existing methods in numerical experiments and real data analysis.

Paper analyzes sparse aggregation in GLMs with Kullback-Leibler risk bounds.

problem Sparse aggregation in GLMs for parameter approximation.
method Exponential weighted aggregation scheme with Kullback-Leibler risk bounds.
result Sharp oracle inequality for Kullback-Leibler risk with leading constant 1 and minimax-optimal rate of aggregation.

Proposes SPCR-glm for generalized linear models combining PCA and regression losses.

problem Lack of response variable information in traditional PCA.
method Combines PCA and regression losses with a sparse penalty for parameter estimation.
result Improves interpretability and classification of principal components.

Bayesian pliable lasso with horseshoe prior models interactions in GLMs with missing data.

problem Modeling interactions in sparse regression problems with missing responses.
method Bayesian pliable lasso with hierarchical horseshoe prior for sparsity and uncertainty quantification.
result Advantages over existing methods in recovering complex interaction patterns under incomplete data.

This paper proposes a new method for GLM estimation using distance penalties to handle constraints.

problem Handling constraints in generalized linear models (GLM) is complicated.
method The approach uses distance penalties to optimize the log-likelihood, avoiding shrinkage.
result Distance penalties provide a flexible and non-shrinking alternative to traditional penalties.

PASS-GLM scales Bayesian GLM inference to large datasets with theoretical guarantees.

problem Inference in Bayesian GLMs is challenging for large datasets.
method Constructs polynomial approximate sufficient statistics for scalable Bayesian GLM inference.
result PASS-GLM provides theoretical guarantees on inference quality.

New methods for quantifying insurance claim cost uncertainty using LightGBM and GLMs.

problem Quantifying prediction uncertainty in insurance claim costs.
method Proposed non-conformity measures for GLMs and GBMs with Tweedie loss.
result Locally weighted Pearson residuals outperform other methods in maintaining nominal coverage with smallest average width.

LR-GLM speeds Bayesian GLM inference for high-dimensional data.

problem Bayesian inference in high-dimensional GLMs is computationally expensive.
method Low-rank data approximation to reduce computational time and memory costs.
result LR-GLM provides a full Bayesian posterior approximation with reduced computational time.

Paper analyzes GLM-tron for high-dimensional ReLU regression, providing upper and lower bounds.

problem Learning a single ReLU neuron in high-dimensional settings with overparameterization.
method Perceptron-type algorithm GLM-tron, with finite-sample analysis.
result Sharp characterization of high-dimensional ReLU regression problems via GLM-tron, contrasting with SGD.

We propose Dirichlet Process mixtures of Generalized Linear Models (DP-GLM), a new method of nonparametric regression that accommodates continuous and categorical inputs, and responses that can be modeled by a generalized linear model. We prove conditions for the asymptotic unbiasedness of the DP-GLM regression mean fu…

2009-09-28abs ↗pdf ↗

New method improves spike count estimation for neural populations.

problem Model overfitting and inaccurate parameter estimates in spike count modeling.
method Hierarchical parametric empirical Bayes method integrating GLMs and empirical Bayes theory.
result Improved accuracy and reliability of parameter estimation compared to existing methods.

DP-GD achieves dimension-independent convergence for unconstrained private GLMs.

problem Differentially private empirical risk minimization for unconstrained GLMs.
method Differentially private gradient descent (DP-GD).
result DP-GD achieves an excess empirical risk of $ ilde O\left(\sqrt{ exttt{rank}}/εn ight)$ for unconstrained GLMs.

Paper improves GLM estimation in NLDP model with public unlabeled data.

problem Estimating smooth GLMs in NLDP model with public unlabeled data.
method Presented (ϵ,δ)(\epsilon, \delta)-NLDP algorithms for GLMs using Stein's lemma and public/unlabeled data.
result Significant improvement in sample complexity for GLM estimation.

Develops a new GLM framework for claims reserving with adaptive estimation.

problem Accurate assessment of claims reserves with dynamic and dependent claim activity.
method Multivariate evolutionary GLM framework with adaptive particle filtering algorithm.
result Adaptive estimation of evolving factors improves claims reserve accuracy.

Paper proposes an alternative to MLE for GLMs with non-canonical link functions.

problem Challenges in MLE for GLMs with non-canonical link functions.
method Variational Inequality (VI) estimation framework.
result Established finite-sample error bounds and asymptotic normality for VI estimator.

The balance property is crucial for insurance pricing, ensuring total actuarial price equals loss. Maximum likelihood GLMs fulfill it, but Lindholm-Wüthrich suggests three methods, with constrained GLM being superior.

problem Ensuring the balance property in insurance pricing models
method Using constrained GLM fitting
result Constrained GLM fitting is superior to the two previously discussed balance correction methods

Paper develops methods for estimating GLMs and SNR under proportional asymptotics.

problem Estimation of regression coefficients and SNR in high-dimensional GLMs.
method Method-of-Moments type estimators that bypass nuisance function estimation.
result Consistent and asymptotically normal estimators derived for targets of inference.

Researchers expand on best subset selection theory, identifying key complexities.

problem Understanding model selection performance in high-dimensional sparse linear regression.
method Analyzing residualized signals, orthogonality, and spurious projections to establish margin conditions.
result Established necessary and sufficient margin conditions for BSS model consistency.

New method simplifies Bayesian analysis for categorical data.

problem Difficulties in scaling GLMs for categorical data due to non-conjugacy or posterior dependencies.
method Defining CB models with binary approximations for tractable inference.
result Fast and scalable inference for thousands of categories, outperforming competitors.

Study improves model fit by transferring info from related datasets.

problem Improving model fit on target data using source data.
method Proposes a transfer learning algorithm for GLMs, derives error bounds, and introduces detection of informative sources.
result Theoretical and practical improvements over classical methods in high-dimensional GLM settings.

BELIEF framework interprets GLMs using binary linear models.

problem Understanding and interpreting generalized linear models (GLMs) with binary outcomes.
method Developed a framework called binary expansion linear effect (BELIEF) to interpret GLMs through transparent linear models.
result BELIEF framework reveals perfect predictors in complete separation scenarios.

Paper connects GLM and LRM for better classification performance.

problem Improving classification performance using statistical inference.
method Derives a statistical test based on SVM and permutation analysis.
result MLE-based inference provides better parameter estimation.

Framework for domain adaptation using pseudo-labels from unlabeled data.

problem Improving prediction accuracy in target domain with covariate shift.
method Kernel GLMs with labeled and pseudo-labeled data, using imputation model for target data.
result Non-asymptotic excess-risk bounds for effective labeled sample size.

Adaptive pricing models for insurance using GLMs and GP regression.

problem Optimizing revenue from new insurance products.
method Developed two adaptive pricing models: GLM and Gaussian Process (GP) regression.
result The adaptive GLM and GP models reduce revenue loss compared to static pricing.

Unified framework for ensemble sampling in nonlinear contextual bandits with provable regret bounds.

problem Efficient exploration in nonlinear contextual bandits with unknown feature dimensions.
method Developed GLM-ES and Neural-ES for generalized linear and neural contextual bandits, respectively, using maximum likelihood estimation on randomly perturbed data.
result Unified high-probability frequentist regret bounds for GLM-ES and Neural-ES, matching state-of-the-art results.

TabPFN doesn't outperform GLM and XGBoost for motor insurance pricing.

problem Improving insurance pricing models using Tabular Foundation Models (TFMs).
method Pre-training on synthetic datasets and in-context learning for inference.
result TabPFN does not consistently outperform established baselines, has longer inference times, and is sensitive to training set size.

A fast, approximate method for variable selection in GLMs tackles correlated data.

problem Variable selection in generalized linear models with correlated data.
method Replica method of statistical mechanics and vector approximate message passing.
result The proposed algorithm provides fast convergence and high approximation accuracy.

A novel multi-objective optimization framework improves insurance pricing fairness.

problem Exacerbated trade-offs between competing fairness criteria in insurance pricing using machine learning.
method Proposes a novel multi-objective optimization framework using NSGA-II to jointly optimize accuracy and fairness criteria.
result Consistently achieves a balanced compromise between accuracy and fairness, outperforming single-model approaches.