PASS-GLM scales Bayesian GLM inference to large datasets with theoretical guarantees.
problem Inference in Bayesian GLMs is challenging for large datasets.
method Constructs polynomial approximate sufficient statistics for scalable Bayesian GLM inference.
result PASS-GLM provides theoretical guarantees on inference quality.
New method improves spike count estimation for neural populations.
problem Model overfitting and inaccurate parameter estimates in spike count modeling.
method Hierarchical parametric empirical Bayes method integrating GLMs and empirical Bayes theory.
result Improved accuracy and reliability of parameter estimation compared to existing methods.
New tools discover latent structure in neural circuits from spike train data.
problem Traditional methods fail to recover neural circuit organization due to noise and temporal dependencies.
method Hierarchical extension of GLM with graph-theoretic priors for latent features and connectivity.
result Reveals latent patterns of neural types and locations from spike trains alone.
A network of spiking agents learns complex tasks using global reward signals.
problem Solving complex reinforcement learning tasks.
method A hierarchical network of GLM spiking agents, each modulating its firing policy based on local and global reward signals.
result A network of spiking agents can learn complex action representations to solve RL tasks.
Bayesian pliable lasso with horseshoe prior models interactions in GLMs with missing data.
problem Modeling interactions in sparse regression problems with missing responses.
method Bayesian pliable lasso with hierarchical horseshoe prior for sparsity and uncertainty quantification.
result Advantages over existing methods in recovering complex interaction patterns under incomplete data.
Hierarchical Federated Learning bounds generalize using Wasserstein distance.
problem Bounding generalization error in Federated Learning with hierarchical sampling.
method Introduced a hierarchical sampling framework and derived generalization bounds using Wasserstein distance.
result Recover and strictly imply existing CMI bounds for bounded losses.
Paper introduces fair GLMs with convex penalty for equalizing GLM outcomes.
problem Achieving fairness in GLMs for practical use.
method Two fairness criteria based on GLM outcomes/log-likelihoods, achieved via a convex penalty on linear components.
result The fair GLM estimator is efficient and can handle various response variables.
We propose the supervised hierarchical Dirichlet process (sHDP), a nonparametric generative model for the joint distribution of a group of observations and a response variable directly associated with that whole group. We compare the sHDP with another leading method for regression on grouped data, the supervised latent…
Two randomized algorithms improve regret bounds for generalized linear bandits.
problem Improving regret bounds for generalized linear bandits.
method Two randomized algorithms: GLM-TSL and GLM-FPL.
result Upper bounds of O ( d n log K ) O(d \sqrt{n \log K}) O ( d n log K ) on regret for both algorithms. Paper proposes robust and sparse GLM regression using stochastic optimization.
problem Sparse GLM's lack robustness against outliers in high-dimensional data.
method Robust and sparse linear regression based on γ γ γ -divergence with stochastic optimization. result The proposed method outperforms existing methods in numerical experiments and real data analysis.
Compression technique recovers interpretability of GLM ensembles without sacrificing predictive power.
problem Loss of interpretability in ensembles of GLMs.
method MDL-motivated compression of GLM ensembles.
result Interpretability recovered without significant loss in predictive performance.
New methods for quantifying insurance claim cost uncertainty using LightGBM and GLMs.
problem Quantifying prediction uncertainty in insurance claim costs.
method Proposed non-conformity measures for GLMs and GBMs with Tweedie loss.
result Locally weighted Pearson residuals outperform other methods in maintaining nominal coverage with smallest average width.
Bayesian models predict Collatz stopping times with high accuracy.
problem Predicting the total stopping time of Collatz sequences.
method Developed two complementary models: a hierarchical Negative Binomial regression and a mechanistic generative approximation.
result Bayesian models outperform generative approximations in predicting Collatz stopping times.
LR-GLM speeds Bayesian GLM inference for high-dimensional data.
problem Bayesian inference in high-dimensional GLMs is computationally expensive.
method Low-rank data approximation to reduce computational time and memory costs.
result LR-GLM provides a full Bayesian posterior approximation with reduced computational time.
Paper analyzes GLM-tron for high-dimensional ReLU regression, providing upper and lower bounds.
problem Learning a single ReLU neuron in high-dimensional settings with overparameterization.
method Perceptron-type algorithm GLM-tron, with finite-sample analysis.
result Sharp characterization of high-dimensional ReLU regression problems via GLM-tron, contrasting with SGD.
A new method connects GLM and MLE for neuroimaging analysis.
problem Limited mathematical elegance and interpretation of MLE for neuroimaging.
method Derives a refined statistical test using SVR-iGLM and RFT.
result MLE and GLM parameter estimations are significantly related to functional tasks.
GLM-PCA simplifies complex data for easier analysis.
problem Non-normally distributed data complicates dimension reduction.
method Derives GLM-PCA, incorporates covariates, and suggests transformations.
result Improves interpretability of latent factors in non-normal data.
We propose Dirichlet Process mixtures of Generalized Linear Models (DP-GLM), a new method of nonparametric regression that accommodates continuous and categorical inputs, and responses that can be modeled by a generalized linear model. We prove conditions for the asymptotic unbiasedness of the DP-GLM regression mean fu…
New algorithm learns sparse GLMs for binary outcomes efficiently.
problem Sparse modeling of binary outcomes in high-dimensional data.
method Iterative hard thresholding algorithm (BIHT) for sparse GLMs.
result BIHT achieves statistical optimality for logistic regression.
New tensor model reduces GLM estimation error and sample complexity.
problem Estimating GLM coefficients with reduced sample complexity.
method Developed LSR tensor model and block coordinate descent algorithm.
result Minimax lower bound on estimation error, suggesting lower sample complexity.
DP-GD achieves dimension-independent convergence for unconstrained private GLMs.
problem Differentially private empirical risk minimization for unconstrained GLMs.
method Differentially private gradient descent (DP-GD).
result DP-GD achieves an excess empirical risk of $ ilde O\left(\sqrt{ exttt{rank}}/εn
ight)$ for unconstrained GLMs.
Generalized Linear Models (GLMs) and Single Index Models (SIMs) provide powerful generalizations of linear regression, where the target variable is assumed to be a (possibly unknown) 1-dimensional function of a linear predictor. In general, these problems entail non-convex estimation procedures, and, in practice, itera…
Paper improves GLM estimation in NLDP model with public unlabeled data.
problem Estimating smooth GLMs in NLDP model with public unlabeled data.
method Presented ( ϵ , δ ) (\epsilon, \delta) ( ϵ , δ ) -NLDP algorithms for GLMs using Stein's lemma and public/unlabeled data. result Significant improvement in sample complexity for GLM estimation.
Genomic models learn DNA sequences to predict functions.
problem Understanding complex genetic interactions.
method Training LLMs on DNA sequences to predict functions.
result gLMs can predict functions of DNA elements.
Paper analyzes sparse aggregation in GLMs with Kullback-Leibler risk bounds.
problem Sparse aggregation in GLMs for parameter approximation.
method Exponential weighted aggregation scheme with Kullback-Leibler risk bounds.
result Sharp oracle inequality for Kullback-Leibler risk with leading constant 1 and minimax-optimal rate of aggregation.
Develops a new GLM framework for claims reserving with adaptive estimation.
problem Accurate assessment of claims reserves with dynamic and dependent claim activity.
method Multivariate evolutionary GLM framework with adaptive particle filtering algorithm.
result Adaptive estimation of evolving factors improves claims reserve accuracy.
Extends matrix factorization for deviance-based losses with GLM theory.
problem Improving data loss models beyond squared error.
method Adapts GLM theory to matrix factorization for deviance losses.
result Strong consistency and robustness of the proposed decomposition.
New method controls FDR for sparse GLMs, identifying positive and negative relationships.
problem Sparse GLMs with high-dimensional data and varying sample size.
method Debiased-Lasso estimator and CLIME method for precision matrix estimation.
result Asymptotically controls directional FDR and FDV for sparse GLMs.
New method for GLMs under DP provides private uncertainty quantification.
problem Private inference for GLMs with uncertainty quantification.
method Noise-aware DP Bayesian inference method for GLMs.
result Posterior uncertainty allows determination of statistically significant coefficients.
Paper proposes an alternative to MLE for GLMs with non-canonical link functions.
problem Challenges in MLE for GLMs with non-canonical link functions.
method Variational Inequality (VI) estimation framework.
result Established finite-sample error bounds and asymptotic normality for VI estimator.
The balance property is crucial for insurance pricing, ensuring total actuarial price equals loss. Maximum likelihood GLMs fulfill it, but Lindholm-Wüthrich suggests three methods, with constrained GLM being superior.
problem Ensuring the balance property in insurance pricing models
method Using constrained GLM fitting
result Constrained GLM fitting is superior to the two previously discussed balance correction methods
Paper develops methods for estimating GLMs and SNR under proportional asymptotics.
problem Estimation of regression coefficients and SNR in high-dimensional GLMs.
method Method-of-Moments type estimators that bypass nuisance function estimation.
result Consistent and asymptotically normal estimators derived for targets of inference.
Holistic GLMs add constraints for better model quality.
problem Improving classical linear regression models.
method Sparsity-inducing, sign-coherence, and linear constraints.
result Holistic GLMs reliably solve GLMs for various responses.
Proposes SPCR-glm for generalized linear models combining PCA and regression losses.
problem Lack of response variable information in traditional PCA.
method Combines PCA and regression losses with a sparse penalty for parameter estimation.
result Improves interpretability and classification of principal components.
New method simplifies Bayesian analysis for categorical data.
problem Difficulties in scaling GLMs for categorical data due to non-conjugacy or posterior dependencies.
method Defining CB models with binary approximations for tractable inference.
result Fast and scalable inference for thousands of categories, outperforming competitors.
New algorithms for private GLM estimation with minimax lower bounds.
problem Privacy in generalized linear models.
method Differentially private algorithms using projected gradient descent.
result Nearly rate-optimal performance with privacy-constrained minimax lower bounds.
Study improves model fit by transferring info from related datasets.
problem Improving model fit on target data using source data.
method Proposes a transfer learning algorithm for GLMs, derives error bounds, and introduces detection of informative sources.
result Theoretical and practical improvements over classical methods in high-dimensional GLM settings.
Automates finding interactions in GLMs using neural networks.
problem Time-consuming and expert-dependent search for GLM interactions.
method Neural networks and model-specific interaction detection method.
result Computational speed improvement over traditional methods.
BELIEF framework interprets GLMs using binary linear models.
problem Understanding and interpreting generalized linear models (GLMs) with binary outcomes.
method Developed a framework called binary expansion linear effect (BELIEF) to interpret GLMs through transparent linear models.
result BELIEF framework reveals perfect predictors in complete separation scenarios.
Improved tensor GLM estimation for complex data.
problem Complex tensor data in GLMs leads to high-dimensional, ill-posed estimation.
method Proposed LSRTR-M algorithm using Muon updates for faster convergence and lower errors.
result LSRTR-M converges faster and achieves lower errors than LSRTR.
Paper connects GLM and LRM for better classification performance.
problem Improving classification performance using statistical inference.
method Derives a statistical test based on SVM and permutation analysis.
result MLE-based inference provides better parameter estimation.
Framework for domain adaptation using pseudo-labels from unlabeled data.
problem Improving prediction accuracy in target domain with covariate shift.
method Kernel GLMs with labeled and pseudo-labeled data, using imputation model for target data.
result Non-asymptotic excess-risk bounds for effective labeled sample size.
Adaptive pricing models for insurance using GLMs and GP regression.
problem Optimizing revenue from new insurance products.
method Developed two adaptive pricing models: GLM and Gaussian Process (GP) regression.
result The adaptive GLM and GP models reduce revenue loss compared to static pricing.
Unified framework for ensemble sampling in nonlinear contextual bandits with provable regret bounds.
problem Efficient exploration in nonlinear contextual bandits with unknown feature dimensions.
method Developed GLM-ES and Neural-ES for generalized linear and neural contextual bandits, respectively, using maximum likelihood estimation on randomly perturbed data.
result Unified high-probability frequentist regret bounds for GLM-ES and Neural-ES, matching state-of-the-art results.
TabPFN doesn't outperform GLM and XGBoost for motor insurance pricing.
problem Improving insurance pricing models using Tabular Foundation Models (TFMs).
method Pre-training on synthetic datasets and in-context learning for inference.
result TabPFN does not consistently outperform established baselines, has longer inference times, and is sensitive to training set size.
Paper trains SNNs for classification using first-to-spike decoding.
problem Training SNNs for classification under GLM model.
method Proposes first-to-spike decoding method for SNNs.
result Improves accuracy and efficiency of SNN classification.
A fast, approximate method for variable selection in GLMs tackles correlated data.
problem Variable selection in generalized linear models with correlated data.
method Replica method of statistical mechanics and vector approximate message passing.
result The proposed algorithm provides fast convergence and high approximation accuracy.
No-regret algorithm for contextual RL with GLM mappings.
problem Learning near-optimal policies in episodic MDPs with contextual information.
method Proposes no-regret online RL algorithm using optimistic and randomized exploration methods.
result Improves previous bounds and provides a lower bound for the setting.