The landscape of empirical risk has been widely studied in a series of machine learning problems, including low-rank matrix factorization, matrix sensing, matrix completion, and phase retrieval. In this work, we focus on the situation where the corresponding population risk is a degenerate non-convex loss function, nam…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper develops a two-population model to assess longevity basis risk.
Develops local population-risk certificates for model updates
We show that model compression can improve the population risk of a pre-trained model, by studying the tradeoff between the decrease in the generalization error and the increase in the empirical risk with model compression. We first prove that model compression reduces an information-theoretic bound on the generalizati…
One primary task of population health analysis is the identification of risk factors that, for some subpopulation, have a significant association with some health condition. Examples include finding lifestyle factors associated with chronic diseases and finding genetic mutations associated with diseases in precision he…
Proposes a federated transfer learning method to improve precision medicine models for underrepresented populations.
Paper tackles efficient risk estimation under dataset shift conditions.
Study models risks for low-carbon economy in Balkan countries, focusing on shadow economy and populism.
This paper studies the landscape of empirical risk of deep neural networks by theoretically analyzing its convergence behavior to the population risk as well as its stationary points and properties. For an -layer linear neural network, we prove its empirical risk uniformly converges to its population risk at the rat…
This work studies a stochastic optimal control problem for a pension scheme which provides an income-drawdown policy to its members after their retirement. To manage the scheme efficiently, the manager and members agree to share the investment risk based on a pre-decided risk-sharing rule. The objective is to maximise …
New bounds for KANs trained with DP-SGD, addressing correlated noise.
We study the Stochastic Gradient Langevin Dynamics (SGLD) algorithm for non-convex optimization. The algorithm performs stochastic gradient descent, where in each step it injects appropriately scaled Gaussian noise to the update. We analyze the algorithm's hitting time to an arbitrary subset of the parameter space. Two…
Population risk is always of primary interest in machine learning; however, learning algorithms only have access to the empirical risk. Even for applications with nonconvex nonsmooth losses (such as modern deep networks), the population risk is generally significantly more well-behaved from an optimization point of vie…
Paper improves privacy-preserving optimization rates for convex functions.
Gradient descent learns a single neuron without knowing the relationship between inputs and labels.
AICov integrates population covariates for better COVID-19 forecasting.
Early recognition of risky trajectories during an Intensive Care Unit (ICU) stay is one of the key steps towards improving patient survival. Learning trajectories from physiological signals continuously measured during an ICU stay requires learning time-series features that are robust and discriminative across diverse …
In stochastic optimization, the population risk is generally approximated by the empirical risk. However, in the large-scale setting, minimization of the empirical risk may be computationally restrictive. In this paper, we design an efficient algorithm to approximate the population risk minimizer in generalized linear …
The curse of dimensionality affects neural network optimization, especially with smooth functions.
The paper analyzes the maximum margin algorithm's performance on noisy data.
Essential to each other, growth and exploration are jointly observed in populations, be it alive such as animals and cells or inanimate such as goods and money. But their ability to move, crucial to cope with uncertainty and optimize returns, is tempered by the space/time properties of the environment. We investigate h…
Develops certificates for local population-risk increments using cross-fitted ridge calibration.
We present a new approach for mitigating unfairness in learned classifiers. In particular, we focus on binary classification tasks over individuals from two populations, where, as our criterion for fairness, we wish to achieve similar false positive rates in both populations, and similar false negative rates in both po…
Most high-dimensional estimation and prediction methods propose to minimize a cost function (empirical risk) that is written as a sum of losses associated to each data point. In this paper we focus on the case of non-convex losses, which is practically important but still poorly understood. Classical empirical process …
Paper proposes robust method to detect risk heterogeneity across ethnic groups.
Bayesian approach clusters survival data for better risk prediction.
Investigates optimal pension policies in PAYG systems with forward utility and ageing population.
U-aggregation combines multiple models without labels for better risk prediction.
Study optimizes data collection from biased, costly sources to minimize risk.
This work establishes uniform convergence of subdifferentials in stochastic optimization.
This work analyzes how users and services adapt to reduce risk, leading to specialization.
A government has to finance a risk for its population. It shares the charges among the population with a fixed scale based on economic criteria. Various organisms have to collect and to redistribute fairly the subsidies. Under these conditions, when the size of the organisms is varied, the distribution's laws of the cr…
Study on price formation in financial markets with a single default event.
Study causal inference under specific sampling methods with monotonicity assumptions.
This work aims to provide understandings on the remarkable success of deep convolutional neural networks (CNNs) by theoretically analyzing their generalization performance and establishing optimization guarantees for gradient descent based training algorithms. Specifically, for a CNN model consisting of convolution…
Method predicts NAFLD risk with high accuracy and distribution-free coverage guarantees.
Lapse-supported life insurance exacerbates adverse selection risks.
Participants enrolled into randomized controlled trials (RCTs) often do not reflect real-world populations. Previous research in how best to translate RCT results to target populations has focused on weighting RCT data to look like the target data. Simulation work, however, has suggested that an outcome model approach …
New algorithm finds approximate stationary points faster under differential privacy constraints.
New algorithms for differentially private optimization in convex and non-convex settings with near-optimal rates.
We develop a personalized real time risk scoring algorithm that provides timely and granular assessments for the clinical acuity of ward patients based on their (temporal) lab tests and vital signs. Heterogeneity of the patients population is captured via a hierarchical latent class model. The proposed algorithm aims t…
Paper develops NN models for diabetes screening using NHANES data.
Estimates population profile from small random samples.
Study creates a multimodal learning framework for CVD risk prediction.
We provide non-asymptotic excess risk guarantees for statistical learning in a setting where the population risk with respect to which we evaluate the target parameter depends on an unknown nuisance parameter that must be estimated from data. We analyze a two-stage sample splitting meta-algorithm that takes as input ar…
New estimates for the population risk are established for two-layer neural networks. These estimates are nearly optimal in the sense that the error rates scale in the same way as the Monte Carlo error rates. They are equally effective in the over-parametrized regime when the network size is much larger than the size of…
The paper analyzes the generalization performance of spectral clustering algorithms and proposes new methods to improve their effectiveness.
The paper addresses differentially private learning for neural networks, focusing on risk bounds and algorithm feasibility.