Boosting framework for vector-valued prediction with geometric stability.
problem Lack of a general theoretical understanding of aggregation for structured prediction.
method Identifies (α,β)-stability property and proposes a boosting framework based on exponential reweighting and geometric-median aggregation. result Obtains exponential decay of empirical divergence error under weak learner condition and (α,β)-stability. DICE learns population dynamics from discrete samples.
problem Learning smooth population dynamics from discrete data.
method Discrete Inverse Continuity Equation (DICE) method.
result DICE models are stable and generate representative samples.
We propose a heterogeneous agent market model (HAM) in continuous time. The market is populated by fundamental traders and chartists, who both use simple linear trading rules. Most of the related literature explores stability, price dynamics and profitability either within deterministic models or by simulation. Our nov…
Develops geometry for Lotka-Volterra model of species competition.
problem Population dynamics of competing species.
method Least squares variational method, Lagrange-Hamilton geometry.
result Jacobi stability discussed for the Lotka-Volterra system.
The paper examines the stability of binary choice models using Gini index and scoring indicators.
problem Stability and discriminatory power of binary choice models.
method Derives the real Gini index and incorporates PSI and KS statistics into the model.
result The real Gini index should be less than the calculated Gini index when the population distribution changes.
Proposes a stability evaluation criterion for learning models using distributional perturbations.
problem Ensuring reliable deployment of learning models in out-of-sample environments.
method Uses optimal transport discrepancy with moment constraints to quantify minimal perturbation required for model deterioration.
result Validates the practical utility of the stability evaluation criterion across various real-world applications.
The paper examines stability of shares in Proof of Stake protocol, identifying different investor behaviors and phase transitions.
problem Stability of shares in Proof of Stake protocol.
method Identification of large, medium, and small investors under various rewarding schemes; dynamical population model analysis.
result Phase transitions and thresholds for stability are characterized; chaotic centralization leads to concentration of shares.
New stability analysis improves generalization of multipass SGD.
problem Improper preconditioning affects generalization in multipass SGD.
method Developed on-average stability analysis for multipass SGD.
result Proper preconditioning yields optimal effective dimension dependence.
Measures of wealth and production have been found to scale superlinearly with the population of a city. Therefore, it makes economic sense for humans to congregate together in dense settlements. A recent model of population dynamics showed that population growth can become superexponential due to the superlinear scalin…
Proposes a method to simulate data for testing credit risk scorecard stability.
problem Ensuring credit risk scorecards remain representative of the population over time.
method Specification of bad ratios to generate parameter values for scorecards.
result Simulated data adheres closely to specified bad ratios.
Clust-PSI-PFL uses PSI to improve accuracy and fairness in federated learning.
problem Non-IID data biases federated learning performance.
method Clust-PSI-PFL uses clustering and PSI to form homogeneous groups of clients.
result Clust-PSI-PFL delivers up to 18% higher global accuracy and improves client fairness.
Mathematical methods of population genetics and framework of exchangeability provide a Markov chain model for analysis and interpretation of stochastic behaviour of equity markets, explaining, in particular, market shape formation, statistical equilibrium and temporal stability of market weights.
Improved online prediction with guaranteed coverage.
problem Creating reliable online predictions for arbitrary sequences.
method Online conformal prediction with decaying step sizes.
result Substantially improved practical properties, including close coverage at every time point.
CLSB models system dynamics from cross-sectional data with population-level regularization.
problem Challenges in modeling system dynamics from limited cross-sectional samples and heterogeneous individual behaviors.
method Introduces CLSB framework for learning dynamics, regularized for population-level temporal variations.
result Empirically superior in single-cell sequencing data analyses, e.g., simulating cell development and drug response.
This work learns models for population dynamics using variational methods and higher-order quadrature.
problem Modeling population dynamics of physical systems with stochastic and mean-field effects.
method Variational problem to infer gradient fields, combining Monte Carlo sampling with higher-order quadrature rules.
result Accurate prediction of population dynamics over a wide range of parameters.
We study differentially private (DP) algorithms for stochastic convex optimization (SCO). In this problem the goal is to approximately minimize the population loss given i.i.d. samples from a distribution over convex and Lipschitz loss functions. A long line of existing work on private convex optimization focuses on th…
The study analyzes the performance of statistical estimators under stability and computational efficiency.
problem Understanding the performance of statistical estimators in relation to stability and computational efficiency.
method Developed a framework to bound statistical accuracy based on the interplay between algorithm convergence rates and stability.
result Unstable algorithms can achieve the same statistical accuracy as stable ones in fewer steps.
New method estimates treatment effects across different populations.
problem Estimating treatment effects across populations with changing distributions.
method SBRL-HAP framework combining balancing and independence regularizers with hierarchical attention.
result Significant improvement in HTE estimation across out-of-distribution populations.
New algorithms help machines forget old data efficiently.
problem Machine learning models can retain old data, hindering new learning.
method Developed TV-stable algorithms based on noisy SGD for convex and non-convex functions.
result Achieved efficient unlearning with upper and lower bounds on risk.
This paper deals with the stability properties of a closed market, where capital and labour force are acting like a predator-prey system in population-dynamics. The spatial movement of the capital and labour force are taken into account by cross-diffusion effect. First, we are showing two possible ways for modeling thi…
While machine learning is rapidly being developed and deployed in health settings such as influenza prediction, there are critical challenges in using data from one environment in another due to variability in features; even within disease labels there can be differences (e.g. "fever" may mean something different repor…
Technological progress is leading to proliferation and diversification of trading venues, thus increasing the relevance of the long-standing question of market fragmentation versus consolidation. To address this issue quantitatively, we analyse systems of adaptive traders that choose where to trade based on their previ…
This paper analyzes stability and generalization of Markov chain stochastic gradient methods.
problem Analyzing stability and generalization of Markov chain stochastic gradient methods.
method Algorithmic stability in statistical learning theory.
result Established optimal generalization bounds for both smooth and non-smooth cases.
Polyak step size GD reaches final radius of convergence after log iterations.
problem Statistical and computational complexities of Polyak step size GD.
method Generalized smoothness and Lojasiewicz conditions, stability of gradients.
result Polyak step size GD reaches final statistical radius of convergence after logarithmic number of iterations.
Databases of electronic health records (EHRs) are increasingly used to inform clinical decisions. Machine learning methods can find patterns in EHRs that are predictive of future adverse outcomes. However, statistical models may be built upon patterns of health-seeking behavior that vary across patient subpopulations, …
This paper analyzes quantiles of heavy-tailed distributions, separating projection direction and quantile threshold effects.
problem Analyzing quantiles of heavy-tailed distributions with estimated parameters.
method Introduces a Q-Q orthogonality formulation to separate projection-direction and quantile-threshold effects.
result Decomposes the difference between empirical and population quantiles into three terms.
New estimator stabilizes higher-order influence functions for stable statistical inference.
problem Numerical instability in estimating inverse population Gram matrix.
method Proposes a new stabilized higher-order estimator without sample splitting.
result Stabilized estimator exhibits more stable performance and similar statistical guarantees.
Income and wealth distribution affect stability of a society to a large extent and high inequality affects it negatively. Moreover, in the case of developed countries, recently has been proven that inequality is closely related to all negative phenomena affecting society. So far, Econophysics papers tried to analyse in…
Proposes a framework to assess model robustness to dataset shifts.
problem Evaluating model robustness to changes in setting or population.
method Derives a debiased estimator for analyzing performance on worst-case distributions.
result Demonstrates the estimator can account for realistic shifts in complex distributions.
Operator calculus for population-based optimization provides a unified framework for analyzing convergence of various methods.
problem Convergence analysis of population-based optimization methods
method Introduce an operator calculus for describing composite mean-field algorithms as compositions of elementary operators acting on probability measures.
result Establish a modular Lyapunov principle for certifying exponential decay of state-space Lyapunov function and search errors.
The paper analyzes generalization bounds for NC-SC/NC-C stochastic minimax optimization.
problem Generalization analysis of nonconvex-(strongly)-concave stochastic minimax optimization.
method Established algorithm-agnostic and algorithm-dependent generalization bounds via uniform convergence and stability arguments.
result Sample complexities and generalization bounds for NC-SC and NC-C settings.
We provide existence, uniqueness and stability results for affine stochastic Volterra equations with L1-kernels and jumps. Such equations arise as scaling limits of branching processes in population genetics and self-exciting Hawkes processes in mathematical finance. The strategy we adopt for the existence part is b…
Geometric stability measures neural network robustness, distinguishing from similarity metrics.
problem Lack of robustness in neural network representations.
method Introduces geometric stability, quantified by Shesha metric measuring self-consistency.
result Stability and similarity are uncorrelated, revealing distinct properties of neural network robustness.
New stability bounds for SGD on nonsmooth convex losses.
problem Understanding stability of SGD on nonsmooth convex losses.
method Sharp upper and lower bounds for SGD and full-batch GD on nonsmooth convex losses.
result SGD can be less stable but still useful for generalization bounds.
New bounds for KANs trained with DP-SGD, addressing correlated noise.
problem Risk bounds for Kolmogorov-Arnold Networks trained by DP-SGD with correlated noise.
method Established new optimization and population risk analysis for KANs trained with DP-SGD, addressing correlated noise.
result First optimization and population risk analysis of correlated-noise mechanisms for DP training in non-convex settings, including neural networks.
Framework selects optimal historical data windows for non-stationary learning.
problem Learning in environments where conditions change over time.
method Stability principle applied to select look-back windows.
result Regret bounds are minimax optimal for strongly convex or Lipschitz population losses.
The stability of statistical analysis is an important indicator for reproducibility, which is one main principle of scientific method. It entails that similar statistical conclusions can be reached based on independent samples from the same underlying population. In this paper, we introduce a general measure of classif…
Proposes φ-balancing for more balanced expert utilization in MoE models.
problem Balanced expert utilization in MoE models to avoid bias.
method Directly targets population-level balance by minimizing a convex potential function.
result Consistently outperforms prior methods in stability and effectiveness.
Proposes BSSP to stabilize predictions in biased data.
problem Distribution shift between training and test data causes prediction instability.
method Balance-subsampled stable prediction (BSSP) algorithm based on fractional factorial design.
result Significantly improves prediction stability across unknown test data.
This paper studies the landscape of empirical risk of deep neural networks by theoretically analyzing its convergence behavior to the population risk as well as its stationary points and properties. For an l-layer linear neural network, we prove its empirical risk uniformly converges to its population risk at the rat…
Online algorithms stabilize in feedback loops of performative prediction.
problem Feedback loops in algorithmic predictions influence data distributions.
method Martingale argument and randomization to avoid distributional assumptions.
result No-regret algorithms converge to performatively stable equilibria.
Deep Reinforcement Learning (DRL) algorithms have been successfully applied to a range of challenging control tasks. However, these methods typically suffer from three core difficulties: temporal credit assignment with sparse rewards, lack of effective exploration, and brittle convergence properties that are extremely …
FHRN uses continuous-time dynamics to stabilize reentrant neural computation.
problem Stabilizing reentrant neural computation.
method Formulated as a continuous-time neural-ODE system, revealing norm-regulated reentry.
result Achieves stable oscillatory trajectories through population-level gain modulation.
Unified framework connects credit risk metrics with information theory.
problem Disconnection between industry-standard metrics and statistical theory.
method Unified information-theoretic framework, proving IV equals PSI, deriving standard errors, formalizing trade-off, automated binning with XGBoost.
result Unified framework connects IV and PSI, providing statistical foundation for metrics.
Study compares statistical properties and power of divergence measures for credit risk monitoring.
problem Detecting distributional shifts in credit risk models.
method Derives statistical properties and chi-square benchmark values for Jensen-Shannon Divergence and Kullback-Leibler Divergence, demonstrating their applicability in credit risk monitoring.
result Jensen-Shannon Divergence and Kullback-Leibler Divergence follow chi-square distributions and reveal practical trade-offs in minimizing false positives vs. detecting changes.
U-aggregation combines multiple models without labels for better risk prediction.
problem Challenges in selecting best model for new populations due to limited data and lack of true labels.
method U-aggregation, an unsupervised model aggregation method that integrates pre-trained models without observed labels.
result U-aggregation improves genetic risk prediction of complex traits using publicly available models.
Resampling outperforms reweighting for correcting biased data in machine learning models.
problem Correcting sampling bias in machine learning models trained on biased data sets.
method Compared resampling and reweighting techniques, focusing on their performance with stochastic gradient algorithms.
result Resampling outperforms reweighting when combined with stochastic gradient algorithms.
The paper studies stability of generative models trained on mixed data.
problem Training generative models on mixed datasets (real and synthetic data).
method Developed a framework to rigorously study stability under specific conditions.
result Proved the stability of iterative training under certain conditions.