New class of heavy-tailed distributions shows weighted averages dominate individual variables.
problem Understanding and comparing risks in heavy-tailed distributions.
method Introducing a new class of heavy-tailed distributions and proving stochastic dominance relations.
result Weighted averages of random variables in this class are stochastically larger than individual variables.
Estimates returns for dollar cost averaging using geometric Brownian motion.
problem Estimating returns for dollar cost averaging investing strategy.
method Uses geometric Brownian motion and log-Normal distribution to construct a lower bound for returns. Computes parameters recursively and in closed form for dollar cost averaging. Compares to lump sum investing for matching wealth distributions.
result Probability of negative returns is less than 2.5% for 40 years of annual dollar cost averaging.
Distributed learning for related tasks using graph weights.
problem Learning multiple tasks on different machines.
method Weighted averaging of messages with skewing or stepsize control.
result Different tasks can be learned on different machines.
Bayesian model averaging improves causal effect estimation by averaging over multiple models.
problem Estimating causal effects under linear Structural Causal Models (SCMs).
method Bayesian model averaging using Gaussian scale mixture distributions for computational efficiency.
result Bayesian model averaging is optimal for causal effect estimation.
Determinantal averaging corrects inversion bias in distributed Newton's method.
problem Inverting a sum of distributed matrices is biased; local averages are incorrect.
method Reweighting local estimates of the Newton's step proportionally to the determinant of the local Hessian estimate, then averaging them.
result Determinantal averaging provides the first known asymptotically consistent distributed Newton step.
Study optimal portfolio selection with Recovery Average Value at Risk, showing better control over liabilities.
problem Optimizing portfolios with a new risk measure under known or uncertain distributions.
method Existence results for mean-risk optimal portfolios under different distributional assumptions.
result Portfolio selection under Recovery Average Value at Risk provides better control over liabilities.
This paper studies distributed linear regression by averaging to reduce communication costs.
problem Reducing communication costs in distributed statistical learning with large datasets.
method One-step and iterative weighted parameter averaging in statistical linear models under data parallelism.
result The method reduces performance loss in estimation, test error, and confidence interval length compared to full data regression in high dimensions.
WASH trains ensembles with shuffled weights to improve accuracy and reduce communication.
problem Training ensembles for weight averaging leads to models converging to different loss basins.
method WASH randomly shuffles a small percentage of weights during training to keep models within the same basin.
result WASH achieves state-of-the-art image classification accuracy with lower communication costs.
Derives pricing formulae for power binary and normal distribution standard options.
problem Developing pricing models for binary and standard options.
method Incorporates Buchen's formulae into power binary options and derives a formula for normal distribution standard options.
result Derives pricing formulae for power binary and normal distribution standard options.
A new distributed SGD algorithm reduces communication by infrequent global reduction.
problem Reducing communication in large-scale machine learning training.
method Hierarchical averaging stochastic gradient descent (Hier-AVG) with infrequent global reduction.
result Hier-AVG achieves comparable training speed with better test accuracy.
Develops unbiased averaging methods for second order optimization in distributed systems.
problem Computing the Hessian is challenging and communication is a bottleneck in distributed optimization.
method Unbiased parameter averaging methods using sampling and sketching of the Hessian.
result Provably minimizes bias for sketched Newton directions.
The study examines how averaging data improves model performance.
problem Understanding the generalization gap in machine learning models.
method Data averaging, covariance analysis, and stochastic gradient descent (SGD) noise modeling.
result A modified generalization gap is always non-negative for a large class of model parameter distributions.
A2SGD reduces distributed SGD communication to O(1) per worker.
problem Heavy communication costs in distributed SGD for large models.
method Two-level gradient averaging to consolidate gradients to two local averages.
result Achieves O(1) communication complexity per worker, significantly reducing traffic and training time.
We consider random vectors drawn from a multivariate normal distribution and compute the sample statistics in the presence of non-stationary correlations. For this purpose, we construct an ensemble of random correlation matrices and average the normal distribution over this ensemble. The resulting distribution contains…
Paper uses averaging from many particle filters to approximate posterior predictive distributions.
problem Approximating posterior predictive distributions efficiently and accurately.
method Particle swarm filter algorithm that averages many particle filter approximations.
result Law of large numbers and central limit theorem support the method's effectiveness.
New bounds for agnostic learning with average smoothness.
problem Distribution-free nonparametric regression with average smoothness.
method Distribution-free uniform convergence bounds and agnostic learning algorithm.
result Distribution-free uniform convergence bounds for average-smoothness classes in the agnostic setting.
Distributed statistical inference has recently attracted enormous attention. Many existing work focuses on the averaging estimator. We propose a one-step approach to enhance a simple-averaging based distributed estimator. We derive the corresponding asymptotic properties of the newly proposed estimator. We find that th…
In this paper we investigate the scaling behavior of the average daily exchange rate returns of the Indian Rupee against four foreign currencies namely US Dollar, Euro, Great Britain Pound and Japanese Yen. Average daily exchange rate return of the Indian Rupee against US Dollar is found to exhibit a persistent scaling…
Sornette et al. claimed that the optimal supply does not agree with the average demand, by analyzing a bakery model where a daily demand fluctuates with a uniform distribution. In this note, we extend the model to general probability distributions, and obtain the formula of the optimal supply for Gaussian distribution,…
COMP-AMS optimizes distributed training with compressed gradients, achieving similar accuracy with less communication.
problem Efficiently training large-scale models in distributed environments with reduced communication costs.
method Distributed optimization framework using gradient averaging and adaptive AMSGrad, with gradient compression and error feedback.
result COMP-AMS achieves the same convergence rate and linear speedup as standard AMSGrad with less communication.
BayesBlend blends multiple models' predictions for better insurance loss predictions.
problem Improving insurance loss predictions by combining multiple models.
method Pseudo-Bayesian model averaging, stacking, and hierarchical stacking.
result BayesBlend provides a user-friendly way to blend model predictions and estimate weights.
This paper examines federated learning from an information-theoretic perspective.
problem Understanding the conditions under which averaging model parameters in federated learning is beneficial.
method Measuring mutual information between representations and inputs/labels in local models and comparing it to the averaged model.
result Empirical results confirm the practical usefulness of averaging for neural networks, even with varying local dataset distributions.
Possible distributions are discussed for intertrade durations and first-passage processes in financial markets. The view-point of renewal theory is assumed. In order to represent market data with relatively long durations, two types of distributions are used, namely, a distribution derived from the so-called Mittag-Lef…
New algorithm for solving minimax problems over distributions converges to Nash equilibrium.
problem Solving minimax problems over probability distributions.
method Symmetric Mean-field Langevin Dynamics (MFL-AG and MFL-ABR) with weighted averaging and best response dynamics.
result Converges to mixed Nash equilibrium with average-iterate and last-iterate convergence.
We consider the problem of sequential sampling from a finite number of independent statistical populations to maximize the expected infinite horizon average outcome per period, under a constraint that the expected average sampling cost does not exceed an upper bound. The outcome distributions are not known. We construc…
In this thesis, we consider the suitability of using the charged cold fluid model in the description of ultra-relativistic beams. The method that we have used is the following. Firstly, the necessary notions of kinetic theory and differential geometry of second order differential equations are explained. Then an averag…
We analyze the question whether sliding window time averages applied to stationary increment processes converge to a limit in probability. The question centers on averages, correlations, and densities constructed via time averages of the increment x(t,T)=x(t+T)-x(t)and the assumption is that the increment is distribute…
We analyze constant step-size and iterate averaging in linear stochastic approximation algorithms.
problem Policy evaluation in reinforcement learning using temporal difference algorithms.
method Constant step-size and Polyak-Ruppert averaging of iterates.
result MSE decays as O(1/t) for a range of constant step-sizes under certain conditions.
The study analyzes when Bayesian averaging over decision trees is reliable.
problem When do Bayesian model averaging weights over decision trees provide reliable information?
method Closed-form solution for Bayesian decision trees with Catalan-exponential priors.
result Established a complete non-asymptotic theory of rational commitment thresholds.
New algorithm reduces communication in distributed SGD, improving efficiency.
problem Slow communication rounds bottleneck synchronous mini-batch SGD convergence.
method Proposes non-asymptotic error analysis for Local-SGD, comparing to averaging methods.
result Local-SGD reduces communication by a factor of O(√T/P^(3/2)) for large step sizes.
New generalization concept considers distribution of errors, not just average error.
problem Classical generalization fails to capture distributional differences in classifier outputs.
method Formal conjectures about distributional generalization based on model architecture, training procedure, and data distribution.
result Distributional generalization can be expected in specific conditions, as evidenced by empirical results.
New analysis shows halting time is predictable for large models, improving optimization efficiency.
problem Understanding the average-case complexity of optimization algorithms for large-scale models.
method Average-case analysis of first-order methods on random least squares and neural networks.
result Halting time is independent of input distribution, leading to tighter convergence rates.
In this paper we present detailed simulation results on the wealth distribution model with quenched saving propensities. Unlike other wealth distribution models where the saving propensities are either zero or constant, this model is not found to be ergodic and self-averaging. The wealth distribution statistics with a …
Paper develops new conformal prediction methods for sum or average of unknown labels.
problem Uncertainty quantification in joint distributions of random variables.
method Introduces novel conformal prediction methods for sum or average of unknown labels.
result Validates the proposed method for sum or average of unknown labels under permutation invariant assumptions.
This paper examines ADMM for network averaging, revealing its convergence rates and network topology impacts.
problem Efficiently averaging local information over a network using ADMM.
method Comparative analysis of ADMM and other algorithms on a canonical distributed averaging problem.
result Characterization of ADMM convergence and optimal parameter tuning based on network spectral properties.
Tree-based model averaging improves CATE estimation from diverse sites.
problem Limited sample size and privacy concerns prevent accurate personalized treatment effect estimation.
method Tree-based model averaging approach to estimate CATEs from multiple heterogeneous sites.
result Improved accuracy in estimating conditional average treatment effects (CATEs) across sites.
A new method compresses conditional distributions of labelled data.
problem No existing method directly compresses the conditional distribution of labelled data.
method Introduce Average Maximum Conditional Mean Discrepancy (AMCMD), derive a closed form estimator, and extend Kernel Herding (KH) to Average Conditional Kernel Herding (ACKH).
result Directly compressing conditional distributions outperforms joint distribution compression and greedy selection.
Unified meta algorithms estimate various distribution functionals in infinite-armed bandits.
problem Estimating various distribution functionals in infinite-armed bandits.
method Unified meta algorithms for offline and online settings, achieving optimal sample complexities.
result Online estimation offers significant advantage for certain distribution functionals.
New aggregation methods improve robustness and efficiency in distributed learning.
problem Outliers and malicious agents compromise traditional averaging in distributed learning.
method Developed statistically efficient and robust aggregation schemes based on median and trimmed mean variations.
result Achieved higher sample efficiency compared to traditional robust aggregation schemes.
A new method speeds up distributed learning while maintaining robustness.
problem Efficient and robust distributed machine learning in high dimensions.
method Multi-Bulyan for gradient aggregation in distributed learning.
result Multi-Bulyan achieves Byzantine resilience and near-averaging speed with linear complexity in d. GANICE improves GAN-based causal inference by minimizing averaged Wasserstein risk.
problem Estimating interventional outcome distributions and quantiles in causal inference.
method GANICE uses extended Wasserstein distance and a cellwise critic to minimize averaged Wasserstein risk.
result GANICE achieves minimax optimality and consistently outperforms existing methods.
New method prunes classification trees for biased data.
problem Pruning classification trees in imbalanced training data.
method Optimal pruning procedure for inhomogeneous data.
result First efficient procedure for optimal pruning under covariate shift.
Study optimizes step size for Metropolis algorithm in non-identifiable cases.
problem Optimizing step size for Metropolis algorithm in non-identifiable models.
method Analytical derivation of average acceptance rate for non-identifiable cases.
result Developed optimization principle for step size based on average acceptance rate.
Quantum circuits are hard to learn on average.
problem Learning the output distributions of quantum circuits is hard.
method Statistical query model analysis.
result Learning quantum circuits requires exponentially many queries.
Meta-learning model predicts intervention effects from uncertain causal graphs.
problem Estimating intervention effects when causal structures are uncertain.
method Model-Averaged Causal Estimation Transformer Neural Process (MACE-TNP) using meta-learning.
result MACE-TNP outperforms Bayesian baselines in predicting intervention distributions.
New method assesses financial and cyber risks under uncertainty.
problem Uncertainty in risk assessment for financial and cyber systems.
method Combines stochastic approximation and distorted mix method to compute worst case average value at risk.
result Efficient algorithm for tail uncertainty in multivariate distributions.
Improved learning rates with new smoothness measure.
problem Learning with noisy data and unknown function class.
method Generalized Hölder smoothness to average smoothness, proving upper and lower bounds.
result Achieved nearly optimal learning rates in realizable and agnostic settings.
Average-case information complexity for learning is bounded, revealing only O(d) bits for most concepts.
problem Understanding the average-case information leakage in learning algorithms for concept classes.
method Developed a learning algorithm that reveals O(d) bits of information for most concepts in a class of VC-dimension d.
result Most concepts in the class do not require large amounts of information leakage, revealing only O(d) bits on average.