The paper examines percentiles of non-identical random variables and provides non-asymptotic bounds.
problem Investigating percentiles of independent but non-identical random variables.
method Analyzing the 100(1−p)%-th percentile X(pn) for a wide class of distributions. result Discovering a connection between the median and the harmonic mean of standard deviations for certain distributions.
New class of heavy-tailed distributions shows weighted averages dominate individual variables.
problem Understanding and comparing risks in heavy-tailed distributions.
method Introducing a new class of heavy-tailed distributions and proving stochastic dominance relations.
result Weighted averages of random variables in this class are stochastically larger than individual variables.
Investigates VaR behavior for sums of one-sided random variables, showing impossibilities and conditions for super-additivity.
problem Investigates the behavior of Value-at-Risk (VaR) for sums of one-sided random variables.
method Analyzes the extremal aggregation behavior of VaR, introduces structural conditions for super-additivity.
result Characterizes when VaR is fully super-additive and provides unified framework for various dependence structures.
The paper analyzes ridge regression with random features for non-identically distributed data.
problem Analyzing ridge regression performance for data with heterogeneous variance profiles.
method Combining linear-plus-chaos approximation and operator-valued free probability.
result Derives asymptotic equivalents for training and test risks under non-identically distributed data.
Study ridge regression for non-identically distributed data with varying variances.
problem Investigate high-dimensional regression with non-identical data variance.
method Propose a random effect model and use tools from random matrix theory.
result Highlight the double descent phenomenon in high-dimensional regression for certain variance profiles.
FedProx tackles heterogeneity in federated learning networks.
problem Significant variability in systems characteristics and non-identically distributed data in federated networks.
method FedProx is a framework that generalizes and re-parametrizes FedAvg, introducing modifications to handle both systems and statistical heterogeneity.
result FedProx demonstrates significantly more stable and accurate convergence behavior than FedAvg, improving test accuracy by 22% on average in highly heterogeneous settings.
The study classifies Riemannian manifolds with specific Hessian properties.
problem Classifying Riemannian manifolds with a special Hessian structure.
method Analyzing manifolds with a non-identically vanishing function f whose Hessian is minus f times the Ricci tensor.
result Partial classification of manifolds with this Hessian structure.
Optimal rates for learning hidden tree structures are determined.
problem Learning hidden tree structures from noisy data.
method Study of the (noisy) information threshold and the Chow-Liu algorithm.
result Optimal rates for structure recovery are inversely proportional to the information threshold squared.
VRL-SGD reduces communication complexity in non-identical data settings.
problem Training machine learning models with non-identical data distribution.
method VRL-SGD, which eliminates gradient variance dependency and achieves linear speedup with lower communication complexity.
result VRL-SGD reduces communication complexity from $O(T^{rac{3}{4}} N^{rac{3}{4}})$ to $O(T^{rac{1}{2}} N^{rac{3}{2}})$.
Random representations of surface groups approach asymptotic freeness in large n limit.
problem Asymptotic freeness of Haar unitary matrices for surface groups.
method Interplay between Dehn's work and classical invariant theory.
result Expected value of trace of a fixed non-identity element is bounded as no∞. This work examines how non-identical data distributions affect Federated Learning performance.
problem The impact of non-identical data distributions on Federated Learning performance.
method Synthesized datasets with varying degrees of data distribution similarity, evaluated Federated Averaging algorithm performance, proposed server momentum mitigation.
result Performance of Federated Learning degrades as data distributions differ more, and a mitigation strategy improves accuracy.
We strengthen the results of \cite{A1}, consequently, we improve the claims of \cite{A2} obtaining the best possible results. Namely, we prove that if a subgroup Γ of Diff+(I) contains a free semigroup on two generators then Γ is not C0-discrete. Using this, we extend the Hölder's Theorem in $\math…
Novel framework for data sharing and coordinated exploration in concurrent RL with non-identical environments.
problem Learning more data-efficient and better policies in concurrent RL with non-identical environments.
method Proposes a novel algorithmic framework that leverages causal inference via ANM-MM to extract model parameters and a new data sharing scheme based on similarity measures.
result Demonstrates superior learning speeds on various tasks and effectiveness of diverse action selection.
Macbeath gave a formula for the number of fixed points for each non-identity element of a cyclic group of automorphisms of a compact Riemann surface in terms of the universal covering transformation group of the cyclic group. We observe that this formula generalizes to determine the fixed-point set of each non-identity…
The paper provides guarantees for learning nonlinear representations from multiple non-identically distributed data sources.
problem Learning from non-identically distributed and dependent data.
method Established statistical guarantees for learning general nonlinear representations from multiple data sources.
result The excess risk of the estimated function decays as a function of the sample complexity and task diversity.
Study introduces new Bernstein inequalities for dependent data in Hilbert spaces.
problem Learning from non-independent and non-identically distributed data.
method Data-dependent Bernstein inequalities tailored for vector-valued processes in Hilbert space.
result Achieved novel risk bounds for covariance operator estimation and operator learning.
In the era of big data, reducing data dimensionality is critical in many areas of science. Widely used Principal Component Analysis (PCA) addresses this problem by computing a low dimensional data embedding that maximally explain variance of the data. However, PCA has two major weaknesses. Firstly, it only considers li…
Meta-analysis improves interpretation and efficiency across similar but non-identical datasets.
problem Meta-analysis of heterogeneous data in high dimensions.
method Integrative sparse regression with a global parameter for adaptability and anonymity.
result Superior identification of global parameter for high-dimensional linear models.
We prove that if Γis subgroup of Diff_{+}^{1+ε}(I) and N is a natural number such that every non-identity element of Γhas at most N fixed points then Γis solvable. If in addition Γis a subgroup of Diff_{+}^{2}(I) then we can claim that Γis metaabelian.
Study resolvent convergence for random matrices with general covariance profiles.
problem Analyzing resolvent convergence for random matrices with non-identically distributed columns.
method Using moments of quadratic forms and deterministic equivalents, the study provides bounds on the trace of matrix products.
result The trace of matrix products is close to the trace of a deterministic equivalent, controlled by matrix norms.
This paper tackles computational bottlenecks in federated learning on mobile devices.
problem Computationally heterogeneous mobile devices hinder federated learning efficiency.
method Proposes efficient algorithms to schedule mobile devices based on data heterogeneity.
result Achieves up to 100x speedup and 7% accuracy gain in federated learning.
fedCI and fedCI-IOD enable federated causal discovery across diverse datasets with privacy and power enhancements.
problem Causal discovery across multiple datasets with privacy constraints and heterogeneity.
method federated conditional independence test (fedCI) and Integration of Overlapping Datasets (IOD) algorithm extension (fedCI-IOD).
result fedCI-IOD achieves comparable performance to fully pooled analyses, enhancing statistical power and privacy.
Study on inequalities for multinomial variables.
problem Understanding concentration inequalities for multinomial variables.
method Investigation of Dirichlet and Multinomial random variables.
result Results on concentration inequalities for multinomial variables.
Sharp concentration results for sums of heavy-tailed random variables.
problem Analyzing sums of independent heavy-tailed random variables.
method Using concentration inequalities and large deviation principles for distributions satisfying specific tail bounds.
result Sharp concentration inequalities and large deviation results for sums of heavy-tailed random variables.
The study examines property testing and estimation under non-identically distributed samples, finding necessary and sufficient sample complexities.
problem Property testing and estimation under non-identically distributed samples.
method Analysis of distributional property testing and estimation in settings with heterogeneous entities.
result Necessary and sufficient sample complexities for property testing and estimation under non-identically distributed samples.
Paper extends stochastic dominance for compound binomial distributions.
problem Stochastic dominance for infinite-mean random variables.
method Investigates properties and inclusion relationships of distribution classes, extends results to compound binomial distributions.
result Establishes necessary and sufficient conditions for first-order stochastic dominance preservation.
Value-at-Risk can be superadditive for sufficiently heavy-tailed losses.
problem Value-at-Risk (VaR) subadditivity failure
method Random vector perspective
result Universal Value-at-Risk superadditivity (UVS)
New tree-structured Markov fields with Poisson marginals for counting variables.
problem Counting variables with complex dependencies.
method Tree-structured Markov random fields with Poisson marginals.
result Straightforward sampling and joint probability calculations.
Undirected graphs are often used to describe high dimensional distributions. Under sparsity conditions, the graph can be estimated using ℓ1 penalization methods. However, current methods assume that the data are independent and identically distributed. If the distribution, and hence the graph, evolves over time t…
In [13], it is proved that any subgroup of Diff+ω(I) (the group of orientation preserving analytic diffeomorphisms of the interval) is either metaabelian or does not satisfy a law. A stronger question is asked whether or not the Girth Alternative holds for subgroups of Diff+ω(I). In th…
Consider an experiment involving a potentially small number of subjects. Some random variables are observed on each subject: a high-dimensional one called the "observed" random variable, and a one-dimensional one called the "outcome" random variable. We are interested in the dependencies between the observed random var…
Diversification improves profits for heavy-tailed investments.
problem Investment portfolios of Pareto-distributed returns.
method Stochastic dominance and majorization order.
result Diversification increases first-order stochastic dominance for heavy-tailed returns.
The paper sets limits on the accuracy of macroeconomic forecasts based on statistical moments and trade volumes.
problem Uncertainty in predicting macroeconomic variables like prices and returns.
method Defines theoretical lower bounds of uncertainty and upper limits on forecast accuracy based on statistical moments and trade volumes.
result Accuracy of forecasts of probabilities of macroeconomic variables doesn't exceed Gaussian approximations.
This study compares machine learning methods for high-cardinality categorical variables.
problem Machine learning struggles with high-cardinality categorical variables.
method Empirical comparison of tree-boosting, deep neural networks, and linear mixed effects models.
result Tree-boosting with random effects outperforms deep neural networks with random effects.
LD-SGD improves communication in decentralized SGD.
problem Efficiently combining local updates and decentralized communication.
method Proposes LD-SGD integrating local updates and decentralized SGD, with a convergence analysis.
result LD-SGD converges to a critical point for non-convex objectives with non-identically distributed data.
Random Forest variable importance is improved by class balancing techniques.
problem Class imbalance problem in machine learning.
method Proposed a variable selection algorithm using RF variable importance and its confidence interval.
result Our algorithm efficiently selects an optimal feature set, leading to improved prediction performance.
Extends results for law-invariant functionals to random variable spaces.
problem Establishing results for a broad class of random variable spaces.
method Using structural results for law-invariant functionals and extending to new spaces.
result Unified perspective on law-invariant functionals, including quantile-based representations.
Investment models account for both probabilistic and possibilistic risks.
problem Investment and background risks are modeled probabilistically and possibilistically.
method Developed three investment models combining probabilistic and possibilistic risks.
result Approximate calculation formula for optimal investment solutions proved.
Random forests can be slow or inconsistent in certain models.
problem Performance issues of random forests in specific data-generating models.
method Intuitive arguments and numerical experiments, combined with variable use and importance statistics.
result Simple methods can create a better predictor using a forced random forest.
We prove new concentration inequalities for random variables.
problem Concentration of random variables in nonlinear functions.
method Efron-Stein inequalities and PAC-Bayesian approach.
result User-friendly concentration bounds for various applications.
A note on extending Chernoff bound for unit interval random variables.
problem Extending the Chernoff bound to random variables in the unit interval.
method Proof of extension of the Chernoff bound.
result A proof provided for the extension of the Chernoff bound.
This paper examines from an experimental perspective random forests, the increasingly used statistical method for classification and regression problems introduced by Leo Breiman in 2001. It first aims at confirming, known but sparse, advice for using random forests and at proposing some complementary remarks for both …
This paper analyzes Mean Decrease Impurity (MDI) variable importance in random forests.
problem Lack of interpretability in random forest variable importances.
method Analysis of Mean Decrease Impurity (MDI) in random forests.
result MDI provides a variance decomposition of the output when variables are independent and there are no interactions.
AugBagg improves random forest accuracy with added noise variables.
problem Improving model accuracy with random forest.
method AugBagg procedure using additional noise variables.
result Out-of-sample predictive accuracy improved with AugBagg.
Paper generalizes bipolar theorems for non-negative random variables.
problem Problems with existing bipolar theorems under stronger assumptions.
method Generalizes existing theorems in a robust probabilistic framework.
result Provides necessary and sufficient conditions for bipolar representation.
New bounds on continuous random variables' right-tail probabilities.
problem Finding precise upper and lower limits for right-tail probabilities of continuous random variables.
method Developed new bounds based on PDF, first derivative, and two parameters.
result The new bounds are tight for various continuous random variables.
Develops a new method for nonlinear dimension reduction using random features.
problem Statistical challenges in generalizing Gaussian process-based latent variable models to non-Gaussian data.
method Random feature latent variable models (RFLVMs) that approximate nonlinear relationships with linear functions of random features.
result RFLVMs produce comparable results to state-of-the-art methods on various data types.
Study large deviations in life insurance portfolios without identical distributions.
problem Large deviations in life insurance portfolios with bounded losses and variances.
method Upper bound from standard large deviations, counterexample for full large deviation principle.
result Exponential bound for average loss exceeding a threshold.