Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

109219328437 · Jun 202019922001200920182026
48 results for non-identical random variables

The paper examines percentiles of non-identical random variables and provides non-asymptotic bounds.

problem Investigating percentiles of independent but non-identical random variables.
method Analyzing the 100(1p)100(1-p)%-th percentile X(pn)X^{(pn)} for a wide class of distributions.
result Discovering a connection between the median and the harmonic mean of standard deviations for certain distributions.

New class of heavy-tailed distributions shows weighted averages dominate individual variables.

problem Understanding and comparing risks in heavy-tailed distributions.
method Introducing a new class of heavy-tailed distributions and proving stochastic dominance relations.
result Weighted averages of random variables in this class are stochastically larger than individual variables.

Investigates VaR behavior for sums of one-sided random variables, showing impossibilities and conditions for super-additivity.

problem Investigates the behavior of Value-at-Risk (VaR) for sums of one-sided random variables.
method Analyzes the extremal aggregation behavior of VaR, introduces structural conditions for super-additivity.
result Characterizes when VaR is fully super-additive and provides unified framework for various dependence structures.

The paper analyzes ridge regression with random features for non-identically distributed data.

problem Analyzing ridge regression performance for data with heterogeneous variance profiles.
method Combining linear-plus-chaos approximation and operator-valued free probability.
result Derives asymptotic equivalents for training and test risks under non-identically distributed data.

Study ridge regression for non-identically distributed data with varying variances.

problem Investigate high-dimensional regression with non-identical data variance.
method Propose a random effect model and use tools from random matrix theory.
result Highlight the double descent phenomenon in high-dimensional regression for certain variance profiles.

FedProx tackles heterogeneity in federated learning networks.

problem Significant variability in systems characteristics and non-identically distributed data in federated networks.
method FedProx is a framework that generalizes and re-parametrizes FedAvg, introducing modifications to handle both systems and statistical heterogeneity.
result FedProx demonstrates significantly more stable and accurate convergence behavior than FedAvg, improving test accuracy by 22% on average in highly heterogeneous settings.

The study classifies Riemannian manifolds with specific Hessian properties.

problem Classifying Riemannian manifolds with a special Hessian structure.
method Analyzing manifolds with a non-identically vanishing function f whose Hessian is minus f times the Ricci tensor.
result Partial classification of manifolds with this Hessian structure.

VRL-SGD reduces communication complexity in non-identical data settings.

problem Training machine learning models with non-identical data distribution.
method VRL-SGD, which eliminates gradient variance dependency and achieves linear speedup with lower communication complexity.
result VRL-SGD reduces communication complexity from $O(T^{ rac{3}{4}} N^{ rac{3}{4}})$ to $O(T^{ rac{1}{2}} N^{ rac{3}{2}})$.

Random representations of surface groups approach asymptotic freeness in large nn limit.

problem Asymptotic freeness of Haar unitary matrices for surface groups.
method Interplay between Dehn's work and classical invariant theory.
result Expected value of trace of a fixed non-identity element is bounded as non o\infty.

This work examines how non-identical data distributions affect Federated Learning performance.

problem The impact of non-identical data distributions on Federated Learning performance.
method Synthesized datasets with varying degrees of data distribution similarity, evaluated Federated Averaging algorithm performance, proposed server momentum mitigation.
result Performance of Federated Learning degrades as data distributions differ more, and a mitigation strategy improves accuracy.

We strengthen the results of \cite{A1}, consequently, we improve the claims of \cite{A2} obtaining the best possible results. Namely, we prove that if a subgroup ΓΓ of Diff+(I)\mathrm{Diff}_{+}(I) contains a free semigroup on two generators then ΓΓ is not C0C_0-discrete. Using this, we extend the Hölder's Theorem in $\math…

2015-03-12abs ↗pdf ↗

Novel framework for data sharing and coordinated exploration in concurrent RL with non-identical environments.

problem Learning more data-efficient and better policies in concurrent RL with non-identical environments.
method Proposes a novel algorithmic framework that leverages causal inference via ANM-MM to extract model parameters and a new data sharing scheme based on similarity measures.
result Demonstrates superior learning speeds on various tasks and effectiveness of diverse action selection.

The paper provides guarantees for learning nonlinear representations from multiple non-identically distributed data sources.

problem Learning from non-identically distributed and dependent data.
method Established statistical guarantees for learning general nonlinear representations from multiple data sources.
result The excess risk of the estimated function decays as a function of the sample complexity and task diversity.

Study introduces new Bernstein inequalities for dependent data in Hilbert spaces.

problem Learning from non-independent and non-identically distributed data.
method Data-dependent Bernstein inequalities tailored for vector-valued processes in Hilbert space.
result Achieved novel risk bounds for covariance operator estimation and operator learning.

In the era of big data, reducing data dimensionality is critical in many areas of science. Widely used Principal Component Analysis (PCA) addresses this problem by computing a low dimensional data embedding that maximally explain variance of the data. However, PCA has two major weaknesses. Firstly, it only considers li…

2017-02-17abs ↗pdf ↗

Meta-analysis improves interpretation and efficiency across similar but non-identical datasets.

problem Meta-analysis of heterogeneous data in high dimensions.
method Integrative sparse regression with a global parameter for adaptability and anonymity.
result Superior identification of global parameter for high-dimensional linear models.

We prove that if Γis subgroup of Diff_{+}^{1+ε}(I) and N is a natural number such that every non-identity element of Γhas at most N fixed points then Γis solvable. If in addition Γis a subgroup of Diff_{+}^{2}(I) then we can claim that Γis metaabelian.

2013-08-01abs ↗pdf ↗

Study resolvent convergence for random matrices with general covariance profiles.

problem Analyzing resolvent convergence for random matrices with non-identically distributed columns.
method Using moments of quadratic forms and deterministic equivalents, the study provides bounds on the trace of matrix products.
result The trace of matrix products is close to the trace of a deterministic equivalent, controlled by matrix norms.

This paper tackles computational bottlenecks in federated learning on mobile devices.

problem Computationally heterogeneous mobile devices hinder federated learning efficiency.
method Proposes efficient algorithms to schedule mobile devices based on data heterogeneity.
result Achieves up to 100x speedup and 7% accuracy gain in federated learning.

fedCI and fedCI-IOD enable federated causal discovery across diverse datasets with privacy and power enhancements.

problem Causal discovery across multiple datasets with privacy constraints and heterogeneity.
method federated conditional independence test (fedCI) and Integration of Overlapping Datasets (IOD) algorithm extension (fedCI-IOD).
result fedCI-IOD achieves comparable performance to fully pooled analyses, enhancing statistical power and privacy.

Sharp concentration results for sums of heavy-tailed random variables.

problem Analyzing sums of independent heavy-tailed random variables.
method Using concentration inequalities and large deviation principles for distributions satisfying specific tail bounds.
result Sharp concentration inequalities and large deviation results for sums of heavy-tailed random variables.

The study examines property testing and estimation under non-identically distributed samples, finding necessary and sufficient sample complexities.

problem Property testing and estimation under non-identically distributed samples.
method Analysis of distributional property testing and estimation in settings with heterogeneous entities.
result Necessary and sufficient sample complexities for property testing and estimation under non-identically distributed samples.

Paper extends stochastic dominance for compound binomial distributions.

problem Stochastic dominance for infinite-mean random variables.
method Investigates properties and inclusion relationships of distribution classes, extends results to compound binomial distributions.
result Establishes necessary and sufficient conditions for first-order stochastic dominance preservation.

Undirected graphs are often used to describe high dimensional distributions. Under sparsity conditions, the graph can be estimated using 1\ell_1 penalization methods. However, current methods assume that the data are independent and identically distributed. If the distribution, and hence the graph, evolves over time t…

2008-02-20abs ↗pdf ↗

In [13], it is proved that any subgroup of Diff+ω(I)\mathrm{Diff}_{+}^{ω}(I) (the group of orientation preserving analytic diffeomorphisms of the interval) is either metaabelian or does not satisfy a law. A stronger question is asked whether or not the Girth Alternative holds for subgroups of Diff+ω(I)\mathrm{Diff}_{+}^{ω}(I). In th…

2015-03-12abs ↗pdf ↗

Consider an experiment involving a potentially small number of subjects. Some random variables are observed on each subject: a high-dimensional one called the "observed" random variable, and a one-dimensional one called the "outcome" random variable. We are interested in the dependencies between the observed random var…

2018-06-13abs ↗pdf ↗

The paper sets limits on the accuracy of macroeconomic forecasts based on statistical moments and trade volumes.

problem Uncertainty in predicting macroeconomic variables like prices and returns.
method Defines theoretical lower bounds of uncertainty and upper limits on forecast accuracy based on statistical moments and trade volumes.
result Accuracy of forecasts of probabilities of macroeconomic variables doesn't exceed Gaussian approximations.

This study compares machine learning methods for high-cardinality categorical variables.

problem Machine learning struggles with high-cardinality categorical variables.
method Empirical comparison of tree-boosting, deep neural networks, and linear mixed effects models.
result Tree-boosting with random effects outperforms deep neural networks with random effects.

LD-SGD improves communication in decentralized SGD.

problem Efficiently combining local updates and decentralized communication.
method Proposes LD-SGD integrating local updates and decentralized SGD, with a convergence analysis.
result LD-SGD converges to a critical point for non-convex objectives with non-identically distributed data.

Random Forest variable importance is improved by class balancing techniques.

problem Class imbalance problem in machine learning.
method Proposed a variable selection algorithm using RF variable importance and its confidence interval.
result Our algorithm efficiently selects an optimal feature set, leading to improved prediction performance.

Extends results for law-invariant functionals to random variable spaces.

problem Establishing results for a broad class of random variable spaces.
method Using structural results for law-invariant functionals and extending to new spaces.
result Unified perspective on law-invariant functionals, including quantile-based representations.

Random forests can be slow or inconsistent in certain models.

problem Performance issues of random forests in specific data-generating models.
method Intuitive arguments and numerical experiments, combined with variable use and importance statistics.
result Simple methods can create a better predictor using a forced random forest.

This paper examines from an experimental perspective random forests, the increasingly used statistical method for classification and regression problems introduced by Leo Breiman in 2001. It first aims at confirming, known but sparse, advice for using random forests and at proposing some complementary remarks for both …

2008-11-21abs ↗pdf ↗

This paper analyzes Mean Decrease Impurity (MDI) variable importance in random forests.

problem Lack of interpretability in random forest variable importances.
method Analysis of Mean Decrease Impurity (MDI) in random forests.
result MDI provides a variance decomposition of the output when variables are independent and there are no interactions.

New bounds on continuous random variables' right-tail probabilities.

problem Finding precise upper and lower limits for right-tail probabilities of continuous random variables.
method Developed new bounds based on PDF, first derivative, and two parameters.
result The new bounds are tight for various continuous random variables.

Develops a new method for nonlinear dimension reduction using random features.

problem Statistical challenges in generalizing Gaussian process-based latent variable models to non-Gaussian data.
method Random feature latent variable models (RFLVMs) that approximate nonlinear relationships with linear functions of random features.
result RFLVMs produce comparable results to state-of-the-art methods on various data types.