Study resolvent convergence for random matrices with general covariance profiles.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New class of heavy-tailed distributions shows weighted averages dominate individual variables.
The paper analyzes ridge regression with random features for non-identically distributed data.
Study ridge regression for non-identically distributed data with varying variances.
This note displays an interesting phenomenon for percentiles of independent but non-identical random variables. Let be independent random variables obeying non-identical continuous distributions and be the corresponding order statistics. For any , we investig…
Federated Learning enables visual models to be trained in a privacy-preserving way using real-world data from mobile devices. Given their distributed nature, the statistics of the data across these devices is likely to differ significantly. In this work, we look at the effect such non-identical data distributions has o…
The paper provides guarantees for learning nonlinear representations from multiple non-identically distributed data sources.
Undirected graphs are often used to describe high dimensional distributions. Under sparsity conditions, the graph can be estimated using penalization methods. However, current methods assume that the data are independent and identically distributed. If the distribution, and hence the graph, evolves over time t…
This paper tackles computational bottlenecks in federated learning on mobile devices.
The study examines property testing and estimation under non-identically distributed samples, finding necessary and sufficient sample complexities.
Federated Learning is a distributed learning paradigm with two key challenges that differentiate it from traditional distributed optimization: (1) significant variability in terms of the systems characteristics on each device in the network (systems heterogeneity), and (2) non-identically distributed data across the ne…
To accelerate the training of machine learning models, distributed stochastic gradient descent (SGD) and its variants have been widely adopted, which apply multiple workers in parallel to speed up training. Among them, Local SGD has gained much attention due to its lower communication cost. Nevertheless, when the data …
Study large deviations in life insurance portfolios without identical distributions.
We consider the problem of large-scale inference on the row or column variables of data in the form of a matrix. Often this data is transposable, meaning that both the row variables and column variables are of potential interest. An example of this scenario is detecting significant genes in microarrays when the samples…
New insights into identifying mixtures of product distributions using Hadamard extensions.
It has been recently shown that numerical semiparametric bounds on the expected payoff of fi- nancial or actuarial instruments can be computed using semidefinite programming. However, this approach has practical limitations. Here we use column generation, a classical optimization technique, to address these limitations…
A common problem in large-scale data analysis is to approximate a matrix using a combination of specifically sampled rows and columns, known as CUR decomposition. Unfortunately, in many real-world environments, the ability to sample specific individual rows or columns of the matrix is limited by either system constrain…
We consider analysis of relational data (a matrix), in which the rows correspond to subjects (e.g., people) and the columns correspond to attributes. The elements of the matrix may be a mix of real and categorical. Each subject and attribute is characterized by a latent binary feature vector, and an inferred matrix map…
A distributed framework for reducing high-dimensional matrix-variate time series data.
We study the column subset selection problem with respect to the entrywise -norm loss. It is known that in the worst case, to obtain a good rank- approximation to a matrix, one needs an arbitrarily large number of columns to obtain a -approximation to the best entrywise -norm low ra…
Correctly detecting the semantic type of data columns is crucial for data science tasks such as automated data cleaning, schema matching, and data discovery. Existing data preparation and analysis systems rely on dictionary lookups and regular expression matching to detect semantic types. However, these matching-based …
We strengthen the results of \cite{A1}, consequently, we improve the claims of \cite{A2} obtaining the best possible results. Namely, we prove that if a subgroup of contains a free semigroup on two generators then is not -discrete. Using this, we extend the Hölder's Theorem in $\math…
Novel framework for data sharing and coordinated exploration in concurrent RL with non-identical environments.
New robust discriminant analysis for non-Gaussian data.
Macbeath gave a formula for the number of fixed points for each non-identity element of a cyclic group of automorphisms of a compact Riemann surface in terms of the universal covering transformation group of the cyclic group. We observe that this formula generalizes to determine the fixed-point set of each non-identity…
A new robust and flexible classification method for non-Gaussian data.
Modeling the probability distribution of rows in tabular data and generating realistic synthetic data is a non-trivial task. Tabular data usually contains a mix of discrete and continuous columns. Continuous columns may have multiple modes whereas discrete columns are sometimes imbalanced making the modeling difficult.…
Client adaptation improves federated learning performance with non-IID data.
In this paper, we study and partially classify those Riemannian man-ifolds carrying a non-identically vanishing function f whose Hessian is minus f times the Ricci-tensor of the manifold.
In this letter, we introduce a distributed Nesterov method, termed as , that does not require doubly-stochastic weight matrices. Instead, the implementation is based on a simultaneous application of both row- and column-stochastic weights that makes this method applicable to arbitrary (strongly-connected…
BIND removes background noise from binary matrices, improving detection accuracy and fairness.
This paper considers the problem of completing a matrix with many missing entries under the assumption that the columns of the matrix belong to a union of multiple low-rank subspaces. This generalizes the standard low-rank matrix completion problem to situations in which the matrix rank can be quite high or even full r…
Paper reinterprets ARP algorithm and improves its analysis and speed.
Fed-TDA augments federated tabular data to improve performance and privacy.
New method recovers matrix column space with active sampling for better results.
Study robust estimation under varying corruption probabilities in data.
We consider the task of meta-analysis in high-dimensional settings in which the data sources are similar but non-identical. To borrow strength across such heterogeneous datasets, we introduce a global parameter that emphasizes interpretability and statistical efficiency in the presence of heterogeneity. We also propose…
Dimensionality reduction is a first step of many machine learning pipelines. Two popular approaches are principal component analysis, which projects onto a small number of well chosen but non-interpretable directions, and feature selection, which selects a small number of the original features. Feature selection can be…
This paper studies quandles with one non-trivial column and their properties.
We consider the problem of matrix column subset selection, which selects a subset of columns from an input matrix such that the input can be well approximated by the span of the selected columns. Column subset selection has been applied to numerous real-world data applications such as population genetics summarization,…
The column group is a subgroup of the symmetric group on the elements of a finite blackboard birack generated by the column permutations in the birack matrix. We use subgroups of the column group associated to birack homomorphisms to define an enhancement of the integral birack counting invariant and give examples whic…
We prove that if Γis subgroup of Diff_{+}^{1+ε}(I) and N is a natural number such that every non-identity element of Γhas at most N fixed points then Γis solvable. If in addition Γis a subgroup of Diff_{+}^{2}(I) then we can claim that Γis metaabelian.
We define a family of probability distributions for random count matrices with a potentially unbounded number of rows and columns. The three distributions we consider are derived from the gamma-Poisson, gamma-negative binomial, and beta-negative binomial processes. Because the models lead to closed-form Gibbs sampling …
In this paper, we consider matrix completion from non-uniformly sampled entries including fully observed and partially observed columns. Specifically, we assume that a small number of columns are randomly selected and fully observed, and each remaining column is partially observed with uniform sampling. To recover the …
Double autoencoder improves missing value imputation in recommender systems.
Neuromorphic column performs online unsupervised clustering.
This paper is concerned with the problem of low rank plus sparse matrix decomposition for big data. Conventional algorithms for matrix decomposition use the entire data to extract the low-rank and sparse components, and are based on optimization problems with complexity that scales with the dimension of the data, which…
In this paper we investigate the feasibility of using synthetic data to augment face datasets. In particular, we propose a novel generative adversarial network (GAN) that can disentangle identity-related attributes from non-identity-related attributes. This is done by training an embedding network that maps discrete id…