The study examines the balancedness of random partition models and finds the rich-get-richer characteristic is a result of model assumptions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
SAM improves deep learning tasks by promoting balancedness, reducing outlier impact.
Regularization leads to balancedness in deep linear networks.
New estimator tackles multi-task linear regression with outliers, avoiding eigenvalue lower bounds.
Distance metric learning (DML), which learns a distance metric from labeled "similar" and "dissimilar" data pairs, is widely utilized. Recently, several works investigate orthogonality-promoting regularization (OPR), which encourages the projection vectors in DML to be close to being orthogonal, to achieve three effect…
Gradient descent solves asymmetric low-rank matrix sensing without balancing.
Wide neural networks with weight decay exhibit neural collapse.
A classical result by Pachner states that two -dimensional combinatorial manifolds with boundary are PL homeomorphic if and only they can be connected by a sequence of shellings and inverse shellings. We prove that for balanced, i.e., properly -colored, manifolds such a sequence can be chosen such that bala…
We analyze the complexity of Gibbs samplers for inference in crossed random effect models used in modern analysis of variance. We demonstrate that for certain designs the plain vanilla Gibbs sampler is not scalable, in the sense that its complexity is worse than proportional to the number of parameters and data. We thu…
This work explores how overparametrization and priors affect Bayesian neural network posteriors.
Gradient descent with large steps leads to chaotic parameter space and unpredictable outcomes.
Many modern clustering methods scale well to a large number of data items, N, but not to a large number of clusters, K. This paper introduces PERCH, a new non-greedy algorithm for online hierarchical clustering that scales to both massive N and K--a problem setting we term extreme clustering. Our algorithm efficiently …
In distributed machine learning, data is dispatched to multiple machines for processing. Motivated by the fact that similar data points often belong to the same or similar classes, and more generally, classification rules of high accuracy tend to be "locally simple but globally complex" (Vapnik & Bottou 1993), we propo…
Gradient descent on ReLU networks with square loss implicitly favors balanced weights.
We study high-dimensional distribution learning in an agnostic setting where an adversary is allowed to arbitrarily corrupt an -fraction of the samples. Such questions have a rich history spanning statistics, machine learning and theoretical computer science. Even in the most basic settings, the only known…