Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

95191286381 · Jun 202019922001200920172026
48 results for group size

Proposes a differentiable hypergeometric distribution for learning group importance.

problem Learning the sizes of subsets in applications like clustering and weakly-supervised learning.
method Introduces a reparameterizable hypergeometric distribution to model group sizes and learn their relative importance.
result Outperforms previous methods in weakly-supervised learning and clustering.

We prove that the rank (that is, the minimal size of a generating set) of lattices in a general connected Lie group is bounded by the co-volume of the projection of the lattice to the semi-simple part of the group. This was proved by Gelander for semi-simple Lie groups and by Mostow for solvable Lie groups. Here we con…

2019-03-12abs ↗pdf ↗

The paper analyzes how contagion affects the survival probability of investment groups in microfinance.

problem The impact of contagion on the survival probability of investment groups in microfinance.
method A probabilistic approach to compute group survival probability with and without contagion effects.
result In homogeneous groups, including more members increases the probability of eventual default to 1.

For nn at least 7 and nn equal to 5, we give generating sets of size 2 for the commutator subgroup of the braid group on nn strands. These generating sets are of the smallest possible cardinality. For nn equal to 4 or 6, we give generating sets of size three. We also prove that the commutator subgroup of the braid …

2019-10-15abs ↗pdf ↗

We present a novel approach which is able to explore the configuration of grouped convolutions within neural networks. Group-size Series (GroSS) decomposition is a mathematical formulation of tensor factorisation into a series of approximations of increasing rank terms. GroSS allows for dynamic and differentiable selec…

2019-12-02abs ↗pdf ↗

Develops a new model to predict training dynamics of large language models.

problem Lack of mechanistic understanding of training dynamics in large language models.
method A first-principles reduced-order model of training dynamics, predicting group-size invariance and stability thresholds.
result Closed-form model predicts training dynamics with high accuracy and provides new diagnostics.

Throwing away data can improve worst-group error in imbalanced datasets.

problem Improving worst-group accuracy in imbalanced datasets.
method Leveraging extreme value theory to analyze the tails of data distributions and their impact on classifier performance.
result Throwing away data restores geometric symmetry in classifiers, improving worst-group generalization.

We derive a lower bound on the size of finite non-cyclic quotients of the braid group that is superexponential in the number of strands. We also derive a similar lower bound for nontrivial finite quotients of the commutator subgroup of the braid group.

2019-10-16abs ↗pdf ↗

Estimates sample size for subgroup analysis in randomized experiments.

problem Determining sample size for accurate subgroup analysis.
method Turns inference problem into simultaneous inference, calculates sample size based on confidence level and margin of error.
result Allows inversion of sample size to feasible number of treatment arms or partition complexity.

Reduced sample complexity for group-invariant distributions.

problem Improving sample complexity for estimating divergences of group-invariant distributions.
method Quantified reduction in sample complexity for Wasserstein-1 metric and Lipschitz-regularized α-divergences under finite and infinite groups.
result Sample complexity reduction proportional to group size for finite groups, and convergence rate depends on intrinsic dimension for infinite groups.

Using the classification of transitive groups we classify indecomposable quandles of size <36. This classification is available in Rig, a GAP package for computations related to racks and quandles. As an application, the list of all indecomposable quandles of size <36 not of type D is computed.

2011-05-26abs ↗pdf ↗

Maximizes filling systems on surfaces with given boundary components.

problem Finding the maximum size of filling systems on surfaces with specific boundary conditions.
method Analyzing the structure of filling systems and their complements.
result The maximum size of a filling system on a surface of genus g with 1 ≤ b ≤ 2g-2 boundary components is 2g + b - 1.

Proving a conjecture of Dennis Johnson, we show that the Torelli subgroup of the mapping class group has a finite generating set whose size grows cubically with respect to the genus of the surface. Our main tool is a new space called the handle graph on which the Torelli group acts cocompactly.

2011-06-16abs ↗pdf ↗

Study shows pooling scores for conformal prediction distorts group coverage.

problem Pooling scores for conformal prediction distorts group coverage.
method Derived conservation law and lower bound, demonstrated tension between fairness definitions, quantified trade-off between policies.
result Pooling scores for conformal prediction distorts group coverage.

The study examines how equivariance in networks affects generalization error using PAC-Bayesian bounds.

problem Understanding how equivariance in networks impacts generalization error.
method Utilized PAC-Bayesian analysis for equivariant networks, deriving norm-based bounds for generalization error.
result The bound indicates that using larger group size in the model improves generalization error.

A new bootstrapping method reduces key sizes and runtime in FHE.

problem Large plaintext evaluation in FHE increases bootstrapping complexity.
method New polynomial vector representation and monic monomial permutation matrices.
result Polynomial factor improvement in key size and constant factor in runtime.

E2GC optimizes energy efficiency in DNNs by balancing computational and data movement costs.

problem Imbalance between computational complexity and data reuse in GConv leads to suboptimal energy efficiency.
method Developed an optimum group size model and proposed E2GC module with constant group size.
result E2GC modules improve energy efficiency by 10.8% and 4.73% on P100 and P4000 GPUs, respectively.

New algorithm groups variables by ancestral relationships to improve causal graph estimation accuracy.

problem Difficulty in estimating causal graphs with small sample sizes relative to variables.
method CAG algorithm groups variables based on ancestral relationships, reducing complexity and improving accuracy.
result CAG outperforms existing methods in estimation accuracy and computation time.

New bounds on sample size for identifying mixture models with grouped samples.

problem Identifying mixture models with minimal sample size.
method Generalized identifiability bounds for mixture models with grouped samples.
result Identifiability with (2m1)/(k1)(2m-1)/(k-1) samples per group, with no improvement possible.

MRI image quality affects statistical and predictive analysis of brain morphology.

problem Impact of MRI image quality on statistical and predictive analysis of brain morphology.
method Systematic testing of image quality on univariate statistics and machine learning classification using three large datasets.
result Low-quality MRI data significantly affects detecting significant sex/gender differences in smaller samples, but not in larger ones.

We present reconstruction algorithms for smooth signals with block sparsity from their compressed measurements. We tackle the issue of varying group size via group-sparse least absolute shrinkage selection operator (LASSO) as well as via latent group LASSO regularizations. We achieve smoothness in the signal via fusion…

2013-09-10abs ↗pdf ↗

Group Shapley evaluates feature groups in business data, improving explainability in AI.

problem Evaluating the importance of feature groups in business and economic data.
method Developed Group Shapley and a significance testing procedure based on chi-square approximation.
result Market-related variables are identified as the most influential feature group.

Reflective of income and wealth distributions, philanthropic gifting appears to follow an approximate power-law size distribution as measured by the size of gifts received by individual institutions. We explore the ecology of gifting by analysing data sets of individual gifts for a diverse group of institutions dedicat…

2013-07-08abs ↗pdf ↗

Study identifies negative data externalities affecting model performance on specific groups.

problem Negative data externalities on group performance in machine learning models.
method Characterized and detected data-model inefficiencies, focusing on specific types of externalities.
result Negative data externalities can lower model performance on specific sub-groups, even with larger datasets.

The paper explores properties of continuous actions on manifolds, proving bounds on subgroup size and fixed points.

problem Properties of continuous finite group actions on topological manifolds.
method Analyzes properties including Jordan property and almost fixed point property, proving bounds on subgroup size.
result Existence of a constant C such that for any continuous action of a finite group G on a manifold X, there is a subgroup H with [G:H] ≤ C and a fixed point.

In this article, we study connections between representation theory and efficient solutions to the conjugacy problem on finitely generated groups. The main focus is on the conjugacy problem in conjugacy separable groups, where we measure efficiency in terms of the size of the quotients required to distinguish a distinc…

2013-12-04abs ↗pdf ↗

The paper improves A/B testing for non-Gaussian data, ensuring reliable results with large sample sizes.

problem Inaccurate A/B testing results due to non-normal data and unequal sample sizes.
method Derives explicit formulas for minimum sample size and introduces an Edgeworth-based correction.
result Corrected method improves reliability of A/B testing in real-world conditions.

We present a plausible micro-founded model for the previously postulated power law finite time singular form of the crash hazard rate in the Johansen-Ledoit-Sornette model of rational expectation bubbles. The model is based on a percolation picture of the network of traders and the concept that clusters of connected tr…

2016-01-28abs ↗pdf ↗

In this paper we consider the problem of grouped variable selection in high-dimensional regression using 1q\ell_1-\ell_q regularization (1q1\leq q \leq \infty), which can be viewed as a natural generalization of the 12\ell_1-\ell_2 regularization (the group Lasso). The key condition is that the dimensionality pnp_n can…

2008-02-11abs ↗pdf ↗

A new method improves AI fairness assessment by estimating performance across intersectional subgroups.

problem Limited evaluation of AI systems across intersectional subgroups due to small sample sizes.
method Structured regression approach to disaggregated evaluation.
result Our method yields more accurate performance estimates, especially for small subgroups.