Reduced sample complexity for group-invariant distributions.
problem Improving sample complexity for estimating divergences of group-invariant distributions.
method Quantified reduction in sample complexity for Wasserstein-1 metric and Lipschitz-regularized α-divergences under finite and infinite groups.
result Sample complexity reduction proportional to group size for finite groups, and convergence rate depends on intrinsic dimension for infinite groups.
Improved multi-group learning with group-realizable concepts.
problem Enhancing multi-group learning efficiency.
method Empirical risk minimization over group-realizable concepts.
result Improved sample complexity in group-realizable settings.
New method for group testing robust to errors in group membership specifications.
problem Errors in specifying group memberships during group testing.
method Debiased Robust Lasso Test Method (DRLT) based on Lasso debiasing.
result Extends LASSO bias mitigation to handle group membership specification errors.
Bayesian design improves accuracy without extra cost.
problem Nested inference in complex systems limits BED accuracy and efficiency.
method Grouped geometric pooled posterior with EKI formulation.
result Improved accuracy and stable estimators at comparable cost.
New bounds on sample size for identifying mixture models with grouped samples.
problem Identifying mixture models with minimal sample size.
method Generalized identifiability bounds for mixture models with grouped samples.
result Identifiability with (2m−1)/(k−1) samples per group, with no improvement possible. A new CVaR test reduces group performance disparity detection complexity.
problem Detecting performance disparities across multiple sensitive groups in ML models.
method Conditional Value-at-Risk (CVaR) testing to reduce sample complexity.
result Sample complexity reduced exponentially to be at most the square root of the number of groups.
Multi-group learners suffer a penalty in transductive learning.
problem The penalty on multi-group learners in transductive learning.
method Analyzing the relationship between the number of groups and the error rate.
result The penalty can increase linearly with the number of groups, up to the square-root of the sample size.
We give a proof of the sublinear tracking property for sample paths of random walks on various groups acting on spaces with hyperbolic-like properties. As an application, we prove sublinear tracking in Teichmueller distance for random walks on mapping class groups, and on Cayley graphs of a large class of finitely gene…
The paper analyzes how minority group imbalance affects neural network performance.
problem The impact of minority group imbalance on neural network performance.
method Formulated group imbalance problem with Gaussian Mixture Model, quantified sample complexity, convergence rate, and testing performance.
result Increasing the minority group fraction does not necessarily improve the generalization performance of the minority group.
New sampling method on Lie groups converges quickly.
problem Sampling on non-Euclidean Lie groups.
method Kinetic Langevin dynamics with noise added.
result Exponential convergence rate proved under W2 distance. FaiREE provides fair classification with guarantees for small datasets.
problem Fairness in classification often requires large sample sizes and distributional assumptions.
method FaiREE offers finite-sample and distribution-free fairness guarantees.
result FaiREE achieves optimal accuracy and satisfies various fairness notions.
With the rapid adoption of machine learning systems in sensitive applications, there is an increasing need to make black-box models explainable. Often we want to identify an influential group of training samples in a particular test prediction for a given machine learning model. Existing influence functions tackle this…
Deep metric learning has yielded impressive results in tasks such as clustering and image retrieval by leveraging neural networks to obtain highly discriminative feature embeddings, which can be used to group samples into different classes. Much research has been devoted to the design of smart loss functions or data mi…
We propose a semantic segmentation model that exploits rotation and reflection symmetries. We demonstrate significant gains in sample efficiency due to increased weight sharing, as well as improvements in robustness to symmetry transformations. The group equivariant CNN framework is extended for segmentation by introdu…
New kernels on symmetric groups enable efficient Gaussian process sampling.
problem Efficiently modeling and sampling on symmetric groups.
method Introduced power sum kernels and methods for efficient calculation and sampling.
result Polynomial computational complexity for sampling Gaussian processes.
Study shows how to reduce data needed for learning under geometric constraints.
problem Learning high-dimensional data with geometric priors.
method Spherical harmonic decompositions and kernel methods for invariance and geometric stability.
result Improvements in sample complexity by leveraging group invariance, with asymptotic behavior depending on spectral properties.
Researchers study fairness-accuracy tradeoffs in predictive models for multiple groups.
problem Understanding the tradeoff between fairness and accuracy in models serving multiple demographic groups.
method Characterizing the fairness-accuracy (FA) Pareto frontier, approximating it from limited data, and bounding the worst-case gap.
result Derivation of worst-case-optimal estimators and uniform finite-sample bounds for the entire FA frontier.
New method GSAT improves robustness against structured perturbations.
problem Structured perturbations in biological data.
method Formulates GSAT as a non-convex concave minimax optimization problem and solves it with GDADMM.
result Improves robustness against group-sparse and rank-constrained perturbations.
Generative model learns from simpler distributions on Lie groups.
problem Learning from complex Lie group data.
method Substituting exponential curves for line segments on Lie groups.
result Simple, intrinsic, and fast implementation for generative modelling.
New algorithms allocate sampling budget to estimate group means without exploration.
problem Allocate sampling budget to estimate means of multiple groups.
method Design exploration-free non-adaptive and adaptive algorithms.
result Prove tighter regret bounds for multi-group mean estimation.
Procedure groups nonparametric regression curves automatically.
problem Determining groups of nonparametric regression curves when curves are numerous.
method Automatic selection of group number through testing procedure.
result Groups of nonparametric regression curves exist in tunnel geometry.
EDGI improves sample efficiency and generalization in tasks with spatial and temporal symmetries.
problem Sample inefficiency and poor generalization in tasks with geometric symmetries.
method Equivariant Diffuser framework, SE(3)xZxSn-equivariant diffusion model.
result EDGI is more sample efficient and generalizes better than non-equivariant models.
Unified framework for fair classification with group-blindness/awareness guarantees.
problem Challenges in enforcing fairness and group-blindness in binary classification.
method Unified framework based on post-processing procedure, applicable to various group fairness notions.
result Minimax rate-optimality of the proposed algorithm with controlled excess risk.
A fast method for discrete OT with group-sparse regularization for class label preservation.
problem Efficiently measuring the distance between two discrete distributions with class labels.
method Fast discrete OT with group-sparse regularizers using gradient-based algorithms.
result Up to 8.6 times faster than original method without degrading accuracy.
Most machine learning algorithms, such as classification or regression, treat the individual data point as the object of interest. Here we consider extending machine learning algorithms to operate on groups of data points. We suggest treating a group of data points as an i.i.d. sample set from an underlying feature dis…
Proposes counterfactual explanations for deep two-sample tests on high-dimensional data.
problem Limited interpretability of deep two-sample tests on high-dimensional data.
method Combines diffusion autoencoder and pretrained deep two-sample test model to generate counterfactuals.
result Counterfactual transformations increase p-values, indicating closer distribution similarity.
Improved SV estimator for efficient data valuation.
problem Computational inefficiency in Shapley value estimation.
method Group Testing-based SV estimator with improvements.
result Enhanced asymptotic sample complexity and insights into challenges.
New method reduces privacy impact on model accuracy for underrepresented groups.
problem Privacy mechanisms disproportionately affect underrepresented groups in machine learning models.
method Proposes DPSGD-F, a modified DPSGD that adjusts group contributions based on clipping bias.
result DPSGD-F removes disparate impact of differential privacy on model accuracy for protected groups.
Datasets containing large samples of time-to-event data arising from several small heterogeneous groups are commonly encountered in statistics. This presents problems as they cannot be pooled directly due to their heterogeneity or analyzed individually because of their small sample size. Bayesian nonparametric modellin…
New group testing method uses Belief Propagation for accurate screening.
problem Efficiently identifying infected samples in large groups with minimal tests.
method Belief Propagation algorithm for inference in group testing schemes.
result Significantly increased accuracy of infection identification with fewer tests.
In multi-label learning, each sample is associated with several labels. Existing works indicate that exploring correlations between labels improve the prediction performance. However, embedding the label correlations into the training process significantly increases the problem size. Moreover, the mapping of the label …
Algorithm samples fair rankings to ensure individual fairness while maintaining group fairness.
problem Fair ranking tasks with group fairness constraints and uncertainty in item utilities.
method Efficient algorithm that samples rankings from an individually-fair distribution ensuring group fairness.
result Expected utility of output ranking is at least α times optimal fair solution, where α depends on utilities and constraints.
Let S=Γ\H be a hyperbolic surface of finite topological type, such that the Fuchsian group Γ≤PSL2(R) is non-elementary, and consider any generating set S of Γ. When sampling by an n-step random walk in π1(S)≅Γ with each step given by an element…
Paper classifies totally symmetric sets in groups and bounds their sizes.
problem Understanding homomorphisms between groups using totally symmetric sets.
method Full classifications and size bounds for totally symmetric sets in various groups.
result Derives restrictions on homomorphisms between certain groups.
Optimizes group testing for COVID-19 to reduce test numbers.
problem Minimizing tests for accurate infection detection.
method Bayesian approach with genetic algorithms and sub-modularity.
result Greedy-adaptive method provides theoretical guarantees.
We study sparse group Lasso for high-dimensional double sparse linear regression, where the parameter of interest is simultaneously element-wise and group-wise sparse. This problem is an important instance of the simultaneously structured model -- an actively studied topic in statistics and machine learning. In the noi…
The paper tackles sampling biases by ensuring minority groups are adequately represented in training data.
problem Sampling biases in training data lead to algorithmic biases in machine learning systems.
method The paper presents adaptive sampling methods to determine if it's possible to assemble a representative dataset from given data sources.
result The methods presented can determine with high confidence if a representative dataset can be assembled from given data sources.
Group factor analysis (GFA) methods have been widely used to infer the common structure and the group-specific signals from multiple related datasets in various fields including systems biology and neuroimaging. To date, most available GFA models require Gibbs sampling or slice sampling to perform inference, which prev…
Proposes active sampling for improving fairness in machine learning.
problem Improving fairness in machine learning models, especially for disadvantaged groups.
method Simple active sampling and reweighting strategies for min-max fairness.
result Proves the rate of convergence to a min-max fair solution for convex problems.
Noisy Pooled PCR tests large groups more efficiently.
problem Efficiently test large populations for viral infections.
method Converts group testing to a linear inverse problem with a message passing algorithm.
result Estimates patient illness status with fewer pooled measurements.
Optimal classification requires choosing the right group symmetries, contrary to intuition.
problem Improving binary classification performance by selecting appropriate group symmetries.
method Developed a theoretical framework for designing group equivariant neural networks.
result Optimal classification performance is achieved by selecting the appropriate subgroups of symmetries, not the largest equivariant groups.
Paper tackles group robustness with partially labeled data.
problem Learning invariant representations from datasets with spurious correlations.
method Constructs a constraint set and derives a high probability bound for group assignment. Proposes an optimization algorithm for worst-off group assignments.
result Improvements in minority group's performance while preserving overall accuracy.
Study learns convolution operators on compact Abelian groups using regularization.
problem Learning convolution operators on compact Abelian groups.
method Regularization-based approach with ridge regression estimator.
result Characterizes the accuracy of the estimator in terms of finite sample bounds.
Proposes a robust optimization method for selecting grouped variables robustly.
problem Selecting grouped variables under data perturbations for regression and classification.
method Distributionally Robust Optimization (DRO) with Wasserstein uncertainty set.
result Coefficients in the same group converge to the same value as sample correlation approaches 1.
Develops non-parametric tests for group symmetry in data.
problem Lack of statistical tests for group symmetry in data.
method Formulates and implements non-parametric tests for distributional symmetry under specified groups.
result Develops tests for conditional invariance/equivariance and applies them to real-world data.
Improved sampling for network community detection.
problem Inefficient sampling from network partition posterior distributions.
method Merge-split Markov chain Monte Carlo for efficient sampling.
result Significantly improved mixing time and correct sampling.
Path signatures adapted for Lie groups improve action recognition in computer vision.
problem Improving action recognition in computer vision with geometric constraints.
method Lifting path signatures to Lie groups and proving universality and characteristic property.
result Path signatures on Lie groups provide comparable performance to shallow learning approaches in action recognition.
We treat the problem of estimation of orientation parameters whose values are invariant to transformations from a spherical symmetry group. Previous work has shown that any such group-invariant distribution must satisfy a restricted finite mixture representation, which allows the orientation parameter to be estimated u…