Algorithm improves binary classification of biased grouped data.
problem Improving binary classification for biased, grouped data.
method Assumes partition-projected class-conditional invariance across groups and derives a semi-supervised algorithm to learn a group-aware classifier.
result Demonstrates improved area under the ROC curve compared to baselines.
FGSV defends against shell company attacks in group data valuation.
problem Shell company attacks on group-level data valuation.
method Developed a provably fast and accurate approximation algorithm for FGSV.
result Empirical results show significant improvement in computational efficiency and accuracy.
Throwing away data can improve worst-group error in imbalanced datasets.
problem Improving worst-group accuracy in imbalanced datasets.
method Leveraging extreme value theory to analyze the tails of data distributions and their impact on classifier performance.
result Throwing away data restores geometric symmetry in classifiers, improving worst-group generalization.
This work provides statistical guarantees for GANs that are invariant to certain group symmetries.
problem Learning group-invariant distributions efficiently.
method Study of group-invariant GANs and their performance guarantees.
result Group-invariant GANs require fewer samples and have a reduced discriminator approximation error.
EquivCNP learns group symmetries for conditional data.
problem Learning conditional models with data symmetries.
method Group equivariant decomposition and Lie group convolutional layers.
result EquivCNP achieves comparable performance and zero-shot generalization.
Paper develops robust methods for panel data with latent groups, improving inference under group separation violations.
problem Inference in latent group panel models under group separation violations.
method Selective conditional inference approach to derive conditional distribution of coefficients given estimated group structure.
result Valid inference under violations of group separation, superior to traditional asymptotic methods.
EbC learns equivariant embeddings from unlabeled group actions.
problem Learning equivariant embeddings from unlabeled group actions.
method Equivariance by Contrast (EbC) method to learn equivariant embeddings from observation pairs (y,g⋅y). result High-fidelity equivariance in latent space for diverse groups.
Flexible framework assesses multilevel data group heterogeneity.
problem Multilevel data structure complicates model selection.
method Flexible framework for assessing differences between levels of grouping variables.
result Framework reliably identifies relevant multilevel components.
Feature noise causes loss discrepancies across groups even with equal data.
problem Loss discrepancies observed in learning procedures across different groups.
method Characterized the effect of feature noise on loss discrepancy in linear regression.
result Feature noise leads to loss discrepancy even when groups have equal data.
Proposes a boosting framework for sparsity in grouped covariates.
problem Sparsity and selection bias in grouped covariates.
method Component-wise and group-wise gradient boosting with adjusted degrees of freedom.
result Reduces bias and improves predictability in variable selection.
Clustering aims to divide a set of points into groups. The current paradigm assumes that the grouping is well-defined (unique) given the probability model from which the data is drawn. Yet, recent experiments have uncovered several high-dimensional datasets that form different binary groupings after projecting the data…
A fast method estimates group-adaptive elastic net penalties using co-data.
problem Computational inefficiency in estimating group-adaptive elastic net penalties.
method Derive low-dimensional representation of Taylor approximation for marginal likelihood and its derivative for group-adaptive ridge penalties; approximate elastic net marginal likelihood by ridge; transform ridge penalties to elastic net penalties.
result Significantly decreases computation time and outperforms other methods.
Sparse Singular Value Decomposition (SVD) models have been proposed for biclustering high dimensional gene expression data to identify block patterns with similar expressions. However, these models do not take into account prior group effects upon variable selection. To this end, we first propose group-sparse SVD model…
Clustering groups similar data points into clusters.
problem Grouping similar data points into coherent clusters.
method Different clustering methods based on similarity and data representations.
result Various clustering methods exist.
This paper improves model robustness to underrepresented groups using ranking metrics and reweighting.
problem Underrepresented groups suffer from low accuracy in models trained via ERM.
method Proposes Discounted Cumulative Gain (DCG) and Discounted Rank Upweighting (DRU) methods.
result Models trained with DRU show superior generalization to unseen groups.
Method estimates group structure in panel data using variance information.
problem Estimating group structure in panel data with unknown groups.
method Proposes a method to estimate unobserved groupings for panel data models using variance information.
result Superior performance compared to existing methods in simulations and empirical applications.
Study of circle configurations in the plane, proving aspherical space and computing fundamental groups.
problem Understanding the space of configurations of circles in the plane.
method Proved the space is aspherical and computed fundamental groups of its components.
result Fundamental groups are iterated semidirect products of braid groups, with structure dictated by a finite rooted tree.
Study identifies negative data externalities affecting model performance on specific groups.
problem Negative data externalities on group performance in machine learning models.
method Characterized and detected data-model inefficiencies, focusing on specific types of externalities.
result Negative data externalities can lower model performance on specific sub-groups, even with larger datasets.
New method groups similar functional covariates for better modeling.
problem Analyzing functional covariates with similar shapes.
method Coefficient shape alignment regularization approach.
result True grouping structure can be accurately identified under certain conditions.
New method identifies differences between groups in low-dimensional data representations.
problem Identifying meaningful differences between groups in low-dimensional data representations.
method Introduce Global Counterfactual Explanation (GCE) and Transitive Global Translations (TGT) for computing GCEs.
result TGT identifies sparse, accurate explanations that match real data patterns.
New algorithm provides robust uncertainty quantification without parameter tuning.
problem Real-world machine learning predictors need reliable uncertainty quantification.
method Parameter-free, group-conditional online prediction algorithm.
result Achieves best group-conditional coverage guarantees.
Study automorphism groups of Inoue surfaces using quadratic number fields.
problem Understanding automorphism groups of Inoue surfaces.
method Construction and description of automorphism groups using quadratic number fields.
result Automorphism groups of Inoue surfaces S(+)/S(−) described in terms of quadratic number fields. New neural networks for non-commutative data.
problem No existing neural networks suitable for non-commutative data.
method Developed compact matrix quantum group equivariant neural networks.
result Characterized weight matrices for easy compact matrix quantum groups.
Proposes a two-stage method for selecting correlated predictors in high-dimensional data.
problem Selecting correlated predictors in high-dimensional data with unknown group structures.
method Two-stage approach: variable clustering followed by group selection.
result The two-stage method improves prediction accuracy and active predictor selection.
Improves fairness in machine learning by adding underrepresented group data.
problem Machine learning biases across subgroups due to under-representation or societal biases.
method Data augmentation via pairwise mixup across subgroups to balance subpopulations.
result Achieves fair outcomes with robust if not improved accuracy.
This work is devoted to elaboration on the idea to use block term decomposition for group data analysis and to raise the possibility of modelling group activity with (Lr, 1) and Tucker blocks. A new generalization of block tensor decomposition was considered in application to group data analysis. Suggested approach was…
Generalizes CNNs for Lie group equivariance across various data types.
problem Equivariance to transformations like rotations for non-image data.
method Constructs equivariant convolutional layers for Lie groups.
result Models conserve linear and angular momentum in Hamiltonian systems.
Paper tackles group robustness with partially labeled data.
problem Learning invariant representations from datasets with spurious correlations.
method Constructs a constraint set and derives a high probability bound for group assignment. Proposes an optimization algorithm for worst-off group assignments.
result Improvements in minority group's performance while preserving overall accuracy.
In this paper we consider the problem of group invariant subspace clustering where the data is assumed to come from a union of group-invariant subspaces of a vector space, i.e. subspaces which are invariant with respect to action of a given group. Algebraically, such group-invariant subspaces are also referred to as su…
Paper introduces a new robust method for estimating Pareto tail index from grouped data.
problem Limited robust methods for estimating Pareto tail index from grouped data.
method Method of Truncated Moments (MTuM)
result Inferential justification and validation of MTuM through simulation study.
A novel framework synthesizes treatment data across sites using optimal transport.
problem Estimating treatment effects across different sites with varying conditions.
method Distributional causal inference, Optimal Transport for alignment of control group distributions.
result Synthetic treatment group data aligns with true target distribution under general conditions.
Study of wild mapping class groups and their cabled braids.
problem Understanding the structure of wild mapping class groups and their cabled versions.
method Define and study generalizations of pure g-braid groups, establish product decompositions, and introduce fission trees. result Obtain cabled versions of braid groups, related to braid operads.
We study a norm for structured sparsity which leads to sparse linear predictors whose supports are unions of prede ned overlapping groups of variables. We call the obtained formulation latent group Lasso, since it is based on applying the usual group Lasso penalty on a set of latent variables. A detailed analysis of th…
We demonstrate how a 3-manifold, a Heegaard diagram, and a group presentation can each be interpreted as a pair of signed permutations in the symmetric group Sd. We demonstrate the power of permutation data in programming and discuss an algorithm we have developed that takes the permutation data as input and determi…
The paper shows how to use proxy attributes for fairness in machine learning models with missing sensitive group data.
problem Measuring and enforcing fairness in machine learning models with incomplete sensitive group data.
method Using proxy-sensitive attributes to derive upper bounds on multiaccuracy and multicalibration violations and adjust models to satisfy these fairness notions.
result Provable upper bounds on multiaccuracy and multicalibration violations can be derived using proxy-sensitive attributes in the absence of sensitive group data.
NFT learns group actions without knowing the data's structure.
problem Learning equivariant representations without knowing the data's structure.
method Neural Fourier Transform (NFT) framework for learning latent linear actions of groups.
result Linear equivariant features are equivalent to group invariants.
This paper develops a theory for group Lasso using a concept called strong group sparsity. Our result shows that group Lasso is superior to standard Lasso for strongly group-sparse signals. This provides a convincing theoretical justification for using group sparse regularization when the underlying group structure is …
G-FIGS uses instance weights to create interpretable models from diverse data.
problem Generalizing to diverse data distributions while maintaining interpretability.
method Estimates group membership probabilities, uses as instance weights in FIGS to grow decision trees.
result Achieves state-of-the-art prediction performance and maintains interpretability.
GCAO improves clustering of high-dimensional data by grouping low-density boundary points.
problem Stability and accuracy of clustering in high-dimensional, non-uniform data.
method Group-level optimization with gravitational attraction and optimization.
result GCAO outperforms 11 clustering methods on multiple datasets.
The paper shows how data augmentation and regularization can enforce group equivariance in machine learning models.
problem Improving model performance by leveraging known symmetries in machine learning tasks.
method Training with data augmentation and regularization to enforce group equivariance.
result Equivariance of the trained model can be achieved through training on augmented data in tandem with regularization.
Selective inference for group lasso estimators across various distributions and covariates.
problem Developing selective inference methods for group lasso estimators.
method Randomized group-regularized optimization problem with post-selection likelihood.
result Selective point estimator and Wald-type confidence regions for regression parameters.
Unified method for CNNs to approximate equivariant maps across various groups.
problem Limited universal approximation theorems for CNNs with specific groups and settings.
method Unified approach to derive universal approximation theorems for equivariant maps by CNNs in diverse settings.
result Ability to handle non-linear equivariant maps between infinite-dimensional spaces for non-compact groups.
DFR reduces the computational cost of sparse-group lasso and adaptive sparse-group lasso.
problem Sparse-group lasso's computational expense and need for tuning.
method Dual Feature Reduction (DFR) using strong screening rules and dual norms.
result DFR drastically reduces computational cost without affecting solution optimality.
New model clusters cells and individuals, revealing genetic influences on cell types.
problem Clustering nested data with group-level and observation-level variables.
method Nested Atoms Model (NAM), Bayesian nonparametric approach.
result Identifies clusters of genetically similar individuals with homogeneous cell-type profiles.
We would like to learn a representation of the data which decomposes an observation into factors of variation which we can independently control. Specifically, we want to use minimal supervision to learn a latent representation that reflects the semantics behind a specific grouping of the data, where within a group the…
PAM models generate dependent random distributions across groups with overlapping clusters.
problem Generating dependent random distributions across multiple groups.
method Atom skipping in an infinite mixture model.
result Interpretable posterior inference of cluster exclusivity and sharing.
Factor analysis provides linear factors that describe relationships between individual variables of a data set. We extend this classical formulation into linear factors that describe relationships between groups of variables, where each group represents either a set of related variables or a data set. The model also na…
We propose the supervised hierarchical Dirichlet process (sHDP), a nonparametric generative model for the joint distribution of a group of observations and a response variable directly associated with that whole group. We compare the sHDP with another leading method for regression on grouped data, the supervised latent…