This work improves group data analysis using modified tensor decompositions.
problem Improving group data analysis models for better signal modeling.
method Introduces a new generalization of block tensor decomposition for group data analysis.
result Demonstrates improved performance in multilabel classification and clustering tasks.
Factor analysis provides linear factors that describe relationships between individual variables of a data set. We extend this classical formulation into linear factors that describe relationships between groups of variables, where each group represents either a set of related variables or a data set. The model also na…
Efficient CVI for NGFA improves GFA inference for large-scale data.
problem Inference limitations in GFA models for large-scale data.
method Collapsed variational inference for nonparametric Bayesian GFA.
result CVI algorithm effectively approximates NGFA posterior in collapsed space.
We introduce a factor analysis model that summarizes the dependencies between observed variable groups, instead of dependencies between individual variables as standard factor analysis does. A group may correspond to one view of the same set of objects, one of many data sets tied by co-occurrence, or a set of alternati…
Study of circle configurations in the plane, proving aspherical space and computing fundamental groups.
problem Understanding the space of configurations of circles in the plane.
method Proved the space is aspherical and computed fundamental groups of its components.
result Fundamental groups are iterated semidirect products of braid groups, with structure dictated by a finite rooted tree.
New method groups similar functional covariates for better modeling.
problem Analyzing functional covariates with similar shapes.
method Coefficient shape alignment regularization approach.
result True grouping structure can be accurately identified under certain conditions.
In this paper we consider the problem of group invariant subspace clustering where the data is assumed to come from a union of group-invariant subspaces of a vector space, i.e. subspaces which are invariant with respect to action of a given group. Algebraically, such group-invariant subspaces are also referred to as su…
BSD is a Bayesian framework for analyzing neural spectral data.
problem Challenges in statistical analysis and group-level comparisons of neural power spectra.
method Bayesian Spectral Decomposition (BSD) for parametric models of neural spectra.
result BSD outperforms existing methods in model selection and parameter estimation.
New Morse theory applied to Vietoris-Rips complexes for topological data analysis and geometric group theory.
problem Understanding homotopy types of Vietoris-Rips complexes for metric spaces.
method Generalization of Bestvina-Brady discrete Morse theory applied to Vietoris-Rips complexes.
result Metric criteria (Morse and Link) to deduce homotopy types of VRt(X). The paper studies geometric properties of group equivariant operators and their Riemannian structure.
problem Understanding the geometric structure of group equivariant operators.
method Endowing the space of group equivariant non-expansive operators with a Riemannian manifold structure and using gradient descent methods.
result Gradient descent methods can be applied to minimize cost functions on the space of group equivariant non-expansive operators.
Proposes Fair Archetypal Analysis to reduce fairness concerns in data representation.
problem Inadvertent encoding of sensitive attributes in Archetypal Analysis.
method Integrates fairness regularization into Archetypal Analysis and its nonlinear extension.
result Reduces group separability without significantly compromising explained variance.
Improved SV estimator for efficient data valuation.
problem Computational inefficiency in Shapley value estimation.
method Group Testing-based SV estimator with improvements.
result Enhanced asymptotic sample complexity and insights into challenges.
A new method clusters mixed-type data tables effectively.
problem Clustering data with mixed types (numerical and categorical).
method Two-step approach: binarize mixed data, then co-cluster.
result Shows improved clustering of mixed-type data compared to MCA.
Data augmentation can achieve the same statistical benefits as full augmentation up to an approximation error.
problem Data augmentation in learning problems
method Using Fourier analysis and representation theory of finite groups
result Partial data augmentation achieves the same minimax rates as full augmentation
Many data-driven approaches exist to extract neural representations of functional magnetic resonance imaging (fMRI) data, but most of them lack a proper probabilistic formulation. We propose a group level scalable probabilistic sparse factor analysis (psFA) allowing spatially sparse maps, component pruning using automa…
Deep generative models have recently yielded encouraging results in producing subjectively realistic samples of complex data. Far less attention has been paid to making these generative models interpretable. In many scenarios, ranging from scientific applications to finance, the observed variables have a natural groupi…
We propose a generative model of a group EEG analysis, based on appropriate kernel assumptions on EEG data. We derive the variational inference update rule using various approximation techniques. The proposed model outperforms the current state-of-the-art algorithms in terms of common pattern extraction. The validity o…
Contrastive ICA identifies features in experimental groups relative to controls.
problem Jointly analyzing experimental and control datasets to identify salient features.
method Developed contrastive ICA (cICA) using tensor decomposition.
result cICA identifies patterns and visualizes data effectively, outperforming existing methods.
Interactive DR framework for comparing datasets.
problem Limited flexibility in existing DR methods for comparative analysis.
method Unified linear comparative analysis (ULCA) with interactive optimization and visualization.
result ULCA and optimization algorithm improve comparative analysis efficiency and flexibility.
Research activities of Kyoto Econophysics Group is reviewed. Strong emphasis has been placed on real economy. While the initial stage of research was a first high-definition data analysis on personal income, it soon progressed to firm dynamics, growth rate distribution and establishment of Pareto's law and Gibrat's law…
Flexible framework assesses multilevel data group heterogeneity.
problem Multilevel data structure complicates model selection.
method Flexible framework for assessing differences between levels of grouping variables.
result Framework reliably identifies relevant multilevel components.
Gen AI improves document understanding but not data analysis in public sector tasks.
problem Understanding the impact of Gen AI on public sector tasks.
method Pre-registered field experiment comparing Gen AI to control group performance.
result Mixed results: Gen AI improves document understanding but not data analysis.
Researchers use shape analysis to recover protein structures from Cryo-EM data.
problem Recovering the three-dimensional backbone structure of single polypeptide proteins from noisy tomographic projections.
method Shape analysis and matrix Lie group actions to deform point clouds to match 2D tomography data.
result Optimal deformations are computed to recover the three-dimensional backbone structure of proteins.
Study presents a method to induce a generalized neural network from joint group invariant functions.
problem Encoding rule of neural network internal data representation.
method Systematic method using joint group invariant function on data-parameter domain.
result Induces a generalized neural network and its inverse operator (ridgelet transform).
We study a norm for structured sparsity which leads to sparse linear predictors whose supports are unions of prede ned overlapping groups of variables. We call the obtained formulation latent group Lasso, since it is based on applying the usual group Lasso penalty on a set of latent variables. A detailed analysis of th…
Unified framework for multi-source data analysis improves network structure identification.
problem High dimensionality and heterogeneity in large-scale network data.
method msLBM framework combining multiple data sources for simultaneous grouping and connectivity analysis.
result Statistically optimal rates achieved for consensus knowledge graph learning.
Due to the escalating growth of big data sets in recent years, new Bayesian Markov chain Monte Carlo (MCMC) parallel computing methods have been developed. These methods partition large data sets by observations into subsets. However, for Bayesian nested hierarchical models, typically only a few parameters are common f…
cMCA uses contrastive learning to identify latent subgroups in political party data.
problem Identifying latent subgroups within political party data.
method Contrastive learning applied to multiple correspondence analysis (MCA).
result cMCA identifies latent subgroups not seen by traditional methods.
Extends functions on symmetric spaces to analytic functions.
problem Extending functions on symmetric spaces to analytic functions.
method Harmonic analysis on symmetric spaces and representation theory of groups.
result Proves Whitney type extension theorems for symmetric spaces.
Bayesian approach groups observations with similar effects for better inference.
problem Estimating heterogeneity across observations in political science.
method Structured sparsity framework integrated into Bayesian regression analysis.
result Method outperforms state-of-the-art methods for heterogeneous effects estimation.
As data collections become larger, exploratory regression analysis becomes more important but more challenging. When observations are hierarchically clustered the problem is even more challenging because model selection with mixed effect models can produce misleading results when nonlinear effects are not included into…
When estimating finite mixture models, it is common to make assumptions on the mixture components, such as parametric assumptions. In this work, we make no distributional assumptions on the mixture components and instead assume that observations from the mixture model are grouped, such that observations in the same gro…
New method groups genetic data into coherent topics for disease insights.
problem Analyzing large, multi-dimensional genetic data sets.
method Conditional Hierarchical Bayesian Tucker Decomposition for genetic data analysis.
result Our models are more coherent than baseline models.
We introduce coroICA, confounding-robust independent component analysis, a novel ICA algorithm which decomposes linearly mixed multivariate observations into independent components that are corrupted (and rendered dependent) by hidden group-wise stationary confounding. It extends the ordinary ICA model in a theoretical…
PCA outperforms random projections in retaining second order signals from latent groups.
problem Preserving second order structure in latent groups under unsupervised linear projections.
method Theoretical framework and quasi-exhaustive enumeration of projections.
result PCA outperforms random projections in retaining second order signals across a broad range of data-generating parameters.
Online selection of dynamic features has attracted intensive interest in recent years. However, existing online feature selection methods evaluate features individually and ignore the underlying structure of feature stream. For instance, in image analysis, features are generated in groups which represent color, texture…
Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying clustering structures. Hence removing noise variables via variable selection is necessary…
A new method speeds up factor analysis for high-dimensional data.
problem Estimating covariance parameters in high-dimensional Gaussian data with limited observations.
method Matrix-free likelihood method using implicitly restarted Lanczos and limited-memory quasi-Newton algorithms.
result Our method is faster than EM without sacrificing accuracy.
Unified analysis of multi-attribute graph learning with non-convex penalties.
problem Graph inference from multi-attribute data.
method Penalized log-likelihood objective function with ADMM and local linear approximation.
result Local consistency in support recovery and precision matrix estimation for non-convex penalties.
Estimates treatment effects in time series data with always-missing controls.
problem Lack of control group in time series data, especially during specific events.
method Recover control group in event period, account for confounders and temporal dependencies.
result Robust estimation of control group's potential outcome and accurate predicted holiday effect.
Method estimates group structure in panel data using variance information.
problem Estimating group structure in panel data with unknown groups.
method Proposes a method to estimate unobserved groupings for panel data models using variance information.
result Superior performance compared to existing methods in simulations and empirical applications.
Paper proposes BSP to find stable bimodules of cross-correlated features.
problem Identify groups of features from two data types with strong cross-correlation.
method Iterative-testing based bimodule search procedure (BSP).
result BSP outperforms existing methods in detecting stable bimodules.
From the homotopy groups of two cubic spherical 3-manifolds we construct the isomorphic groups of deck transformations acting on the 3-sphere. These groups become the cyclic group of order eight and the quaternion group respectively. By reduction of representations from the orthogonal group to the identity representati…
Develops MGQDA for multi-group classification with theoretical guarantees and practical applications.
problem Complex multi-group classification problems with nonlinear decision boundaries and group-specific covariance patterns.
method MGQDA, a method based on quadratic discriminant analysis that projects predictors onto a lower-dimensional subspace.
result MGQDA achieves competitive or improved predictive performance compared to existing methods.
Motivation: Modelling methods that find structure in data are necessary with the current large volumes of genomic data, and there have been various efforts to find subsets of genes exhibiting consistent patterns over subsets of treatments. These biclustering techniques have focused on one data source, often gene expres…
A novel method for estimating group-representative functional networks from multi-subject fMRI data.
problem Estimating common neuronal characteristics in a population from multi-subject fMRI data.
method Two-phase approach: clustering-based ICA for component maps, MAP-MRF labeling for group-representative map estimation.
result Demonstrated the viability of the proposed method in extracting group-representative functional networks from simulated fMRI data.
The study visualizes Spanish fish and meat processing companies using financial, environmental, and social ratios.
problem Mapping financial, environmental, and social performance of Spanish processing companies.
method Used compositional data and principal-component analysis biplot for statistical analysis.
result Identified clusters of companies with similar financial, environmental, and social performance.
The paper develops a new mathematical framework for group-equivariant operators in machine learning.
problem Developing a robust mathematical framework for group-equivariant operators in machine learning.
method The paper introduces group-equivariant non-expansive operators (GENEOs) and studies their topological and metric properties.
result The space of GENEOs is compact and convex, providing fundamental guarantees for machine learning.