Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

58117175233 · Jun 202019922001200920182026
48 results for group-level labels

We present a new approach for transferring knowledge from groups to individuals that comprise them. We evaluate our method in text, by inferring the ratings of individual sentences using full-review ratings. This approach, which combines ideas from transfer learning, deep learning and multi-instance learning, reduces t…

2014-11-12abs ↗pdf ↗

GROOVE learns representations for weakly paired multimodal data.

problem Learning representations for high-content perturbation data with weakly paired samples.
method GroupCLIP contrastive loss integrated with an autoencoder framework.
result GROOVE performs on par with or outperforms existing approaches for cross-modal tasks.

New model clusters cells and individuals, revealing genetic influences on cell types.

problem Clustering nested data with group-level and observation-level variables.
method Nested Atoms Model (NAM), Bayesian nonparametric approach.
result Identifies clusters of genetically similar individuals with homogeneous cell-type profiles.

BSD is a Bayesian framework for analyzing neural spectral data.

problem Challenges in statistical analysis and group-level comparisons of neural power spectra.
method Bayesian Spectral Decomposition (BSD) for parametric models of neural spectra.
result BSD outperforms existing methods in model selection and parameter estimation.

The paper analyzes how minority group imbalance affects neural network performance.

problem The impact of minority group imbalance on neural network performance.
method Formulated group imbalance problem with Gaussian Mixture Model, quantified sample complexity, convergence rate, and testing performance.
result Increasing the minority group fraction does not necessarily improve the generalization performance of the minority group.

Develops a method to estimate uncertainty for group-level recommendations in matrix completion.

problem Uncertainty estimation for group-level recommendations in matrix completion.
method Structured conformal inference method combining any matrix completion algorithm.
result Stronger group-level guarantees through structured calibration.

Constellation learns group-level visual relationships for abstract reasoning.

problem Learning configurational properties of entire groups of objects.
method Introduces Constellation, a network that learns relational abstractions over static visual scenes.
result Offers a basis for abstract relational reasoning and sensory imagination.

We present a Bayesian nonparametric framework for multilevel clustering which utilizes group-level context information to simultaneously discover low-dimensional structures of the group contents and partitions groups into clusters. Using the Dirichlet process as the building block, our model constructs a product base-m…

2014-01-09abs ↗pdf ↗

Scalable psFA for fMRI data extracts sparse components.

problem Extracting neural representations from fMRI data with probabilistic formulation.
method Group level scalable probabilistic sparse factor analysis (psFA) with spatial sparsity, component pruning, and heteroscedastic noise modeling.
result Sparse components similar to group ICA and reduced noise in activated areas.

Proposes a new RNN model for grouped sequential data with varying time intervals.

problem Implicitly models fixed time intervals between observations and lacks group-level effects.
method Mixed membership framework for RNN, learning group-level base parameter.
result Demonstrates dynamic topic modeling with evolving topic distributions over time.

Framework infers coordination strategies from movement data.

problem Inferring individual movement strategies from group data.
method Formalizes Coordination Strategy Inference Problem; provides methodology to infer strategies.
result Framework accurately infers strategies in simulated and real-world datasets.

GCAO improves clustering of high-dimensional data by grouping low-density boundary points.

problem Stability and accuracy of clustering in high-dimensional, non-uniform data.
method Group-level optimization with gravitational attraction and optimization.
result GCAO outperforms 11 clustering methods on multiple datasets.

Sparse modeling is a powerful framework for data analysis and processing. Traditionally, encoding in this framework is performed by solving an L1-regularized linear regression problem, commonly referred to as Lasso or Basis Pursuit. In this work we combine the sparsity-inducing property of the Lasso model at the indivi…

2010-06-07abs ↗pdf ↗

Identifies patient-specific root causes of disease using structural equation models.

problem Detecting significant variables in complex diseases that differ between patients.
method Defining patient-specific root causes as exogenous errors in a structural equation model, quantifying predictivity using Shapley values, and developing a fast algorithm called Root Causal Inference.
result Significant improvements in accuracy by uncovering root causes with large effect sizes at the individual level but clinically insignificant effect sizes at the group level.

Review of central extensions for symplectic and divergence-free vector fields.

problem Comparing central extensions of symplectic and divergence-free vector fields.
method Analyzing universal central extensions and integrability.
result Similar results for both types of vector fields.

A new model for multi-level non-parametric admixture modeling.

problem Non-parametric topic modeling with shared distributions across levels.
method Nested Hierarchical Dirichlet Processes (nHDP) with a multi-level Chinese Restaurant Franchise (nCRF) representation.
result Significantly better generalization and detection of missing author entities.

We present the discrete infinite logistic normal distribution (DILN), a Bayesian nonparametric prior for mixed membership models. DILN is a generalization of the hierarchical Dirichlet process (HDP) that models correlation structure between the weights of the atoms at the group level. We derive a representation of DILN…

2011-03-24abs ↗pdf ↗

Study shows multisite fMRI data can be used reliably, slightly reducing detection power but not prediction accuracy.

problem Systematic biases in connectivity measures across multiple sites in fMRI studies.
method Empirical analysis of real multisite fMRI datasets and Monte-Carlo simulations.
result Multisite fMRI data can be used reliably, but slightly decreases detection power in statistical tests and prediction accuracy.

Develops a new criterion for subgroup fairness in algorithmic decision support.

problem Identifying fair recommendations in algorithms despite group-level differences.
method IJDI criterion and IJDI-Scan approach to detect and mitigate disparities.
result Identifies significant disparities in recommendations across subpopulations.

An anologue of the Calabi invariant for Poisson manifolds is considered. For any Poisson manifold PP, the Poisson bracket on C(P)C^{\infty}(P) extends to a Lie bracket on the space Ω1(P)Ω^{1}(P) of all differential one-forms, under which the space Z1(P)Z^{1}(P) of closed one-forms and the space B1(P)B^{1}(P) of exact one-forms a…

1996-05-05abs ↗pdf ↗

The paper compares different fairness definitions under various worldviews.

problem Avoiding disparity amplification under different worldviews.
method Mathematical comparison of four fairness definitions using a theoretical framework.
result Different worldviews require different fairness definitions to avoid disparity amplification.

We consider the problem of sparse variable selection in nonparametric additive models, with the prior knowledge of the structure among the covariates to encourage those variables within a group to be selected jointly. Previous works either study the group sparsity in the parametric setting (e.g., group lasso), or addre…

2012-06-18abs ↗pdf ↗

The paper proposes a method to measure fairness through equality of effort using algorithmic recourse.

problem Measuring fairness through equality of effort in automated systems.
method Applying algorithmic recourse to quantify equality of effort, overcoming previous limitations.
result An algorithm for assessing equality of effort has been developed and validated.

Personalized models using group attributes reduce performance, study finds.

problem Reducing performance of models using group attributes like race or gender.
method Formal conditions and collective preference guarantees to ensure fair use.
result Models personalized with group attributes reduce performance at a group level.

ABROCA assesses algorithmic bias, revealing skewed distributions that inflate results.

problem Detecting nuanced performance differences in classifier fairness.
method Study of ABROCA metric's statistical properties under various conditions.
result ABROCA distributions are skewed, inflating results by chance in imbalanced classes.

LCMQR improves prediction intervals by adapting to local heteroscedasticity.

problem Efficient and adaptive prediction intervals for local heteroscedasticity.
method LCMQR combines multi-quantile information with kernel-based localization.
result LCMQR constructs tighter intervals than prior methods, especially in heterogeneous environments.

Bayesian Additive Distribution Regression (DistBART) predicts distributions from grouped data.

problem Predicting distributions from grouped data with varying characteristics.
method Bayesian nonparametric approach using BART for modeling the regression function.
result Empirical and theoretical evidence supports DistBART's effectiveness in learning from low-dimensional marginals.

Proposes a model for multi-agent reinforcement learning with hierarchical graph attention network.

problem Limited transferability of trained policies to new multi-agent tasks.
method Uses hierarchical graph attention network for representation learning and multi-agent actor-critic for policy learning.
result Demonstrates superior performance in mixed cooperative and competitive tasks compared to existing methods.