New mechanism detects overlap density for weak-to-strong generalization.
problem Understanding what aspects of data enable weak-to-strong generalization.
method Data-centric mechanism and overlap detection algorithm.
result Overlap density is a key factor in weak-to-strong generalization.
As research into community finding in social networks progresses, there is a need for algorithms capable of detecting overlapping community structure. Many algorithms have been proposed in recent years that are capable of assigning each node to more than a single community. The performance of these algorithms tends to …
New metrics assess class overlap and imbalance in datasets.
problem Class overlap and imbalance make datasets hard to classify.
method Developed new metrics based on ball coverage by classes.
result Metrics correlate well with classifier performance.
ION-C solves overlapping network integration problems efficiently.
problem Integrating overlapping networks with different datasets.
method Formulated as an ASP problem and solved with clingo.
result Significantly improved efficiency in runtime and solution graphs.
A new method optimizes anomaly scoring from score distribution to improve AD performance.
problem Vulnerability to anomaly contamination and lack of adaptability in existing AD methods.
method Optimizes anomaly scoring function from score distribution perspective, using Overlap loss.
result Overlap loss-based AD models significantly outperform state-of-the-art methods.
We calculate eigenvector overlaps between intersecting time periods of covariance matrices.
problem Analyzing overlapping time periods in covariance matrices.
method Girko linearisation and extended local laws.
result Computed eigenvector overlaps for intersecting time intervals.
Estimates overlap in observational studies for causal effect estimation.
problem Overlap between treatment groups is crucial for causal effect estimation.
method Formalizes overlap estimation as a binary classification problem with Boolean rule classifiers.
result Rules provide interpretable explanations for causal conclusions.
The paper addresses causal estimation for text data with apparent overlap violations.
problem Estimating causal effects from text data with unknown confounders and apparent overlap.
method Uses supervised representation learning to create a representation that preserves confounding information while eliminating predictive information, satisfying overlap assumptions.
result Shows how to obtain robust causal estimation in the presence of apparent overlap violations.
RISA improves VFL by using imputed samples with low uncertainty.
problem Limited overlapping samples constrain VFL performance.
method Imputing non-overlapping samples and using evidence theory to select reliable imputed samples.
result Significant performance gains achieved, especially with limited overlapping samples.
New clustering method avoids chaining issues and refines single-linkage clusters.
problem Clustering overlapping data with chaining issues.
method Functorial constraints on overlapping clustering in metric spaces.
result Any clustering functor is constrained to refine single-linkage clusters.
The paper tackles the problem of identifying same-cluster elements in overlapping clusters with minimum queries.
problem Identifying same-cluster elements in overlapping clusters with minimum queries.
method The paper provides algorithms for identifying same-cluster elements in overlapping clusters with minimum queries, under both arbitrary and statistical modeling assumptions.
result The algorithms are order optimal, parameter free, efficient, and work in the presence of random noise.
Paper proposes a new co-clustering method for overlapping clusters and outliers.
problem Real-world datasets often contain overlaps and outliers in co-clusters.
method Formulated Non-Exhaustive, Overlapping Co-Clustering problem and developed NEO-CC algorithm.
result NEO-CC algorithm effectively captures underlying co-clustering structure of real-world data.
New survival learners estimate heterogeneous treatment effects from time-to-event data.
problem Estimating HTEs from time-to-event data with censoring outcomes.
method Orthogonal survival learners with theoretical guarantees and custom weighting functions.
result Orthogonal survival learners provide robust and model-agnostic HTE estimation.
Improved off-policy evaluation for MDPs with weak distributional overlap.
problem Evaluation of policies when target and data-collection distributions are not strongly overlapping.
method Truncated Doubly Robust (TDR) estimators for off-policy evaluation in MDPs under weak distributional overlap.
result TDR estimators can recover large-sample behavior and are consistent even when distribution ratios are not square-integrable.
This paper presents a novel spectral algorithm with additive clustering designed to identify overlapping communities in networks. The algorithm is based on geometric properties of the spectrum of the expected adjacency matrix in a random graph model that we call stochastic blockmodel with overlap (SBMO). An adaptive ve…
Subject Cross Validation improves Human Activity Recognition performance by up to 16%.
problem Overestimation of Human Activity Recognition performance using k-fold cross validation.
method Investigated Subject Cross Validation vs. k-fold cross validation for Human Activity Recognition.
result Subject Cross Validation increases performance by up to 16%.
Efficiently solves high-dimensional regression with overlapping groups using greedy hard-thresholding.
problem High-dimensional regression problems with overlapping groups of relevant features.
method Greedy hard-thresholding combined with submodular optimization to avoid NP-hard projections.
result Strong theoretical guarantees even with poorly conditioned data and overlapping features.
New group-sparse SVD models improve biclustering of gene expression data.
problem Identifying block patterns with similar expressions in high-dimensional gene expression data.
method Proposed GL1-SVD, GL0-SVD, OGL1-SVD, and OGL0-SVD models with group Lasso and L0-norm penalties, using alternating iterative strategies and ADMM.
result Effective in identifying biologically interpretable gene modules with gene prior group knowledge.
Deconfounding scores improve causal effect estimation with weak overlap.
problem Poor overlap in treatment and control groups makes causal effect estimators brittle.
method Introduces feature representations that improve overlap without introducing bias.
result Deconfounding scores satisfy a zero-covariance condition that is identifiable in observed data.
Quantum theory improves counting overlapping clusters.
problem Counting overlapping clusters in machine learning.
method Applied quantum theory using path integral technique.
result Quantum theory provides a robust statistical method for counting clusters.
Proposes a new framework for detecting overlapping and non-overlapping communities.
problem Lack of methods for both overlapping and non-overlapping community detection.
method Integrated framework based on primary node criteria of internal and external association degrees.
result Outperforms existing methods on evaluation criteria.
This study quantifies systemic risk from overlapping portfolios in the Mexican financial system.
problem Systemic risk from indirect interconnections between financial institutions.
method Represented the Mexican financial system as a bipartite network of securities and financial institutions; quantified systemic risk from overlapping portfolios.
result Total systemic risk levels underestimated by up to 50% when only direct exposures are considered.
PAM models generate dependent random distributions across groups with overlapping clusters.
problem Generating dependent random distributions across multiple groups.
method Atom skipping in an infinite mixture model.
result Interpretable posterior inference of cluster exclusivity and sharing.
Automated image segmentation distinguishes overlapping human chromosomes.
problem Distinguishing overlapping human chromosomes for medical diagnostics.
method Customized convolutional neural network for image segmentation.
result IOU scores of 94.7% for overlapping regions, 88-94% for non-overlapping regions.
A new overlapping space solves the configuration search problem for graph embeddings.
problem Configuring product spaces for graph embeddings is resource-intensive and impractical.
method Introducing overlapping spaces that share subsets of coordinates between different types of spaces (Euclidean, hyperbolic, spherical).
result Overlapping spaces achieve nearly optimal results without configuration tuning, reducing training time.
Syncytial clustering merges groups from standard algorithms to reveal complex data structures.
problem Challenges in finding clusters with irregular structures.
method Estimates nonparametric overlap between clusters and merges groups with high overlap.
result Always a top performer in identifying groups with regular and irregular structures.
Overlapping clustering problem is an important learning issue in which clusters are not mutually exclusive and each object may belongs simultaneously to several clusters. This paper presents a kernel based method that produces overlapping clusters on a high feature space using mercer kernel techniques to improve separa…
CausalMix generates synthetic data with causal controls for mixed-type tables.
problem Synthetic data for causal inference with mixed-type and multimodal tabular data.
method CausalMix combines Gaussian latent priors with data-type-specific decoders for control over overlap, confounding, and treatment effect heterogeneity.
result CausalMix achieves state-of-the-art distributional metrics and stable causal control.
Deconfounding scores improve causal effect estimation with weak overlap.
problem Challenges in causal treatment effect estimation due to weak overlap in high-dimensional data.
method Propose deconfounding scores to preserve identification and target estimation while improving overlap.
result Prognostic scores are overlap-optimal under a broad family of generalized linear models with Gaussian features.
New algorithm learns policies without uniform overlap assumption.
problem Learning optimal policies from non-uniformly collected data.
method Pessimistic Policy Learning (PPL) using lower confidence bounds.
result Efficient policy learning for adaptively collected data.
A new method speeds up overlapping group lasso computations.
problem Time-consuming optimization of overlapping group lasso on large-scale problems.
method Non-overlapping statistical approximation to overlapping group lasso.
result The proposed penalty is statistically equivalent to overlapping group lasso.
New method improves CATE estimation in low overlap regions.
problem Low overlap in CATE estimation leads to poor performance of meta-learners.
method Overlap-Adaptive Regularization (OAR) that regularizes models proportionally to overlap weights.
result OAR significantly improves CATE estimation in low-overlap settings.
Proposes a sensitivity framework to handle limited overlap in causal inference.
problem Limited overlap between treated and control groups in observational studies.
method Sensitivity framework based on worst-case confidence bounds on bias introduced by trimming.
result Protects against spurious findings by quantifying uncertainty in regions with limited overlap.
Producing overlapping schemes is a major issue in clustering. Recent proposed overlapping methods relies on the search of an optimal covering and are based on different metrics, such as Euclidean distance and I-Divergence, used to measure closeness between observations. In this paper, we propose the use of another meas…
Emergent misalignment is influenced by training dynamics, model priors, and data.
problem Emergent misalignment in models
method Exploring training dynamics, model priors, and data
result Activation deltas before and after narrow fine-tuning correlate with their similarities when measured with the last prompt-token activations.
Effective multilingual search with instance-based transfer learning.
problem Search in multilingual setting, especially next-sentence prediction and inverse cloze.
method Instance-based transfer learning, analyzing vocabulary overlaps and transitive overlaps.
result Positive transfer on all 35 target languages and two tasks, even with no vocabulary overlap.
New theory explains contrastive learning via overlapping augmented views.
problem Lack of theoretical understanding of contrastive learning.
method Augmentation overlap perspective to improve downstream performance.
result Asymptotically closed bounds for downstream performance under weaker assumptions.
The study simplifies assessing overlap in logistic regression models using empirical likelihood.
problem Assessing overlap in multidimensional logistic regression models.
method Translation of Silvapulle's condition to empirical likelihood maximization, mechanized with R code.
result Minimal overlapping structures are cataloged in dimensions less than four, providing rules for higher dimensions.
Method leverages data transfer for estimating CATE with KRR.
problem Leveraging findings from one study to estimate CATE in a different population.
method Overlap-adaptive transfer learning of CATE using kernel ridge regression.
result The method achieves superior efficiency and adaptability in estimating CATE.
The paper tackles CF in CL by analyzing NTK overlap matrix and proposing OGD.
problem Catastrophic Forgetting in continual learning.
method Analysis of NTK overlap matrix, OGD with PCA.
result Proposes OGD to mitigate CF, supported by experiments.
Novel unsupervised scheme for highly imbalanced and overlapping datasets.
problem Highly imbalanced and overlapping classes in medical datasets.
method Unsupervised domain adaptation scheme based on Quantification.
result High quality results for Quantification and Domain Adaptation.
Proposes a meta-algorithm for classification with overlapping classes in high-energy physics.
problem Challenges of class overlap in binary classification.
method Combines bagging and boosting techniques with a randomization trick.
result Improves statistical significance of Higgs discovery.
A new model for detecting overlapping communities in weighted networks.
problem Community detection in overlapping weighted networks with mixed membership and edge weights.
method Mixed membership distribution-free (MMDF) model with an efficient spectral algorithm and fuzzy weighted modularity.
result The MMDF model can estimate community memberships and evaluate community quality for weighted networks.
Proposes efficient estimators for weighted cumulative treatment effects in observational studies.
problem Inconsistent and inefficient estimators due to model misspecification and lack of overlap.
method Double/debiased machine learning for weighted cumulative causal effects.
result Proposed estimators are consistent, asymptotically linear, and reach semiparametric efficiency bounds.
Temperature scaling fails for distributions with class overlaps, while Mixup improves calibration.
problem Temperature scaling's performance degrades with class overlaps, leading to poor calibration.
method Identified temperature scaling's limitations and compared it with Mixup for calibration.
result Mixup significantly outperforms temperature scaling in calibration metrics with class overlaps.
Unified clustering comparison framework for overlapping and hierarchical structures.
problem Critical biases in existing clustering comparison measures.
method Element-centric framework comparing relationships induced by cluster structure.
result Framework does not suffer from biases and provides unique insights.
MODWST improves classification tasks with wavelet scattering.
problem Signal classification challenges.
method Combines MODWT and WST for feature extraction.
result MODWST outperforms CNNs in limited data scenarios.
Paper extends causal inference methods beyond unconfoundedness and overlap assumptions.
problem Treatment effect identification in studies violating unconfoundedness and overlap.
method Statistical learning theory approach to identify ATE and ATT.
result General conditions for identifying ATE and ATT, including scenarios like Regression Discontinuity designs.