New methods identify concepts in trained embeddings reliably without human labels.
problem Identifying interpretable concepts in trained embedding spaces without human labels.
method Explicitly connecting concept discovery to PCA and ICA, proposing novel approaches for dependent concepts.
result Proven methods outperform competitors on a variety of experiments, achieving up to 29% better alignment with ground truth.
New method discovers concepts in hidden feature layers using sparse subspace clustering.
problem Local attribution methods fail to identify coherent model behavior across samples.
method Sparse Subspace Clustering (SSCC) for concept discovery.
result Empirically validated method for various image classification tasks.
LORL learns object-centric representations from vision and language.
problem Learning disentangled, object-centric scene representations from vision and language.
method LORL integrates unsupervised object discovery and segmentation with language input to learn object-centric concepts.
result LORL improves unsupervised object discovery methods and aids downstream tasks.
New method extracts biological concepts from cell microscopy images.
problem Extracting meaningful concepts from vision foundation models trained on cell microscopy images.
method Sparse dictionary learning (DL) combined with PCA whitening pre-processing.
result Successfully retrieved biologically meaningful concepts like cell types and genetic perturbations.
Efron et al. (2001) proposed empirical Bayes formulation of the frequentist Benjamini and Hochbergs False Discovery Rate method (Benjamini and Hochberg,1995). This article attempts to unify the `two cultures' using concepts of comparison density and distribution function. We have also shown how almost all of the existi…
DFF detects similar concepts in images, visualized as heat maps.
problem Localizing similar semantic concepts within images.
method Deep Feature Factorization (DFF) to detect hierarchical cluster structures in feature space.
result Visualizes semantically matching regions across images, revealing network perception.
This work improves interpretability in deep learning models by introducing a two-level concept discovery framework.
problem High complexity and lack of interpretability in deep learning models, especially for safety-critical tasks.
method Concept Bottleneck Models (CBMs) framework combining vision-language models and data-driven coarse-to-fine concept selection.
result The proposed framework outperforms recent CBM approaches and provides a principled interpretability.
CB-SLICE identifies concept-based error slices in deep learning models.
problem Systematic errors in deep learning models on specific groups.
method Concept Bottleneck Models (CBMs) and concept representations.
result CB-SLICE outperforms state-of-the-art methods in error slice identification.
MCD offers a complete model understanding for high-stake decisions.
problem Local model understanding in XAI methods is not sufficient for high-stake decisions.
method MCD extends concept-based methods to ensure global model understanding via multi-dimensional subspaces.
result MCD provides a complete model understanding, ensuring the model reasoning is related to the actual model.
Algorithm transfers visual concepts to answer out-of-vocabulary questions.
problem Leveraging off-the-shelf visual and linguistic data for out-of-vocabulary answers in visual question answering.
method Unsupervised task discovery for learning task conditional visual classifier, then transferring to visual question answering models.
result Algorithm generalizes to out-of-vocabulary answers successfully.
Causal discovery algorithms can help generate legal arguments.
problem Leveraging causal discovery algorithms in legal decision-making.
method Prepared a legal dataset, annotated with 17 legal concepts, applied causal discovery algorithms, and quantified degrees of belief.
result Some causal relationships help generate viable legal arguments.
New algorithm discovers causal relationships in complex data.
problem Discovering causal relationships in data with cycles, latent confounders, and non-linearities.
method Introducing σ-connection graphs and extending σ-separation to handle these complexities.
result First algorithm capable of handling non-linear, cyclic, and latent confounders.
AlphaZero reveals new chess concepts learnable by top experts.
problem Extracting and understanding hidden knowledge from AI systems.
method Proposed method to extract new chess concepts from AlphaZero.
result Top chess grandmasters show improvements in learning new concepts.
AI helps in drug discovery with understandable explanations.
problem Understanding the complex models behind AI-generated drugs.
method Explainable AI methods to interpret deep learning models.
result Improved interpretability of AI-generated drug properties.
Concept Relation Discovery and Innovation Enabling Technology (CORDIET), is a toolbox for gaining new knowledge from unstructured text data. At the core of CORDIET is the C-K theory which captures the essential elements of innovation. The tool uses Formal Concept Analysis (FCA), Emergent Self Organizing Maps (ESOM) and…
New algorithm for clustering categorical data without using distance measures.
problem Challenges in clustering categorical data due to its unordered structure.
method Developed a Matching based clustering algorithm using a similarity matrix and feature importance criteria.
result The algorithm can serve as an alternative to existing clustering methods for categorical data.
GPR enhances materials discovery by automating parameter space exploration.
problem Automating exploration of large, high-dimensional parameter spaces in materials science.
method Gaussian process regression with inhomogeneous measurement noise and anisotropic kernels.
result Importance and benefits of tuning GPR for materials science experiments.
This paper is a tutorial on Formal Concept Analysis (FCA) and its applications. FCA is an applied branch of Lattice Theory, a mathematical discipline which enables formalisation of concepts as basic units of human thinking and analysing data in the object-attribute form. Originated in early 80s, during the last three d…
This paper introduces a method to find complete and interpretable concept-based explanations for deep neural networks.
problem Lack of complete and interpretable concept-based explanations in deep neural networks.
method Definition of completeness, concept discovery method, and importance score calculation using game-theoretic notions.
result The proposed method finds complete and interpretable concept-based explanations for deep neural networks.
Paper discovers differential equations from data using neural networks and Bayesian methods.
problem Discovering differential equations from datasets using machine learning.
method Integrates neural network-based surrogates with Sparse Bayesian Learning (SBL).
result Proposes a robust model discovery algorithm and a Physics Informed Normalizing Flow (PINF).
A new framework for interpretable models using sparse linear layers.
problem Performance degradation and lower interpretability in concept bottleneck models.
method Contrastive Language Image models and a single sparse linear layer with Bayesian inference.
result Our framework outperforms recent CBM approaches in accuracy and concept sparsity.
PACC Discovery improves causal inference from limited data.
problem Inferring causal relationships from finite data.
method Extends PAC learning principles to causal inference.
result Theoretical guarantees for various causal methods.
Geometric framework detects concept frustration between human concepts and machine representations.
problem Aligning human concepts with machine learning representations.
method Geometric framework and similarity measures for detecting concept frustration.
result Concept frustration affects machine learning model performance and reorganizes learned concept representations.
Paper relaxes faithfulness assumption for MB discovery in k-order Markov blanket.
problem Learning graphical Markov blanket from data under faithfulness assumption violations.
method k-order relaxation of faithfulness assumption; k-OMB algorithm.
result k-OMB recovers MB under true and empirical faithfulness violations.
Proposes a weaker faithfulness assumption for causal discovery.
problem Violation of the faithfulness assumption in causal discovery.
method Proposes a new assumption called 2-adjacency faithfulness and a modified Grow and Shrink algorithm.
result Proves the correctness of the modified algorithm under weaker assumptions.
Automates organizing diverse web data into a hierarchical topic model.
problem Manual classification of all scientific and popular scientific knowledge is impractical.
method Proposes an algorithm to aggregate multiple collections into a single hierarchical topic model.
result Demonstrates a web service for topical exploratory search.
Study uses LLMs to automate data insights discovery.
problem Extracting relevant insights from large data sets.
method Capture the Flag principle, LLMs, reasoning, code generation.
result LLMs can recognize meaningful data insights.
New method improves bivariate causal discovery by accurately estimating cause variable complexity.
problem Improper estimation of cause variable complexity in current MDL-based methods.
method Rate-distortion MDL (RDMDL) using information dimension for cause variable complexity estimation.
result RDMDL achieves competitive performance on Tübingen dataset.
Recursive causal discovery reduces errors and complexity in causal graph learning.
problem Challenges in causal discovery from limited data and computational complexity.
method Removable variables for recursive causal discovery, reducing problem size and CI tests.
result Worst-case performances nearly match lower bound, with state-of-the-art efficiency.
This paper uses group theory to create data-free, feature-free clustering.
problem Creating data-free, feature-free hierarchical clustering.
method Symmetry-driven hierarchical clustering using group theory.
result A new clustering framework that is globally hierarchical and data-free.
The paper develops online methods to control familywise error rate in growing hypothesis testing sequences.
problem Controlling familywise error rate in a growing sequence of hypotheses over time.
method Unified algorithmic concepts for offline and online FWER control, including new adaptive online algorithms.
result Substantial gains in power demonstrated and formally proved in a Gaussian sequence model.
Unified framework controls false discovery rate in bandit multiple testing.
problem Designing adaptive algorithms to identify true discoveries in multiple hypothesis testing.
method Unified modular framework using e-processes for FDR control in arbitrary settings.
result Unified framework ensures FDR control for dependent and simultaneous arm queries.
Proposes a method for time-evolving and difficulty-level topic discovery.
problem Discovering evolving and advanced topics in dynamic corpora.
method Constrained Coupled Matrix-Tensor Factorization with expertise-level constraints.
result Identifies evolving and difficulty-level topics in community-contributed content.
This study creates a new conceptual framework for news aggregation.
problem Aggregating and presenting diverse news sources in a unified and accessible way.
method Developed a mobile app that analyzes unstructured data patterns to create a conceptual framework for news.
result Users can easily find and navigate through updated news using a new conceptual multilevel structure.
New framework learns physics from output measurements only.
problem Learning governing physics from only output measurements.
method Stochastic calculus, sparse learning, Bayesian statistics, Euler Maruyama scheme.
result Potential to identify governing physics from sparse, noisy, incomplete data.
The paper tests semantic importance in opaque models using betting.
problem Precise statistical guarantees for semantic concepts in black-box models.
method Formalizes global and local statistical importance via conditional independence and SKIT.
result Shows effectiveness and flexibility of the framework on various models.
This paper presents a framework for exact discovery of the top-k sequential patterns under Leverage. It combines (1) a novel definition of the expected support for a sequential pattern - a concept on which most interestingness measures directly rely - with (2) SkOPUS: a new branch-and-bound algorithm for the exact disc…
Deep convolutional neural networks comprise a subclass of deep neural networks (DNN) with a constrained architecture that leverages the spatial and temporal structure of the domain they model. Convolutional networks achieve the best predictive performance in areas such as speech and image recognition by hierarchically …
Study examines XAI methods for ECG analysis to improve model transparency.
problem Lack of transparency in deep learning models for ECG analysis.
method Investigates post-hoc XAI methods for local and global perspectives, establishes sanity checks, and demonstrates knowledge discovery.
result Quantitative evidence supports expert rules for sensible attribution methods and demonstrates XAI's utility for knowledge discovery.
Automated discovery of early visual concepts from raw image data is a major open challenge in AI research. Addressing this problem, we propose an unsupervised approach for learning disentangled representations of the underlying factors of variation. We draw inspiration from neuroscience, and show how this can be achiev…
Deep networks respond to specific linguistic units, not arbitrary patterns.
problem Understanding how deep convolutional networks interpret natural language.
method Concept alignment method based on unit responsiveness to replicated text.
result Deep networks selectively respond to morphemes, words, and phrases, not arbitrary patterns.
DL-PDE discovers PDEs from noisy, sparse data using neural networks and sparse regressions.
problem Discovering PDEs from noisy, sparse data.
method Combines neural networks and sparse regressions to discover PDEs from meta-data generated by a neural network.
result Achieves satisfactory results in real-world engineering settings with noisy and limited data.
Dynamic Structural Causal Models handle time-dependent systems with cycles and latent confounding.
problem Representing and analyzing systems of Stochastic Differential Equations (SDEs) with DSCMs.
method Define time-splitting and subsampling operations to analyze DSCMs of SDEs, and apply existing causal discovery algorithms to time-series data.
result DSCMs provide a graphical Markov property for SDEs and enable identification of time-dependent causal effects.
New algorithms discover and utilize 'voids' in data to improve machine learning models.
problem Improving machine learning models by considering the unknown aspects of data.
method Developed algorithms to discover and utilize 'voids' in data, creating ignorance-aware prototypes.
result Improved performance of nearest neighbor classifiers through ignorance-aware prototype selection.
LFD method improves text classification by making features clearer and less label-leaking.
problem Creating interpretable text representations that are both predictive and understandable.
method LFD method: proposes lexical and semantic features from contrastive text pairs, screens candidates using κ, and selects features by residual gain. result LFD features achieve higher human-human and human-LLM agreement than baseline concepts and are less label-leaking.
This paper uses spectrum analysis to understand price behavior in the Indian stock market.
problem Understanding price formation and discovery in the Indian stock market.
method Adapting mathematical physics theories and spectrum analysis to decompose price cycles.
result Decomposing price cycles helps in understanding the effect of information on price formation and discovery.
This paper examines how institutional liquidity affects prediction markets.
problem How institutional liquidity impacts prediction markets and their quality.
method Defines a market-quality lens, separates channels, and uses synthetic microstructure lab.
result Institutional liquidity does not necessarily translate to equal gains for all traders.
Paper introduces a new identifiability criterion for DAGs using conditional variances.
problem Challenges in discovering causal relationships from observational data.
method Introduces a novel identifiability criterion for DAGs using conditional variances. Uses weak majorization on Cholesky factor of covariance matrix for learning DAGs.
result Demonstrates effectiveness of the new approach in recovering DAGs through simulations and real data analysis.