RFA-LCF improves clustering accuracy by robustly handling noise and errors.
problem Inaccurate representation and clustering results due to noise and hard constraints.
method Integrates robust flexible CF, sparse local-coordinate coding, and adaptive weighting into a unified model.
result Delivers state-of-the-art clustering results on public databases.
A neural network model tackles high-dimensional data with latent structures.
problem Modeling high-dimensional data with latent low-dimensional structures.
method Integrates PCA and Soft PCA layers into neural network architecture for factor modeling and non-linear transformations.
result Demonstrates improved performance in forecasting and nowcasting with real-world data.
The study compares different models for predicting factor premiums and finds neural networks perform better but have unstable weights.
problem Predicting and timing the CMA factor premium using machine learning models.
method Compared regression models (OLS, Ridge, Random Forest, Neural Network) and tested factor timing strategies.
result Neural networks outperform linear models in explaining factor premium variance, but weights are unstable.
Flexible framework for CMTF with ADMM for various constraints and couplings.
problem Challenges in data fusion from multiple sources with varying characteristics.
method Flexible algorithmic framework using AO and ADMM for various constraints, loss functions, and couplings.
result Accurate and computationally efficient results for various loss functions, including KL divergence.
We develop a new model and algorithms for machine learning-based learning analytics, which estimate a learner's knowledge of the concepts underlying a domain, and content analytics, which estimate the relationships among a collection of questions and those concepts. Our model represents the probability that a learner p…
The recently proposed SPARse Factor Analysis (SPARFA) framework for personalized learning performs factor analysis on ordinal or binary-valued (e.g., correct/incorrect) graded learner responses to questions. The underlying factors are termed "concepts" (or knowledge components) and are used for learning analytics (LA),…
DSCF-Net learns deep features for clustering with robustness and locality preservation.
problem Unsupervised deep representation learning for clustering.
method Integrates robust deep concept factorization, deep self-expressive representation, and adaptive locality preserving feature learning.
result Delivers state-of-the-art performance on public databases.
A drift detection method for large datasets without labels.
problem Early detection of concept drift in large, unlabeled datasets.
method Classical statistical process control in a label-less setting.
result Better statistical power than previous methods under computational constraints.
Unified framework for generating data by modeling causal and correlational dependencies.
problem Modeling both causal and correlational dependencies among latent factors.
method Causal-Correlation Variational Autoencoder (C2VAE) framework.
result Improves generation quality, disentanglement, and intervention fidelity.
Diffusion models enhance SBI with flexible parameter and observation learning.
problem Efficient and accurate estimation of latent parameters from simulations and real data.
method Score-based diffusion models, guidance, score composition, flow matching, consistency models, joint modeling.
result Flexibility and versatility in modeling various problems.
In information retrieval, a fundamental goal is to transform a document into concepts that are representative of its content. The term "representative" is in itself challenging to define, and various tasks require different granularities of concepts. In this paper, we aim to model concepts that are sparse over the voca…
Levy copulas are the most general concept to capture jump dependence in multivariate Levy processes. They translate the intuition and many features of the copula concept into a time series setting. A challenge faced by both, distributional and Levy copulas, is to find flexible but still applicable models for higher dim…
We present a class of flexible and tractable static factor models for the term structure of joint default probabilities, the factor copula models. These high-dimensional models remain parsimonious with pair-copula constructions, and nest many standard models as special cases. The loss distribution of a portfolio of con…
Survey on concept factorization methods for better feature learning.
problem Understanding the quality of learned features in representation learning.
method Categorizing and summarizing recent advances in concept factorization methods.
result Provides insights into theoretical basis and current developments in concept factorization.
The study identifies latent concepts from diverse observations without assuming specific models.
problem Lack of general theoretical support for concept learning.
method Develops a nonparametric framework for identifying latent concepts from multiple classes of observations.
result Correctness guarantees for concept identification without parametric assumptions.
CREAM models enable concept-grounded predictions and interpretability.
problem Designing models that can encode and extend prior knowledge about concept-concept and concept-task relationships.
method Proposes a flexible and efficient framework (CREAMs) that encodes arbitrary C−C and CoY relationships, incorporating a side-channel for incomplete concept sets. result CREAM models achieve competitive task performance while encouraging concept-grounded predictions, avoiding concept leakage and achieving black-box-level performance.
Flexible models cluster RNA sequencing data.
problem Clustering discrete data from RNA sequencing studies.
method Finite mixtures of multivariate Poisson-log normal factor analyzers with constraints.
result Models give favorable clustering performance on real and simulated data.
Machine learning offers novel ways and means to design personalized learning systems wherein each student's educational experience is customized in real time depending on their background, learning goals, and performance to date. SPARse Factor Analysis (SPARFA) is a novel framework for machine learning-based learning a…
A new matrix factorization method for high-dimensional data.
problem Exploiting sparse structures in complex data for better interpretability.
method Bayesian shrinkage priors and flexible sparse patterns modeled through row and column dependencies.
result Demonstrated practical advantages through simulation and soccer heatmap analysis.
We propose Deep Feature Factorization (DFF), a method capable of localizing similar semantic concepts within an image or a set of images. We use DFF to gain insight into a deep convolutional neural network's learned features, where we detect hierarchical cluster structures in feature space. This is visualized as heat m…
Paper formalizes Simon's satisficing through FFSD, proving its equivalence to expected utility theory.
problem Formalizing Herbert Simon's bounded rationality concept in economic decision-making.
method Developed FFSD framework using Lean 4 theorem prover, proving equivalence to expected utility theory.
result Equivalence theorem linking FFSD to expected utility maximization for approximate indicator functions.
Bayesian non-parametric model adapts to concept drifts in streaming data.
problem Inference under concept drift phenomenon for non-stationary data streams.
method Variational inference algorithm for Dirichlet process mixture models with exponential forgetting.
result The proposed model outperforms state-of-the-art algorithms in clustering problems.
Variational Bayesian neural networks combine the flexibility of deep learning with Bayesian uncertainty estimation. However, inference procedures for flexible variational posteriors are computationally expensive. A recently proposed method, noisy natural gradient, is a surprisingly simple method to fit expressive poste…
A new framework for deep matrix factorizations improves model consistency and flexibility.
problem Inconsistent loss functions in deep matrix factorizations.
method Introduces two new loss functions and a generic optimization framework.
result Demonstrates improved model performance on synthetic and real data.
Proposes a flexible feature allocation model for sparse factor analysis.
problem Sparse data and rigid assumptions in traditional exploratory tools.
method Adaptive latent feature sharing with control over feature sparsity.
result Derives a novel adaptive Factor analysis (aFA) and aPPCA for flexible dimensionality reduction.
CAVs reveal latent concept distributions, but are vulnerable to adversarial attacks.
problem Understanding latent concept encodings in AI models.
method Probabilistic perspective on CAVs, deriving mean and covariance.
result CAVs can be adversarially manipulated, highlighting a vulnerability.
Study uses MLP models to predict large-cap US stocks, finding 2-3 hidden layers more flexible.
problem Predicting asset prices for large-cap US stocks.
method Applied MLP models with dynamic structure to factor models, focusing on firm characteristics.
result MLP models with 2-3 hidden layers more flexible in modeling factors, better for downside risk control.
Study finds relevance of exchange and inflation rates to economic factors.
problem Determining the relevance of exchange and inflation rates to economic factors.
method Introduced concept of adequacy, established positive relation between exchange and inflation rates and other economic factors.
result Close positive relation found between exchange and inflation rates and other economic factors.
A new prior is proposed for learning representations of high-level concepts of the kind we manipulate with language. This prior can be combined with other priors in order to help disentangling abstract factors from each other. It is inspired by cognitive neuroscience theories of consciousness, seen as a bottleneck thro…
New methods identify concepts in trained embeddings reliably without human labels.
problem Identifying interpretable concepts in trained embedding spaces without human labels.
method Explicitly connecting concept discovery to PCA and ICA, proposing novel approaches for dependent concepts.
result Proven methods outperform competitors on a variety of experiments, achieving up to 29% better alignment with ground truth.
MGLM models all possible language channel factorizations for improved multilingual generation.
problem Generating multilingual text with flexibility and quality.
method Generative joint distribution model over language channels, marginalizing all possible factorizations.
result MGLM outperforms traditional models in multilingual generation tasks.
We propose SPARFA-Trace, a new machine learning-based framework for time-varying learning and content analytics for education applications. We develop a novel message passing-based, blind, approximate Kalman filter for sparse factor analysis (SPARFA), that jointly (i) traces learner concept knowledge over time, (ii) an…
The nonnegative matrix factorization is a widely used, flexible matrix decomposition, finding applications in biology, image and signal processing and information retrieval, among other areas. Here we present a related matrix factorization. A multi-objective optimization problem finds conical combinations of templates …
Natural graph networks are a new class of graph neural networks that are more flexible and scalable.
problem Traditional graph neural networks are limited by equivariance to node permutations.
method Introduced natural graph networks, which are more flexible and scalable than conventional graph neural networks.
result Natural graph networks are as scalable as conventional message passing graph neural networks but more flexible.
Nonnegative Matrix Factorization (NMF) aims to factorize a matrix into two optimized nonnegative matrices appropriate for the intended applications. The method has been widely used for unsupervised learning tasks, including recommender systems (rating matrix of users by items) and document clustering (weighting matrix …
Meta-causal states group equivalent qualitative causal dynamics, useful for analyzing system changes.
problem Qualitative changes in causal relationships due to agent actions or environmental tipping points.
method Propose meta-causal states to group causal models based on equivalent qualitative behavior and parameterize specific mechanisms.
result Meta-causal states can be inferred from observed agent behavior and disentangled from unlabeled data.
Finite translation surfaces can be classified by the order of their singularities. When generalizing to infinite translation surfaces, however, the notion of order of a singularity is no longer well-defined and has to be replaced by new concepts. This article discusses the nature of two such concepts, recently introduc…
New ADMM method for PARAFAC2 tensor decomposition with flexible regularization.
problem Challenges in applying regularisation to the evolving mode of PARAFAC2.
method Alternating Direction Method of Multipliers (AO-ADMM) for PARAFAC2 tensor fitting.
result The proposed ADMM-based approach accurately recovers underlying components from simulated data.
A new tensor decomposition method for fMRI data captures both spatial and temporal variability.
problem Challenges in modeling shared and subject-specific structure in multisubject spatiotemporal data, especially in neuroimaging.
method Introduces a spatiotemporal variational tensor decomposition (ST-VTD) framework combining tensor factorization with structured priors for flexible representation of spatial and temporal dynamics.
result Significantly improves latent factor recovery in fMRI data compared to classical and probabilistic decomposition benchmarks.
The paper analyzes the convergence of CART under a SID condition, improving previous results.
problem Investigating the convergence rate of CART under a sufficient impurity decrease condition.
method Established an upper bound on prediction error under SID condition, introduced easily verifiable conditions.
result Improved convergence rate of CART under SID condition, demonstrated examples of error bound limitations.
During the past few years Boolean matrix factorization (BMF) has become an important direction in data analysis. The minimum description length principle (MDL) was successfully adapted in BMF for the model order selection. Nevertheless, a BMF algorithm performing good results from the standpoint of standard measures in…
Tensor factorization has been demonstrated as an efficient approach for computational phenotyping, where massive electronic health records (EHRs) are converted to concise and meaningful clinical concepts. While distributing the tensor factorization tasks to local sites can avoid direct data sharing, it still requires t…
Tensor factorization is a powerful tool to analyse multi-way data. Compared with traditional multi-linear methods, nonlinear tensor factorization models are capable of capturing more complex relationships in the data. However, they are computationally expensive and may suffer severe learning bias in case of extreme dat…
GD with large init shows incremental learning in matrix factorization.
problem Understanding GD's behavior with large initial values in matrix factorization.
method Signal-to-noise ratio concepts and inductive arguments.
result Uncovering an incremental learning phenomenon in GD with large initialization.
Extends matrix factorization for deviance-based losses with GLM theory.
problem Improving data loss models beyond squared error.
method Adapts GLM theory to matrix factorization for deviance losses.
result Strong consistency and robustness of the proposed decomposition.
Unified model learns concepts across domains like left and right.
problem Limited generalization of language concepts in inference-only models.
method Logic-Enhanced Foundation Model (LEFT) with a differentiable, domain-independent program executor.
result LEFT flexibly learns and reasons with concepts across 2D images, 3D scenes, human motions, and robotic manipulation.
New model analyzes dynamic correlations in stock returns.
problem Analyzing time-varying correlations in high-dimensional data.
method Dynamic factor correlation model with novel parametrization.
result Model accurately captures heterogeneous heavy-tailed distributions and dependent shocks.
Enhancing spectral embedding for low-dimensional embeddings in rare disease cohorts
problem Representing clinical concepts and patients in electronic health records
method Spectral-based unsupervised learning with flexible knowledge transfer
result Outperforms competing approaches in challenging scenarios