The paper proposes a method to analyze categorical feature interactions in large datasets using graph covariance and LLMs.
problem Analyzing complex datasets with numerous categorical features and timestamps.
method Binarization of categorical features using one-hot encoding, computation of graph covariance, identifying significant feature pairs, and using LLMs to generate explanations.
result The method identifies meaningful feature pairs and potential data stories underlying categorical feature interactions.
Bayesian model improves categorization of explosions from sparse data.
problem Challenges in categorizing explosions from limited data.
method Bayesian update to Event Categorization Matrix model with Bayesian Decision Theory.
result Consistent gains in overall accuracy and lower false negative rates.
Reduces high granularity and dimensionality in hierarchical categorical variables.
problem Overfitting and estimation issues in predictive models due to high granularity and dimensionality.
method Entity embedding and top-down clustering algorithm to reduce granularity and dimensionality.
result The reduced hierarchy improves model fit and complexity balance.
Probabilistic learning for binary classification with categorical variables.
problem Binary classification with categorical covariates.
method Probabilistic analysis and two algorithms for learning boolean functions.
result Effective learning of boolean functions from binary data.
DPERC efficiently estimates covariance matrices for mixed data with missing values.
problem Estimating covariance matrices for datasets with missing values and mixed features.
method Direct Parameter Estimation for Randomly Missing Data with Categorical Features (DPERC).
result DPERC outperforms other methods in estimating covariance matrices for mixed data with missing values.
A new method for finding significant patterns in stratified data.
problem Statistical challenges in finding enriched itemsets in stratified data.
method Proposes a strategy and efficient algorithm for significant pattern mining in the presence of categorical covariates.
result Efficiently corrects for multiple testing in pattern mining while retaining statistical power.
lgpr interprets longitudinal data to find covariate effects.
problem Inferring covariate effects from longitudinal data with complex structures and interactions.
method Additive Gaussian processes for nonparametric analysis.
result lgpr outperforms previous methods in identifying relevant covariates.
A new model integrates covariates with grade of membership analysis for better latent structure recovery.
problem Improving latent structure recovery in multivariate categorical data analysis.
method Covariate-assisted grade of membership model exploiting shared low-rank simplex geometry.
result Auxiliary covariates can provably improve latent structure recovery, leading to faster convergence rates.
A method for representing and comparing categorical trajectories using multivariate functional principal components.
problem Statistical description and comparison of categorical trajectories.
method Transforming categorical trajectories into binary indicator functions and applying multivariate functional principal components analysis.
result Consistent estimators of mean trajectories and covariance functions are obtained under weak regularity assumptions.
A new Bayesian multinomial regression model using permuted and augmented stick-breaking.
problem Modeling categorical response variables given covariates.
method Permuted and augmented stick-breaking (paSB) construction.
result Transforms multinomial regression into regression of stick-specific binary variables.
Two novel methods estimate multiple FDR directions for binary categorical responses.
problem Estimating multiple FDR directions for categorical responses.
method Information maximization and square loss mutual information.
result Statistical consistency of the proposed methods established.
A new matching method for causal inference that handles irrelevant variables and missing data.
problem Creating high-quality treatment-control matches for categorical data in social sciences.
method A weighted Hamming distance matching method that considers covariate importance and creates a hierarchy of covariate combinations.
result The method produces high-quality matches and handles irrelevant variables and missing data.
CardiCat generates synthetic data for high-cardinality tabular datasets.
problem Learning complexities of high-cardinality categorical features in tabular data.
method Substitutes one-hot encoding with regularized dual encoder-decoder embedding layers.
result Generates high-quality synthetic data with a smaller parameter space.
Standardization boosts sparse group lasso performance for categorical variables.
problem Improving performance of sparse group lasso with categorical variables.
method Column-wise scaling and re-scaling of coefficients, without orthonormalization.
result Column-wise scaling and re-scaling of coefficients achieves the same effect as orthonormalization.
GCVAE uses Gaussian copula for mixed data, improving over standard VAE.
problem Handling mixed categorical and continuous data.
method Employed Gaussian copula to model local dependency in mixed data, using rank-one approximation for covariance.
result GCVAE better captures the data manifold compared to standard VAE.
MCPCA improves data dimensionality by maximizing nonlinear correlations.
problem PCA's limitations in handling nonlinearity and categorical data.
method MCPCA computes nonlinear transformations of variables to maximize covariance matrix Ky Fan norm.
result MCPCA outperforms other methods in dimensionality reduction tasks.
FLAME efficiently matches high-dimensional categorical datasets for causal inference.
problem Matching treatment and control units based on covariate information in causal inference.
method FLAME learns a distance metric using a hold-out training set and uses query processing techniques for large datasets.
result FLAME achieves significantly better performance than other matching methods, scaling to huge datasets.
This study improves hyperparameter optimization for categorical and non-normal data.
problem Bayesian hyperparameter optimization struggles with categorical hyperparameters and non-normal data.
method Integrates conformalized quantile regression to address estimation weaknesses and provides robust calibration guarantees.
result Quantile surrogate architectures and acquisition functions yield superior performance compared to existing methods.
Global covariance pooling improves deep CNNs' representation and generalization.
problem Capturing richer statistics of deep features for better representation and generalization.
method Integrates global covariance pooling into deep CNNs, addressing challenges with robust covariance estimation and geometry exploitation.
result Proposes MPN-COV Pooling and a Gaussian embedding network, achieving state-of-the-art performance.
A new method for domain generalization using source-specific classifiers.
problem Generalizing across multiple sources for any target domain.
method Multiple domain-specific classifiers and a domain agnostic component.
result Improved performance on public benchmarks.
Efficiently models categorical data with low to medium class overlap, improving accuracy over standard distributions.
problem Poor parameter estimates and accuracy in multinomial and Dirichlet multinomial distributions when assumptions are violated.
method Introduces Beta-Liouville multinomial distribution and efficient estimation methods.
result Beta-Liouville multinomial outperforms standard distributions on two out of four datasets.
Categorical Co-Frequency Analysis clusters diagnoses to predict hospital readmissions.
problem Predicting patients' risk of 30-day hospital readmission.
method Categorical Co-Frequency Analysis (CoFA) measures diagnosis similarity using random forests.
result Identified three groups of diagnoses with varying readmission risk.
We consider analysis of relational data (a matrix), in which the rows correspond to subjects (e.g., people) and the columns correspond to attributes. The elements of the matrix may be a mix of real and categorical. Each subject and attribute is characterized by a latent binary feature vector, and an inferred matrix map…
We propose Dirichlet Process mixtures of Generalized Linear Models (DP-GLM), a new method of nonparametric regression that accommodates continuous and categorical inputs, and responses that can be modeled by a generalized linear model. We prove conditions for the asymptotic unbiasedness of the DP-GLM regression mean fu…
Extends FJS analysis to general label spaces, including classification and regression.
problem Distribution shift in general label spaces, including covariate and label shifts.
method Proposes a framework for analyzing FJS in general label spaces and generalizes existing results.
result Generalizes FJS analysis to general label spaces, including classification and regression.
Develops multiscale covariance tensor fields for data shape analysis.
problem Quantifying variation of data at all scales.
method Localized covariance tensor fields (CTF) and strong stability theorems.
result CTFs are robust to sampling, noise, and outliers.
Paper proposes Gini distance statistics for estimating feature-label dependence.
problem Identifying statistical dependence between features and categorical labels.
method Generalized Gini distance in RKHS for feature-label dependence estimation.
result Gini distance statistics converge faster and have tighter error bounds than distance covariance.
Bayesian model detects communities in networks with covariates.
problem Extracting meaningful community clusters from network data with covariates.
method Proposes a Bayesian stochastic block model with a covariate-dependent random partition prior.
result Can learn the number of communities via posterior inference.
Selective inference for group lasso estimators across various distributions and covariates.
problem Developing selective inference methods for group lasso estimators.
method Randomized group-regularized optimization problem with post-selection likelihood.
result Selective point estimator and Wald-type confidence regions for regression parameters.
Combines boosting with Gaussian process and mixed effects models.
problem Model misspecifications and independence assumptions in boosting.
method Relaxes zero or linearity assumption in Gaussian process and mixed effects models, and independence assumption in boosting.
result Increased prediction accuracy compared to existing approaches.
Enhances Gaussian process models for handling variable error variances and multiple responses.
problem Limited ability of Gaussian process models to capture abrupt changes and heteroscedastic errors.
method Introduces a novel heteroscedastic Gaussian process (HeGP) framework coupled with variational inference and EM algorithm.
result Effective modeling of multivariate responses with varying error variances.
Study on categorical bundles for gauge theories with new product concept.
problem Developing a new framework for gauge theories involving multiple gauge groups.
method Investigate product bundles in the categorical sense, construct cocycles, and introduce twisted-product bundles.
result Established a new concept of local triviality for categorical bundles.
New combinatorial method for sparse PCA works beyond spiked identity model.
problem Sparse PCA under general covariance matrices.
method Combinatorial truncated power method with global convergence guarantee.
result First combinatorial sparse PCA method provably successful for general covariance matrices.
StructureBoost improves gradient boosting for complex categorical variables efficiently.
problem Efficiently handling complex categorical variables with known structure.
method Two methods to overcome computational obstacles in SCDT enumeration for structured categorical variables.
result StructureBoost outperforms existing packages on complex categorical problems.
Study compares different covariance estimation methods for portfolio allocation.
problem Comparing methods for estimating covariance and precision matrices in portfolio allocation.
method Gaussian Graphical Model (GGM), Shrinkage, Thresholding, Random Matrix Theory (RMT) methods.
result GGM methods outperform other methods in predictive ability for portfolio allocation.
This paper aims to incorporate passive symmetries in machine learning for better generalization.
problem Machine learning's reliance on arbitrary choices leads to passive symmetries that can limit generalization.
method Translation among physics, mathematics, and machine learning to understand and implement passive symmetries.
result Respecting passive symmetries can improve machine learning's ability to generalize.
A new method for backpropagating through categorical distributions.
problem Difficulty in backpropagating through categorical latent variables in neural networks.
method Introducing Gumbel-Softmax distribution for differentiable sampling.
result Gumbel-Softmax estimator outperforms existing methods on tasks with categorical latent variables.
UNTIE learns representations of coupled categorical data.
problem Challenges in learning from unlabeled categorical data with complex couplings.
method UNTIE approach for unsupervised representation learning of heterogeneous couplings.
result UNTIE significantly improves categorical data representations on 25 diverse datasets.
Paper introduces Categorical Normalizing Flows for better handling of categorical data.
problem Limited application of normalizing flows on categorical data due to lack of intrinsic order.
method Categorical Normalizing Flows use continuous transformations to model latent relations in categorical data, optimizing both continuous representation and model likelihood.
result GraphCNF, a permutation-invariant generative model, outperforms state-of-the-art on molecule generation.
Authors conjecture categorically diagonalizable complex for full twists.
problem Categorically diagonalizing the complex of Soergel bimodules for full twists.
method Utilizes categorical diagonalization theory.
result Proves conjecture in type A, categorifies Young idempotents.
This paper proposes a method to reduce complexity in GLMs with categorical predictors.
problem Wasteful, hard-to-interpret, and prone to overfitting of traditional one-hot encoding for high-cardinality categorical predictors.
method Clustering categories of categorical predictors through a numerical method that preserves or improves accuracy while reducing the number of coefficients.
result Clustering categories of categorical predictors reduces complexity substantially without harming accuracy.
Gaussian-Dirichlet posterior dominance proven for sequential categorical data.
problem Sequential learning from categorical observations bounded in [0,1]
method Establishing an ordering between Dirichlet and Gaussian posteriors under N(0,1) noise
result Posterior mean of categorical distribution stochastically dominates Gaussian distribution
CARD models predict the distribution of continuous or categorical responses.
problem Uncertainty in predicting continuous or categorical responses.
method Combines denoising diffusion and conditional mean estimation.
result Outperforms state-of-the-art methods in conditional distribution prediction.
A new method optimises problems with both continuous and categorical inputs.
problem Optimising black-box problems with mixed continuous and categorical inputs.
method Continuous and Categorical Bayesian Optimisation (CoCaBO) combining multi-armed bandits and Bayesian optimisation.
result CoCaBO outperforms existing methods on synthetic and real-world tasks.
The paper shows how integrating categorical semantics can enhance unsupervised domain translation.
problem Improving unsupervised domain translation between perceptually different domains.
method Learning invariant categorical semantic features in an unsupervised manner and conditioning them on the style encoder.
result Conditioning the style encoder on learned categorical semantics improves translation and stylization.
New connections found for quantum flag manifolds modules.
problem Unique connections for relative line modules over quantum flag manifolds.
method Applied general results on quantum principal bundles to Heckenberger-Kolb calculi.
result Found bimodule connections with invertible maps.
New algorithm recovers labels from noisy categorical data.
problem Recovering latent labels from noisy observations in structured instances.
method Approximate algorithm for graphs with categorical variables.
result Logarithmic dependency of Hamming error to the number of categories.
Method trains GANs on categorical data, outperforming existing models.
problem GANs struggle with discrete data, especially categorical values.
method Proposes architectures with multiple Gumbel softmax output layers.
result Proposed architecture outperforms existing models on various datasets.