Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

306089119 · May 202619922001200920182026
48 results for categorical covariates

The paper proposes a method to analyze categorical feature interactions in large datasets using graph covariance and LLMs.

problem Analyzing complex datasets with numerous categorical features and timestamps.
method Binarization of categorical features using one-hot encoding, computation of graph covariance, identifying significant feature pairs, and using LLMs to generate explanations.
result The method identifies meaningful feature pairs and potential data stories underlying categorical feature interactions.

Bayesian model improves categorization of explosions from sparse data.

problem Challenges in categorizing explosions from limited data.
method Bayesian update to Event Categorization Matrix model with Bayesian Decision Theory.
result Consistent gains in overall accuracy and lower false negative rates.

Reduces high granularity and dimensionality in hierarchical categorical variables.

problem Overfitting and estimation issues in predictive models due to high granularity and dimensionality.
method Entity embedding and top-down clustering algorithm to reduce granularity and dimensionality.
result The reduced hierarchy improves model fit and complexity balance.

Probabilistic learning for binary classification with categorical variables.

problem Binary classification with categorical covariates.
method Probabilistic analysis and two algorithms for learning boolean functions.
result Effective learning of boolean functions from binary data.

DPERC efficiently estimates covariance matrices for mixed data with missing values.

problem Estimating covariance matrices for datasets with missing values and mixed features.
method Direct Parameter Estimation for Randomly Missing Data with Categorical Features (DPERC).
result DPERC outperforms other methods in estimating covariance matrices for mixed data with missing values.

A new method for finding significant patterns in stratified data.

problem Statistical challenges in finding enriched itemsets in stratified data.
method Proposes a strategy and efficient algorithm for significant pattern mining in the presence of categorical covariates.
result Efficiently corrects for multiple testing in pattern mining while retaining statistical power.

lgpr interprets longitudinal data to find covariate effects.

problem Inferring covariate effects from longitudinal data with complex structures and interactions.
method Additive Gaussian processes for nonparametric analysis.
result lgpr outperforms previous methods in identifying relevant covariates.

A new model integrates covariates with grade of membership analysis for better latent structure recovery.

problem Improving latent structure recovery in multivariate categorical data analysis.
method Covariate-assisted grade of membership model exploiting shared low-rank simplex geometry.
result Auxiliary covariates can provably improve latent structure recovery, leading to faster convergence rates.

A method for representing and comparing categorical trajectories using multivariate functional principal components.

problem Statistical description and comparison of categorical trajectories.
method Transforming categorical trajectories into binary indicator functions and applying multivariate functional principal components analysis.
result Consistent estimators of mean trajectories and covariance functions are obtained under weak regularity assumptions.

A new Bayesian multinomial regression model using permuted and augmented stick-breaking.

problem Modeling categorical response variables given covariates.
method Permuted and augmented stick-breaking (paSB) construction.
result Transforms multinomial regression into regression of stick-specific binary variables.

Two novel methods estimate multiple FDR directions for binary categorical responses.

problem Estimating multiple FDR directions for categorical responses.
method Information maximization and square loss mutual information.
result Statistical consistency of the proposed methods established.

A new matching method for causal inference that handles irrelevant variables and missing data.

problem Creating high-quality treatment-control matches for categorical data in social sciences.
method A weighted Hamming distance matching method that considers covariate importance and creates a hierarchy of covariate combinations.
result The method produces high-quality matches and handles irrelevant variables and missing data.

CardiCat generates synthetic data for high-cardinality tabular datasets.

problem Learning complexities of high-cardinality categorical features in tabular data.
method Substitutes one-hot encoding with regularized dual encoder-decoder embedding layers.
result Generates high-quality synthetic data with a smaller parameter space.

Standardization boosts sparse group lasso performance for categorical variables.

problem Improving performance of sparse group lasso with categorical variables.
method Column-wise scaling and re-scaling of coefficients, without orthonormalization.
result Column-wise scaling and re-scaling of coefficients achieves the same effect as orthonormalization.

MCPCA improves data dimensionality by maximizing nonlinear correlations.

problem PCA's limitations in handling nonlinearity and categorical data.
method MCPCA computes nonlinear transformations of variables to maximize covariance matrix Ky Fan norm.
result MCPCA outperforms other methods in dimensionality reduction tasks.

FLAME efficiently matches high-dimensional categorical datasets for causal inference.

problem Matching treatment and control units based on covariate information in causal inference.
method FLAME learns a distance metric using a hold-out training set and uses query processing techniques for large datasets.
result FLAME achieves significantly better performance than other matching methods, scaling to huge datasets.

This study improves hyperparameter optimization for categorical and non-normal data.

problem Bayesian hyperparameter optimization struggles with categorical hyperparameters and non-normal data.
method Integrates conformalized quantile regression to address estimation weaknesses and provides robust calibration guarantees.
result Quantile surrogate architectures and acquisition functions yield superior performance compared to existing methods.

Global covariance pooling improves deep CNNs' representation and generalization.

problem Capturing richer statistics of deep features for better representation and generalization.
method Integrates global covariance pooling into deep CNNs, addressing challenges with robust covariance estimation and geometry exploitation.
result Proposes MPN-COV Pooling and a Gaussian embedding network, achieving state-of-the-art performance.

Efficiently models categorical data with low to medium class overlap, improving accuracy over standard distributions.

problem Poor parameter estimates and accuracy in multinomial and Dirichlet multinomial distributions when assumptions are violated.
method Introduces Beta-Liouville multinomial distribution and efficient estimation methods.
result Beta-Liouville multinomial outperforms standard distributions on two out of four datasets.

Categorical Co-Frequency Analysis clusters diagnoses to predict hospital readmissions.

problem Predicting patients' risk of 30-day hospital readmission.
method Categorical Co-Frequency Analysis (CoFA) measures diagnosis similarity using random forests.
result Identified three groups of diagnoses with varying readmission risk.

We propose Dirichlet Process mixtures of Generalized Linear Models (DP-GLM), a new method of nonparametric regression that accommodates continuous and categorical inputs, and responses that can be modeled by a generalized linear model. We prove conditions for the asymptotic unbiasedness of the DP-GLM regression mean fu…

2009-09-28abs ↗pdf ↗

Extends FJS analysis to general label spaces, including classification and regression.

problem Distribution shift in general label spaces, including covariate and label shifts.
method Proposes a framework for analyzing FJS in general label spaces and generalizes existing results.
result Generalizes FJS analysis to general label spaces, including classification and regression.

Paper proposes Gini distance statistics for estimating feature-label dependence.

problem Identifying statistical dependence between features and categorical labels.
method Generalized Gini distance in RKHS for feature-label dependence estimation.
result Gini distance statistics converge faster and have tighter error bounds than distance covariance.

Selective inference for group lasso estimators across various distributions and covariates.

problem Developing selective inference methods for group lasso estimators.
method Randomized group-regularized optimization problem with post-selection likelihood.
result Selective point estimator and Wald-type confidence regions for regression parameters.

Enhances Gaussian process models for handling variable error variances and multiple responses.

problem Limited ability of Gaussian process models to capture abrupt changes and heteroscedastic errors.
method Introduces a novel heteroscedastic Gaussian process (HeGP) framework coupled with variational inference and EM algorithm.
result Effective modeling of multivariate responses with varying error variances.

Study on categorical bundles for gauge theories with new product concept.

problem Developing a new framework for gauge theories involving multiple gauge groups.
method Investigate product bundles in the categorical sense, construct cocycles, and introduce twisted-product bundles.
result Established a new concept of local triviality for categorical bundles.

StructureBoost improves gradient boosting for complex categorical variables efficiently.

problem Efficiently handling complex categorical variables with known structure.
method Two methods to overcome computational obstacles in SCDT enumeration for structured categorical variables.
result StructureBoost outperforms existing packages on complex categorical problems.

Study compares different covariance estimation methods for portfolio allocation.

problem Comparing methods for estimating covariance and precision matrices in portfolio allocation.
method Gaussian Graphical Model (GGM), Shrinkage, Thresholding, Random Matrix Theory (RMT) methods.
result GGM methods outperform other methods in predictive ability for portfolio allocation.

This paper aims to incorporate passive symmetries in machine learning for better generalization.

problem Machine learning's reliance on arbitrary choices leads to passive symmetries that can limit generalization.
method Translation among physics, mathematics, and machine learning to understand and implement passive symmetries.
result Respecting passive symmetries can improve machine learning's ability to generalize.

A new method for backpropagating through categorical distributions.

problem Difficulty in backpropagating through categorical latent variables in neural networks.
method Introducing Gumbel-Softmax distribution for differentiable sampling.
result Gumbel-Softmax estimator outperforms existing methods on tasks with categorical latent variables.

UNTIE learns representations of coupled categorical data.

problem Challenges in learning from unlabeled categorical data with complex couplings.
method UNTIE approach for unsupervised representation learning of heterogeneous couplings.
result UNTIE significantly improves categorical data representations on 25 diverse datasets.

Paper introduces Categorical Normalizing Flows for better handling of categorical data.

problem Limited application of normalizing flows on categorical data due to lack of intrinsic order.
method Categorical Normalizing Flows use continuous transformations to model latent relations in categorical data, optimizing both continuous representation and model likelihood.
result GraphCNF, a permutation-invariant generative model, outperforms state-of-the-art on molecule generation.

This paper proposes a method to reduce complexity in GLMs with categorical predictors.

problem Wasteful, hard-to-interpret, and prone to overfitting of traditional one-hot encoding for high-cardinality categorical predictors.
method Clustering categories of categorical predictors through a numerical method that preserves or improves accuracy while reducing the number of coefficients.
result Clustering categories of categorical predictors reduces complexity substantially without harming accuracy.

Gaussian-Dirichlet posterior dominance proven for sequential categorical data.

problem Sequential learning from categorical observations bounded in [0,1]
method Establishing an ordering between Dirichlet and Gaussian posteriors under N(0,1) noise
result Posterior mean of categorical distribution stochastically dominates Gaussian distribution

A new method optimises problems with both continuous and categorical inputs.

problem Optimising black-box problems with mixed continuous and categorical inputs.
method Continuous and Categorical Bayesian Optimisation (CoCaBO) combining multi-armed bandits and Bayesian optimisation.
result CoCaBO outperforms existing methods on synthetic and real-world tasks.

The paper shows how integrating categorical semantics can enhance unsupervised domain translation.

problem Improving unsupervised domain translation between perceptually different domains.
method Learning invariant categorical semantic features in an unsupervised manner and conditioning them on the style encoder.
result Conditioning the style encoder on learned categorical semantics improves translation and stylization.