Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

3.5%7.1%10.6%14.2% · May 202619922001200920182026
48 results for heterogeneous variables

New method quantifies variable importance in causal forests for treatment effect heterogeneity.

problem Lack of understanding how input variables affect treatment effect heterogeneity in causal forests.
method Developed a new importance variable algorithm for causal forests based on the drop and relearn principle.
result Shows how to handle forest retraining without a confounding variable and introduces a corrective term for confounders.

Framework assesses variable importance for heterogeneous treatment effects.

problem High-risk domains need reliable methods to assess treatment effect heterogeneity.
method Inferential framework based on Shapley values and semiparametric theory.
result Valid inference on variable importance for heterogeneous treatment effects.

Method estimates heterogeneous causal effects using instrumental variables.

problem Estimating heterogeneity in causal effects using instrumental variables.
method Two-part method: 1) Discover effect modifiers using machine learning, 2) Test for significance using closed testing.
result Evidence of heterogeneity found in specific subgroups.

Study estimates heterogeneous principal causal effects with binary treatments and intermediate variables.

problem Estimating subgroup effects within strata defined by potential values of an intermediate variable.
method Proposes a framework for estimating and forming confidence intervals for heterogeneous principal causal effects under principal ignorability assumption. Develops several estimators with varying robustness properties.
result Established large-sample theory and analyzed bias contributions of each approach.

Develops HCQRF for estimating heterogeneous treatment effects with censored data.

problem Estimating heterogeneous treatment effects on censored responses with high-dimensional variables.
method Hybrid Censored Quantile Regression Forest (HCQRF) combining random forests and censored quantile regression.
result Demonstrates the effectiveness and stability of HCQRF through simulation studies and real-world application.

Paper identifies and estimates CAPCEs in continuous treatment settings.

problem Estimating heterogeneous causal effects of continuous treatments.
method Instrumental variable approach to identify CAPCEs under weaker conditions.
result Developed three families of CAPCE estimators with statistical properties analyzed.

A new distance for mixed-variable, hierarchical datasets with meta variables.

problem Heterogeneous datasets limit generalizability and performance in machine learning and optimization.
method Developed a modeling framework for mixed-variable and hierarchical domains with meta variables, and a novel distance function.
result The novel distance function allows comparison of heterogeneous datasets, improving model performance.

New method removes hidden confounders for unbiased treatment effect estimation.

problem Bias in treatment effect estimation due to unobserved confounders.
method Proposes a new debiased estimation approach via SVD to handle heterogeneous confounding.
result Established rate of convergence for the estimator under different noise conditions.

FedProx tackles heterogeneity in federated learning networks.

problem Significant variability in systems characteristics and non-identically distributed data in federated networks.
method FedProx is a framework that generalizes and re-parametrizes FedAvg, introducing modifications to handle both systems and statistical heterogeneity.
result FedProx demonstrates significantly more stable and accurate convergence behavior than FedAvg, improving test accuracy by 22% on average in highly heterogeneous settings.

Bayesian machine learning algorithm for causal effects with imperfect compliance.

problem Heterogeneous causal effects in imperfect compliance scenarios.
method Bayesian Causal Forest with Instrumental Variable (BCF-IV) methodology.
result BCF-IV outperforms other techniques in discovering and estimating heterogeneous causal effects.

FAIR-NN finds invariant variables for causal inference across diverse environments.

problem Nonparametric invariance and causal learning in regression models with varying joint distributions.
method FAIR-NN framework using adversarial optimization and neural networks.
result FAIR-NN identifies invariant variables and quasi-causal variables under minimal conditions.

EBM reduces dimensionality for estimating heterogeneous CATEs.

problem Estimating CATEs requires many confounding variables, increasing sample complexity.
method Proposes an EBM that learns a low-dimensional representation of variables.
result EBM representations keep CATE estimates consistent and perform better than other methods.

Paper proposes a method to optimize policies for diverse individuals using heterogeneous data.

problem Learning optimal policies for a heterogeneous population from pre-collected data.
method Individualized offline policy optimization framework for heterogeneous MDPs.
result The proposed P4L algorithm achieves a fast rate of average regret.

VSAE learns from missing heterogeneous data by modeling latent dependencies.

problem Learning from partially-observed heterogeneous data with missingness.
method Variational selective autoencoder (VSAE) models joint distribution of observed, unobserved, and missing data.
result VSAE improves over state-of-the-art models in data generation and imputation tasks.

Proposes a framework to fuse heterogeneous data sources for better modeling.

problem Heterogeneous data sources with different input parameter spaces.
method Input mapping calibration (IMC) and latent variable Gaussian process (LVGP).
result Improved predictive accuracy over single source models.

The paper tackles misspecification in contextual bandits by incorporating arm-specific variables.

problem Misspecification in contextual bandits due to unexplained inter-arm heterogeneity.
method Develops robust contextual bandit algorithms (RoLinUCB and RoLinTS) that incorporate arm-specific variables to address misspecification.
result The developed algorithms bound the nn-round Bayes regret and show superior performance in various misspecification scenarios.

fedCI and fedCI-IOD enable federated causal discovery across diverse datasets with privacy and power enhancements.

problem Causal discovery across multiple datasets with privacy constraints and heterogeneity.
method federated conditional independence test (fedCI) and Integration of Overlapping Datasets (IOD) algorithm extension (fedCI-IOD).
result fedCI-IOD achieves comparable performance to fully pooled analyses, enhancing statistical power and privacy.

This work explores variably scaled kernels to improve non-stationary Gaussian processes.

problem Limited ability of stationary kernels to represent heterogeneous correlation structures.
method Introduces variably scaled kernels to modify correlation structures explicitly.
result Improved reconstruction accuracy and better uncertainty estimates for non-stationary data.

MTHetGNN models complex relations in multivariate time series forecasting.

problem Complex relations among variables in multivariate time series forecasting.
method Designs a relation embedding module and a temporal embedding module, using graph neural networks and CNNs.
result Achieves state-of-the-art results in multivariate time series forecasting.

New algorithms for clustering and synthetic data generation of heterogeneous tabular datasets.

problem Clustering and generating synthetic data from heterogeneous tabular datasets with hidden cluster structure.
method Developed MMM and MMMsynth algorithms for clustering and synthetic data generation.
result MMMsynth algorithm outperforms other literature tabular-data generators and approaches real data performance.

Proposes a flexible deep learning model for complex distributions.

problem Complex shapes, strong skews, and multiple modes in output variable distributions.
method Uncountable Mixture of Asymmetric Laplacians (UMAL) deep learning framework.
result UMAL can estimate heterogeneous distributions without strong assumptions.

Model learns low-dimensional representation from heterogeneous data with missing values.

problem Handling high-dimensional, noisy, and missing data from clinical records.
method Latent Gaussian process with composite likelihoods and numerical quadrature.
result Improves upon existing GPLVM methods for heterogeneous data.

The paper addresses misspecification in econometric models of discrete unobserved heterogeneity.

problem Misspecification in econometric models of discrete unobserved heterogeneity.
method Generalizing previous approaches to allow multiple latent variables, developing inference results for a k-means style estimator, and proposing information criteria for model selection.
result Over-fitting can be severe in k-means style estimators when the number of clusters is over-specified.

A new model decomposes market variability into interpretable components.

problem Understanding the factors driving market variability and predicting future movements.
method H-SGDLM framework with HAR-RV model for GPU-scalable multivariate volatility estimation.
result Superior performance in predicting large moves and longer-term market variability.

Efficiently solves heterogeneous QPs by reducing variables using instance-specific projections.

problem Solving high-dimensional quadratic programming problems efficiently.
method Data-driven framework with a graph neural network generating projections tailored to each QP instance.
result Produces high-quality solutions with reduced computation time, outperforming existing methods.

The paper introduces a new volatility model for state heterogeneous financial markets using high-frequency data.

problem State heterogeneity in financial volatility processes.
method Developed a state heterogeneous GARCH-Ito (SG-Ito) model based on continuous Ito diffusion process.
result Empirical studies reveal various state heterogeneities in S&P 500 index volatility.

Framework for causal discovery from changing data.

problem Challenges of causal discovery in heterogeneous or nonstationary data.
method Constraint-based CD-NOD framework for causal skeleton and orientation recovery, independent changes detection.
result Efficient estimation of causal mechanism changes and low-dimensional representation of nonstationarity.

New methods for estimating complex causal effects in econometrics.

problem Estimating causal parameters in short panel data models using nested nonparametric instrumental variable regression.
method Introducing techniques to limit ill-posedness in nested NPIV, providing explicit mean square rates and efficient inference.
result Explicit mean square rates for nested NPIV and efficient inference for causal parameters.

This paper introduces a general Bayesian non- parametric latent feature model suitable to per- form automatic exploratory analysis of heterogeneous datasets, where the attributes describing each object can be either discrete, continuous or mixed variables. The proposed model presents several important properties. First…

2017-07-26abs ↗pdf ↗

New model clusters cells and individuals, revealing genetic influences on cell types.

problem Clustering nested data with group-level and observation-level variables.
method Nested Atoms Model (NAM), Bayesian nonparametric approach.
result Identifies clusters of genetically similar individuals with homogeneous cell-type profiles.

MISTR improves HTE estimation in survival data with heavy censoring and instrumental variables.

problem Estimating HTE in survival data with censoring and unobserved confounders.
method MISTR uses recursively imputed survival trees to handle censoring and instrumental variables.
result MISTR outperforms existing methods under heavy censoring and instrumental variable settings.

Paper uses CT-IV to estimate causal effects in non-randomized settings.

problem Estimating causal effects in non-randomized observational studies.
method Modified Causal Tree (CT-IV) algorithm combining CART and IV framework.
result Demonstrates efficiency in handling heterogeneity of causal effects.

Modern datasets are becoming heterogeneous. To this end, we present in this paper Mixed-Variate Restricted Boltzmann Machines for simultaneously modelling variables of multiple types and modalities, including binary and continuous responses, categorical options, multicategorical choices, ordinal assessment and category…

2014-08-06abs ↗pdf ↗

New model identifies microbial subcommunities robustly, accounting for cross-sample heterogeneity.

problem Inference in LDA is sensitive to the number of subcommunities and often creates artificial ones.
method Incorporates logistic-tree normal (LTN) model into LDA to account for cross-sample heterogeneity.
result Restores robustness of inference and identifies meaningful subcommunities.

Tree ensemble method tackles multi-objective constrained optimization in energy systems.

problem Complex, multi-objective, and constrained optimization problems in energy systems.
method Data-driven tree ensemble approach for black-box problems with heterogeneous variable spaces.
result Competitive performance and sampling efficiency compared to state-of-the-art tools.

HL-VAE extends VAE for heterogeneous temporal and longitudinal data.

problem Handling heterogeneous data in temporal and longitudinal datasets.
method Proposes HL-VAE, an extension of existing VAEs for temporal and longitudinal data, incorporating likelihood models for various data types.
result HL-VAE achieves competitive performance in missing value imputation and predictive accuracy.