Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

3517021,0521,403 · Jun 202019922001200920172026
48 results for Covariate-based models

Novel approach for SEM in small samples with p>np>n.

problem Small sample size and p>np>n issues in factor-based SEM.
method Reformulates covariance structure into self-covariance and cross-covariance, defines a feasible set with relative error constraint.
result Improved stability and directional information in small-sample settings.

Study predicts when ALS patients will lose speech, swallowing, etc. based on covariates.

problem Predicting when ALS patients will experience significant functional decline.
method Multi-event survival analysis, covariate-based models.
result Covariate-based models outperform Kaplan-Meier estimator in predicting time-to-event outcomes.

Kernel measures similarity of nonlinear causal structures in heterogeneous populations.

problem Learning causal structure in populations with diverse underlying structures.
method Distance covariance-based kernel for measuring similarity of causal structures.
result Kernel enables clustering of homogeneous subpopulations for causal structure learning.

The study forecasts portfolio volatility using cointegrated asset dynamics.

problem Forecasting volatility in portfolios with high accuracy.
method Developed HVR/DVR ratios and used Vector Error Correction Model (VECM) to forecast volatility.
result VECM forecasts of portfolio volatility have lower MAPE than covariance-based forecasts.

This paper addresses two crucial problems of learning disentangled image representations, namely controlling the degree of disentanglement during image editing, and balancing the disentanglement strength and the reconstruction quality. To encourage disentanglement, we devise a distance covariance based decorrelation re…

2020-01-22abs ↗pdf ↗

Method uses random forest with distance covariance for transfer learning in healthcare.

problem Transfer learning in random forests with sparse differences between source and target.
method Distance covariance-based feature weights in residual random forest.
result Upper bound on mean square error rate for transfer learning in RF.

Proposes a new method to estimate variable importance in black box models, mitigating correlation effects.

problem Correlation between covariates affects the interpretation of variable importance parameters.
method Develops a modified LOCO (Leave Out COvariates) method and uses semiparametric models for estimation.
result Shows how to estimate a modified LOCO method that mitigates correlation effects.

TraCeR uses transformers to analyze survival data with longitudinal covariates.

problem Handling longitudinal covariates and assessing model calibration in survival analysis.
method Transformer-based survival analysis framework with factorized self-attention architecture.
result TraCeR achieves significant performance improvements over state-of-the-art methods.

BICauseTree improves causal effect estimation by identifying clusters and balancing treatment allocation.

problem Improving interpretability and transparency in causal effect models from observational data.
method Hierarchical bias-driven stratification using decision trees with a customized objective function.
result BICauseTree provides interpretable causal effect estimation and is comparable to existing methods.

New method links covariates to CTMCs using RKHS, improving state transitions modeling.

problem Traditional multistate models rely on linear relationships, limiting flexibility.
method Nonparametric approach using RKHS, with Frequentist and Bayesian versions.
result Effective in identifying nonlinear transition functions and predicting long-term behaviors.

We study the value of information in sequential compressed sensing by characterizing the performance of sequential information guided sensing in practical scenarios when information is inaccurate. In particular, we assume the signal distribution is parameterized through Gaussian or Gaussian mixtures with estimated mean…

2015-09-01abs ↗pdf ↗

We solve a high-dimensional model where nonlinear autoencoders detect hidden structure missed by PCA.

problem Hidden structure in high-dimensional data not detected by PCA.
method Tractable spiked model with two latent factors, one visible and one uncorrelated.
result Nonlinear autoencoders can extract hidden structure missed by PCA, even if reconstruction loss is higher.

This article proposes a method to consistently estimate functionals 1pi=1pf(λi(C1C2))\frac1p\sum_{i=1}^pf(λ_i(C_1C_2)) of the eigenvalues of the product of two covariance matrices C1,C2Rp×pC_1,C_2\in\mathbb{R}^{p\times p} based on the empirical estimates λi(C^1C^2)λ_i(\hat C_1\hat C_2) ($\hat C_a=\frac1{n_a}\sum_{i=1}^{n_a} x_i^{(a)}x_i^{(a){\sf T}…

2019-03-08abs ↗pdf ↗

We study the problem of structured output learning from a regression perspective. We first provide a general formulation of the kernel dependency estimation (KDE) problem using operator-valued kernels. We show that some of the existing formulations of this problem are special cases of our framework. We then propose a c…

2012-05-10abs ↗pdf ↗

Cryptocurrency markets exhibit violent, synchronised drawdowns, challenging diversification claims.

problem Cryptocurrency markets' violent drawdowns challenge diversification claims.
method Dynamic conditional tail dependence analysis
result Near-complete and stable lower-tail graph, upper tail that thins over time, dissolution of token categories into a core.

STVNN models spatiotemporal data using covariance matrices.

problem Challenges in modeling spatiotemporal interactions in multivariate time series.
method Introduces SpatioTemporal coVariance Neural Network (STVNN) that operates on sample covariance matrix and uses joint spatiotemporal convolutions.
result STVNN is stable to online estimation uncertainties and outperforms temporal PCA.

A new algorithm COVA-FC improves subgroup-fair clustering efficiency.

problem Challenges in making cluster assignments independent of sensitive attributes in subgroups.
method Defining a subgroup-fairness gap, deriving a covariance-based surrogate, and introducing a continuous relaxation for efficient optimization.
result COVA-FC achieves competitive cost-fairness trade-offs and improves computational efficiency.

A scalable algorithm for GP regression selects relevant covariates efficiently.

problem Scalable variable selection in large GP regression models.
method VGPR algorithm using Vecchia approximation for sparse precision matrix, mini-batch subsampling.
result Improved scalability and accuracy in selecting relevant covariates.

Unsupervised method detects earthquakes from raw waveforms, generalizing across datasets.

problem Lack of labeled data for earthquake detection.
method Uses deep autoencoders with cross-covariance triggering at bottleneck.
result Performance comparable to supervised methods, with strong cross-dataset generalization.

FVNNs use graph convolutions on fair covariance estimates to improve fairness in machine learning.

problem Data-driven methods can encode biases in sample covariance matrices, leading to unfair treatment of different subpopulations.
method FVNNs perform graph convolutions on fair covariance estimates and use a fairness regularizer in the loss function.
result FVNNs provide a flexible model that is intrinsically fairer than PCA approaches and can handle low sample regimes.

Estimates Gaussian location model with ridge regularization, comparing variational and spectral methods.

problem Estimating parameters in Gaussian location model with regularization.
method Ridge-regularized log-density-ratio estimation, variational and spectral approaches.
result Regularized variational estimator has lower risk with many observations, spectral estimator with fewer observations.

This research tackles sample complexity in causal graph recovery with temporal heterogeneity.

problem Recovering a unique causal graph from observational data with temporal heterogeneity.
method Integrates time-series dynamics and multi-environment heterogeneity to constrain the problem, enabling a rigorous analysis of statistical limits.
result Unified necessary identifiability conditions and explicit information-theoretic bounds quantify the sample complexity under different noise distributions.

Detection of power-law behavior and studies of scaling exponents uncover the characteristics of complexity in many real world phenomena. The complexity of financial markets has always presented challenging issues and provided interesting findings, such as the inverse cubic law in the tails of stock price fluctuation di…

2018-03-22abs ↗pdf ↗

Flexible spatial models improve predictive performance over nonstationary alternatives.

problem Improving predictive performance in nonstationary spatial modeling.
method Introduces a modular parametric covariance function that extends nonstationary spatial models.
result The proposed covariance function outperforms nonparametric methods in predictive performance.

STICC clusters geographic objects considering both spatial contiguity and attributes.

problem Discovering repeated geographic patterns with spatial contiguity.
method Spatial Toeplitz Inverse Covariance-Based Clustering (STICC) method.
result STICC significantly outperforms baseline methods in adjusted rand index and macro-F1 score.

The paper proposes DEA to make graph neural networks fairer in link prediction.

problem Graph neural networks can unfairly prioritize certain social groups in link prediction.
method Drop Edges and Adapt (DEA) fine-tuning strategy with covariance constraints.
result DEA improves fairness and accuracy in link prediction tasks.

The paper tackles temporal coverage bias in financial panel data, proposing a structuring framework to correct for incomplete histories.

problem Incomplete histories of financial instruments lead to biased panel data.
method Formalizes the problem and proposes a coverage-aware structuring framework using structured metadata and an availability matrix.
result The framework reveals substantial distortions in return dynamics and volatility when naive temporal alignment is used.

Proposes PEMI for online selective conformal prediction with asymmetric rules.

problem Challenges of handling asymmetric selection mechanisms in online selective conformal prediction.
method PEMI: permutation-based framework for selective conformal prediction with arbitrary asymmetric selection rules.
result Achieves exact selection-conditional coverage for any asymmetric selection mechanism and any prediction model.

Study uses detrended cross-correlation to analyze cryptocurrency market, revealing robust collective modes and distinguishing interdependencies.

problem Nonstationarity, long-range memory, and heavy-tailed fluctuations obscure traditional correlations in complex systems.
method Constructs detrended correlation matrices using multifractal detrended cross-correlation coefficient ρrρ_r to emphasize different fluctuations.
result Detrending and fluctuation analysis reveal distinct spectral properties from random case, identifying market and sectoral components.

Given i.i.d. observations of a random vector XRpX \in \mathbb{R}^p, we study the problem of estimating both its covariance matrix ΣΣ^*, and its inverse covariance or concentration matrix {Θ=(Σ)1Θ^* = (Σ^*)^{-1}.} We estimate ΘΘ^* by minimizing an 1\ell_1-penalized log-determinant Bregman divergence; in the multivariate G…

2008-11-21abs ↗pdf ↗

The paper introduces BCART models for aggregate claim amount, improving frequency-severity and joint modeling.

problem Modeling aggregate claim amount with frequency-severity and joint dependencies.
method Developed three types of BCART models: frequency-severity, sequential, and joint models. Used various distributions for claim severity data.
result Weibull distribution outperforms gamma and lognormal for right-skewed, heavy-tailed claim severity data.

The paper uses model-based trees to create interpretable surrogate models for complex machine learning models.

problem Interpreting complex machine learning models.
method Using model-based trees to partition feature space and create interpretable models.
result Model-based trees generate optimal surrogate models that balance interpretability and performance.