New Krylov subspace methods speed up mixed-effects models with crossed random effects.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A fast bootstrap method estimates cross-validation standard error.
Proposes a new cross-validation method to estimate model performance.
We study the problem of treatment effect estimation in randomized experiments with high-dimensional covariate information, and show that essentially any risk-consistent regression adjustment can be used to obtain efficient estimates of the average treatment effect. Our results considerably extend the range of settings …
Study compares Islamic banks' accounting and market performance.
Improved accuracy in machine learning with Cross-Cluster Weighted Forests.
We derive streamlined mean field variational Bayes algorithms for fitting linear mixed models with crossed random effects. In the most general situation, where the dimensions of the crossed groups are arbitrarily large, streamlining is hindered by lack of sparseness in the underlying least squares system. Because of th…
We confirm universal behaviors such as eigenvalue distribution and spacings predicted by Random Matrix Theory (RMT) for the cross correlation matrix of the daily stock prices of Tokyo Stock Exchange from 1993 to 2001, which have been reported for New York Stock Exchange in previous studies. It is shown that the random …
We analyze the complexity of Gibbs samplers for inference in crossed random effect models used in modern analysis of variance. We demonstrate that for certain designs the plain vanilla Gibbs sampler is not scalable, in the sense that its complexity is worse than proportional to the number of parameters and data. We thu…
Paper proposes a novel method to assess treatment effect estimators using cross-validation.
Develops a new model for cross-currency derivatives pricing.
The cross-correlations between the exchange rate fluctuations of 74 currencies over the period 1995-2012 are analyzed in this paper. The eigenvalue distribution of the cross-correlation matrix exhibits a bulk which approximately matches the bounds predicted from random matrices constructed using mutually uncorrelated t…
Introduces intrinsic Riemannian cross-covariance for manifold-valued random objects.
We analyse the dependence of stock return cross-correlations on the sampling frequency of the data known as the Epps effect: For high resolution data the cross-correlations are significantly smaller than their asymptotic value as observed on daily data. The former description implies that changing trading frequency sho…
Paper uses ML to improve A/B testing for complex treatment effects.
Estimates causal effect of managed care plans on NYC Medicaid spending.
The paper compares methods for estimating heterogeneous treatment effects using multiple randomized trials.
Paper adapts DML for panel data, addressing unobserved heterogeneity.
MOMENT selects and estimates mixed-effects models using moment identities.
Develops statistical inference for ML-discovered heterogeneous treatment effects.
Statistical machine learning models should be evaluated and validated before putting to work. Conventional k-fold Monte Carlo Cross-Validation (MCCV) procedure uses a pseudo-random sequence to partition instances into k subsets, which usually causes subsampling bias, inflates generalization errors and jeopardizes the r…
This research shows how to learn shared representations from unpaired data.
New method cleans cross-covariance matrices for better financial forecasting.
Random forests perform bootstrap-aggregation by sampling the training samples with replacement. This enables the evaluation of out-of-bag error which serves as a internal cross-validation mechanism. Our motivation lies in using the unsampled training samples to improve each decision tree in the ensemble. We study the e…
This paper improves cross-domain learning using random forests for manifold alignment.
We consider a natural model of random knotting- choose a knot diagram at random from the finite set of diagrams with n crossings. We tabulate diagrams with 10 and fewer crossings and classify the diagrams by knot type, allowing us to compute exact probabilities for knots in this model. As expected, most diagrams with 1…
The cross-correlation matrix of daily returns of stock market indices in a diverse set of 37 countries worldwide was analyzed. Comparison of the spectrum of this matrix with predictions of random matrix theory provides an empirical evidence of strong interactions between individual economies, as manifested by three lar…
Linear-cost unbiased estimates for complex models via couplings.
New method stabilizes machine learning predictions across random seeds.
Machine learning reduces variance in online experiment results.
We use the Chebyshev knot diagram model of Koseleff and Pecker in order to introduce a random knot diagram model by assigning the crossings to be positive or negative uniformly at random. We give a formula for the probability of choosing a knot at random among all knots with bridge index at most 2. Restricted to this c…
Study relaxes identification assumptions for natural direct effects in non-randomized settings.
ECV method optimizes ensemble parameters for randomized ensembles.
We review the decomposition method of stock return cross-correlations, presented previously for studying the dependence of the correlation coefficient on the resolution of data (Epps effect). Through a toy model of random walk/Brownian motion and memoryless renewal process (i.e. Poisson point process) of observation ti…
Machine learning models predict depression risk based on various factors.
The study assesses external validity by evaluating worst-case treatment effects across subpopulations.
We describe a model of random links based on random 4-valent maps, which can be sampled due to the work of Schaeffer. We will look at the relationship between the combinatorial information in the diagram and the hyperbolic volume. Specifically, we show that for random alternating diagrams, the expected hyperbolic volum…
In a previous work, the first and third authors studied a random knot model for all two-bridge knots using billiard table diagrams. Here we present a closed formula for the distribution of the crossing numbers of such random knots. We also show that the probability of any given knot appearing in this model decays to ze…
In a very high-dimensional vector space, two randomly-chosen vectors are almost orthogonal with high probability. Starting from this observation, we develop a statistical factor model, the random factor model, in which factors are chosen at random based on the random projection method. Randomness of factors has the con…
In this paper we study the problems of estimating heterogeneity in causal effects in experimental or observational studies and conducting inference about the magnitude of the differences in treatment effects across subsets of the population. In applications, our method provides a data-driven approach to determine which…
We develop importance sampling based efficient simulation techniques for three commonly encountered rare event probabilities associated with random walks having i.i.d. regularly varying increments; namely, 1) the large deviation probabilities, 2) the level crossing probabilities, and 3) the level crossing probabilities…
This paper studies a class of optimal multiple stopping problems driven by Lévy processes. Our model allows for a negative effective discount rate, which arises in a number of financial applications, including stock loans and real options, where the strike price can potentially grow at a higher rate than the original d…
We present an original and novel method based on random matrix approach that enables to distinguish the respective role of temporal autocorrelations inside given time series and cross correlations between various time series. The proposed algorithm is based on properties of Wigner eigenspectrum of random matrices inste…
Study improves prediction accuracy and uncertainty for mobile sensor data using randomized neural networks.
We build a simple diagnostic criterion for approximate factor structure in large cross-sectional equity datasets. Given a model for asset returns with observable factors, the criterion checks whether the error terms are weakly cross-sectionally correlated or share at least one unobservable common factor. It only requir…
PLS-SVD struggles with missing data in multimodal datasets, showing a phase transition in performance.
XTNet estimates complex cross-treatment effects in multi-category, multi-valued settings.
We investigate the accuracy of the two most common estimators for the maximum expected value of a general set of random variables: a generalization of the maximum sample average, and cross validation. No unbiased estimator exists and we show that it is non-trivial to select a good estimator without knowledge about the …