Unified Bayesian framework improves clinical trial hypothesis testing.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Efficiently estimates Cox model coefficients without sharing data.
Unified framework for fair decision-making across diverse groups.
We investigate the volatility return intervals in the NYSE and FOREX markets. We explain previous empirical findings using a model based on the interacting agent hypothesis instead of the widely-used efficient market hypothesis. We derive macroscopic equations based on the microscopic herding interactions of agents and…
The paper improves confidence intervals for test error using cross-validation.
The upsilon distribution, the sum of independent chi random variates and a normal, is introduced. As a special case, the upsilon distribution includes Lecoutre's lambda-prime distribution. The upsilon distribution finds application in Frequentist inference on the Sharpe ratio, including hypothesis tests on independent …
A new method for adaptive experiments improves inference.
A note on extending Chernoff bound for unit interval random variables.
While statistics focusses on hypothesis testing and on estimating (properties of) the true sampling distribution, in machine learning the performance of learning algorithms on future data is the primary issue. In this paper we bridge the gap with a general principle (PHI) that identifies hypotheses with best predictive…
Active inference uses machine learning to prioritize data labeling for more efficient statistical inference.
Long-term relative arbitrage exists in markets where the excess growth rate of the market portfolio is bounded away from zero. Here it is shown that under a time-homogeneity hypothesis this condition will also imply the existence of relative arbitrage over arbitrarily short intervals.
Hypothesis testing in the linear regression model is a fundamental statistical problem. We consider linear regression in the high-dimensional regime where the number of parameters exceeds the number of samples (). In order to make informative inference, we assume that the model is approximately sparse, that is th…
A new approach to the understanding of complex behavior of financial markets index using tools from thermodynamics and statistical physics is developed. Physical complexity, a magnitude rooted in Kolmogorov-Chaitin theory is applied to binary sequences built up from real time series of financial markets indexes. The st…
A new approach to the understanding of the complex behavior of financial markets index using tools from thermodynamics and statistical physics is developed. Physical complexity, a magnitude rooted in the Kolmogorov-Chaitin theory is applied to binary sequences built up from real time series of financial markets indices…
The paper discusses methods for interval estimation of coefficients in penalized regression models for insurance data.
For analysis of a high-dimensional dataset, a common approach is to test a null hypothesis of statistical independence on all variable pairs using a non-parametric measure of dependence. However, because this approach attempts to identify any non-trivial relationship no matter how weak, it often identifies too many rel…
We develop a framework for post model selection inference, via marginal screening, in linear regression. At the core of this framework is a result that characterizes the exact distribution of linear functions of the response , conditional on the model being selected (``condition on selection" framework). This allows…
This chapter reviews statistical tools for reinforcement learning.
Develops a hypothesis testing framework for generalized Thurstone models.
The paper proposes an efficient method for estimating ATEs using adaptive experiments.
Paper develops PAC verification for hypothesis classes and statistical algorithms.
Paper proposes a new UCB approach for estimating maximum mean.
Proposes a method for inference in high-dimensional classification with non-differentiable surrogate losses.
We present an empirical study of the subordination hypothesis for a stochastic time series of a stock price. The fluctuating rate of trading is identified with the stochastic variance of the stock price, as in the continuous-time random walk (CTRW) framework. The probability distribution of the stock price changes (log…
Modeling disease progression in healthcare administrative databases is complicated by the fact that patients are observed only at irregular intervals when they seek healthcare services. In a longitudinal cohort of 76,888 patients with chronic obstructive pulmonary disease (COPD), we used a continuous-time hidden Markov…
A universal method for hypothesis tests and confidence sets without regularity conditions.
A method for efficient statistical inference from online algorithms.
Study geodesics on random hyperbolic surfaces, finding variance similar to prime number theory.
The paper tackles high-dimensional mixed linear regression with unknown parameters and proposes methods for estimation, confidence intervals, and hypothesis testing.
This paper develops embeddings that preserve likelihood-based statistical inference.
Method constructs confidence regions for linear models with arbitrary predictors.
Proposes online debiasing estimators for adaptive linear regression.
Develops tools to audit ML models for bias and unfairness.
The efficient market hypothesis has been considered one of the most controversial arguments in finance, with the academia divided between who claims the impossibility of beating the market and who believes that it is possible to gain over the average profits. If the hypothesis holds, it means, as suggested by Burton Ma…
Unified framework for statistical inference in gradient boosting regression.
New methods estimate interventional effects with multiple mediators using machine learning.
For many causal effect parameters of interest, doubly robust machine learning (DRML) estimators are the state-of-the-art, incorporating the good prediction performance of machine learning; the decreased bias of doubly robust estimators; and the analytic tractability and bias reduction of sample splitting wi…
This paper optimizes predicting support and resistance levels in financial markets.
The paper tackles individual fairness in ML models, developing statistical methods to detect bias.
This work develops formal statistical inference procedures for machine learning ensemble methods. Ensemble methods based on bootstrapping, such as bagging and random forests, have improved the predictive accuracy of individual trees, but fail to provide a framework in which distributional results can be easily determin…
e-LOND algorithm controls FDR in online testing with arbitrary dependencies.
We propose a statistical approach to tornadoes modeling for predicting and simulating occurrences of tornadoes and accumulated cost distributions over a time interval. This is achieved by modeling the tornadoes intensity, measured with the Fujita scale, as a stochastic process. Since the Fujita scale divides tornadoes …
Virtual links are generalizations of classical links that can be represented by links embedded in a ``thickened'' surface , product of a Riemann surface of genus with an interval. In this paper, we show that virtual alternating links and tangles are naturally associated with the expansion of an i…
Blockwise bootstrap improves ASR performance testing for correlated data.
This paper presents a unified geometric framework for the statistical analysis of a general ill-posed linear inverse model which includes as special cases noisy compressed sensing, sign vector recovery, trace regression, orthogonal matrix estimation, and noisy matrix completion. We propose computationally feasible conv…
The stochastic gradient descent (SGD) algorithm has been widely used in statistical estimation for large-scale data due to its computational and memory efficiency. While most existing works focus on the convergence of the objective function or the error of the obtained solution, we investigate the problem of statistica…
This paper addresses privacy concerns in ratio statistics using differential privacy.
The determinants of the velocity of money have been examined based on life-cycle hypothesis. The velocity of money can be expressed by reciprocal of the average value of holding time which is defined as interval between participating exchanges for one unit of money. This expression indicates that the velocity is govern…