Enhances power of covariance matrix tests for high-dimensional data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Novel power transform unifies various mathematical functions.
We introduce a new statistical tool (the TP-statistic and TE-statistic) designed specifically to compare the behavior of the sample tail of distributions with power-law and exponential tails as a function of the lower threshold u. One important property of these statistics is that they converge to zero for power laws o…
A new family of nonparametric statistics, the r-statistics, is introduced. It consists of counting the number of records of the cumulative sum of the sample. The single-sample r-statistic is almost as powerful as Student's t-statistic for Gaussian and uniformly distributed variables, and more powerful than the sign and…
Testing independence is of significant interest in many important areas of large-scale inference. Using extreme-value form statistics to test against sparse alternatives and using quadratic form statistics to test against dense alternatives are two important testing procedures for high-dimensional independence. However…
Estimates statistical power for cluster analysis in biomedical research.
Develops algorithms to balance personalization and statistical power in mobile health studies.
PAS improves estimation of multiple means using ML predictions and shrinkage.
The paper integrates statistical significance and discriminative power in pattern discovery.
Signals consisting of a sequence of pulses show that inherent origin of the 1/f noise is a Brownian fluctuation of the average interevent time between subsequent pulses of the pulse sequence. In this paper we generalize the model of interevent time to reproduce a variety of self-affine time series exhibiting power spec…
A new test method improves goodness-of-fit tests for copulas.
Nonparametric two sample testing is a decision theoretic problem that involves identifying differences between two random variables without making parametric assumptions about their underlying distributions. We refer to the most common settings as mean difference alternatives (MDA), for testing differences only in firs…
Study improves statistical power for detecting algorithmic bias in educational data.
Auto-regressive conditionally heteroskedastic (ARCH) family models are still used, by practitioners in business and economic policy making, as a conditional volatility forecasting models. Furthermore ARCH models still are attracting an interest of the researchers. In this contribution we consider the well known GARCH(1…
The role of kernels is central to machine learning. Motivated by the importance of power-law distributions in statistical modeling, in this paper, we propose the notion of power-law kernels to investigate power-laws in learning problem. We propose two power-law kernels by generalizing Gaussian and Laplacian kernels. Th…
The paper analyzes the power of MX CI tests and finds likelihood-based statistics most powerful.
This work improves independence tests for high-dimensional data.
New metrics boost A/B-test power by up to 210%.
Factor models are a class of powerful statistical models that have been widely used to deal with dependent measurements that arise frequently from various applications from genomics and neuroscience to economics and finance. As data are collected at an ever-growing scale, statistical machine learning faces some new cha…
Superposition accelerates training to a universal power-law exponent.
E-C2ST uses E-values for high-dimensional data two-sample tests.
Study spectral learning for odeco tensors, addressing initialization bottlenecks.
We study the problem of nonparametric dependence detection. Many existing methods may suffer severe power loss due to non-uniform consistency, which we illustrate with a paradox. To avoid such power loss, we approach the nonparametric test of independence through the new framework of binary expansion statistics (BEStat…
GAIF enhances online multiple testing with feedback, improving statistical power.
Predictive e-values enhance statistical inference across various tasks.
Value functions struggle to represent transition dynamics, impacting statistical efficiency.
Improves normalizing flows by incorporating data dependencies.
The paper analyzes a private likelihood-ratio test for frequency tables under differential privacy constraints.
Pronounced variability due to the growth of renewable energy sources, flexible loads, and distributed generation is challenging residential distribution systems. This context, motivates well fast, efficient, and robust reactive power control. Real-time optimal reactive power control is possible in theory by solving a n…
Enhances weak lensing inference with neural summaries.
The paper proposes a method to identify power system oscillation modes using blind source separation.
Cheap permutation tests speed up distribution testing without sacrificing accuracy.
Improves A/B testing by detecting minor treatment effects.
We investigate relationship between annual electric power consumption per capita and gross domestic production (GDP) per capita for 131 countries. We found that the relationship can be fitted with a power-law function. We examine the relationship for 47 prefectures in Japan. Furthermore, we investigate values of annual…
Paper improves power of conditional randomization tests.
A permutation-based SW test achieves minimax-optimal power for two-sample testing.
Earlier we proposed the stochastic point process model, which reproduces a variety of self-affine time series exhibiting power spectral density S(f) scaling as power of the frequency f and derived a stochastic differential equation with the same long range memory properties. Here we present a stochastic differential eq…
Enhances statistical inference using synthetic data.
Paper detects and estimates breaks in high-dimensional functional time series.
The problem of finding itemsets that are statistically significantly enriched in a class of transactions is complicated by the need to correct for multiple hypothesis testing. Pruning untestable hypotheses was recently proposed as a strategy for this task of significant itemset mining. It was shown to lead to greater s…
Enhances selective inference for generalized lasso using parametric programming.
In this study, we investigate the statistical properties of the returns and the trading volume. We show a typical example of power-law distributions of the return and of the trading volume. Next, we propose an interacting agent model of stock markets inspired from statistical mechanics [24] to explore the empirical fin…
Following the work of Okuyama, Takayasu and Takayasu [Okuyama, Takayasu and Takayasu 1999] we analyze huge databases of Japanese companies' financial figures and confirm that the Zipf's law, a power law distribution with the exponent -1, has been maintained over 30 years in the income distribution of Japanese companies…
Study compares statistical properties and power of divergence measures for credit risk monitoring.
QAOA matches classical tensor power iteration in spiked tensor model recovery.
There is a vast body of literature related to methods for detecting changepoints (CP). However, less attention has been paid to assessing the statistical reliability of the detected CPs. In this paper, we introduce a novel method to perform statistical inference on the significance of the CPs, estimated by a Dynamic Pr…
We introduce preferential behavior into the study on statistical mechanics of money circulation. The computer simulation results show that the preferential behavior can lead to power laws on distributions over both holding time and amount of money held by agents. However, some constraints are needed in generation mecha…
FPPI selectively uses predictions to improve inference efficiency.