Enhances power of covariance matrix tests for high-dimensional data.
problem Testing large covariance matrices in high-dimensional data.
method Proposes a new Fisher's combined probability test for quadratic form and maximum form statistics.
result Boosts power against more general alternatives.
Novel power transform unifies various mathematical functions.
problem Normalizing and standardizing datasets.
method Presented a novel power transform.
result Unified various mathematical functions.
We introduce a new statistical tool (the TP-statistic and TE-statistic) designed specifically to compare the behavior of the sample tail of distributions with power-law and exponential tails as a function of the lower threshold u. One important property of these statistics is that they converge to zero for power laws o…
A new family of nonparametric statistics, the r-statistics, is introduced. It consists of counting the number of records of the cumulative sum of the sample. The single-sample r-statistic is almost as powerful as Student's t-statistic for Gaussian and uniformly distributed variables, and more powerful than the sign and…
Testing independence is of significant interest in many important areas of large-scale inference. Using extreme-value form statistics to test against sparse alternatives and using quadratic form statistics to test against dense alternatives are two important testing procedures for high-dimensional independence. However…
Estimates statistical power for cluster analysis in biomedical research.
problem Lack of established methods to compute a priori statistical power for cluster analysis.
method Simulation studies varying subgroup size, number, separation, and covariance structure.
result Sufficient statistical power achieved with small samples (N=20-30) for large effect sizes.
Develops algorithms to balance personalization and statistical power in mobile health studies.
problem Balancing personalization and statistical power in mobile health studies.
method Develops general meta-algorithms to modify existing bandit algorithms.
result Guarantees sufficient power while improving user well-being.
PAS improves estimation of multiple means using ML predictions and shrinkage.
problem Improving statistical estimates with limited gold-standard data and noisy ML predictions.
method Prediction-Powered Adaptive Shrinkage (PAS) that combines PPI with empirical Bayes shrinkage.
result PAS adapts to the reliability of ML predictions and outperforms traditional methods in large-scale applications.
The paper integrates statistical significance and discriminative power in pattern discovery.
problem Discovering actionable patterns that meet rigorous statistical significance and discriminative power criteria.
method Integrates statistical significance and discriminative power criteria into state-of-the-art algorithms.
result Improves discriminative power and statistical significance of discovered patterns without quality deterioration.
Signals consisting of a sequence of pulses show that inherent origin of the 1/f noise is a Brownian fluctuation of the average interevent time between subsequent pulses of the pulse sequence. In this paper we generalize the model of interevent time to reproduce a variety of self-affine time series exhibiting power spec…
A new test method improves goodness-of-fit tests for copulas.
problem Developing robust tests for copula goodness-of-fit.
method Binary Expansion Approximation of UniformiTY (BEAUTY) and Binary Expansion Adaptive Symmetry Test (BEAST).
result The BEAST method improves empirical power against various alternatives.
Nonparametric two sample testing is a decision theoretic problem that involves identifying differences between two random variables without making parametric assumptions about their underlying distributions. We refer to the most common settings as mean difference alternatives (MDA), for testing differences only in firs…
Study improves statistical power for detecting algorithmic bias in educational data.
problem Challenges in measuring algorithmic bias using ABROCA due to skewed distribution.
method Investigates ABROCA's distributional properties and proposes nonparametric randomization tests.
result ABROCA-based bias assessments are underpowered in typical EDM sample sizes.
Auto-regressive conditionally heteroskedastic (ARCH) family models are still used, by practitioners in business and economic policy making, as a conditional volatility forecasting models. Furthermore ARCH models still are attracting an interest of the researchers. In this contribution we consider the well known GARCH(1…
The role of kernels is central to machine learning. Motivated by the importance of power-law distributions in statistical modeling, in this paper, we propose the notion of power-law kernels to investigate power-laws in learning problem. We propose two power-law kernels by generalizing Gaussian and Laplacian kernels. Th…
The paper analyzes the power of MX CI tests and finds likelihood-based statistics most powerful.
problem Testing conditional independence under model-X assumptions.
method Conditional randomization test (CRT) and MX knockoffs.
result Likelihood-based statistics are most powerful in MX CI tests.
This work improves independence tests for high-dimensional data.
problem Detecting subtle dependencies between high-dimensional random variables with complex distributions.
method Develops two approaches to learn powerful independence tests using variational mutual information and HSIC.
result Optimized HSIC tests generally outperform other approaches on detecting structured dependence.
New metrics boost A/B-test power by up to 210%.
problem High cost and type-II errors in A/B-tests.
method Learn metrics from short-term signals to maximize power.
result Statistical power increased by up to 210%.
Factor models are a class of powerful statistical models that have been widely used to deal with dependent measurements that arise frequently from various applications from genomics and neuroscience to economics and finance. As data are collected at an ever-growing scale, statistical machine learning faces some new cha…
Superposition accelerates training to a universal power-law exponent.
problem Training dynamics in neural networks.
method Teacher-student framework and analytic theory.
result Superposition leads to a universal power-law exponent of ~1, independent of data and channel statistics.
E-C2ST uses E-values for high-dimensional data two-sample tests.
problem Statistical testing for high-dimensional data.
method Combines split likelihood ratio tests and predictive independence tests, using E-values for anytime-valid sequential tests.
result E-C2ST achieves enhanced statistical power by partitioning datasets into multiple batches.
Study spectral learning for odeco tensors, addressing initialization bottlenecks.
problem Recovering orthogonally decomposable tensors under noise.
method Investigates perturbation bounds, non-convex optimization, and initialization strategies.
result Initialization is the main bottleneck for efficient algorithms.
We study the problem of nonparametric dependence detection. Many existing methods may suffer severe power loss due to non-uniform consistency, which we illustrate with a paradox. To avoid such power loss, we approach the nonparametric test of independence through the new framework of binary expansion statistics (BEStat…
GAIF enhances online multiple testing with feedback, improving statistical power.
problem Sequential online multiple testing with delayed feedback.
method GAIF framework using dynamic threshold adjustment and feedback-driven model selection.
result Improves statistical power through feedback-driven model selection.
Predictive e-values enhance statistical inference across various tasks.
problem Insufficient data limits traditional statistical inference.
method Apply prediction-powered inference to e-values.
result Every e-value-based inference has a prediction-powered counterpart.
Value functions struggle to represent transition dynamics, impacting statistical efficiency.
problem Limited representational power of value functions in capturing transition dynamics.
method Case studies of various reinforcement learning problems to explore the limitations of value-based methods.
result Value-based methods can be as efficient as model-based ones in some cases but severely underperform in others due to information loss.
Improves normalizing flows by incorporating data dependencies.
problem Current normalizing flow learning assumes independent data, leading to errors.
method Proposes a likelihood objective with dependencies and efficient learning algorithm.
result Improves density estimation and data generation on real-world data.
The paper analyzes a private likelihood-ratio test for frequency tables under differential privacy constraints.
problem Achieving privacy in statistical data analysis while maintaining statistical utility.
method A rigorous analysis of a private likelihood-ratio (LR) test for goodness-of-fit in frequency tables, considering (ε,δ)-differential privacy. result Characterization of the trade-off between differential privacy parameters (ε,δ) and statistical power of the private LR test. Pronounced variability due to the growth of renewable energy sources, flexible loads, and distributed generation is challenging residential distribution systems. This context, motivates well fast, efficient, and robust reactive power control. Real-time optimal reactive power control is possible in theory by solving a n…
Enhances weak lensing inference with neural summaries.
problem Extracting additional information from weak lensing convergence maps.
method Hybrid approach combining physics-based and neural summaries.
result Neural summaries extract up to 8 times more information than angular power spectra.
The paper proposes a method to identify power system oscillation modes using blind source separation.
problem Accurately identifying oscillation modes in power systems with renewable energy sources.
method A high-order blind source identification (HOBI) algorithm based on copula statistic combined with Hilbert transform and iteration procedure.
result The method can identify all oscillation modes and model order from a single channel of observation signals, outperforming state-of-the-art methods.
Cheap permutation tests speed up distribution testing without sacrificing accuracy.
problem Efficiently testing distribution differences and independence.
method Group datapoints into bins and permute only these bins, using stored sufficient statistics.
result Cheap permutation tests maintain the accuracy and optimality of standard tests but are significantly faster.
Improves A/B testing by detecting minor treatment effects.
problem Challenges in identifying small average treatment effects.
method Maximum probability-driven two-armed bandit (TAB) process with weighted mean volatility statistic.
result Significant improvement in A/B testing with reduced experimental costs.
We investigate relationship between annual electric power consumption per capita and gross domestic production (GDP) per capita for 131 countries. We found that the relationship can be fitted with a power-law function. We examine the relationship for 47 prefectures in Japan. Furthermore, we investigate values of annual…
Paper improves power of conditional randomization tests.
problem Improving power of conditional randomization tests.
method Introducing a new cost function to maximize test statistic power.
result Consistently increases the number of correct discoveries.
A permutation-based SW test achieves minimax-optimal power for two-sample testing.
problem Nonparametric two-sample testing using the sliced Wasserstein distance.
method Proposes a permutation-based SW test and analyzes its performance.
result Achieves minimax separation rate n−1/2 over multinomial and bounded-support alternatives. Earlier we proposed the stochastic point process model, which reproduces a variety of self-affine time series exhibiting power spectral density S(f) scaling as power of the frequency f and derived a stochastic differential equation with the same long range memory properties. Here we present a stochastic differential eq…
Paper introduces a method to assess the statistical reliability of changepoints using selective inference and dynamic programming.
problem Assessing the statistical reliability of detected changepoints.
method Selective inference framework combined with dynamic programming for exact p-value computation.
result Proposes a method with high statistical power and decent computational efficiency.
Enhances statistical inference using synthetic data.
problem Limited labeled data for statistical inference.
method GESPI framework that combines synthetic and real data.
result Error rate remains below a user-specified bound and decreases with synthetic data quality.
Paper detects and estimates breaks in high-dimensional functional time series.
problem Detecting and estimating structural breaks in heterogeneous mean functions of high-dimensional functional time series.
method Proposes a new test statistic combining functional CUSUM and power enhancement components, with a clustering algorithm for group structure estimation.
result The proposed techniques have satisfactory performance in finite samples, detecting and estimating breaks effectively.
The problem of finding itemsets that are statistically significantly enriched in a class of transactions is complicated by the need to correct for multiple hypothesis testing. Pruning untestable hypotheses was recently proposed as a strategy for this task of significant itemset mining. It was shown to lead to greater s…
Enhances selective inference for generalized lasso using parametric programming.
problem Low statistical power in selective inference for generalized lasso.
method Parametric programming to compute solution paths and identify model selection events.
result Improves selective inference power and practicality for various problems.
In this study, we investigate the statistical properties of the returns and the trading volume. We show a typical example of power-law distributions of the return and of the trading volume. Next, we propose an interacting agent model of stock markets inspired from statistical mechanics [24] to explore the empirical fin…
Following the work of Okuyama, Takayasu and Takayasu [Okuyama, Takayasu and Takayasu 1999] we analyze huge databases of Japanese companies' financial figures and confirm that the Zipf's law, a power law distribution with the exponent -1, has been maintained over 30 years in the income distribution of Japanese companies…
Study compares statistical properties and power of divergence measures for credit risk monitoring.
problem Detecting distributional shifts in credit risk models.
method Derives statistical properties and chi-square benchmark values for Jensen-Shannon Divergence and Kullback-Leibler Divergence, demonstrating their applicability in credit risk monitoring.
result Jensen-Shannon Divergence and Kullback-Leibler Divergence follow chi-square distributions and reveal practical trade-offs in minimizing false positives vs. detecting changes.
QAOA matches classical tensor power iteration in spiked tensor model recovery.
problem Statistical estimation in spiked tensor model with computational gap.
method Analysis of QAOA performance on spiked tensor model.
result QAOA weak recovery threshold matches tensor power iteration.
We introduce preferential behavior into the study on statistical mechanics of money circulation. The computer simulation results show that the preferential behavior can lead to power laws on distributions over both holding time and amount of money held by agents. However, some constraints are needed in generation mecha…
FPPI selectively uses predictions to improve inference efficiency.
problem Improving statistical inference with limited labeled data and heterogeneous prediction quality.
method Filtered Prediction-Powered Inference (FPPI) framework.
result FPPI achieves strictly improved asymptotic efficiency compared to existing methods.