Distance correlation has gained much recent attention in the data science community: the sample statistic is straightforward to compute and asymptotically equals zero if and only if independence, making it an ideal choice to discover any type of dependency structure given sufficient sample size. One major bottleneck is…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We introduce the chi-square test neural network: a single hidden layer backpropagation neural network using chi-square test theorem to redefine the cost function and the error function. The weights and thresholds are modified using standard backpropagation algorithm. The proposed approach has the advantage of making co…
USP test improves on Pearson's chi-squared and -test for independence.
Study tests uniformity of categorical data against missing-ball alternatives, finding chi-squared test outperforms.
Study compares chi-squared divergence and KL-divergence posteriors for PAC-Bayesian bounds.
A goodness-of-fit test for DCSBM improves scalability and power for large sparse networks.
PQMass assesses generative model quality using chi-squared tests.
Study robust hypothesis testing under Hellinger distance, proving lower bounds and providing tests.
SDYNA is a general framework designed to address large stochastic reinforcement learning problems. Unlike previous model based methods in FMDPs, it incrementally learns the structure and the parameters of a RL problem using supervised learning techniques. Then, it integrates decision-theoric planning algorithms based o…
Transformers learn to adapt to different task difficulties and resist distribution shifts.
Tests if vertices in graphs have the same latent positions.
Objectives: Text categorization has been used in biomedical informatics for identifying documents containing relevant topics of interest. We developed a simple method that uses a chi-square-based scoring function to determine the likelihood of MEDLINE citations containing genetic relevant topic. Methods: Our procedure …
This paper tests the multivariate normality of node degrees in Erdős-Rényi graphs.
Proposes DR-ME test for interpretable distributional treatment effects.
Group Shapley evaluates feature groups in business data, improving explainability in AI.
RENAL test evaluates generative models for time series data.
We develop a pivotal test to assess the statistical significance of the feature variables in a single-layer feedforward neural network regression model. We propose a gradient-based test statistic and study its asymptotics using nonparametric techniques. Under technical conditions, the limiting distribution is given by …
The transition probability of a Cox-Ingersoll-Ross process can be represented by a non-central chi-square density. First we prove a new representation for the central chi-square density based on sums of powers of generalized Gaussian random variables. Second we prove Marsaglia's polar method extends to this distributio…
Four new methods for computing generalized chi-square distribution.
We consider the problem of comparing probability densities between two groups. A new probabilistic tensor product smoothing spline framework is developed to model the joint density of two variables. Under such a framework, the probability density comparison is equivalent to testing the presence/absence of interactions.…
Logistic regression is used thousands of times a day to fit data, predict future outcomes, and assess the statistical significance of explanatory variables. When used for the purpose of statistical inference, logistic models produce p-values for the regression coefficients by using an approximation to the distribution …
Directly simulates squared Bessel processes efficiently.
This paper investigates the utilization of maximum and average distance correlations for multivariate independence testing. We characterize their consistency properties in high-dimensional settings with respect to the number of marginally dependent dimensions, compare the advantages of each test statistic, examine thei…
AES scheme improves Bermudan and American option pricing for Heston models.
SBI provides more accurate pole positions than chi-squared minimization in model misspecification.
Network data is prevalent in many contemporary big data applications in which a common interest is to unveil important latent links between different pairs of nodes. Yet a simple fundamental question of how to precisely quantify the statistical uncertainty associated with the identification of latent links still remain…
The paper develops approximations for Pearson's chi-square statistic and applies them to confidence intervals.
A new method detects concept drift in streaming data using k-means space partitioning.
LMC algorithm converges to target in Chi-squared and Renyi divergence.
Test partial effects in Frechet regression on Bures-Wasserstein manifolds.
Constraint-based (CB) learning is a formalism for learning a causal network with a database D by performing a series of conditional-independence tests to infer structural information. This paper considers a new test of independence that combines ideas from Bayesian learning, Bayesian network inference, and classical hy…
Share price returns on different time scales can be well modelled by a superstatistical dynamics. Here we provide an investigation which type of superstatistics is most suitable to properly describe share price dynamics on various time scales. It is shown that while chi-square superstatistics works well on a time scale…
A test for neural networks identifies genetic associations.
Kernel two-sample testing is a useful statistical tool in determining whether data samples arise from different distributions without imposing any parametric assumptions on those distributions. However, raw data samples can expose sensitive information about individuals who participate in scientific studies, which make…
New SQ lower bounds for NGCA without requiring chi-squared condition.
A test for comparing networks using stochastic block models.
Efficient ANN search for sparse embeddings in ads targeting.
A theory which describes the share price evolution at financial markets as a continuous-time random walk has been generalized in order to take into account the dependence of waiting times t on price returns x. A joint probability density function (pdf) which uses the concept of a Lévy stable distribution is worked out.…
New adaptive test for NPIV models controls size and has superior power.
Fast nonparametric conditional independence testing via two-stage regression
The scaled complex Wishart distribution is a widely used model for multilook full polarimetric SAR data whose adequacy has been attested in the literature. Classification, segmentation, and image analysis techniques which depend on this model have been devised, and many of them employ some type of dissimilarity measure…
This paper aims to explore models based on the extreme gradient boosting (XGBoost) approach for business risk classification. Feature selection (FS) algorithms and hyper-parameter optimizations are simultaneously considered during model training. The five most commonly used FS methods including weight by Gini, weight b…
We consider the goodness-of-fit testing problem of distinguishing whether the data are drawn from a specified distribution, versus a composite alternative separated from the null in the total variation metric. In the discrete case, we consider goodness-of-fit testing when the null distribution has a possibly growing or…
Study finds tax avoidance and IT issues hinder revenue in Gombe state.
Paper develops statistical tests for covariance matrix regression on manifold.
Much work has been done on feature selection. Existing methods are based on document frequency, such as Chi-Square Statistic, Information Gain etc. However, these methods have two shortcomings: one is that they are not reliable for low-frequency terms, and the other is that they only count whether one term occurs in a …
Develops an empirical likelihood framework for random forests and ensembles.
Previous algorithms for constructing regression tree models for longitudinal and multiresponse data have mostly followed the CART approach. Consequently, they inherit the same selection biases and computational difficulties as CART. We propose an alternative, based on the GUIDE approach, that treats each longitudinal d…