New estimator stabilizes higher-order influence functions for stable statistical inference.
problem Numerical instability in estimating inverse population Gram matrix.
method Proposes a new stabilized higher-order estimator without sample splitting.
result Stabilized estimator exhibits more stable performance and similar statistical guarantees.
New estimator stabilizes higher-order influence functions for bilinear forms.
problem Stability issues in estimating bilinear forms using higher-order influence functions.
method Proposes a new stabilized higher-order estimator for a class of bilinear forms without sample splitting.
result New estimator exhibits more stable finite-sample performance compared to the empirical higher-order estimator.
This paper simplifies computing higher-order U-statistics efficiently.
problem The inefficiency of computing higher-order U-statistics in practice. method Decomposition, connection to Einstein summation, and treewidth-based complexity estimate.
result A new, more efficient algorithm to compute U-statistics. Study improves BN TTA under distribution shift using higher-order asymptotics.
problem Improving BN TTA for changing data distributions.
method Integrates Edgeworth expansion and saddlepoint approximation with one-step M-estimation.
result Derives optimal weighting parameter for minimized mean-squared error.
Diffusion models learn simple statistics before complex ones, revealing a sample complexity exponent.
problem Understanding the learning dynamics of diffusion models.
method Empirical observations and theoretical analysis of diffusion models and denoisers.
result Diffusion models learn simple statistics (pair-wise correlations) at linear sample complexity, while higher-order statistics (e.g., fourth cumulant) require cubic sample complexity.
Study quantifies how LLMs capture higher-order statistical structure using cumulant expansion.
problem Understanding how LLMs internalize statistical structure during next-token prediction.
method Cumulant-expansion framework treating softmax entropy as perturbation around center distribution.
result Cumulants reveal distinct signatures for mathematical vs. general text prompts, quantifying feature-learning dynamics.
Bayesian method reconstructs hidden higher-order interactions from network data.
problem Lack of explicit higher-order interactions in pairwise network data.
method Bayesian approach based on parsimony, infers higher-order structures when statistically supported.
result Demonstrated applicability to various datasets, synthetic and empirical.
Neural networks can learn from higher-order cumulants efficiently, requiring quadratic samples.
problem Learning from higher-order cumulants in high-dimensional data.
method Spiked cumulant model, polynomial time algorithms, neural networks, random features.
result Neural networks require quadratic samples to learn from higher-order cumulants efficiently, while random features require more samples.
Estimates hypergraphons for modeling complex interactions efficiently.
problem Modeling higher-order interactions using hypergraphons.
method Restricted class of Simple Lipschitz Hypergraphons (SLH) for efficient estimation.
result Optimal rates of convergence for SLH estimator.
We develop a theory of higher-order feature attribution for complex models.
problem Interpreting feature contributions in models with interactions is challenging.
method We extend Integrated Gradients (IG) to higher-order feature attributions.
result We establish natural connections to statistics and topological signal processing.
Improved estimation of higher order integrals using shrinkage techniques.
problem Estimating higher order Bochner integrals in non-parametric settings.
method Shrinkage of U-statistic towards a target element, considering kernel degeneracy.
result Consistent shrinkage estimators with fast rates of convergence, even for non-degenerate kernels.
New quantum states capture more information, enabling advanced processing tasks.
problem Quantum information processing challenges with limited statistical information.
method Introducing Random-Coefficient Pure States (RCPS) and exploiting their higher-order statistics.
result RCPS provide richer information than density operators, enabling new quantum tasks.
New neural networks model complex phenomena with fewer parameters.
problem Challenges in studying higher-order interactions in neural networks.
method Introducing curved neural networks using the maximum entropy principle.
result Curved neural networks accelerate memory retrieval and exhibit explosive phase transitions.
Study higher-order spin glass models for social network behavior with peer-group effects.
problem Modeling correlation phenomena on social networks with peer-group effects.
method Inference in higher-order Ising models to recover coefficients and peer-group effects.
result Strong concavity of log pseudo-likelihood implies statistical error rate of sqrt(d/n) for MPLE.
We present an extension of the Kolmogorov-Smirnov (KS) two-sample test, which can be more sensitive to differences in the tails. Our test statistic is an integral probability metric (IPM) defined over a higher-order total variation ball, recovering the original KS test as its simplest case. We give an exact representer…
A framework infers hyperedges and overlapping communities in hypergraphs.
problem Characterizing the structural organization of hypergraphs with higher-order interactions.
method Statistical inference to infer missing hyperedges and detect overlapping communities.
result Efficient numerical implementation and strong performance on real-world systems.
The parametric complexity is the key quantity in the minimum description length (MDL) approach to statistical model selection. Rissanen and others have shown that the parametric complexity of a statistical model approaches a simple function of the Fisher information volume of the model as the sample size n goes to in…
New model estimates higher-order interactions in stochastic processes using lower-dimensional projections.
problem Estimating higher-order interaction effects in stochastic processes with limited data.
method Additive Poisson Process (APP) combines information geometry and generalized additive models to model intensity functions in lower dimensions.
result The model can estimate higher-order intensity functions with sparse data.
In this paper, we investigate the popular deep learning optimization routine, Adam, from the perspective of statistical moments. While Adam is an adaptive lower-order moment based (of the stochastic gradient) method, we propose an extension namely, HAdam, which uses higher order moments of the stochastic gradient. Our …
New stable HOIF estimators for statistical functionals.
problem Constructing numerically stable HOIF estimators for statistical functionals.
method Developed new sHOIF estimators with provable guarantees.
result 2nd order sHOIF estimators were validated in synthetic experiments.
Paper provides Edgeworth expansions for network moments, improving accuracy of sampling distributions.
problem Accurate descriptions of sampling distributions of network moment statistics.
method Edgeworth expansion applied to studentized network moment statistics.
result Higher-order accurate approximation to sampling CDF of network moment statistics.
New method bypasses assumptions for unbiased estimation of complex system interactions.
problem Inferring pair-wise and higher-order interactions from observational data.
method Cross-disciplinary approach using Targeted Learning for unbiased estimation.
result Universal estimator of all-order symmetric interactions without parametric assumptions.
TGCCA analyzes higher-order tensors using orthogonal rank-R CP decomposition.
problem Handling higher-order structures in multi-block data analysis.
method Tensor Generalized Canonical Correlation Analysis (TGCCA) with orthogonal rank-R CP decomposition.
result TGCCA outperforms state-of-the-art methods on simulated and real data.
The search for higher-order feature interactions that are statistically significantly associated with a class variable is of high relevance in fields such as Genetics or Healthcare, but the combinatorial explosion of the candidate space makes this problem extremely challenging in terms of computational efficiency and p…
Predicts node sequences in graphs using multi-order network models.
problem Predicting sequences of node traversals in graphs.
method Combines multiple higher-order network models into a multi-order model, fitting and selecting the optimal maximum order.
result Outperforms state-of-the-art algorithms for next-element and full sequence prediction.
Tensor train (TT) decomposition provides a space-efficient representation for higher-order tensors. Despite its advantage, we face two crucial limitations when we apply the TT decomposition to machine learning problems: the lack of statistical theory and of scalable algorithms. In this paper, we address the limitations…
We apply information-based complexity analysis to support vector machine (SVM) algorithms, with the goal of a comprehensive continuous algorithmic analysis of such algorithms. This involves complexity measures in which some higher order operations (e.g., certain optimizations) are considered primitive for the purposes …
We develop and implement a novel fast bootstrap for dependent data. Our scheme is based on the i.i.d. resampling of the smoothed moment indicators. We characterize the class of parametric and semi-parametric estimation problems for which the method is valid. We show the asymptotic refinements of the proposed procedure,…
Cross validation (CV) and the bootstrap are ubiquitous model-agnostic tools for assessing the error or variability of machine learning and statistical estimators. However, these methods require repeatedly re-fitting the model with different weighted versions of the original dataset, which can be prohibitively time-cons…
New method shows fully-connected networks can learn convolutional structures from data.
problem How to learn convolutional structures from translation-invariant data.
method Data-driven emergence of convolutional structure in neural networks.
result Initially fully-connected networks can learn convolutional structures from their inputs.
iLOCO measures feature interactions without assumptions, providing statistical inference.
problem Lack of methods to statistically infer feature interactions.
method iLOCO metric and LOCO inference for distribution-free, efficient computation.
result First inferential approach to detecting feature interactions.
Betas are possibly the most frequently applied tool to analyze how securities relate to the market. While in very widespread use, betas only express dynamics derived from second moment statistics. Financial returns data often deviate from normal assumptions in the sense that they have significant third and fourth order…
Paper proposes a new tensor model for mixed memberships and provides error bounds.
problem Estimating mixed memberships in higher-order multiway data.
method Tensor mixed-membership blockmodel, higher-order orthogonal iteration algorithm (HOOI), simplex corner-finding algorithm.
result Consistency of estimation procedure with error bounds under specific conditions.
A model predicts influential nodes in complex networks by considering indirect interactions.
problem Identifying influential nodes in complex networks using indirect interactions.
method Proposes MOGen, a multi-order generative model that considers all indirect influences up to a maximum distance.
result MOGen consistently outperforms network models and path-based approaches in predicting influential nodes.
Generalizes SMCI to improve Boltzmann machine learning accuracy.
problem Limitation in applying higher-order SMCI to dense systems.
method Generalized SMCI (GSMCI) and new PBM learning method.
result GSMCI allows higher-order approximations for dense systems.
Tensor decompositions have rich applications in statistics and machine learning, and developing efficient, accurate algorithms for the problem has received much attention recently. Here, we present a new method built on Kruskal's uniqueness theorem to decompose symmetric, nearly orthogonally decomposable tensors. Unlik…
Deterministic bounds for tensor singular values and vectors, differing from matrix cases.
problem Spectral learning of higher-order orthogonally decomposable tensors.
method Deterministic perturbation bounds for singular values and vectors of orthogonally decomposable tensors.
result Perturbation affects each essential singular value/vector in isolation, independent of multiplicity and distance from other singular values.
New Gini indices capture more nuanced income inequality.
problem Measuring joint dispersion across multiple observations.
method Axiomatic approach to define and characterize n-th order Gini deviations.
result Higher-order Gini coefficients reveal more extreme income disparities.
Unified framework for higher-order network analysis.
problem Complex structure of space of networks.
method Measure-theoretic formalism, Gromov-Wasserstein distance, co-optimal transport distance.
result Unified theoretical treatment of generalized networks.
In topic modeling, many algorithms that guarantee identifiability of the topics have been developed under the premise that there exist anchor words -- i.e., words that only appear (with positive probability) in one topic. Follow-up work has resorted to three or higher-order statistics of the data corpus to relax the an…
Generative graph models create instances of graphs that mimic the properties of real-world networks. Generative models are successful at retaining pairwise associations in the underlying networks but often fail to capture higher-order connectivity patterns known as network motifs. Different types of graphs contain diff…
Four algorithms improve sparse tensor BR1Approx with theoretical guarantees.
problem Sparse tensor best rank-1 approximation.
method Four approximation algorithms exploiting multilinearity and sparsity.
result Theoretical worst-case approximation lower bounds for all algorithms.
Finding statistically significant interactions between binary variables is computationally and statistically challenging in high-dimensional settings, due to the combinatorial explosion in the number of hypotheses. Terada et al. recently showed how to elegantly address this multiple testing problem by excluding non-tes…
Study privacy vs. utility in estimating network parameters with aggregated data.
problem Privacy-preserving estimation of network parameters from aggregated node degrees.
method β model, local and central differential privacy, minimax lower bounds, simple estimators.
result Achieved minimax-optimal risk bounds for parameter estimation under privacy constraints.
For certain classes of knots we define geometric invariants called higher-order genera. Each of these invariants is a refinement of the slice genus of a knot. We find lower bounds for the higher-order genera in terms of certain von Neumann ρ-invariants, which we call higher-order signatures. The higher-order genera o…
A fundamental property of complex networks is the tendency for edges to cluster. The extent of the clustering is typically quantified by the clustering coefficient, which is the probability that a length-2 path is closed, i.e., induces a triangle in the network. However, higher-order cliques beyond triangles are crucia…
Stability of capillary hypersurfaces with higher order mean curvature.
problem Stability of capillary hypersurfaces with constant higher order mean curvature.
method Generalization of classical stability theory for capillary hypersurfaces.
result Results on stability for capillary hypersurfaces with higher order mean curvature.
A key feature of inductive logic programming (ILP) is its ability to learn first-order programs, which are intrinsically more expressive than propositional programs. In this paper, we introduce techniques to learn higher-order programs. Specifically, we extend meta-interpretive learning (MIL) to support learning higher…