Survey of statistical queries and their applications.
problem Understanding statistical queries and their applications.
method Exploration of statistical queries model, definitions, and connections to learnability.
result Connections to learnability and applications in optimization, evolvability, and differential privacy.
Information geometry offers new tools for statistical analysis.
problem Statistical analysis of probability distributions.
method Geometric perspective on statistical manifolds.
result New applications in radar sensing, signal processing, etc.
New framework controls statistical dispersion for high-stakes applications.
problem Understanding and controlling the dispersion of loss distributions in high-stakes applications.
method Simple yet flexible framework for distribution-free control of statistical dispersion measures.
result Proposed methods control statistical dispersion measures with societal implications.
Factor models are a class of powerful statistical models that have been widely used to deal with dependent measurements that arise frequently from various applications from genomics and neuroscience to economics and finance. As data are collected at an ever-growing scale, statistical machine learning faces some new cha…
Statistical and probabilistic characteristics of locally free group with growing number of generators are defined and their application to statistics of braid groups is given.
Recently we reported on an application of the Tsallis non-extensive statistics to the S&P500 stock index. There we argued that the statistics are applicable to a broad range of markets and exchanges where anamolous (super) diffusion and 'heavy' tails of the distribution are present, as they are in the S&P500. We have c…
Machine learning improves official statistics but needs rigorous validation.
problem Lack of methodological robustness in machine learning for official statistics.
method Total Machine Learning Error (TMLE) framework to validate ML models.
result TMLE addresses representativeness and measurement errors in ML models.
We review basic notions in the field of information geometry such as Fisher metric on statistical manifold, α-connection and corresponding curvature following Amari's work . We show application of information geometry to asymptotic statistical inference.
This paper deals with the applications of an optimization method on submanifolds, that is, geometric inequalities can be considered as optimization problems. In this regard, we obtain optimal Casorati inequalities and Chen-Ricci inequality for a statistical submanifold in a statistical warped product manifold of type $…
Risk statistic is a critical factor not only for risk analysis but also for financial application. However, the traditional risk statistics may fail to describe the characteristics of regulator-based risk. In this paper, we consider the regulator-based risk statistics for portfolios. By further developing the propertie…
A new test statistic measures discrepancy between conditional distributions.
problem Measuring the discrepancy between two conditional distributions.
method Proposes a Bregman matrix divergence-based statistic that avoids explicit distribution estimation.
result The new statistic inherits high-order statistics and demonstrates utility in multi-task learning, concept drift detection, and feature selection.
Kenmotsu geometry is a valuable part of contact geometry with nice applications in other fields such as theoretical physics. In this article, we study the statistical counterpart of a Kenmotsu manifold, that is, Kenmotsu statistical manifold with some related examples. We investigate some statistical curvature properti…
Survey on importance weighting in machine learning applications.
problem Distribution shift in supervised learning.
method Weighting objective function or probability distribution based on instance importance.
result Importance weighting can guarantee desirable statistical properties in distribution shift scenarios.
Survey on factor models and their applications in econometrics.
problem Estimating low-rank structures in high-dimensional models.
method Low-rank recovery techniques for factor model estimation.
result New insights into factor model applications in econometrics.
Defines vector Laplacian on statistical manifolds.
problem No specific problem stated; focuses on mathematical definition.
method Defines and derives vector Laplacian formula.
result Derives formula for vector Laplacian.
The Bonnet theorem is proven for statistical manifolds.
problem Locally embeddable statistical manifolds in flat spaces.
method Using statistical embedding and the Gauss--Codazzi--Ricci equations.
result Statistical manifolds with specific tensor properties are locally embeddable to flat statistical manifolds.
The paper examines extreme value statistics of high-dimensional sample covariances, with applications in finance and image analysis.
problem Statistical validation of normal conditions in high-dimensional time series data.
method Generalizes the maximal deviation of sample autocovariances to high dimensions and applies Gumbel-type extreme value asymptotics.
result Gumbel-type extreme value asymptotics holds true for high-dimensional sample covariances.
Paper reproduces a kernel-based scan B-statistic for online change-point detection.
problem Continuous detection of distribution changes in online data streams.
method Efficient kernel-based scan B-statistic for online change-point detection.
result Scan B-statistic outperforms parametric methods in challenging scenarios.
The paper derives Einstein tensors for a family of α-connections on quasi-statistical manifolds.
problem Deriving Einstein tensors for a new family of connections.
method Developed mathematical foundations of statistical and quasi-statistical manifolds, including dual and equiaffine connections.
result Explicit expressions for curvatures and Einstein tensors of the α-connections.
Sharp statistical theory for conditional diffusion models.
problem Lack of theoretical foundation for conditional diffusion models.
method Sharp statistical theory with approximation of conditional score function.
result Sample complexity bound that adapts to data distribution smoothness.
LLMs struggle to generate random numbers from statistical distributions, leading to biased results in applications.
problem LLMs' inability to generate random numbers accurately from specified distributions.
method Dual-protocol design: Batch Generation and Independent Requests, benchmarking 11 models across 15 distributions.
result Sampling fidelity degrades with distributional complexity and horizon, leading to systematic biases in downstream applications.
The need for new methods to deal with big data is a common theme in most scientific fields, although its definition tends to vary with the context. Statistical ideas are an essential part of this, and as a partial response, a thematic program on statistical inference, learning, and models in big data was held in 2015 i…
A textbook on statistical machine learning for astronomy.
problem Uncertainty quantification in astronomical data analysis.
method Bayesian inference and classical statistical methods.
result Unified framework connecting modern and traditional methods.
When the data are stored in a distributed manner, direct application of traditional statistical inference procedures is often prohibitive due to communication cost and privacy concerns. This paper develops and investigates two Communication-Efficient Accurate Statistical Estimators (CEASE), implemented through iterativ…
Abstract mathematical formulas for statistical structures and curvatures.
problem Developing formulas for statistical structures and curvatures.
method Proving new formulas and theorems for statistical structures and curvatures.
result Generalized formulas for statistical structures and curvatures.
The paper provides statistical guarantees for sparse deep learning.
problem Understanding the potential and limitations of sparse deep learning.
method Develops statistical guarantees for different types of sparsity in sparse deep learning.
result Statistical guarantees for sparse deep learning with mild dependence on network widths and depths.
Cross-sectional "Information Coefficient" (IC) is a widely and deeply accepted measure in portfolio management. The paper gives an insight into IC in view of high-dimensional directional statistics: IC is a linear operator on the components of a centralizing-unitizing standardized random vector of next-period cross-sec…
Discusses new probabilistic morphisms and geometric methods in machine and statistical learning.
problem Addressing challenges in statistical, machine, and manifold learning.
method Introduces category of probabilistic morphisms and geometric methods.
result New insights and applications in various learning fields.
Contrastive learning simplifies statistical inference for complex models.
problem Computational intractability of likelihood functions for certain models.
method Contrastive learning as an alternative for parameter estimation and inference.
result Contrastive learning enables practical methods for diverse statistical problems.
This paper quantifies uncertainty in Data Shapley using statistical inference.
problem Uncertainty in data valuation due to dynamic data distribution.
method Established relationship with U-statistics and quantified uncertainty using statistical inference.
result Confidence intervals for Data Shapley estimations are provided.
In this paper, we introduce the concept of principal bundles on statistical manifolds. After necessary preliminaries on information geometry and principal bundles on manifolds, we study the α-structure of frame bundles over statistical manifolds with respect to α-connections, by giving geometric structures. The man…
Active inference framework improves U-statistic estimation efficiency.
problem Costly acquisition of labels for U-statistics. method Active inference framework with optimal sampling rule.
result Substantial gains in estimation efficiency over baseline methods.
Modern technologies are generating ever-increasing amounts of data. Making use of these data requires methods that are both statistically sound and computationally efficient. Typically, the statistical and computational aspects are treated separately. In this paper, we propose an approach to entangle these two aspects …
Develops methods for GWAS of high dimensional phenotypes using summary statistics.
problem Lack of methods to model pleiotropy in multi-phenotype GWAS.
method Bayesian inference model using summary statistics, fast computation, and biologically informed priors.
result Demonstrates utility in metabolite GWAS with interpretable pathway-level inference.
The time value of money is a critical factor not only in risk analysis, but also in insurance and financial applications. In this paper, we consider a special class of set-valued risk statistics by introducing the time value of money. In fact, the risk statistics established by this method is closer to financial realit…
Overview of high-dimensional time series regression methods.
problem Estimation and inference with high-dimensional time series data.
method Limit theory for high-dimensional dependent data, asymptotic theory for time series regression, statistical learning methods.
result Main limit theory results and asymptotic theory for high-dimensional time series regression.
Establishes statistical and computational bounds for influence diagnostics.
problem Identifying influential datapoints or subsets in machine learning models.
method Finite-sample statistical bounds and computational complexity for influence functions and approximate maximum influence perturbations.
result Established statistical and computational guarantees for influence diagnostics.
Paper develops approximation and statistical theory for signature-based path regression.
problem Understanding how fast signatures approximate continuous path functionals.
method Develops \(L^2\) approximation rate for smooth functionals of Itô diffusions and establishes consistency of statistical learning procedures.
result Signature-based methods improve prediction over handcrafted features in various real-data applications.
Transform non-private e-values into differentially private ones.
problem Leaking sensitive data through non-private e-values.
method Developed a novel biased multiplicative noise mechanism.
result Differentially private e-values maintain strong statistical power and asymptotic equivalence to non-private ones.
Theoretical guarantees for neural estimators in parametric statistics are derived.
problem Lack of theoretical guarantees for neural estimators in parametric statistics.
method Decompose risk into terms and verify assumptions for convergence.
result Derive theoretical guarantees for neural estimators.
High-dimensional U-statistics show surprising phase transitions, impacting kernel-based tests.
problem Understanding phase transitions in high-dimensional U-statistics.
method Proved a convergence theorem for U-statistics of degree two in high dimensions.
result High-dimensional U-statistics can have non-Gaussian limits with larger variance and asymmetry.
New formulae identify discrete probability laws without needing normalization constants.
problem Characterizing non-normalized discrete probability distributions.
method Derive explicit formulae for mass functions using Stein's method.
result Developed tools for solving statistical problems without normalization constants.
Learning representations of data is an important problem in statistics and machine learning. While the origin of learning representations can be traced back to factor analysis and multidimensional scaling in statistics, it has become a central theme in deep learning with important applications in computer vision and co…
The differential geometry of Kenmotsu manifold is a valuable part of contact geometry with nice applications in other fields such as theoretical physics. In fact, its statistical counterpart, that is, Kenmotsu statistical manifold also has same importance as that of Kenmotsu manifold. Theoretical physicists have also b…
Unified approach for quantum and classical learning from evaluation oracles.
problem Learning from evaluation oracles in quantum and classical settings.
method Inspired by Kearns' SQ and Valiant's weak evaluation oracle, a unified framework is established.
result Characterizes query complexity for learning linear function classes and extends learnability results for quantum circuits.
Paper stabilizes persistent homology rank functions for statistical inference.
problem Stability issues in persistent homology rank functions.
method Derive stability results for rank functions under FDA metrics.
result Rank functions stabilize, improving statistical inference.
FNNs can be made more interpretable with statistical methods.
problem FNNs lack interpretability and are often used as black-box models.
method Supplement FNNs with statistical inference and covariate-effect visualizations.
result FNNs can be made more like traditional statistical models.
New statistical inference method for high-dimensional Hawkes processes.
problem Uncertainty evaluation of network estimates in high-dimensional point process data.
method Develops a new statistical inference procedure using concentration inequalities and martingale central limit theory.
result Characterizes the convergence rate of test statistics for high-dimensional Hawkes processes.