Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

187375562749 · Jun 202019922001200920172026
48 results for statistical applications

New framework controls statistical dispersion for high-stakes applications.

problem Understanding and controlling the dispersion of loss distributions in high-stakes applications.
method Simple yet flexible framework for distribution-free control of statistical dispersion measures.
result Proposed methods control statistical dispersion measures with societal implications.

Recently we reported on an application of the Tsallis non-extensive statistics to the S&P500 stock index. There we argued that the statistics are applicable to a broad range of markets and exchanges where anamolous (super) diffusion and 'heavy' tails of the distribution are present, as they are in the S&P500. We have c…

2002-07-16abs ↗pdf ↗

Machine learning improves official statistics but needs rigorous validation.

problem Lack of methodological robustness in machine learning for official statistics.
method Total Machine Learning Error (TMLE) framework to validate ML models.
result TMLE addresses representativeness and measurement errors in ML models.

We review basic notions in the field of information geometry such as Fisher metric on statistical manifold, αα-connection and corresponding curvature following Amari's work . We show application of information geometry to asymptotic statistical inference.

2014-10-09abs ↗pdf ↗

Risk statistic is a critical factor not only for risk analysis but also for financial application. However, the traditional risk statistics may fail to describe the characteristics of regulator-based risk. In this paper, we consider the regulator-based risk statistics for portfolios. By further developing the propertie…

2019-04-16abs ↗pdf ↗

A new test statistic measures discrepancy between conditional distributions.

problem Measuring the discrepancy between two conditional distributions.
method Proposes a Bregman matrix divergence-based statistic that avoids explicit distribution estimation.
result The new statistic inherits high-order statistics and demonstrates utility in multi-task learning, concept drift detection, and feature selection.

The paper examines extreme value statistics of high-dimensional sample covariances, with applications in finance and image analysis.

problem Statistical validation of normal conditions in high-dimensional time series data.
method Generalizes the maximal deviation of sample autocovariances to high dimensions and applies Gumbel-type extreme value asymptotics.
result Gumbel-type extreme value asymptotics holds true for high-dimensional sample covariances.

Paper reproduces a kernel-based scan B-statistic for online change-point detection.

problem Continuous detection of distribution changes in online data streams.
method Efficient kernel-based scan B-statistic for online change-point detection.
result Scan B-statistic outperforms parametric methods in challenging scenarios.

The paper derives Einstein tensors for a family of α-connections on quasi-statistical manifolds.

problem Deriving Einstein tensors for a new family of connections.
method Developed mathematical foundations of statistical and quasi-statistical manifolds, including dual and equiaffine connections.
result Explicit expressions for curvatures and Einstein tensors of the α-connections.

Sharp statistical theory for conditional diffusion models.

problem Lack of theoretical foundation for conditional diffusion models.
method Sharp statistical theory with approximation of conditional score function.
result Sample complexity bound that adapts to data distribution smoothness.

LLMs struggle to generate random numbers from statistical distributions, leading to biased results in applications.

problem LLMs' inability to generate random numbers accurately from specified distributions.
method Dual-protocol design: Batch Generation and Independent Requests, benchmarking 11 models across 15 distributions.
result Sampling fidelity degrades with distributional complexity and horizon, leading to systematic biases in downstream applications.

The need for new methods to deal with big data is a common theme in most scientific fields, although its definition tends to vary with the context. Statistical ideas are an essential part of this, and as a partial response, a thematic program on statistical inference, learning, and models in big data was held in 2015 i…

2015-09-09abs ↗pdf ↗

When the data are stored in a distributed manner, direct application of traditional statistical inference procedures is often prohibitive due to communication cost and privacy concerns. This paper develops and investigates two Communication-Efficient Accurate Statistical Estimators (CEASE), implemented through iterativ…

2019-06-12abs ↗pdf ↗

Abstract mathematical formulas for statistical structures and curvatures.

problem Developing formulas for statistical structures and curvatures.
method Proving new formulas and theorems for statistical structures and curvatures.
result Generalized formulas for statistical structures and curvatures.

Discusses new probabilistic morphisms and geometric methods in machine and statistical learning.

problem Addressing challenges in statistical, machine, and manifold learning.
method Introduces category of probabilistic morphisms and geometric methods.
result New insights and applications in various learning fields.

This paper quantifies uncertainty in Data Shapley using statistical inference.

problem Uncertainty in data valuation due to dynamic data distribution.
method Established relationship with U-statistics and quantified uncertainty using statistical inference.
result Confidence intervals for Data Shapley estimations are provided.

In this paper, we introduce the concept of principal bundles on statistical manifolds. After necessary preliminaries on information geometry and principal bundles on manifolds, we study the αα-structure of frame bundles over statistical manifolds with respect to αα-connections, by giving geometric structures. The man…

2014-03-18abs ↗pdf ↗

Develops methods for GWAS of high dimensional phenotypes using summary statistics.

problem Lack of methods to model pleiotropy in multi-phenotype GWAS.
method Bayesian inference model using summary statistics, fast computation, and biologically informed priors.
result Demonstrates utility in metabolite GWAS with interpretable pathway-level inference.

The time value of money is a critical factor not only in risk analysis, but also in insurance and financial applications. In this paper, we consider a special class of set-valued risk statistics by introducing the time value of money. In fact, the risk statistics established by this method is closer to financial realit…

2019-04-16abs ↗pdf ↗

Overview of high-dimensional time series regression methods.

problem Estimation and inference with high-dimensional time series data.
method Limit theory for high-dimensional dependent data, asymptotic theory for time series regression, statistical learning methods.
result Main limit theory results and asymptotic theory for high-dimensional time series regression.

Establishes statistical and computational bounds for influence diagnostics.

problem Identifying influential datapoints or subsets in machine learning models.
method Finite-sample statistical bounds and computational complexity for influence functions and approximate maximum influence perturbations.
result Established statistical and computational guarantees for influence diagnostics.

Paper develops approximation and statistical theory for signature-based path regression.

problem Understanding how fast signatures approximate continuous path functionals.
method Develops \(L^2\) approximation rate for smooth functionals of Itô diffusions and establishes consistency of statistical learning procedures.
result Signature-based methods improve prediction over handcrafted features in various real-data applications.

High-dimensional U-statistics show surprising phase transitions, impacting kernel-based tests.

problem Understanding phase transitions in high-dimensional U-statistics.
method Proved a convergence theorem for U-statistics of degree two in high dimensions.
result High-dimensional U-statistics can have non-Gaussian limits with larger variance and asymmetry.

New formulae identify discrete probability laws without needing normalization constants.

problem Characterizing non-normalized discrete probability distributions.
method Derive explicit formulae for mass functions using Stein's method.
result Developed tools for solving statistical problems without normalization constants.

Learning representations of data is an important problem in statistics and machine learning. While the origin of learning representations can be traced back to factor analysis and multidimensional scaling in statistics, it has become a central theme in deep learning with important applications in computer vision and co…

2019-11-26abs ↗pdf ↗

The differential geometry of Kenmotsu manifold is a valuable part of contact geometry with nice applications in other fields such as theoretical physics. In fact, its statistical counterpart, that is, Kenmotsu statistical manifold also has same importance as that of Kenmotsu manifold. Theoretical physicists have also b…

2019-05-30abs ↗pdf ↗

Unified approach for quantum and classical learning from evaluation oracles.

problem Learning from evaluation oracles in quantum and classical settings.
method Inspired by Kearns' SQ and Valiant's weak evaluation oracle, a unified framework is established.
result Characterizes query complexity for learning linear function classes and extends learnability results for quantum circuits.

FNNs can be made more interpretable with statistical methods.

problem FNNs lack interpretability and are often used as black-box models.
method Supplement FNNs with statistical inference and covariate-effect visualizations.
result FNNs can be made more like traditional statistical models.

New statistical inference method for high-dimensional Hawkes processes.

problem Uncertainty evaluation of network estimates in high-dimensional point process data.
method Develops a new statistical inference procedure using concentration inequalities and martingale central limit theory.
result Characterizes the convergence rate of test statistics for high-dimensional Hawkes processes.