This paper connects functional data analysis with machine learning techniques.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes a statistical test for feature selection pipelines using selective inference.
Study proposes a statistical testing framework for evaluating clustering pipelines.
Unified platform for statistical and machine learning in bioinformatics.
DPpack offers R tools for private data analysis and machine learning.
Computational topology has recently known an important development toward data analysis, giving birth to the field of topological data analysis. Topological persistence, or persistent homology, appears as a fundamental tool in this field. In this paper, we study topological persistence in general metric spaces, with a …
The paper proposes methods to infer from privacy-protected data using simulation-based techniques.
Tensor analysis tackles complex multidimensional data across fields.
Data science enhances knot theory by analyzing invariant relations.
Making sense of a dataset in an automatic and unsupervised fashion is a challenging problem in statistics and AI. Classical approaches for {exploratory data analysis} are usually not flexible enough to deal with the uncertainty inherent to real-world data: they are often restricted to fixed latent interaction models an…
CEDA improves understanding of data fit to models.
This paper quantifies privacy loss in exploratory data analysis.
Paper stabilizes persistent homology rank functions for statistical inference.
RobPy offers robust statistical methods in Python.
As regulators pay more attentions to losses rather than gains, we are able to derive a new class of risk statistics, named regulator-based risk statistics with scenario analysis in this paper. This new class of risk statistics can be considered as a kind of risk extension of risk statistics introduced by Kou et al. \ci…
The paper examines extreme value statistics of high-dimensional sample covariances, with applications in finance and image analysis.
A new method for analyzing shapes using FDA techniques.
The book chapter discusses tail risk analysis for financial data using extreme value statistics.
New method robustifies topological data analysis against outliers.
In recent years, ideas from statistics and scientific computing have begun to interact in increasingly sophisticated and fruitful ways with ideas from computer science and the theory of algorithms to aid in the development of improved worst-case algorithms that are useful for large-scale scientific and Internet data an…
We introduce RSE to measure robustness in estimation problems.
Survey of techniques for diagnosing pediatric sleep apnea from inexpensive data.
This describes a statistical technique called "tonsuring" for exploratory data analysis in finance. Instead of rejecting "outlier" data that conflicts with the model, this strips out "inlier" data to get a clearer picture of how the market changes for larger moves.
Adapts data analysis for growing data, improving generalization guarantees.
The paper gives picture of enrichment to economic and financial system analysis using agent-based models as a form of advanced study for financial economic data post-statistical-data analysis and micro-simulation analysis. Theoretical exploration is carried out by using comparisons of some usual financial economy syste…
This paper describes a new approach to time series modeling that combines subject-matter knowledge of the system dynamics with statistical techniques in time series analysis and regression. Applications to American option pricing and the Canadian lynx data are given to illustrate this approach.
Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data intro…
This paper introduces compositional data analysis for financial ratios, improving industry-level analysis.
Neuroimaging research has predominantly drawn conclusions based on classical statistics, including null-hypothesis testing, t-tests, and ANOVA. Throughout recent years, statistical learning methods enjoy increasing popularity, including cross-validation, pattern classification, and sparsity-inducing regression. These t…
Develops methods for causal inference in compositional data using instrumental variables.
The problem of complex data analysis is a central topic of modern statistical science and learning systems and is becoming of broader interest with the increasing prevalence of high-dimensional data. The challenge is to develop statistical models and autonomous algorithms that are able to acquire knowledge from raw dat…
Traditional statistical theory assumes that the analysis to be performed on a given data set is selected independently of the data themselves. This assumption breaks downs when data are re-used across analyses and the analysis to be performed at a given stage depends on the results of earlier stages. Such dependency ca…
Hypothesis testing is one of the most common types of data analysis and forms the backbone of scientific research in many disciplines. Analysis of variance (ANOVA) in particular is used to detect dependence between a categorical and a numerical variable. Here we show how one can carry out this hypothesis test under the…
Framework for efficient statistical estimation with privacy guarantees.
Kernel methods are powerful learning methodologies that allow to perform non-linear data analysis. Despite their popularity, they suffer from poor scalability in big data scenarios. Various approximation methods, including random feature approximation, have been proposed to alleviate the problem. However, the statistic…
Real data often contain anomalous cases, also known as outliers. These may spoil the resulting analysis but they may also contain valuable information. In either case, the ability to detect such anomalies is essential. A useful tool for this purpose is robust statistics, which aims to detect the outliers by first fitti…
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
This paper investigates asymptotic behaviors of gradient descent algorithms (particularly accelerated gradient descent and stochastic gradient descent) in the context of stochastic optimization arising in statistics and machine learning where objective functions are estimated from available data. We show that these alg…
The statistical analysis of data lying on a differentiable, locally Euclidean, manifold introduces a variety of challenges because the analogous measures to standard Euclidean statistics are local, that is only defined within a neighbourhood of each datapoint. This is because the curvature of the space means that the c…
Symbolic data analysis (SDA) is an emerging area of statistics concerned with understanding and modelling data that takes distributional form (i.e. symbols), such as random lists, intervals and histograms. It was developed under the premise that the statistical unit of interest is the symbol, and that inference is requ…
Neuroscience is undergoing faster changes than ever before. Over 100 years our field qualitatively described and invasively manipulated single or few organisms to gain anatomical, physiological, and pharmacological insights. In the last 10 years neuroscience spawned quantitative big-sample datasets on microanatomy, syn…
Machine learning and statistical modeling complement each other in healthcare analytics.
Divide-and-conquer method splits large data sets for efficient analysis.
Fine-grained atlases improve fMRI analysis of brain activity.
Novel framework for ML-assisted inference valid for any statistical task.
For high dimensional data, some of the standard statistical techniques do not work well. So modification or further development of statistical methods are necessary. In this paper, we explore these modifications. We start with the important problem of estimating high dimensional covariance matrix. Then we explore some …
This paper proposes a statistical mechanics approach to the analysis of income distribution and inequality. A new distribution function, having its roots in the framework of k-generalized statistics, is derived that is particularly suitable to describe the whole spectrum of incomes, from the low-middle income region up…
Overview of high-dimensional time series regression methods.