We highlight a very simple statistical tool for the analysis of financial bubbles, which has already been studied in [1]. We provide extensive empirical tests of this statistical tool and investigate analytically its link with stocks correlation structure.
This chapter reviews statistical tools for reinforcement learning.
problem Applying RL algorithms in healthcare and ride-sharing platforms.
method Statistical inference tools for RL, including hypothesis testing and confidence interval construction.
result Highlighting the value of statistical inference in RL for both communities.
Statistical physics helps solve complex machine learning problems.
problem Large dimensional inference problems in machine learning.
method Replica symmetric level analysis and cavity methods.
result General framework for solving various problems with weak long-range interactions.
Python tools for 3D shape analysis on Kendall's space.
problem Lack of practical utilities for advanced 3D shape analysis.
method Developed Python tools for 3D shape analysis on Kendall's 3D Shape Space.
result Efficient, accessible software solutions for researchers.
The paper uses statistics to improve the explainability of models.
problem Subjective human assessment of explanations and lack of theoretical guarantees.
method Leveraging statistical estimators for proper definition and evaluation of explanations.
result Statistical tools provide theoretical guarantees and evaluation metrics for explanations.
The paper tackles individual fairness in ML models, developing statistical methods to detect bias.
problem Detecting and measuring violations of individual fairness in machine learning models.
method Formalizing the problem as adversarial attack, developing inference tools for the adversarial cost function.
result Statistical methods to assess and test hypotheses of model fairness with non-coverage error rate control.
A review of ML and DL for ecological data analysis.
problem Understanding the strengths and limitations of ML and DL in ecological research.
method Historical overview, algorithm families, differences, universal principles, and emerging trends.
result ML and DL excel in prediction tasks but are still debated for causal inference.
Enhances U-statistics for semi-supervised datasets using unlabeled data.
problem Efficiently utilizing unlabeled data in semi-supervised settings.
method Semi-supervised U-statistics enhanced by unlabeled data.
result Proposed method is asymptotically Normal and more efficient than classical U-statistics.
Survey on statistical learning theory for control, focusing on linear systems.
problem Applying machine learning techniques to control systems, especially linear ones.
method Adapting tools from modern high-dimensional statistics and learning theory.
result Recent advances in statistical learning theory for control, particularly for linear systems.
TML package uses tropical geometry for machine learning tasks.
problem Statistical learning problems.
method Tropical convexity computations, Hit and Run sampler, tropical metrics.
result First R package for tropical geometric machine learning.
Sliced Optimal Transport simplifies OT for fast computation.
problem Efficient computation of distances and barycenters for probability measures.
method Combines OT, integral geometry, and statistics for fast computation.
result Retains rich geometric structure while speeding up computations.
Information theory plays an indispensable role in the development of algorithm-independent impossibility results, both for communication problems and for seemingly distinct areas such as statistics and machine learning. While numerous information-theoretic tools have been proposed for this purpose, the oldest one remai…
RobPy offers robust statistical methods in Python.
problem Lack of robust statistical methods in Python.
method Built on NumPy, SciPy, and scikit-learn, RobPy includes robust tools for various statistical tasks.
result RobPy enables more users to perform robust data analysis in Python.
In these notes we describe heuristics to predict computational-to-statistical gaps in certain statistical problems. These are regimes in which the underlying statistical problem is information-theoretically possible although no efficient algorithm exists, rendering the problem essentially unsolvable for large instances…
New statistical inference method for high-dimensional Hawkes processes.
problem Uncertainty evaluation of network estimates in high-dimensional point process data.
method Develops a new statistical inference procedure using concentration inequalities and martingale central limit theory.
result Characterizes the convergence rate of test statistics for high-dimensional Hawkes processes.
Neural networks speed up statistical inference.
problem Efficient statistical inference for complex models.
method Neural networks for learning complex mappings.
result Amortized inference speeds up inference processes.
Machine learning and statistical modeling complement each other in healthcare analytics.
problem Choosing between machine learning and statistical modeling for analytics challenges.
method Choosing based on problem, data, and desired outcomes.
result Machine learning and statistical modeling are complementary, using similar principles but different tools.
Introduces m-connecting imset and factorization for ADMG models.
problem Handling latent confounding in DAG models.
method Introduces m-connecting imset and m-connecting factorization criterion for ADMG models.
result Equivalence of m-connecting factorization criterion to global Markov property.
Establishes statistical and computational bounds for influence diagnostics.
problem Identifying influential datapoints or subsets in machine learning models.
method Finite-sample statistical bounds and computational complexity for influence functions and approximate maximum influence perturbations.
result Established statistical and computational guarantees for influence diagnostics.
Sparklen is a Python toolkit for high-dimensional Hawkes processes.
problem Efficiently analyzing high-dimensional Hawkes processes.
method Combines Python for ease of use with C++ for performance.
result Provides state-of-the-art tools for estimating and classifying Hawkes processes.
Develops algorithms to balance personalization and statistical power in mobile health studies.
problem Balancing personalization and statistical power in mobile health studies.
method Develops general meta-algorithms to modify existing bandit algorithms.
result Guarantees sufficient power while improving user well-being.
Real data often contain anomalous cases, also known as outliers. These may spoil the resulting analysis but they may also contain valuable information. In either case, the ability to detect such anomalies is essential. A useful tool for this purpose is robust statistics, which aims to detect the outliers by first fitti…
Socio-economic inequalities are manifested in different aspects of our social life. We discuss various aspects, beginning with the evolutionary and historical origins, and discussing the major issues from the social and economic point of view. The subject has attracted scholars from across various disciplines, includin…
Using the tools developed for statistical physics, we simultaneously analyze statistical properties of the Jakarta and Kuala Lumpur Stock Exchange indices. In spite of the small number of data used in the analysis, the result shows the universal behavior of complex systems previously found in the leading stock indices.…
Probabilistic models can handle causal inference without special tools.
problem Confusion over necessary tools for causal inference.
method Demonstrated through concrete examples that causal questions can be answered using standard probabilistic models.
result Causal questions can be addressed using standard probabilistic modelling and inference.
I introduce a new geometrical approach to thermo--statistical mechanics. Here I highlight the main physical ideas, and how do they translate into geometrical language. I contrast the present approach with previous thermo--statistical--geometrical formalisms, (pseudo-)Riemannian [Weinhold 1975; Ruppeiner 1979] as well a…
An algorithm reduces breast cancer detection data complexity using effect sizes.
problem Improving accuracy in breast cancer detection.
method Statistical feature selection and SVM classifier with linear kernel.
result SVM classifier achieved over 90% accuracy.
I present in this paper some tools in Symplectic and Poisson Geometry in view of their applications in Geometric mechanics and Mathematical Physics. After a short discussion of the Lagrangian and Hamiltonian formalisms, including the use of symmetry groups, and a presentation of the Tulczyjew's isomorphisms (which expl…
A new model improves analysis of neural activity from calcium imaging.
problem Statistical modeling of deconvolved calcium signals for neural activity interpretation.
method Proposed a zero-inflated gamma (ZIG) model to characterize calcium responses as a mixture of a gamma distribution and a point mass.
result The ZIG model outperforms simpler models in neural encoding and decoding problems.
New tools quantify deep generative models' performance.
problem Measuring the quality-diversity trade-off in deep generative models.
method Established non-asymptotic bounds on sample complexity and introduced frontier integrals.
result Smoothed estimators improve convergence rates of divergence frontiers.
Generative model improves intraday electricity price forecasting.
problem Intraday electricity price forecasting for improved trading strategies.
method Generative neural network model for probabilistic path forecasts.
result Generative model leads to higher profit gains than benchmark methods.
Paper analyzes shapes of brain arterial networks using statistical methods.
problem Quantifying and comparing shapes of brain arterial networks.
method Mathematical representation of BAN shapes as elastic shape graphs, development of Riemannian metrics and geometrical tools.
result Age has a clear, quantifiable effect on BAN shapes, with increased variance in shapes as age increases.
A new statistical model uses Orlicz-Sobolev spaces with Gaussian weight.
problem Statistical modeling of infinite-dimensional probability measures.
method Affine statistical bundle on Gaussian Orlicz-Sobolev space.
result Provides tools for solving infinite-dimensional evolution problems.
The paper gives picture of enrichment to economic and financial system analysis using agent-based models as a form of advanced study for financial economic data post-statistical-data analysis and micro-simulation analysis. Theoretical exploration is carried out by using comparisons of some usual financial economy syste…
MM (majorization--minimization) algorithms are an increasingly popular tool for solving optimization problems in machine learning and statistical estimation. This article introduces the MM algorithm framework in general and via three popular example applications: Gaussian mixture regressions, multinomial logistic regre…
Over the past decades, statisticians and machine-learning researchers have developed literally thousands of new tools for the reduction of high-dimensional data in order to identify the variables most responsible for a particular trait. These tools have applications in a plethora of settings, including data analysis in…
Computational topology has recently known an important development toward data analysis, giving birth to the field of topological data analysis. Topological persistence, or persistent homology, appears as a fundamental tool in this field. In this paper, we study topological persistence in general metric spaces, with a …
In this paper, we consider formal series associated with events, profiles derived from events, and statistical models that make predictions about events. We prove theorems about realizations for these formal series using the language and tools of Hopf algebras.
New statistical framework for coresets in density estimation.
problem Improving computational efficiency in density estimation.
method Developed a statistical framework for coresets in nonparametric density estimation.
result Practical coreset kernel density estimators are near-minimax optimal.
Using AI predictions as data can mislead inference, study shows.
problem Misleading inference when using AI predictions instead of real data.
method Characterized statistical challenges and reviewed methods for IPD.
result High predictive accuracy doesn't ensure valid inference.
New arithmetic phenomenon 'murmurations' detected using AI.
problem Detecting new arithmetic patterns in large datasets.
method Machine learning interpretability tools (PCA, saliency, convolutional filters).
result Murmurations encode Frobenius traces and connect to number theory.
Lecture notes on linear neural networks for deep learning optimization and generalization.
problem Understanding optimization and generalization in deep learning models.
method Mathematical tools and dynamical systems theory.
result Potential of mathematical tools to enhance understanding of deep learning.
One aim of data mining is the identification of interesting structures in data. For better analytical results, the basic properties of an empirical distribution, such as skewness and eventual clipping, i.e. hard limits in value ranges, need to be assessed. Of particular interest is the question of whether the data orig…
Hypothesis tests are a crucial statistical tool for data mining and are the workhorse of scientific research in many fields. Here we present a differentially private analogue of the classic Wilcoxon signed-rank hypothesis test, which is used when comparing sets of paired (e.g., before-and-after) data values. We present…
Information geometry offers new tools for statistical analysis.
problem Statistical analysis of probability distributions.
method Geometric perspective on statistical manifolds.
result New applications in radar sensing, signal processing, etc.
This paper provides a dictionary of closed-form kernel mean embeddings.
problem Challenges in deriving closed-form kernel mean embeddings.
method Comprehensive dictionary and practical tools for deriving new embeddings.
result Provides a Python library with minimal implementations of embeddings.
sig-MMD tests compare path distributions using kernel methods.
problem Comparing path distributions in stochastic processes.
method Signature kernel for path space valued distributions.
result sig-MMD can lead to Type 2 errors in limited data settings.
StatEcoNet models species distribution using neural networks to correct observation errors.
problem Correcting observation errors in wildlife surveys for accurate species distribution modeling.
method StatEcoNet integrates a graphical generative model with neural networks to address SDM challenges.
result StatEcoNet outperforms traditional methods on simulated and real datasets.