This work uses statistical mechanics to explain AI learning.
problem Understanding the statistical principles behind AI learning.
method Starting from sample concentration behaviors, the study applies statistical mechanics principles to AI and machine learning.
result Exponential families and statistical quantities are key in AI and machine learning.
AI uses statistical PQRS workflow for human-machine collaboration.
problem Integrating AI with statistical principles for better results.
method PQRS workflow integrating statistical concepts with AI.
result Reproducibility and interpretability of AI algorithms and results.
A quantum circuit designed for efficient statistical model preparation and training.
problem Challenges in preparing and learning statistical models on quantum processors.
method Utilizes the maximum entropy principle to design a statistics-informed parameterized quantum circuit (SI-PQC).
result Improves trainability and interpretability for learning quantum states and classical model parameters.
Statistical methods remain relevant for ODE inverse problems, especially with sparse data.
problem The relevance of statistical methods in the era of deep learning for ODE inverse problems.
method Employed physics-informed neural networks (PINN) and manifold-constrained Gaussian process inference (MAGI) to compare statistical and deep learning approaches.
result Statistically principled methods outperform deep learning models in tasks like parameter inference and trajectory reconstruction.
We consider the concept of equilibrium in economic systems from statistical mechanics viewpoint. A new method is suggested for computing the premium on this basis. The Bühlmann economic premium principle is derived as a special case of our method.
Study on U-statistics with heavy-tailed samples, providing tail bounds and LDP.
problem Deviation of U-statistics with heavy-tailed samples.
method Exponential tail bounds and Large Deviation Principle (LDP) for U-statistics.
result Obtained an exponential upper bound for U-statistics tail decay, showing two regions of decay.
A new principle for extrapolating regression outside training data.
problem Regression extrapolation when predictions are outside training data range.
method Data-adaptive marginal transformation and simple relationship assumption.
result Progression method offers guarantees on approximation error beyond training data range.
A new principle and method improve out-of-distribution detection in generative models.
problem Out-of-distribution detection in deep generative models often fails due to poor likelihood estimates.
method Introducing the Likelihood Path (LPath) principle and new theoretical tools for OOD detection.
result Non-asymptotic provable OOD detection guarantees for variational autoencoders (VAEs).
We present a new statistical learning paradigm for Boltzmann machines based on a new inference principle we have proposed: the latent maximum entropy principle (LME). LME is different both from Jaynes maximum entropy principle and from standard maximum likelihood estimation.We demonstrate the LME principle BY deriving …
Paper justifies regularization in machine learning using Occam's razor.
problem Justifying Occam's razor in machine learning.
method Statistical learning theory to justify preference for simplicity over fit.
result Justification for regularization in machine learning.
These notes review six lectures given by Prof. Andrea Montanari on the topic of statistical estimation for linear models. The first two lectures cover the principles of signal recovery from linear measurements in terms of minimax risk. Subsequent lectures demonstrate the application of these principles to several pract…
New approach uses physics principles to improve business analytics.
problem Current business analytics methods fail with new data.
method Divide KPIs into controllable and uncontrollable groups; apply physics principles to controllable ones.
result Improves understanding and optimization of controllable KPI dynamics.
Abstract notes on robust statistical learning theory.
problem Developing robust estimators for statistical learning.
method Stressing principles of robust estimators construction and analysis.
result Emphasizes main principles of robust estimators construction and analysis.
This paper explores probabilistic numerical methods for integrating statistical computations.
problem Handling numerical error as epistemic uncertainty in statistical computation.
method Probabilistic integrators that model numerical error as a distribution.
result Probabilistic integrators can achieve posterior contraction rates similar to Monte Carlo methods.
Foundation models alter medical data science workflow, challenging veridical data science principles.
problem Foundation models disrupt traditional data science practices in medicine.
method Critically examined the medical foundation model lifecycle and its deviation from veridical data science principles.
result Foundation models challenge veridical data science principles of predictability, computability, and stability.
Bayesian framework for choosing number of blocks in stochastic block models.
problem Lack of principled statistical model selection criteria for stochastic block models.
method Bayesian framework for choosing the number of blocks and comparing to degree-corrected block models.
result Universal model selection framework capable of comparing multiple modeling combinations.
Estimates peeking effects in p-values to correct bias.
problem Data peeking biases reported p-values downward.
method Develops mechanisms to estimate running extrema of test statistics.
result Corrects bias in p-values due to peeking.
New principle in online learning: Regret can be expressed using sufficient statistics and a Burkholder function.
problem Achieving optimal online learning performance with limited memory.
method Introducing a Burkholder function that depends only on sufficient statistics, not the entire data sequence.
result Developed novel online strategies for matrix prediction and parameter-free supervised learning.
A textbook on statistical machine learning for astronomy.
problem Uncertainty quantification in astronomical data analysis.
method Bayesian inference and classical statistical methods.
result Unified framework connecting modern and traditional methods.
Simplifies matrix completion with statistical models.
problem Matrix completion under MCAR assumption.
method Statistical models and missing data analysis.
result Matrix completion valid without MCAR assumption.
Paper proves large deviation principle for stochastic approximations.
problem Asymptotic estimates of learning algorithm deviations.
method Weak convergence approach to large deviations.
result Identifies appropriate scaling sequence and new representation for rate function.
Bayesian method improves star location and flux estimation from coadded images.
problem Statistical analysis of coadded astronomical images is complicated by pixel dependence.
method Bayesian approach that implicitly marginalizes single-exposure pixel intensities.
result Method outperforms single-exposure image training for star parameter estimation.
Unified geometric approach to quantum indeterminacy.
problem Quantum indeterminacy and uncertainty principles.
method Geometric formulation using convex geometry and symplectic topology.
result Robertson-Schrodinger inequalities emerge as geometric principles.
Model uses statistical physics principles to predict financial market volatility and returns.
problem Predicting price volatility and expected returns in financial markets.
method Inspired by statistical physics, the study introduces a physical model using Level 3 order book data to measure kinetic energy and momentum.
result The model outperforms traditional and machine learning approaches in forecasting volatility and expected returns.
Global analysis of EM for mixtures of two Gaussians provides convergence insights.
problem Discrepancy between EM's theoretical guarantees and practical behavior.
method Analyzes EM's behavior in infinite sample limit and actual sequence of parameters.
result Characterizes limit points of EM sequence and establishes statistical consistency.
Paper reviews synthetic data from AI models for statistical inference.
problem When can synthetic data be used reliably in statistical inference?
method Survey of generative models, statistical analysis of pitfalls.
result Principled use of synthetic data requires careful model specification.
A new strategy selects k in k-NN regression without hold-out data.
problem Choosing optimal k in k-NN regression without hold-out data.
method Iterative procedure over k, minimum discrepancy principle.
result Minimax-optimal over smoothness function classes.
Study on discrepancy principle for learning algorithms in nonparametric regression.
problem Determining optimal iteration number in nonparametric regression with unknown optimal iteration.
method Investigates discrepancy principle and modified principles for kernelized spectral filters, using deviation inequalities and change-of-norm arguments.
result Classical discrepancy principle is adaptive for slow rates, while modified principles are adaptive for faster rates.
New method improves data-driven optimization by first-order statistical gains.
problem Unclear statistical benefits of data-driven optimization methods.
method Directionally perturbed empirical optimization (EO+) framework.
result First-order statistical improvements possible with geometrically effective side information.
New method identifies common organizational principles in networks.
problem Challenging to cluster networks of different size and density.
method Introduces a new network comparison methodology.
result Identifies common organizational principles in networks.
FALKON efficiently processes large datasets using kernel methods.
problem Limited applicability of kernel methods in large scale scenarios.
method Combining stochastic subsampling, iterative solvers, and preconditioning.
result Optimal statistical accuracy achieved with O ( n ) O(n) O ( n ) memory and O ( n n ) O(n\sqrt{n}) O ( n n ) time. New strategies for efficient distributed ML on big data.
problem Running ML algorithms on big data requires significant engineering effort.
method Discussion of principles and strategies for efficient, scalable ML systems.
result Principles and strategies for designing efficient, scalable ML systems.
Proposes LatLapMED for detecting high-utility anomalies.
problem Detecting statistically rare instances with real-world significance.
method Uses EM algorithm to combine entropy minimization and maximum entropy discrimination.
result Superior performance over existing anomaly detection methods.
This paper provides a method for noise-calibrated inference from DP synthetic data.
problem Inference from DP synthetic data is often miscalibrated and lacks principled uncertainty quantification.
method Release DP sufficient statistics, perform noise-calibrated likelihood-based inference, and optional synthetic data generation.
result Asymptotic normality and valid confidence intervals for the plug-in DP MLE.
New method improves neural network robustness to adversarial attacks.
problem Vulnerability of neural networks to adversarial examples.
method Distributionally robust optimization with Wasserstein ball penalties.
result Provable robustness with minimal computational cost.
A novel feature selection method using noise-based hypothesis testing improves feature selection accuracy.
problem Challenges in feature selection for complex, high-dimensional datasets.
method Introduces multiple random noise features and evaluates feature importance against noise feature maxima using non-parametric bootstrap-based hypothesis testing.
result Outperforms existing methods in simulated and real-world datasets.
Prove non-asymptotic bounds for minimal risk in statistical learning
problem Estimating minimal risk in statistical learning
method Using concentration inequalities
result Non-asymptotic bounds for minimal risk
The paper explores new learning principles based on least cognitive action.
problem Formulating learning problems in a way that incorporates time and perception.
method Introduces the principle of least cognitive action and proves its well-posedness.
result Proves the existence of a minimum for a special form of cognitive action, leading to fourth-order differential equations of learning.
Deep neural networks predict summary statistics for ABC methods.
problem Constructing effective summary statistics for ABC methods in models with unknown likelihoods.
method Train deep neural networks to predict parameters from generated data.
result Automated summary statistics improve posterior accuracy in Ising and moving-average models.
Designs new functionals for ranking joint probability distributions based on correlations.
problem Ranking joint probability distributions based on their correlations.
method Using first principles from inference, a set of functionals are designed with the Principle of Constant Correlations (PCC) guiding the construction.
result The n n n -partite information (NPI) uniquely determines whether inferential transformations preserve, destroy, or create correlations. A comprehensive guide to MDL principle, a statistical inference theory.
problem Statistical inference and pattern recognition.
method MDL principle, model selection, hypothesis testing, averaging.
result Unified perspective on various statistical methods.
Paper explores SVGD for Bayesian inference, linking deterministic and stochastic dynamics.
problem Bayesian inference and Markov chain Monte Carlo methods.
method Stein variational gradient descent (SVGD) with deterministic and stochastic dynamics.
result Identifies Stein-Fisher information as the leading order contribution in the long-time and many-particle regime.
This Chapter is written for the Festschrift celebrating the 70th birthday of the distinguished economist Duncan Foley from the New School for Social Research in New York. This Chapter reviews applications of statistical physics methods, such as the principle of entropy maximization, to the probability distributions of …
We show how to incorporate ethical principles into machine learning models.
problem Machine-learned systems often violate deontological ethical principles.
method Add shape constraints to models to ensure positive responses to relevant inputs.
result Shape constraints help produce more ethical AI by avoiding penalties on good attributes.
This paper analyzes the Shattering coefficient for supervised learning algorithms.
problem Ensuring uniform convergence of empirical risk to expected risk.
method Analyzes the Shattering coefficient for Hilbert spaces containing input spaces.
result Proves the Shattering coefficient's polynomial growth for any Hilbert space.
Develops a statistical framework for coherent risk estimation.
problem Constructing coherent risk estimators with sound financial and statistical properties.
method Inspired by axiomatic risk measure theory, defines coherent risk estimators through robust representations linked to L L L -estimators. result Demonstrates that coherence of a risk measure does not necessarily carry over to its estimators and shows alternative weight structures can lead to different outcomes.
Transformers recall from long distributions with statistical guarantees.
problem Designing Transformers that can recall from arbitrarily long, distributional contexts.
method Recast associative memory as probability measures, decomposing the task into recall and prediction.
result A shallow measure-theoretic Transformer learns the recall-and-predict map under spectral assumptions.
Paper develops a statistical model for summarizing event sequences.
problem Discovering frequent serial episodes from sequential data.
method Minimum Description Length (MDL) principle with modifications.
result Reduces dictionary size by more than four-fold without losing accuracy.