Deep learning's success is puzzling from a statistical perspective.
problem Deep learning's success is puzzling from a statistical perspective.
method Physics-informed investigation of deep learning features and surprises.
result Neural scaling laws and their interplay with inductive biases.
Statistical field theory aids in understanding deep learning complexities.
problem Complexity and lack of theoretical understanding in deep learning.
method Statistical field theory as a theoretical framework.
result Field theory provides insights into generalization, bias, and feature learning.
Teaches deep learning to statisticians.
problem Statisticians lack expertise in deep learning.
method Developed a program and taught DL to statistics graduate students.
result Provided tips and resources for teaching DL.
The paper provides statistical guarantees for sparse deep learning.
problem Understanding the potential and limitations of sparse deep learning.
method Develops statistical guarantees for different types of sparsity in sparse deep learning.
result Statistical guarantees for sparse deep learning with mild dependence on network widths and depths.
Deep learning uncovers patterns between knot types.
problem Discovering connections between combinatorial and hyperbolic knot invariants.
method Statistical approach using linear regression and deep learning.
result Revealed empirical connections between knot types.
Deep learning used for parameter estimation in hard-to-infer models.
problem Parameter estimation in intractable models like max-stable processes.
method Train deep neural networks on simulated data to estimate parameters.
result Deep learning provides accurate and faster parameter estimation.
Statistical methods remain relevant for ODE inverse problems, especially with sparse data.
problem The relevance of statistical methods in the era of deep learning for ODE inverse problems.
method Employed physics-informed neural networks (PINN) and manifold-constrained Gaussian process inference (MAGI) to compare statistical and deep learning approaches.
result Statistically principled methods outperform deep learning models in tasks like parameter inference and trajectory reconstruction.
New framework tackles deep learning issues like local traps and miscalibration.
problem Local traps and miscalibration in deep neural networks.
method Sparse deep learning framework with prior annealing algorithms.
result Proposed method successfully addresses local traps and miscalibration.
Deep learning has arguably achieved tremendous success in recent years. In simple words, deep learning uses the composition of many nonlinear functions to model the complex dependency between input features and labels. While neural networks have a long history, recent advances have greatly improved their performance in…
Flexible framework for deep distributional regression models.
problem Learning conditional distributions from semi-structured data.
method Combines additive regression models with deep networks using TensorFlow.
result State-of-the-art predictive performance with interpretability.
Develops a deep learning approach for statistical arbitrage.
problem Temporal price differences between similar assets.
method Constructs arbitrage portfolios using latent asset pricing factors and a convolutional transformer for time series signals.
result High risk-adjusted returns and Sharpe ratios with optimal trading policy.
Survey of deep learning methods for time series forecasting.
problem Improving accuracy in time series predictions across various domains.
method Analysis of common encoder and decoder designs, hybrid models, and decision support.
result Advancements in deep learning for time series forecasting.
This study compares deep learning and statistical models for stock price forecasting.
problem Accurate stock price prediction is challenging due to market volatility.
method Used deep learning (LSTM, RNN, CNN, FULL CNN) and statistical models (ARIMA, Moving Averages) on S&P 500 data.
result LSTM model showed the lowest Mean Absolute Error (MAE), indicating highest accuracy.
Bayesian Neural Networks help quantify uncertainty in deep learning predictions.
problem Uncertainty quantification in deep learning predictions.
method Bayesian statistics applied to neural networks.
result Design, implementation, training, and evaluation of Bayesian Neural Networks.
This paper introduces a novel measure-theoretic theory for machine learning that does not require statistical assumptions. Based on this theory, a new regularization method in deep learning is derived and shown to outperform previous methods in CIFAR-10, CIFAR-100, and SVHN. Moreover, the proposed theory provides a the…
Deep learning uses layers of transformations to predict structured data with uncertainty.
problem Predicting structured high-dimensional data efficiently and with uncertainty.
method Applying layers of semi-affine input transformations to find features for probabilistic statistical methods.
result Achieves scalable prediction rules with uncertainty quantification and feature selection.
In this paper, we present a statistical-mechanical analysis of deep learning. We elucidate some of the essential components of deep learning---pre-training by unsupervised learning and fine tuning by supervised learning. We formulate the extraction of features from the training data as a margin criterion in a high-dime…
In this paper we develop a statistical theory and an implementation of deep learning models. We show that an elegant variable splitting scheme for the alternating direction method of multipliers optimises a deep learning objective. We allow for non-smooth non-convex regularisation penalties to induce sparsity in parame…
Lectures on deep learning properties in infinite and large-width networks.
problem Understanding deep neural networks in extreme width conditions.
method Analysis of random deep neural networks, connections to linear models, kernels, and Gaussian processes, perturbative and non-perturbative treatments.
result Properties and behaviors of deep neural networks in the infinite-width limit and large-width regime.
MegazordNet combines stats and ML for better financial time series forecasting.
problem Forecasting financial time series is challenging due to its chaotic nature.
method MegazordNet integrates statistical features with a deep learning model.
result MegazordNet outperforms single statistical and machine learning methods in S&P 500 stock price prediction.
Lecture notes on linear neural networks for deep learning optimization and generalization.
problem Understanding optimization and generalization in deep learning models.
method Mathematical tools and dynamical systems theory.
result Potential of mathematical tools to enhance understanding of deep learning.
Lectures on deep learning from a learning theory perspective.
problem Understanding how deep learning architectures lead to inductive bias.
method Statistical learning theory and stochastic optimization.
result Gradient descent on linear diagonal networks can lead to various forms of implicit bias.
si4onnx enables selective inference on deep learning models.
problem Establishing the reliability of AI systems through statistical significance of identified regions.
method Selective inference techniques implemented through a Python package.
result Controlled type I error rates for hypothesis testing on deep learning models.
New deep learning model interprets tabular data with variable selection and explainability.
problem Deep learning models lack interpretability and variable selection.
method Proposes a new network architecture that combines deep learning with generalized linear models.
result The model provides superior predictive power and interpretable results.
DRIFT uses neural flows to replace distributional regression models.
problem Lack of neural network representations for distributional regression models.
method Inverse flow transformations (DRIFT) for distributional regression.
result Neural representations in DRIFT match classical statistical methods in performance.
A theory of deep learning is emerging, focusing on training dynamics and statistics.
problem Develop a scientific theory to understand deep learning.
method Synthesize research into five areas: idealized settings, tractable limits, mathematical laws, hyperparameters, and universal behaviors.
result The emerging theory is a mechanics of the learning process, named learning mechanics.
Deep-learning method improves hypothesis testing for independence.
problem Improving hypothesis testing for independence using deep learning.
method Proposes deep-testing, a novel procedure that uses a deep neural network to distinguish between data generated under and outside a given statistical model.
result Deep-testing achieves the highest overall power against nineteen competing methods across various dependence structures.
Paper introduces a new gradient statistic to improve deep learning convergence.
problem Fluctuation effect of gradient updates between iterations.
method Introduces an unbiased stratified statistic \(\bar{G}_{mst}\) and a new algorithm MSSG.
result MSSG algorithm outperforms other sgd-like algorithms in training deep models.
GenFormer uses deep learning to generate complex stochastic data.
problem Creating synthetic stochastic data that matches real-world statistical properties.
method Transformer-based deep learning model that maps Markov state sequences to time series values.
result GenFormer preserves target marginal distributions and other statistical properties in multivariate spatio-temporal data.
MES-LSTM hybrid method improves multivariate time series forecasting and mortality modeling.
problem Challenges in applying hybrid forecast methods to multivariate data.
method Generalized multivariate extension of ES-RNN, utilizing vectorized implementation.
result MES-LSTM shows significant improvement over pure statistical and deep learning methods in forecast accuracy and prediction interval construction.
CRL uses causality to build interpretable AI models from complex data.
problem Interpreting deep neural networks' implicit representations.
method Causal representation learning (CRL) synthesizing latent variable models, causal graphical models, and nonparametric statistics.
result CRL can improve interpretability of generative AI models.
Deep learning architectures have demonstrated state-of-the-art performance for object classification and have become ubiquitous in commercial products. These methods are often applied without understanding (a) the difficulty of a classification task given the input data, and (b) how a specific deep learning architectur…
Paper examines the structure of stochastic gradients in deep learning.
problem Exploring the structure and heavy tails of stochastic gradients in deep learning.
method Conducted formal statistical tests on stochastic gradients and gradient noise.
result Stochastic gradients and gradient noise do not exhibit power-law heavy tails, but their covariance spectra do.
The burgeoning success of deep learning has raised the security and privacy concerns as more and more tasks are accompanied with sensitive data. Adversarial attacks in deep learning have emerged as one of the dominating security threat to a range of mission-critical deep learning systems and applications. This paper ta…
Deep learning depends on tuning layers near critical points.
problem Understanding how deep learning architectures depend on tuning parameters.
method Random energy approach to analyze statistical dependence in deep belief networks.
result Statistical dependence can propagate only if layers are tuned near critical points.
DeRisk improves credit risk prediction using deep learning.
problem Challenges in training deep neural networks with real-world financial data.
method DeRisk, an effective deep learning framework for credit risk prediction.
result DeRisk outperforms statistical learning methods in credit risk prediction.
Simplified neural network EFTs reveal a single critical condition.
problem Understanding neuron statistics in neural networks at initialization.
method Diagrammatic approach to effective field theories (EFTs).
result A single condition governs criticality of all neuron preactivations.
Improved deep learning performance in financial markets by using rank space.
problem High volatility and low signal-to-noise ratio in equity market dynamics.
method Transformed equity market data from name space to rank space, enabling better learning by DNNs.
result DNNs achieve superior performance in statistical arbitrage in rank space compared to name space.
Deep learning models predict mutual funds' performance better than traditional methods.
problem Predicting mutual funds' performance accurately.
method Deep learning models (LSTM, GRUs) trained with Bayesian optimization and ensemble methods.
result Ensemble method of LSTM and GRUs achieves the highest accuracy in forecasting mutual funds' Sharpe ratios.
Efficient method for training deep learning models with human validation and statistical analysis.
problem Challenges in labeling medical images for deep learning, including time and cost.
method Four-step method using automated data and human visual checks for iterative refinement and statistical validation.
result Initial model accuracy improved from 92% to 98% with statistical validation.
Study excess risk in statistical inference with transformations.
problem Excess risk in estimating random variables from feature vectors and transformations.
method Characterize lossless transformations, develop test statistics, and information-theoretic bounds.
result Strongly consistent partitioning test statistic for lossless transformations.
Automated rock fragmentation assessment using deep learning and spatial statistics.
problem Assessing post-blast rock fragmentation in real-time.
method Fine-tuned YOLO12l-seg model for instance segmentation, followed by spatial statistics.
result Framework accurately assesses rock fragmentation patterns in real-time.
This research improves model interpretability and uncertainty estimation for deep learning models on non-iid data.
problem Improving interpretability and uncertainty estimation for deep learning models on non-iid data.
method 4 UQ approaches (BNN, SWAG, MC dropout, ensemble) applied to ARMED MEDL models.
result Ensemble approaches, especially with 90% subsampling, provide best performance in prediction and uncertainty estimation.
Study shows depth improves generalization in deep learning models.
problem Understanding why and when depth improves generalization in deep learning.
method Implementation-agnostic state-transition model to analyze depth and generalization.
result Identifies geometric and semigroup mechanisms that keep entropy contribution saturated or polynomial, clarifying depth's statistical advantage.
Deep learning struggles with out-of-distribution data, so this paper tackles domain generalization.
problem Deep learning models fail with out-of-distribution data.
method Formulates domain generalization as a constrained statistical learning problem, then uses nonconvex duality theory to develop an algorithm with convergence guarantees.
result Improves domain generalization by up to 30 percentage points on various benchmarks.
This paper examines error bounds for deep learning classifiers with noisy labels.
problem Understanding the performance of classifiers trained on noisy data.
method Derives error bounds for excess risk, decomposing it into statistical and approximation errors. Uses independent block construction for statistical dependencies and vector-valued setting for approximation error.
result Established theoretical results for error bounds in deep learning with noisy labels, mitigating the impact of high-dimensional input spaces.
Researchers use DL and XAI to evaluate climate downscaling models.
problem Evaluating complex DL models for climate downscaling.
method Intercompare DL models, expand standard evaluation methods with XAI.
result XAI techniques provide new evaluation dimensions and model insights.
Statistical physics explains deep learning's feature learning capacity.
problem Understanding neural networks' ability to learn complex features.
method Study of a multi-layer perceptron in the interpolation regime.
result Optimal learning requires specialization across layers and neurons.