Proposes mGBDTs for learning hierarchical representations in gradient boosting decision trees.
problem Inability of gradient boosting decision trees to learn hierarchical representations.
method Introduces multi-layered GBDT forest (mGBDTs) with explicit emphasis on hierarchical learning.
result Jointly trained mGBDTs can learn hierarchical representations effectively without backpropagation.
The multi-layer IB problem optimizes relevance and compression rates.
problem Optimizing relevance and compression rates in multi-layer information propagation.
method Single-letter characterization of the rate-relevance region, conditions for successive refinability, and counterexamples.
result Successive refinability of binary and Gaussian models, counterexample provided.
SuperSpike learns spiking neural networks to perform complex computations.
problem Training spiking neural networks to perform nonlinear computations.
method Derive SuperSpike, a learning rule for multi-layer spiking neural networks.
result SuperSpike enables training of multi-layer spiking neural networks to solve complex tasks.
New methods for uncertainty in neural networks with leaky ReLU activations.
problem Uncertainty in feed-forward neural networks with random input perturbations.
method Analytical expressions for PDF and moments of neural network output, linearization of leaky ReLU, Gaussian copula surrogate models.
result Accurate statistical results for large input perturbations, excellent agreement with Monte Carlo simulations.
Deep Gaussian processes (DGPs) are multi-layer hierarchical generalisations of Gaussian processes (GPs) and are formally equivalent to neural networks with multiple, infinitely wide hidden layers. DGPs are probabilistic and non-parametric and as such are arguably more flexible, have a greater capacity to generalise, an…
This review tackles long horizon forecasting in time series analysis using deep learning.
problem Long horizon forecasting in time series analysis.
method Incorporates deep learning techniques such as trend, seasonality, Fourier and wavelet transforms, and various model architectures.
result LHF is an error propagation problem, with models like xLSTM and Triformer showing better performance.
Adapts BP-based algorithms for deep learning, improving performance and accuracy.
problem Training deep neural networks with discrete weights and activations.
method Message-passing algorithms based on Belief Propagation, with reinforcement field.
result Comparable performance to SGD-inspired heuristics (BinaryNet) and higher accuracy in predictions.
DeepAM optimizes deep neural networks for image super-resolution.
problem Image super-resolution using deep neural networks.
method L-layer analysis dictionary model with IPAD and CAD.
result DeepAM outperforms deep neural networks with back-propagation.
A new neural network initialization method is proposed for faster and more accurate training.
problem Efficient initialization for training multi-layer feedforward neural networks.
method Initialization based on Stein's identity, using eigenvectors of cross-moment matrix.
result The SteinGLM method is faster and more accurate than other initialization methods.
Noise stability improves understanding of Transformer models.
problem Lack of robustness metrics for real-valued domains and junta-like input dependence in modern LLMs.
method Proposed noise stability as a new metric and developed a practical regularization method.
result Noise stability regularization method accelerates training by 35-75%.
Deep Gaussian processes (DGPs) are multi-layer hierarchical generalisations of Gaussian processes (GPs) and are formally equivalent to neural networks with multiple, infinitely wide hidden layers. DGPs are nonparametric probabilistic models and as such are arguably more flexible, have a greater capacity to generalise, …
Model infers diffusion networks from heterogeneous cascade data.
problem Understanding and predicting diffusion processes in interconnected populations.
method Double mixture directed graph model with layer-specific constraints.
result Convex formulation allows for statistical and computational guarantees.
Our paper explains deep neural collapse in multiple layers.
problem Understanding deep neural collapse in multi-layered neural networks.
method Generalized unconstrained features model for deep networks.
result Deep unconstrained features model exhibits deep neural collapse.
New multi-layer transform models improve image denoising.
problem Improving image denoising techniques.
method Developed multi-layer transform learning algorithms.
result Multi-layer models outperform single-layer schemes in image denoising.
Two spectral clustering methods for multi-layer networks are analyzed and compared.
problem Community detection in multi-layer networks.
method Sum and debiased sum of squared adjacency matrices for spectral clustering.
result Debiased sum of squared adjacency matrices outperforms sum of adjacency matrices.
New deep network derived from rate reduction principles, explaining features and efficiency.
problem Understanding and optimizing deep learning architectures.
method Gradient ascent scheme for rate reduction leading to multi-layer deep network.
result Explicitly constructed multi-layer network with precise optimization and interpretation.
MGCN improves multi-layer graph classification using node attributes and relations.
problem Lack of comprehensive multi-layer graph embedding methods considering node attributes and different types of edges.
method Proposes MGCN, a method that combines GCN for multi-layer graphs, incorporating node attributes and both within and between layer relations.
result MGCN outperforms other multi-layer and single-layer methods in semi-supervised node classification tasks.
New methods estimate mixed memberships in multi-layer networks.
problem Complex community structure in multi-layer networks.
method Spectral methods using eigen-decomposition of aggregate matrices.
result Theoretical guarantees and empirical validation for mixed membership estimation.
New algorithm for signal reconstruction from multi-layered measurements.
problem Reconstructing signals from multi-layered non-linear measurements.
method Multi-Layer Approximate Message Passing (ML-AMP) algorithm and state evolution equations.
result Asymptotic free energy and minimal achievable error derived.
TWIST algorithm detects communities in multi-layer networks with tensor decomposition.
problem Community detection in multi-layer networks with multiple node-modality relationships.
method Tensor-based TWIST algorithm for global/local node and layer memberships.
result Accurate community detection with small misclassification error as network size increases.
New models explain residual and dilated dense neural networks using sparse coding.
problem Lack of theoretical understanding of residual and dilated dense neural networks.
method Proposed Res-CSC and MSD-CSC models, derived mathematical relationships, implemented ISTA.
result Mathematical understanding of residual and dilated dense neural networks.
This work tackles asymmetric community estimation in multi-layer directed networks.
problem Estimating different numbers of sender and receiver communities in multi-layer directed networks.
method Proposes a goodness-of-fit test based on the largest singular value of an aggregated normalized residual matrix.
result Develops sequential and ratio-based testing procedures to consistently determine true sender and receiver community numbers.
Proposes a balanced multi-component and multi-layer neural network for efficient function approximation.
problem Accurately and efficiently approximating complex functions with high degrees of freedom and computational cost.
method Inspired by a multi-component approach, MMNN combines single-layer networks with a multi-layer decomposition strategy.
result Significant reduction in training parameters, more efficient training process, and improved accuracy compared to FCNNs or MLPs.
New AMP algorithms improve multi-layer signal reconstruction.
problem Reconstructing signals and hidden variables from multi-layer networks with rotationally invariant weights.
method Developed multi-layer rotationally invariant generalized AMP (ML-RI-GAMP) algorithms and state evolution recursion.
result ML-RI-GAMP outperforms existing methods in terms of lower complexity and similar performance.
New model for multi-layer categorical data improves latent class analysis.
problem Traditional latent class analysis for single-layer categorical data is insufficient for multi-layer data.
method Developed a multi-layer latent class model (multi-layer LCM) and three spectral methods for estimation.
result The debiased sum of Gram matrices method performs best in estimating latent classes.
Proposes a method for multi-layered network embeddings that improves performance.
problem Mapping multi-layered network nodes into a vector space while preserving relationships.
method Jointly embed nodes in all layers via DeepWalk on a supra graph, then fine-tune embeddings for cohesive structure.
result Outperforms existing single- and multi-layered network embedding algorithms on benchmarks.
This paper optimizes multi-layer reinsurance policies to minimize risk measures.
problem Minimizing risk for insurance companies with multiple layers of reinsurance.
method Generalizes optimal stop-loss reinsurance to multi-layer policies using conditional tail expectation (CTE) risk measure.
result An optimal multi-layer reinsurance policy can be derived and estimated.
Develops a framework to assess systemic risk in the economy using bank-firm network data.
problem Measuring systemic risk in the economy using multilayer network data.
method Unified framework combining techniques to reconstruct multilayer economy structure from bank and firm balance sheets, and dynamics of shock propagation.
result Identifies systemically important firms and banks, and assesses systemic risk determinants.
Paper proposes a new ML approach using only additions and thresholding.
problem Energy efficiency and reduced complexity for IoT ML devices.
method Margin-Propagation (MP) network for inference and learning without MVMs.
result MP-based classifiers achieve comparable results to traditional ML methods with energy savings.
Proposes GrAMME for semi-supervised learning with multi-layered graphs.
problem Semi-supervised learning with multi-layered graphs.
method Uses attention models for feature learning in multi-layered graphs.
result Significant performance improvements over state-of-the-art network embedding strategies.
Paper presents ML-VAMP for efficient multi-layer inference with exact performance analysis.
problem Inference in multi-layer deep neural networks with non-convex optimization.
method ML-VAMP algorithm for MAP and MMSE estimates, with performance predictions in high dimensions.
result ML-VAMP achieves Bayes-optimal MSE under certain conditions, providing exact performance characterization.
New methods for community detection in multi-layer networks improve upon existing techniques.
problem Estimating a consensus community structure in multi-layer networks.
method Spectral clustering and matrix factorization methods for low-rank matrix optimization.
result Consistency properties of intermediate fusion techniques under multi-layer stochastic blockmodel.
Enhances relational reasoning with multi-layer architecture.
problem Limited relational reasoning with shallow architectures.
method Multi-layer relation network architecture.
result Solved all 20 tasks in bAbI 20 QA dataset.
New AI algorithm improves multi-layer optical film design efficiency.
problem Traditional algorithms converge to local optima, limiting global optimal solutions.
method Deep Q-learning for global optimal multi-layer optical film design.
result Deep Q-learning model converges global optimum of optical thin film structure.
In recent years there has been an increased interest in statistical analysis of data with multiple types of relations among a set of entities. Such multi-relational data can be represented as multi-layer graphs where the set of vertices represents the entities and multiple types of edges represent the different relatio…
Algorithm predicts performance of learning in multi-layer networks with matrix-valued hidden variables.
problem Signal recovery and learning in multi-layer neural networks with matrix-valued hidden variables.
method Unified approximation algorithm for MAP and MMSE inference, extending ML-VAMP to handle matrix-valued unknowns.
result Performance of ML-Mat-VAMP algorithm can be predicted in a random large-system limit.
New spectral clustering method for multi-layer networks improves accuracy.
problem Detecting community structure in multi-layer networks.
method Integrative spectral clustering based on adaptive layer aggregation.
result Our methods minimize mis-clustering error and outperform existing methods.
Paper trains multi-layer SNNs using NormAD for spatio-temporal error backpropagation.
problem Training multi-layer SNNs with non-linear integrate-and-fire dynamics.
method Formulates training as optimization, uses NormAD for iterative synaptic weight update.
result Validated on 2- and 3-layer SNNs solving spike-based XOR and generic problems.
New multi-layer algorithm improves CNN performance.
problem Efficiently modeling and processing information with parsimonious representations.
method Generalized Basis Pursuit to multi-layer setting, proposing ML-ISTA and ML-FISTA algorithms.
result Nested first order algorithms converge to solve the multi-layer problem.
Combines multi-layer graphs to infer global network structure.
problem Leveraging domain knowledge in structure inference for multi-layer graphs.
method Mask combination of multi-layer graphs using optimization.
result Enhanced structure inference through multi-layer graph integration.
Paper optimizes clustering for multi-layer networks and discrete mixtures.
problem Optimizing clustering in multi-layer networks and discrete mixtures.
method Two-stage method: tensor-based initialization and likelihood-based refinement.
result Achieves minimax optimal error rate for multi-layer networks and discrete mixtures.
FMMNN combines sine activations with multi-component, multi-layer structure for high-frequency function approximation.
problem Effective representation and learning of high-frequency features in neural networks.
method Introduces FMMNN with sine-type activations and multi-component, multi-layer structure.
result FMMNN achieves strong accuracy and favorable convergence on oscillatory function-approximation benchmarks.
SVGP KAN integrates uncertainty quantification into Kolmogorov-Arnold networks.
problem Uncertainty quantification in scientific machine learning models.
method Sparse variational Gaussian process inference with Kolmogorov-Arnold topology.
result Demonstrated ability to distinguish aleatoric and epistemic uncertainty in various scientific applications.
Statistical physics explains deep learning's feature learning capacity.
problem Understanding neural networks' ability to learn complex features.
method Study of a multi-layer perceptron in the interpolation regime.
result Optimal learning requires specialization across layers and neurons.
This work defines a new function space for multi-layer neural networks.
problem Characterizing the function space of multi-layer neural networks.
method Defining a neural Hilbert ladder (NHL) as an infinite union of reproducing kernel Hilbert spaces (RKHSs).
result Established theoretical properties of the new function space, including generalization guarantees and dynamics of random fields.
In this paper, we present a statistical-mechanical analysis of deep learning. We elucidate some of the essential components of deep learning---pre-training by unsupervised learning and fine tuning by supervised learning. We formulate the extraction of features from the training data as a margin criterion in a high-dime…
ReduNet optimizes data compression by maximizing rate reduction in deep networks.
problem Optimizing deep networks for high-dimensional multi-class data.
method Maximizing rate reduction through iterative gradient ascent, leading to a multi-layer deep network.
result ReduNet achieves optimal linear discriminative representation and is more efficient in the spectral domain.
A novel multi-layer architecture for one-class classification using graph-embedded kernel ridge regression.
problem Outlier detection in one-class classification using only normal samples.
method Stacking various Graph-Embedded Kernel Ridge Regression (KRR) based Auto-Encoders in a hierarchical fashion.
result The proposed method outperforms existing one-class classifiers on 21 benchmark datasets.