Parsimonious neural networks discover interpretable physical laws from data.
problem Discovering interpretable physical laws from data using machine learning.
method Combining neural networks with evolutionary optimization to balance accuracy and parsimony.
result Developed models for classical mechanics and materials melting temperature prediction.
Combining Bayesian nonparametrics and a forward model selection strategy, we construct parsimonious Bayesian deep networks (PBDNs) that infer capacity-regularized network architectures from the data and require neither cross-validation nor fine-tuning when training the model. One of the two essential components of a PB…
Neural networks can model chaos efficiently by becoming geometrically chaotic.
problem Lack of theoretical understanding of how neural networks learn chaos.
method Employed a geometric perspective to show neural networks can model chaotic dynamics.
result Neural networks can reconstruct strange attractors and accurately predict local divergence rates.
A neural network approach solves dynamic portfolio optimization without dynamic programming.
problem Dynamic portfolio optimization with multiple constraints and high rebalancing frequency.
method Parsimonious neural network without dynamic programming, avoiding high-dimensional expectations.
result Proves convergence to theoretical optimal solution under general conditions.
Machine Learning algorithms are increasingly being used in recent years due to their flexibility in model fitting and increased predictive performance. However, the complexity of the models makes them hard for the data analyst to interpret the results and explain them without additional tools. This has led to much rese…
New neural network models for complex functional data analysis.
problem Complex relations between functional predictors and responses.
method Function-on-Function regression models using neural networks with continuous hidden layers.
result Demonstrated power and flexibility in handling complex functional models.
Proposes a new framework for predicting stock market movements using sparse neural architectures.
problem Challenging problem of predicting stock market movements using technical indicators.
method Multi-criteria optimization approach to evolve sparse neural architectures.
result Evolved parsimonious networks with better generalization capabilities.
Solving for adversarial examples with projected gradient descent has been demonstrated to be highly effective in fooling the neural network based classifiers. However, in the black-box setting, the attacker is limited only to the query access to the network and solving for a successful adversarial example becomes much …
GIT-Net uses neural networks to approximate PDE operators efficiently.
problem Approximating PDE operators for complex geometries.
method Parametrizes adaptive generalized integral transforms with deep neural networks.
result GIT-Net outperforms existing neural network operators in multiple areas.
New neural networks model complex phenomena with fewer parameters.
problem Challenges in studying higher-order interactions in neural networks.
method Introducing curved neural networks using the maximum entropy principle.
result Curved neural networks accelerate memory retrieval and exhibit explosive phase transitions.
Novel method combines neural network features with survival models for ICU infections.
problem Improving predictive models of ICU infections while maintaining interpretability.
method Semi-parametric approach combining low-resolution and high-resolution data.
result Improved predictive power with interpretability maintained.
Extracts coarse-grained PDEs from microscopic simulations.
problem Discovering effective PDEs for macro-scale processes from microscopic data.
method Combining neural networks with equation-free numerics and data-driven approaches.
result Efficiently discovers macro-scale PDEs from microscopic simulations.
Method selects most useful network model for various tasks.
problem Impact of translating raw data to network models is unexamined.
method Proposes a network model selection methodology focusing on utility and parsimony.
result Demonstrates the importance of network definition for system behavior.
Overparameterisation can limit the benefits of curriculum learning in neural networks.
problem The ineffectiveness of curriculum learning in deep learning applications.
method Analytical study connecting curriculum learning and overparameterisation in an online learning setting for a 2-layer network in the XOR-like Gaussian Mixture problem.
result High degree of overparameterisation can limit the benefit from curricula.
MDL principle aids in learning neural network-based causal structures.
problem Learning causal relationships from observations with neural networks.
method Prequential minimum description length (MDL) principle.
result Competitive results on synthetic and real-world data, often recovering correct structure.
Adaptive neural networks learn functional data bases for improved performance.
problem Applying deep learning to functional data is challenging due to high dimensionality.
method Proposes adaptive neural networks with Basis Layers that learn relevant basis functions.
result Empirically outperforms other neural network approaches across various tasks.
A new model captures financial asset returns' tail behaviors and outperforms GARCH family.
problem Capturing the dynamic tail behaviors of financial asset returns.
method Combines LSTM with a novel parametric quantile function.
result Out-of-sample forecasts of conditional quantiles or VaR outperform GARCH family.
New fusion blocks improve equivariant neural networks for molecular dynamics.
problem Designing equivariant neural networks for tasks with global symmetries.
method Using fusion diagrams from tensor networks to design novel equivariant components.
result Improved performance with fewer parameters on chemical problems.
Proposes a method for neural networks to learn causal relationships and humans to contest and modify them.
problem Neural networks learn relevant causal relationships unclearly and are black-box, making them hard to debug.
method Two-way interaction between neural networks and humans, allowing contestation and modification of causal graphs.
result Improves predictive performance up to 2.4x and produces smaller networks up to 7x compared to SOTA.
The parsimonious Gaussian mixture models, which exploit an eigenvalue decomposition of the group covariance matrices of the Gaussian mixture, have shown their success in particular in cluster analysis. Their estimation is in general performed by maximum likelihood estimation and has also been considered from a parametr…
Bayesian method reconstructs hidden higher-order interactions from network data.
problem Lack of explicit higher-order interactions in pairwise network data.
method Bayesian approach based on parsimony, infers higher-order structures when statistically supported.
result Demonstrated applicability to various datasets, synthetic and empirical.
This paper introduces hierarchical Gaussian process priors for neural networks to capture weight correlations and inductive biases.
problem Capturing weight correlations and inductive biases in neural networks.
method Hierarchical Gaussian process priors with unit embeddings and input-dependent kernels.
result Hierarchical Gaussian process priors provide competitive predictive performance and desirable uncertainty estimates.
Path regularization reveals convex optimization in deep ReLU networks.
problem Understanding the optimization landscape of deep neural networks.
method Introducing path regularization to make the training problem convex and sparsity-inducing.
result Path regularized parallel ReLU networks are a parsimonious convex model in high dimensions.
Investigates deep hedging under rough volatility models.
problem Performance of deep hedging framework under non-Markovian conditions.
method Analysis of rough volatility models, use of parsimonious network architectures.
result Parsimonious network architectures can capture non-Markovian time-series.
Deep neural network generates symbolic equations from data.
problem Lack of insight into underlying mappings from traditional deep learning.
method Combines deep learning flexibility with symbolic solutions.
result Accurately generates governing equations for dynamical systems.
GAMI-Net improves neural network interpretability while maintaining accuracy.
problem Lack of interpretability in neural network models.
method GAMI-Net is a disentangled feedforward network with multiple additive subnetworks designed for capturing main effects and pairwise interactions, considering sparsity, heredity, and marginal clarity.
result GAMI-Net achieves superior interpretability and competitive prediction accuracy compared to explainable boosting machine and other models.
This work reduces DIM computation costs by training neural networks on single MC paths.
problem Training neural networks for Dynamic Initial Margin (DIM) computation in counterparty credit risk.
method Constructing a training dataset with noisy but unbiased DIM samples from single MC paths, employing a multi-output neural network structure.
result The approach reduces dataset generation cost to a single MC execution and validates its general applicability and efficiency.
Proposes a differentiable LSE-ICNN for modeling multi-well potentials.
problem Modeling multi-well potentials in various scientific domains.
method Log-sum-exponential (LSE) mixture of input convex neural network (ICNN) modes.
result Smooth surrogate that retains convexity within basins and allows gradient-based learning.
Improved SINDy autoencoder for identifying noisy dynamical systems.
problem Robust identification of noisy dynamical systems from data.
method Incorporates noise-separating neural network structures into SINDy autoencoder architecture.
result Accurately recovers latent dynamics and estimates measurement noise from noisy observations.
Kolmogorov-Arnold Networks offer improved interpretability and parsimony in science tasks.
problem Improving interpretability and parsimony in science-oriented tasks.
method Theoretical analysis of Kolmogorov-Arnold Networks (KAN) with generalization bounds and model complexity.
result Generalization bounds for KAN with various activation functions, scaling with the l1 norm of coefficient matrices and Lipschitz constants. New method for learning multidimensional CDFs using Archimedean copulas.
problem Learning multidimensional CDFs in high dimensions.
method Generative modeling technique using Archimedean copulas as mixture models with latent variables from neural networks.
result Efficacy and computational efficiency compared to existing methods.
Implantable, closed-loop devices for automated early detection and stimulation of epileptic seizures are promising treatment options for patients with severe epilepsy that cannot be treated with traditional means. Most approaches for early seizure detection in the literature are, however, not optimized for implementati…
Introduces LLC, a new complexity measure for DNNs based on SLT.
problem Lack of effective complexity measures for DNNs.
method Uses Singular Learning Theory to define LLC and proposes scalable estimator.
result Empirical evidence shows LLC provides valuable insights into DNN complexity.
This paper proposes a parsimoniously time varying parameter vector autoregressive model (with exogenous variables, VARX) and studies the properties of the Lasso and adaptive Lasso as estimators of this model. The parameters of the model are assumed to follow parsimonious random walks, where parsimony stems from the ass…
We consider the problem of modeling multivariate time series with parsimonious dynamical models which can be represented as sparse dynamic Bayesian networks with few latent nodes. This structure translates into a sparse plus low rank model. In this paper, we propose a Gaussian regression approach to identify such a mod…
Compact neural networks improve reinforcement learning performance with less data.
problem Sample inefficiency in deep reinforcement learning.
method Tensor factorization and wavelet scattering for compact representation.
result Achieved 2-10 times fewer total coefficients with comparable performance.
This paper presents a new causal network learning algorithm (FSNN, Feedback System Neural Network) based on the construction and analysis of a non-linear system of Ordinary Differential Equations (ODEs). The constructed system provides insight into the mechanisms responsible for generating the past and potential future…
A parsimonious model reduces over-parameterization in skewed matrix variate mixtures.
problem Over-parameterization in skewed matrix variate mixtures.
method Parsimonious family of 256 models using bilinear factor analyzers constrained over clusters, with AECM algorithm for estimation.
result Extensive simulations and real-world datasets (MNIST, Olivetti faces) demonstrate the method's effectiveness.
Enhances KANs for accuracy and interpretability with multi-exit architecture.
problem Unclear optimal depth for KANs and difficulty in optimization and interpretation.
method Introduces multi-exit KANs with each layer having its own prediction branch.
result Multi-exit KANs outperform single-exit versions on various datasets.
GD-VAEs learn dynamics from observations using geometric and topological information.
problem Learning parsimonious representations of nonlinear dynamics from observations.
method Develops data-driven methods incorporating geometric and topological information using Variational Autoencoders (VAEs).
result GD-VAEs provide methods for learning reduced dimensional representations of nonlinear dynamics.
A new method reduces computational cost for gene expression inference in large microarray data sets.
problem Efficiently predicting gene expression in large datasets with limited resources.
method Adaptive Lipschitz constant inspired learning rate, random sub-sampling, and A-ReLU activation function.
result Remarkable improvement in saving computational cost while maintaining prediction accuracy.
New method measures generalizability of deep neural networks based on decision boundary complexity.
problem Lack of generalization methods for deep neural networks.
method Created Decision Boundary Complexity (DBC) score to measure DNN complexity.
result Simpler decision boundaries lead to better generalizability, supporting Occam's Razor.
Unified framework for intermittent demand forecasting using renewal processes.
problem Intermittency in demand forecasting.
method Unified framework based on extensions of discrete-time renewal processes.
result Efficacy demonstrated in forecasting practice with favorable predictive accuracy.
Superior performance and ease of implementation have fostered the adoption of Convolutional Neural Networks (CNNs) for a wide array of inference and reconstruction tasks. CNNs implement three basic blocks: convolution, pooling and pointwise nonlinearity. Since the two first operations are well-defined only on regular-s…
New model combines neural networks and embeddings for better choice modeling interpretability.
problem Limited behavioral insights in embedding representations for categorical variables.
method Combines discrete choice models and neural networks using embeddings for interpretability.
result Proposed models deliver state-of-the-art predictive performance while preserving interpretability.
Tensor network decomposition, originated from quantum physics to model entangled many-particle quantum systems, turns out to be a promising mathematical technique to efficiently represent and process big data in parsimonious manner. In this study, we show that tensor networks can systematically partition structured dat…
A family of parsimonious shifted asymmetric Laplace mixture models is introduced. We extend the mixture of factor analyzers model to the shifted asymmetric Laplace distribution. Imposing constraints on the constitute parts of the resulting decomposed component scale matrices leads to a family of parsimonious models. An…
Bayesian optimisation method targets graph classification models against adversarial attacks.
problem Adversarial attacks on graph classification models, especially for graph-level tasks.
method Bayesian optimisation-based attack method for graph classification models.
result Effectiveness and flexibility of the proposed method validated on various graph classification tasks.