Paper analyzes and proves convergence of a new method for solving complex PDEs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Combining Bayesian deep learning and split conformal prediction affects out-of-distribution coverage.
TSSM splits neural networks for parallel training with minimal accuracy loss.
Unified predictive uncertainty disentangled using deep split ensembles.
New method escapes local optima in neural architecture optimization.
In this paper we develop a statistical theory and an implementation of deep learning models. We show that an elegant variable splitting scheme for the alternating direction method of multipliers optimises a deep learning objective. We allow for non-smooth non-convex regularisation penalties to induce sparsity in parame…
Memory split advantage: thinner networks outperform a single wide network.
This work investigates power laws in deep neural network ensembles and predicts their performance.
Split learning preserves privacy in 1D CNN models for detecting heart abnormalities.
A new method for causal inference in high-dimensional data using machine learning.
Shortage of labeled data has been holding the surge of deep learning in healthcare back, as sample sizes are often small, patient information cannot be shared openly, and multi-center collaborative studies are a burden to set up. Distributed machine learning methods promise to mitigate these problems. We argue for a sp…
In this paper we introduce a numerical method for nonlinear parabolic PDEs that combines operator splitting with deep learning. It divides the PDE approximation problem into a sequence of separate learning problems. Since the computational graph for each of the subproblems is comparatively small, the approach can handl…
The paper presents a method for generating well-calibrated prediction intervals using quality-driven deep ensembles.
A deep learning method solves nonlinear filtering problems efficiently.
This paper introduces a generic method which enables to use conventional deep neural networks as end-to-end one-class classifiers. The method is based on splitting given data from one class into two subsets. In one-class classification, only samples of one normal class are available for training. During inference, a cl…
A new numerical scheme approximates nonlinear filtering densities for noisy and partial measurements.
Deep neural networks have been proven powerful at processing perceptual data, such as images and audio. However for tabular data, tree-based models are more popular. A nice property of tree-based models is their natural interpretability. In this work, we present Deep Neural Decision Trees (DNDT) -- tree models realised…
Study numerical methods for singular FBSDEs with degenerate forward component.
Paper introduces a new sampling method for Bayesian inference.
Develops significance tests for neural networks without strong assumptions or excessive computation.
Improves DPGMM sampler by better initializing subclusters for more effective clustering.
SplitEasy trains ML models on mobile devices without server data transfer.
New approach makes survival analysis fairer without specifying sensitive features.
Designing energy-efficient networks is of critical importance for enabling state-of-the-art deep learning in mobile and edge settings where the computation and energy budgets are highly limited. Recently, Liu et al. (2019) framed the search of efficient neural architectures into a continuous splitting process: it itera…
In this work, we propose a deep neural network architecture motivated by primal-dual splitting methods from convex optimization. We show theoretically that there exists a close relation between the derived architecture and residual networks, and further investigate this connection in numerical experiments. Moreover, we…
Batch-splitting (data-parallelism) is the dominant distributed Deep Neural Network (DNN) training strategy, due to its universal applicability and its amenability to Single-Program-Multiple-Data (SPMD) programming. However, batch-splitting suffers from problems including the inability to train very large models (due to…
The training of Deep Neural Networks usually needs tremendous computing resources. Therefore many deep models are trained in large cluster instead of single machine or GPU. Though major researchs at present try to run whole model on all machines by using asynchronous asynchronous stochastic gradient descent (ASGD), we …
DeepDPM clusters images without knowing the number of clusters.
New method uses symmetric splitting for efficient HMC inference in large neural networks.
This work approximates full conformal prediction for neural networks without sample splitting.
This paper proposes a method to use deep neural networks as end-to-end open-set classifiers. It is based on intra-class data splitting. In open-set recognition, only samples from a limited number of known classes are available for training. During inference, an open-set classifier must reject samples from unknown class…
Wide neural networks' last hidden layers split into groups of redundant neurons.
The paper uses deep neural networks to estimate and infer ATE without needing to know the dimension of the data.
Adaptive Multilevel Splitting improves rare event pricing for financial derivatives.
A novel variational inference based resampling framework is proposed to evaluate the robustness and generalization capability of deep learning models with respect to distribution shift. We use Auto Encoding Variational Bayes to find a latent representation of the data, on which a Variational Gaussian Mixture Model is a…
SVHN dataset's split affects generative models but not digit classification.
The emergence of various intelligent mobile applications demands the deployment of powerful deep learning models at resource-constrained mobile devices. The device-edge co-inference framework provides a promising solution by splitting a neural network at a mobile device and an edge computing server. In order to balance…
Can health entities collaboratively train deep learning models without sharing sensitive raw data? This paper proposes several configurations of a distributed deep learning method called SplitNN to facilitate such collaborations. SplitNN does not share raw data or model details with collaborating institutions. The prop…
We survey distributed deep learning models for training or inference without accessing raw data from clients. These methods aim to protect confidential patterns in data while still allowing servers to train models. The distributed deep learning methods of federated learning, split learning and large batch stochastic gr…
SplitNN-driven Vertical Partitioning enables distributed learning from diverse data sources.
This paper proposes a novel Stochastic Split Linearized Bregman Iteration (-LBI) algorithm to efficiently train the deep network. The -LBI introduces an iterative regularization path with structural sparsity. Our -LBI combines the computational efficiency of the LBI, and model selection consistency…
Backpropagation-free trunk training improves model performance on various benchmarks.
Handel and Mosher have proved that the free splitting complex FS for the free group is Gromov hyperbolic. This is a deep and much sought-after result, since it establishes FS as a good analogue of the curve complex for surfaces. We give a shorter alternative proof of this theorem, using surgery paths in Hatcher's spher…
SPlit optimizes dataset splitting for better model performance.
Despite recent advances in architectures for mobile devices, deep learning computational requirements remains prohibitive for most embedded devices. To address that issue, we envision sharing the computational costs of inference between local devices and the cloud, taking advantage of the compression performed by the f…
This paper proposes a novel model for the rating prediction task in recommender systems which significantly outperforms previous state-of-the art models on a time-split Netflix data set. Our model is based on deep autoencoder with 6 layers and is trained end-to-end without any layer-wise pre-training. We empirically de…
A new MMD-based test combines kernels for two-sample testing without splitting data.
The paper extends keenness concept to bridge splittings and finds conditions for existence.