Improved sleep apnea detection using sensor fusion and backward shortcut connections.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Due to the success of residual networks (resnets) and related architectures, shortcut connections have quickly become standard tools for building convolutional neural networks. The explanations in the literature for the apparent effectiveness of shortcuts are varied and often contradictory. We hypothesize that shortcut…
In established network architectures, shortcut connections are often used to take the outputs of earlier layers as additional inputs to later layers. Despite the extraordinary effectiveness of shortcuts, there remain open questions on the mechanism and characteristics. For example, why are shortcuts powerful? Why do sh…
The Residual Network (ResNet), proposed in He et al. (2015), utilized shortcut connections to significantly reduce the difficulty of training, which resulted in great performance boosts in terms of both training and generalization error. It was empirically observed in He et al. (2015) that stacking more layers of resid…
We propose a neural architecture search (NAS) algorithm, Petridish, to iteratively add shortcut connections to existing network layers. The added shortcut connections effectively perform gradient boosting on the augmented layers. The proposed algorithm is motivated by the feature selection algorithm forward stage-wise …
New method trains deep vanilla networks as fast as ResNets without shortcut connections.
New method flattens decision boundary by targeting shortcut-aligned axes in disentangled latent space.
Study on shortcuts in deep networks, revealing their layer-wise distribution and impact.
Default-ERM shortcut learning persists even without additional information.
Regularization methods can overregulate, suppressing causal features.
Residual Network (ResNet) is undoubtedly a milestone in deep learning. ResNet is equipped with shortcut connections between layers, and exhibits efficient training using simple first order algorithms. Despite of the great empirical success, the reason behind is far from being well understood. In this paper, we study a …
This work identifies and mitigates reasoning shortcuts in Neuro-Symbolic models.
Very deep CNNs achieve state-of-the-art results in both computer vision and speech recognition, but are difficult to train. The most popular way to train very deep CNNs is to use shortcut connections (SC) together with batch normalization (BN). Inspired by Self- Normalizing Neural Networks, we propose the self-normaliz…
A network supporting deep unsupervised learning is presented. The network is an autoencoder with lateral shortcut connections from the encoder to decoder at each level of the hierarchy. The lateral shortcut connections allow the higher levels of the hierarchy to focus on abstract invariant features. While standard auto…
Matrix Product States (MPS), also known as Tensor Train (TT) decomposition in mathematics, has been proposed originally for describing an (especially one-dimensional) quantum system, and recently has found applications in various applications such as compressing high-dimensional data, supervised kernel linear classifie…
Transformers learn to use induction heads or shortcuts based on data diversity.
Transformers simulate finite-state automata with fewer layers.
To have a superior generalization, a deep learning neural network often involves a large size of training sample. With increase of hidden layers in order to increase learning ability, neural network has potential degradation in accuracy. Both could seriously limit applicability of deep learning in some domains particul…
Attention weights may not accurately highlight important parts due to combinatorial shortcuts.
Study reveals DNNs prefer easy-to-learn cues over essential ones in image recognition.
Materials discovery is crucial for making scientific advances in many domains. Collections of data from experiments and first-principle computations have spurred interest in applying machine learning methods to create predictive models capable of mapping from composition and crystal structures to materials properties. …
Neurosymbolic predictors fail to model uncertainty under independence assumption.
Deep vanilla transformers trained without shortcuts achieve similar performance to standard models.
One pixel modification can make deep models unlearnable.
The paper develops methods to price options under rough volatility models using BSPDEs.
Recently, we proposed to transform the outputs of each hidden neuron in a multi-layer perceptron network to have zero output and zero slope on average, and use separate shortcut connections to model the linear dependencies instead. We continue the work by firstly introducing a third transformation to normalize the scal…
This paper presents a novel approach to numerically solve stochastic differential games for nonlinear systems. The proposed approach relies on the nonlinear Feynman-Kac theorem that establishes a connection between parabolic deterministic partial differential equations and forward-backward stochastic differential equat…
In this paper, we propose an implicit gradient descent algorithm for the classic -means problem. The implicit gradient step or backward Euler is solved via stochastic fixed-point iteration, in which we randomly sample a mini-batch gradient in every iteration. It is the average of the fixed-point trajectory that is c…
Accelerates Birkhoff projection for manifold-constrained hyper-connections with high accuracy and speed.
In that paper, we provide a new characterization of the solutions of specific reflected backward stochastic differential equations (or RBSDEs) whose driver is convex and has quadratic growth in its second variable: this is done by introducing the extended notion of -Snell enveloppe. Then, in a second step, we re…
Backward stochastic partial differential equations of parabolic type in bounded domains are studied in the setting where the coercivity condition is not necessary satisfied and the equation can be degenerate. Some generalized solutions based on the representation theorem are suggested. In addition to problems with a st…
New method for dynamic valuation in markets with random endowments.
Measures financial resilience using BSDEs and their properties.
Recent theoretical results establish that time-consistent valuations (i.e. pricing operators) can be created by backward iteration of one-period valuations. In this paper we investigate the continuous-time limits of well-known actuarial premium principles when such backward iteration procedures are applied. We show tha…
We prove that when n >= 5, the Dehn function of SL(n;Z) is quadratic. The proof involves decomposing a disc in SL(n;R)/SO(n) into triangles of varying sizes. By mapping these triangles into SL(n;Z) and replacing large elementary matrices by "shortcuts," we obtain words of a particular form, and we use combinatorial tec…
Residual networks (ResNet) and weight normalization play an important role in various deep learning applications. However, parameter initialization strategies have not been studied previously for weight normalized networks and, in practice, initialization methods designed for un-normalized networks are used as a proxy.…
Paper finds a new principle for optimizing consumption and wealth using Tsallis entropy.
In this paper, we first establish the reflected backward stochastic difference equations with finite state (FS-RBSDEs for short). Then we explore the Existence and Uniqueness Theorem as well as the Comparison Theorem by "one step" method. The connections between FS-RBSDEs and optimal stopping time problems are investig…
We show that any two non-conjugate points on a forward or backward complete connected Finsler manifold can be joined by infinitely many geodesics which are not covered by finitely many closed ones, provided that the Betti numbers of the based loop space grow unbounded.
Approximate Incremental Value-at-Risk formulae provide an easy-to-use preliminary guideline for risk allocation. Both the cases of risk adding and risk pooling are examined and beta-based formulae achieved. Results highlight how much the conditions for adding new risky positions are stronger than those required for ris…
In this paper, we provide conceptional explanations for the geodesic and Jacobi field correspondences for homothetic navigation, and then let them guide us to the shortcuts to some well known flag curvature and S-curvature formulas. They also help us directly see the local correspondence between isoparametric functions…
ACI identifies cause-effect relationships and causal influence ranges in dynamical systems.
Paper presents a new backward deep BSDE method for solving nonlinear FBSDE problems.
This paper extends ResNet theory to infinitely deep networks, linking them to diffusion processes.
We show some computations related to the motion by mean curvature flow of a submanifold inside an ambient Riemannian manifold evolving by Ricci or backward Ricci flow. Special emphasis is given to the possible generalization of Huisken's monotonicity formula and its connection with the validity of some Li--Yau--Hamilto…
Machine learning reveals hidden features in knot classification.
Study connects curvature bounds to map existence and flow solutions.
Gradient-based framework for optimizing text prompts in diffusion models.