Due to the success of residual networks (resnets) and related architectures, shortcut connections have quickly become standard tools for building convolutional neural networks. The explanations in the literature for the apparent effectiveness of shortcuts are varied and often contradictory. We hypothesize that shortcut…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Shortcut connections in ResNet help avoid local optima, leading to efficient training.
In established network architectures, shortcut connections are often used to take the outputs of earlier layers as additional inputs to later layers. Despite the extraordinary effectiveness of shortcuts, there remain open questions on the mechanism and characteristics. For example, why are shortcuts powerful? Why do sh…
Improved sleep apnea detection using sensor fusion and backward shortcut connections.
The Residual Network (ResNet), proposed in He et al. (2015), utilized shortcut connections to significantly reduce the difficulty of training, which resulted in great performance boosts in terms of both training and generalization error. It was empirically observed in He et al. (2015) that stacking more layers of resid…
We propose a neural architecture search (NAS) algorithm, Petridish, to iteratively add shortcut connections to existing network layers. The added shortcut connections effectively perform gradient boosting on the augmented layers. The proposed algorithm is motivated by the feature selection algorithm forward stage-wise …
New method trains deep vanilla networks as fast as ResNets without shortcut connections.
Improved speech recognition with faster training and inference.
New method flattens decision boundary by targeting shortcut-aligned axes in disentangled latent space.
Study on shortcuts in deep networks, revealing their layer-wise distribution and impact.
Default-ERM shortcut learning persists even without additional information.
Regularization methods can overregulate, suppressing causal features.
This work identifies and mitigates reasoning shortcuts in Neuro-Symbolic models.
A network supporting deep unsupervised learning is presented. The network is an autoencoder with lateral shortcut connections from the encoder to decoder at each level of the hierarchy. The lateral shortcut connections allow the higher levels of the hierarchy to focus on abstract invariant features. While standard auto…
Matrix Product States (MPS), also known as Tensor Train (TT) decomposition in mathematics, has been proposed originally for describing an (especially one-dimensional) quantum system, and recently has found applications in various applications such as compressing high-dimensional data, supervised kernel linear classifie…
Transformers learn to use induction heads or shortcuts based on data diversity.
Transformers simulate finite-state automata with fewer layers.
To have a superior generalization, a deep learning neural network often involves a large size of training sample. With increase of hidden layers in order to increase learning ability, neural network has potential degradation in accuracy. Both could seriously limit applicability of deep learning in some domains particul…
Attention weights may not accurately highlight important parts due to combinatorial shortcuts.
Study reveals DNNs prefer easy-to-learn cues over essential ones in image recognition.
Materials discovery is crucial for making scientific advances in many domains. Collections of data from experiments and first-principle computations have spurred interest in applying machine learning methods to create predictive models capable of mapping from composition and crystal structures to materials properties. …
Neurosymbolic predictors fail to model uncertainty under independence assumption.
Deep vanilla transformers trained without shortcuts achieve similar performance to standard models.
One pixel modification can make deep models unlearnable.
Recently, we proposed to transform the outputs of each hidden neuron in a multi-layer perceptron network to have zero output and zero slope on average, and use separate shortcut connections to model the linear dependencies instead. We continue the work by firstly introducing a third transformation to normalize the scal…
New method discourages models from using bias shortcuts for better generalization.
We prove that when n >= 5, the Dehn function of SL(n;Z) is quadratic. The proof involves decomposing a disc in SL(n;R)/SO(n) into triangles of varying sizes. By mapping these triangles into SL(n;Z) and replacing large elementary matrices by "shortcuts," we obtain words of a particular form, and we use combinatorial tec…
Residual networks (ResNet) and weight normalization play an important role in various deep learning applications. However, parameter initialization strategies have not been studied previously for weight normalized networks and, in practice, initialization methods designed for un-normalized networks are used as a proxy.…
Approximate Incremental Value-at-Risk formulae provide an easy-to-use preliminary guideline for risk allocation. Both the cases of risk adding and risk pooling are examined and beta-based formulae achieved. Results highlight how much the conditions for adding new risky positions are stronger than those required for ris…
Machine learning reveals hidden features in knot classification.
Gradient-based framework for optimizing text prompts in diffusion models.
A result of M. Ledoux is that a complete Riemannian manifold with non negative Ricci curvature satisfying the Euclidean Sobolev inequality is the Euclidean space. We present a shortcut of the proof. We also give a refinement of a result of B-L. Chen et X-P. Zhu about locally conformally flat manifolds with non negative…
This paper proposes a deep speech enhancement method which exploits the high potential of residual connections in a wide neural network architecture, a topology known as Wide Residual Network. This is supported on single dimensional convolutions computed alongside the time domain, which is a powerful approach to proces…
We give a self-contained introduction to the theory of Turaev's shadows as a tool to study 3 and 4-manifolds. The goal of the present paper twofold: on one side it is intended to be a shortcut to a basic use of the theory of shadows, on the other side it gives a sketchy overview of some of the recent results on shadows…
Ozsvath and Szabo have defined a knot concordance invariant tau that bounds the 4-ball genus of a knot. Here we discuss shortcuts to its computation. We include examples of Alexander polynomial one knots for which the invariant is nontrivial, including all iterated untwisted positive doubles of knots with nonnegative T…
Recent empirical results on long-term dependency tasks have shown that neural networks augmented with an external memory can learn the long-term dependency tasks more easily and achieve better generalization than vanilla recurrent neural networks (RNN). We suggest that memory augmented neural networks can reduce the ef…
The paper explains geometric correspondences for homothetic navigation.
In this paper, we study the Hausdorff dimension of the Floyd and Bowditch boundaries of a relatively hyperbolic group, and show that for the Floyd metric and shortcut metrics respectively, they are are both equal to a constant times the growth rate of the group. In the proof, we study a special class of conical points …
Hybrid model combines LSTM and ETS for mid-term electric load forecasting.
New diagonal knots found with non-torus structure.
This work explores non-negative low-rank matrix factorization based on regularized Poisson models (PF or "Poisson factorization" for short) for recommender systems with implicit-feedback data. The properties of Poisson likelihood allow a shortcut for very fast computations over zero-valued inputs, and oftentimes result…
Study highlights robustness issues in healthcare diagnostic models due to distribution shifts.
New models explain residual and dilated dense neural networks using sparse coding.
SGD quickly learns a spurious XOR feature before the signal feature, revealing learning dynamics.
Martingale Doppelgänger-Eval benchmarks VLMs on candlestick evidence vs. trend extrapolation
We define and study a discrete process that generalizes the convex-layer decomposition of a planar point set. Our process, which we call "homotopic curve shortening" (HCS), starts with a closed curve (which might self-intersect) in the presence of a set of point obstacles, and evolves in discrete…
SurvNAM explains survival model predictions using machine learning.
New geometric regularizers improve deep learning generalization.