New method flattens decision boundary by targeting shortcut-aligned axes in disentangled latent space.
problem Shortcut learning in neural networks, leading to poor out-of-distribution generalization.
method Injects targeted anisotropic noise to regularize classifier sensitivity along shortcut-aligned axes.
result Achieves state-of-the-art OOD performance without shortcut labels or conflicting samples.
Study on shortcuts in deep networks, revealing their layer-wise distribution and impact.
problem Understudied impact of shortcuts on feature representations in deep networks.
method Layer-wise localization through counterfactual training on clean and skewed datasets.
result Shortcuts are distributed throughout the network, not localized in specific layers.
This paper explores a novel sparse shortcut topology in neural networks.
problem Understanding the effectiveness and characteristics of shortcut connections in neural networks.
method Investigates a novel sparse shortcut topology, demonstrating its expressivity and generalizability.
result The proposed topology enables a one-neuron-wide deep network to approximate any univariate continuous function and shows excellent generalizability.
Default-ERM shortcut learning persists even without additional information.
problem Default-ERM shortcut learning in perception tasks despite stable feature sufficiency.
method Studied linear perception task; developed margin control (MARG-CTRL) loss functions.
result Margin control mitigates shortcut learning on various tasks.
The Residual Network (ResNet), proposed in He et al. (2015), utilized shortcut connections to significantly reduce the difficulty of training, which resulted in great performance boosts in terms of both training and generalization error. It was empirically observed in He et al. (2015) that stacking more layers of resid…
Regularization methods can overregulate, suppressing causal features.
problem Mitigating shortcuts in models exploiting spurious correlations.
method Analysis of regularization methods and their effects on causal features.
result Regularization can overregulate, suppressing causal features.
This work identifies and mitigates reasoning shortcuts in Neuro-Symbolic models.
problem Neuro-Symbolic models can achieve high accuracy by using unintended concepts.
method Characterized reasoning shortcuts as unintended optima of the learning objective and identified four key conditions.
result Reasoning shortcuts are difficult to mitigate, casting doubt on NeSy solutions' trustworthiness and interpretability.
Due to the success of residual networks (resnets) and related architectures, shortcut connections have quickly become standard tools for building convolutional neural networks. The explanations in the literature for the apparent effectiveness of shortcuts are varied and often contradictory. We hypothesize that shortcut…
Petridish efficiently searches neural architectures by iteratively adding shortcut connections.
problem Finding efficient neural architectures for various tasks.
method Iteratively adds shortcut connections to existing network layers, motivated by feature selection.
result Petridish efficiently finds competitive models with few GPU days.
A new model SMPS alleviates the exponential decay of correlations in MPS.
problem Exponential decay of correlations in Matrix Product States (MPS) limits their power in capturing long-range dependences.
method Introducing long-range interactions (shortcuts) to MPS to decrease correlation length while preserving computational efficiency.
result SMPS can decrease significantly the correlation length of MPS, improving its ability to capture long-range dependences.
Shortcut connections in ResNet help avoid local optima, leading to efficient training.
problem Understanding why shortcut connections in ResNet lead to efficient training.
method Two-layer non-overlapping convolutional ResNet, gradient descent with proper normalization.
result Gradient descent avoids spurious local optima, converging to a global optimum.
Improved sleep apnea detection using sensor fusion and backward shortcut connections.
problem Untreated sleep apnea leads to severe health consequences; automated detection is needed.
method Late sensor fusion using backward shortcut connections to improve deep learning models.
result Significant improvement in predictive performance over single sensor methods.
Transformers learn to use induction heads or shortcuts based on data diversity.
problem How data diversity influences the behavior of transformers.
method Gradient-based training of a single-layer transformer on a minimal task.
result Data diversity steers transformers toward induction heads or shortcuts.
Transformers simulate finite-state automata with fewer layers.
problem How do shallow, non-recurrent Transformers simulate complex computations?
method Hierarchical reparameterization of recurrent dynamics to simulate automata.
result Polynomial-sized, O(logT)-depth solutions exist and are common. Attention weights may not accurately highlight important parts due to combinatorial shortcuts.
problem Inaccurate interpretation of attention weights in models.
method Theoretical analysis and design of experiments to show combinatorial shortcuts. Proposed two methods to mitigate this issue.
result Proposed methods improve the interpretability of attention mechanisms.
Study reveals DNNs prefer easy-to-learn cues over essential ones in image recognition.
problem DNNs learn easy-to-learn features that aren't essential to the task.
method WCST-ML training setup with shortcut cues on synthetic and face datasets.
result DNNs converge to solutions focusing on preferred cues, leading to flat minima.
New method trains deep vanilla networks as fast as ResNets without shortcut connections.
problem Training very deep neural networks is challenging.
method Developed a new type of transformation compatible with Leaky ReLUs.
result Validation accuracies with deep vanilla networks are competitive with ResNets and significantly higher.
Neurosymbolic predictors fail to model uncertainty under independence assumption.
problem Neurosymbolic predictors' reliance on independence assumption limits their ability to model uncertainty.
method Formal analysis of NeSy predictors under independence assumption.
result Assuming independence among symbolic concepts prevents NeSy predictors from representing uncertainty.
Improved speech recognition with faster training and inference.
problem Training very deep CNNs for speech recognition is difficult.
method Proposed SNDCNN using SELU activations instead of RELU and shortcut connections/BN.
result Achieved similar or lower WER with faster training and inference.
One pixel modification can make deep models unlearnable.
problem Protecting data from unauthorized training of deep neural networks.
method Perturbing only one pixel in each image to degrade model accuracy.
result Generated One-Pixel Shortcut (OPS) cannot be erased by adversarial training and strong augmentations.
New method discourages models from using bias shortcuts for better generalization.
problem Training models on biased data can lead to poor generalization when bias shifts.
method Train a de-biased representation by encouraging it to differ from biased representations.
result Improved generalization across synthetic and real-world biases.
A network supporting deep unsupervised learning is presented. The network is an autoencoder with lateral shortcut connections from the encoder to decoder at each level of the hierarchy. The lateral shortcut connections allow the higher levels of the hierarchy to focus on abstract invariant features. While standard auto…
We prove that when n >= 5, the Dehn function of SL(n;Z) is quadratic. The proof involves decomposing a disc in SL(n;R)/SO(n) into triangles of varying sizes. By mapping these triangles into SL(n;Z) and replacing large elementary matrices by "shortcuts," we obtain words of a particular form, and we use combinatorial tec…
Approximate Incremental Value-at-Risk formulae provide an easy-to-use preliminary guideline for risk allocation. Both the cases of risk adding and risk pooling are examined and beta-based formulae achieved. Results highlight how much the conditions for adding new risky positions are stronger than those required for ris…
A new deep learning model improves robustness and efficiency in predicting continuous variables.
problem Limited applicability of deep learning in domains with small sample sizes.
method Autoencoder-based residual deep network with shortcut connections.
result Achieves cutting-edge accuracy and efficiency in multiple datasets.
Machine learning reveals hidden features in knot classification.
problem Classifying the topology of closed curves.
method Investigating shortcut methods used by ML for knot classification.
result Developed a dataset and code to remove non-topological features.
Gradient-based framework for optimizing text prompts in diffusion models.
problem Efficiently optimizing prompts in text-to-image diffusion models with large domain space and non-differentiable embeddings.
method Formulated as discrete optimization over language space, designed compact subspaces, and introduced shortcut text gradient.
result Empirically discovered prompts that enhance or destroy image faithfulness.
A result of M. Ledoux is that a complete Riemannian manifold with non negative Ricci curvature satisfying the Euclidean Sobolev inequality is the Euclidean space. We present a shortcut of the proof. We also give a refinement of a result of B-L. Chen et X-P. Zhu about locally conformally flat manifolds with non negative…
Deep vanilla transformers trained without shortcuts achieve similar performance to standard models.
problem Training deep vanilla transformers without shortcuts and normalizations.
method Parameter initializations, bias matrices, and location-dependent rescaling.
result Deep vanilla transformers can train at similar speeds and performance to standard models.
We give a self-contained introduction to the theory of Turaev's shadows as a tool to study 3 and 4-manifolds. The goal of the present paper twofold: on one side it is intended to be a shortcut to a basic use of the theory of shadows, on the other side it gives a sketchy overview of some of the recent results on shadows…
Ozsvath and Szabo have defined a knot concordance invariant tau that bounds the 4-ball genus of a knot. Here we discuss shortcuts to its computation. We include examples of Alexander polynomial one knots for which the invariant is nontrivial, including all iterated untwisted positive doubles of knots with nonnegative T…
The paper explains geometric correspondences for homothetic navigation.
problem Understanding geodesic and Jacobi field correspondences in homothetic navigation.
method Providing conceptual explanations and shortcuts to formulas.
result Directly seeing local correspondence between isoparametric functions or hypersurfaces.
In this paper, we study the Hausdorff dimension of the Floyd and Bowditch boundaries of a relatively hyperbolic group, and show that for the Floyd metric and shortcut metrics respectively, they are are both equal to a constant times the growth rate of the group. In the proof, we study a special class of conical points …
New diagonal knots found with non-torus structure.
problem Identifying knots with diagonal grid diagrams.
method Analysis of knots represented by diagonal grid diagrams.
result All diagonal knots are positive, and a new non-torus example is found.
Study highlights robustness issues in healthcare diagnostic models due to distribution shifts.
problem Robustness of diagnostic models in healthcare is compromised by distribution shifts.
method Theoretical analysis and simulation studies to understand and mitigate shortcuts learned by models.
result Ignoring covariates or using invariant learning approaches leads to non-robust predictors.
SGD quickly learns a spurious XOR feature before the signal feature, revealing learning dynamics.
problem Over-reliance on spurious correlations in neural networks trained by SGD.
method Theoretical analysis of SGD on two-layer ReLU networks trained on XOR data.
result SGD learns the spurious feature first and exponentially fast, dominating the signal feature.
Martingale Doppelgänger-Eval benchmarks VLMs on candlestick evidence vs. trend extrapolation
problem Auditing whether VLMs use chart evidence or trend extrapolation
method Proving formal limitations and designing controlled mechanisms
result Identifying regression coefficients for evidence vs. trend
IRNet improves material property prediction from composition and crystal structure.
problem Predicting material properties from composition and crystal structure.
method Deep residual regression network with individual residual learning.
result IRNet outperforms state-of-the-art machine learning approaches in predicting material properties.
This study improves fast non-Bayesian Poisson factorization for implicit-feedback recommendation systems.
problem Improving recommendation quality and speed for implicit-feedback data.
method Regularized Poisson models, frequentist optimization, sparse solutions.
result Frequentist approach yields better top-N recommendations with shorter fitting times.
The paper analyzes how generated data improves adversarial training in high-dimensional regression.
problem Improving adversarial training in high-dimensional regression.
method Theoretical analysis of a two-stage training approach with generated data and pseudo-labels.
result Two-stage adversarial training achieves better performance than ridgeless training in high-dimensional linear regression.
Sarkar and Wang proved that the hat version of Heegaard Floer homology group of a closed oriented 3-manifold is combinatorial starting from an arbitrary nice Heegaard diagram and in fact every closed oriented 3-manifold admits such a Heegaard diagram. Plamenevskaya showed that the contact Ozsvath-Szabo invariant is com…
Graph neural controlled differential equations learn graph dynamics from vertex observations.
problem Predicting future states of dynamical systems on graphs with limited vertex data.
method Incorporates graph topology information into NCDE to predict graph dynamics.
result Informed NCDE requires fewer parameters and lower MAE compared to previous methods.
Recently, we proposed to transform the outputs of each hidden neuron in a multi-layer perceptron network to have zero output and zero slope on average, and use separate shortcut connections to model the linear dependencies instead. We continue the work by firstly introducing a third transformation to normalize the scal…
Generative classifiers show surprising human-like performance.
problem Comparing generative and discriminative models for object recognition.
method Built on recent advances in generative modeling to create classifiers and compared them to discriminative models.
result Generative classifiers outperform discriminative models in several key areas, including shape bias and out-of-distribution accuracy.
New method reduces memory usage in deep HRNNs by replacing gradient backpropagation with local losses.
problem Memory constraints in training deep hierarchical RNNs.
method Replace gradient backpropagation with locally computable losses in deep HRNNs.
result Memory requirements reduced by a factor exponential in hierarchy depth.
This paper is a new step in the project of systematic description of colored knot polynomials started in arXiv:1506.00339. In this paper, we managed to explicitly find the inclusive Racah matrix, i.e. the whole set of mixing matrices in channels R^3->Q with all possible Q, for R=[3,1]. The calculation is made possible …
Paper explores anchoring for vision models, improving generalization and safety.
problem Anchoring can lead to undesirable shortcuts, limiting generalization.
method Introduced a new anchored training protocol with a regularizer to mitigate undesirable shortcuts.
result Significant performance gains in generalization and safety metrics.
We introduce a new parameterization method for deep learning layers using spectral tensor train decomposition.
problem Efficiency and stability in deep learning models with weight matrix compression.
method Spectral Tensor Train Parameterization (STTP) of weight matrices.
result Improved compression and training stability in neural networks.