New research finds deep networks often settle in asymmetric valleys, improving generalization.
problem Understanding the nature of local minima in deep neural networks.
method Formal analysis and empirical investigation of SGD trajectories and BN effects.
result SGD solutions biased towards flat sides of asymmetric valleys generalize better.
Looped transformers outperform standard transformers in complex reasoning tasks due to a specific loss landscape geometry.
problem Understanding why looped transformers outperform standard transformers in complex reasoning tasks.
method Explained through loss landscape geometry, distinguishing between U-shaped and V-shaped valleys, and proposing SHIFT training strategy.
result Looped transformers' recursive architecture induces a River-V-Valley landscape, leading to better loss convergence and complex pattern learning.
Deep networks avoid bad local minima with no bad valleys.
problem Finding sub-optimal local minima in deep neural networks.
method Analyzing over-parameterized networks with specific activation and loss functions.
result No bad local valleys, implying no sub-optimal strict local minima.
We present novel empirical observations regarding how stochastic gradient descent (SGD) navigates the loss landscape of over-parametrized deep neural networks (DNNs). These observations expose the qualitatively different roles of learning rate and batch-size in DNN optimization and generalization. Specifically we study…
Neural networks provide a rich class of high-dimensional, non-convex optimization problems. Despite their non-convexity, gradient-descent methods often successfully optimize these models. This has motivated a recent spur in research attempting to characterize properties of their loss surface that may explain such succe…
This paper explores loss landscapes of sparse neural networks, finding unique characteristics compared to dense networks.
problem Understanding the loss landscape of sparse neural networks, especially one-hidden-layer networks.
method Analyzes sparse networks with dense and sparse final layers, focusing on linear and non-linear models.
result Sparse networks can have no spurious valleys under certain conditions, but spurious valleys and minima can exist for wide sparse networks.
LoRA-Curve connects independent LoRA optima through continuous low-loss valleys, improving Bayesian model averaging.
problem Challenges in estimating epistemic uncertainty in LoRA-based Bayesian inference.
method Introduces LoRA-Curve, a segmented Bézier curve parameterization in the LoRA space, with free and anchored configurations.
result Empirically shows that connecting independent LoRA optima through continuous low-loss valleys improves mutual information of the predictive distribution.
A new pruning method reduces neural network computation without retraining.
problem Efficiently reduce neural network computation while maintaining accuracy.
method Structured directional pruning via perturbation orthogonal projection.
result Achieves state-of-the-art pruned accuracy without retraining.
A new Kolmogorov-Arnold network improves function approximation and optimization.
problem Approximating potentially irregular functions in high dimensions.
method Proposes a new Kolmogorov-Arnold network (KAN) and provides error bounds and universal approximation theorems.
result Outperforms multilayer perceptrons in accuracy and convergence speed for irregular functions.
Quantum annealer improves classification of handwritten digits.
problem Classical MCMC struggles with missing labels in RBM models.
method Embedding RBM into D-Wave quantum annealer for classification.
result D-Wave QA reduces classification error by more than two times.
Deep learning networks have connected sublevel sets, avoiding local minima.
problem Finding local minima in deep learning networks.
method Analyzing sublevel sets of loss functions in over-parameterized neural nets.
result Sublevel sets are connected and unbounded, ensuring all global minima are accessible.
Reservoir computing minimizes prediction error in a spectral radius interval.
problem Lack of guiding principles for neural network parameters.
method Model-free prediction of spatiotemporal dynamical systems using recurrent neural networks.
result A spectral radius interval minimizes prediction error for nonlinear dynamical systems.
This paper examines SVB's failure and its impact on bank stocks.
problem SVB failure and its contagion effects on bank stocks.
method Analyzed bank-specific vulnerabilities and stock performance.
result Uninsured deposits and unrealized losses were key factors in SVB's impact.
Proposes a method to partition univariate data into unimodal subsets.
problem Partitioning univariate multimodal data into unimodal subsets.
method Recursive splitting around valley points of the data density using properties of critical points on the convex hull of the ecdf plot.
result Obtains a hierarchical statistical model of the initial dataset as a mixture of UMMs.
New framework reveals thermodynamic principles for LLM training.
problem Understanding the training dynamics of large language models.
method Introducing Neural Thermodynamic Laws (NTL) under river-valley loss landscape assumptions.
result Key thermodynamic quantities and principles naturally emerge in LLM training.
WSD schedule improves model training efficiency by adapting learning rates dynamically.
problem Fixed compute budgets limit training efficiency of language models.
method Introduces a WSD schedule that uses a constant learning rate followed by a rapid decay phase.
result WSD schedule generates a non-traditional loss curve with stable and decay phases.
Deep networks exhibit permutation saddles and valleys between equivalent minima.
problem Understanding the structure of loss landscapes in deep neural networks.
method Geometric approach to constructing paths between equivalent minima and saddle points.
result Existence of permutation saddles and valleys in deep neural networks.
HyPV-LEAD detects cryptocurrency anomalies proactively, improving financial security.
problem Cryptocurrency anomalies like mixing, fraud, and pump-and-dump operations are hard to detect due to class imbalance and temporal volatility.
method HyPV-LEAD integrates lead time into anomaly detection through window-horizon modeling, Peak-Valley sampling, and hyperbolic embedding.
result HyPV-LEAD achieves a PR-AUC of 0.9624 on Bitcoin transaction data, significantly outperforming state-of-the-art methods.
Quantization-aware training can recover accuracy lost by post-training quantization.
problem Post-training quantization (PTQ) can fail sharply at aggressive bitwidths.
method A unified geometric framework that explains PTQ failure and QAT recovery.
result QAT has a useful bias that steers iterates back into the basin.
This paper proposes a new optimization algorithm called Entropy-SGD for training deep neural networks that is motivated by the local geometry of the energy landscape. Local extrema with low generalization error have a large proportion of almost-zero eigenvalues in the Hessian with very few positive or negative eigenval…
Study of geometric analysis on asymmetric metric spaces, including heat flow and Sobolev spaces.
problem Analysis of geometric properties on asymmetric metric measure spaces.
method Introduction of upper gradients, q-Laplacian, and q-heat flow in asymmetric settings. result Extension of concepts from symmetric to asymmetric metric measure spaces.
A new asymmetric correntropy method improves robust adaptive filtering for asymmetric error distributions.
problem Inadequate handling of asymmetric error distributions in adaptive filtering.
method Proposes asymmetric correntropy using an asymmetric Gaussian kernel and develops a robust adaptive filtering algorithm.
result The proposed algorithm shows better steady-state convergence performance for asymmetric error distributions.
Solves local minima problems on smooth manifolds.
problem Local minima issues on smooth manifolds.
method Introducing valley functions and applying Morse's lemma.
result Eliminates critical points and reduces to 1D.
New metrics for Anosov representations defined from Thurston's asymmetric metrics.
problem Defining metrics for Anosov representations.
method Generalizing Thurston's asymmetric metric to Anosov representations.
result Provides a (possibly asymmetric) Finsler distance in some cases.
Unbalanced data arises in many learning tasks such as clustering of multi-class data, hierarchical divisive clustering and semisupervised learning. Graph-based approaches are popular tools for these problems. Graph construction is an important aspect of graph-based learning. We show that graph-based algorithms can fail…
Generalizes Thurston's asymmetric metric to flat metrics.
problem Defining an asymmetric metric on flat metrics.
method Defined an asymmetric metric on the space of unit-area flat metrics.
result Discussed two different topologies from the asymmetry.
This study examines asymmetric cross-correlations in cryptocurrency markets using fractal analysis.
problem Exploring asymmetric multifractal cross-correlations in cryptocurrency markets.
method Fractal analysis and MF-ADCCA method to investigate asymmetric volatility dynamics.
result Cross-correlations are stronger in downtrend markets than in uptrend markets for maturing BTC and ETH.
Theoretical justification for asymmetric actor-critic algorithms in reinforcement learning.
problem Lack of precise theoretical justification for asymmetric actor-critic algorithms in reinforcement learning.
method Adapting a finite-time convergence analysis to the asymmetric actor-critic setting with linear function approximators.
result A finite-time bound reveals that the asymmetric critic eliminates aliasing errors in the agent state.
This work presents deep asymmetric networks with a set of node-wise variant activation functions. The nodes' sensitivities are affected by activation function selections such that the nodes with smaller indices become increasingly more sensitive. As a result, features learned by the nodes are sorted by the node indices…
We consider the problem of designing locality sensitive hashes (LSH) for inner product similarity, and of the power of asymmetric hashes in this context. Shrivastava and Li argue that there is no symmetric LSH for the problem and propose an asymmetric LSH based on different mappings for query and database points. Howev…
New asymmetric kernel methods improve feature learning.
problem Improving feature learning with asymmetric kernels.
method Coupled covariance eigenproblem and Nyström method.
result Empirical evaluations show benefits of KSVD.
A new algorithm improves model-based reinforcement learning by guiding latent representations.
problem Improving model-based reinforcement learning through additional supervision.
method Proposed a novel asymmetric representation learning objective using latent guidance.
result Significantly improved performance over previous asymmetric approaches.
The article confirms two quasi-alternating surgeries for 9 asymmetric L-space knots.
problem Understanding quasi-alternating surgeries on asymmetric L-space knots.
method Using the Montesinos trick to confirm known surgeries.
result Confirmation of two quasi-alternating surgeries for each of 9 asymmetric L-space knots.
The paper improves asymmetric causality tests by addressing inefficiencies and statistical significance issues.
problem Inefficiencies and statistical significance issues in asymmetric causality tests.
method Improved asymmetric causality tests via partial cumulative sums for positive and negative components, explicitly testing differences between causal parameters.
result Efficiently tested hypotheses on asymmetric causal interaction between financial markets.
Designs chiral photonic structures using machine learning for efficient optical properties.
problem Optimizing chiral photonic nanostructures for light-matter interactions.
method Evolutionary algorithm and neural network approach for rapid optimization.
result Frequency-dependent modification in reflected light's degree of circular polarization.
Asymmetric expansion preserves convexity in hyperbolic geometry.
problem Maintaining convexity in hyperbolic geometry under asymmetric expansions.
method Generalizing earlier results on radial expansion to asymmetric expansion.
result Asymmetric expansion of hyperbolic convex sets remains convex.
Tilting loss functions improves machine learning performance.
problem Improving machine learning models, especially in under- and over-parameterized networks.
method Using evolving loss functions that emphasize different classes cyclically.
result Dynamical loss functions lead to better generalization and stability in training.
We propose Deep Asymmetric Multitask Feature Learning (Deep-AMTFL) which can learn deep representations shared across multiple tasks while effectively preventing negative transfer that may happen in the feature sharing process. Specifically, we introduce an asymmetric autoencoder term that allows reliable predictors fo…
Study of classification in asymmetric quasi-metric spaces.
problem Classification in asymmetric quasi-metric spaces.
method Sample compression and nearest neighbor algorithm.
result Algorithm has favorable statistical properties.
In this paper we show how the study of asymmetric R&D alliances, that are those between young and small firms and large and MNEs firms for knowledge exploration and/or exploitation, requires the adoption of a coopetitive framework which consider both collaboration and competition. We draw upon the literature on asymmet…
Stablecoin liquidity was affected by the SVB collapse, with USDC's transparency leading to market reactions.
problem Impact of stablecoin transparency on liquidity during market turmoil.
method Adapted MCI measure to Uniswap, Difference-in-Differences analysis on MCI and TVL, measured liquidity concentration.
result USDC's transparency led to swift market reactions, while USDT's opacity provided a safety net.
This research explains why SGD generalizes better than ADAM in deep learning.
problem Understanding the generalization gap between SGD and ADAM in deep learning.
method Analyzing local convergence behaviors through Levy-driven stochastic differential equations (SDEs).
result SGD is more locally unstable and better escapes from sharp minima to flatter ones, leading to better generalization.
Bayesian VI copula models capture asymmetric intraday equity dependence.
problem Modeling asymmetric and extreme tail dependence in financial data.
method Bayesian variational inference for skew-t copula models in high dimensions.
result The copula captures substantial heterogeneity in asymmetric dependence over equity pairs and time.
This work describes compactifications of metric spaces and vector spaces using asymmetric norms.
problem Compactifying metric spaces and vector spaces using asymmetric norms.
method Nonstandard methods, ultrapowers of the spaces at hand.
result Polyhedral compactifications of vector spaces with stratified structure.
Extends multidimensional scaling to analyze three-way asymmetric proximities.
problem Analyzing asymmetric and three-way proximities in a Euclidean space.
method Unified h-plot methodology for three-way asymmetric proximities, including symmetric and conditional frameworks.
result Identification of archetypal profiles and clustering structures.
Extends metric to Margulis spacetimes for convex properties.
problem No specific problem stated; extends metric.
method Extends Thurston's asymmetric metric to Margulis spacetimes and proves convex properties.
result Established convex properties of the extended metric.
This paper introduces constrained mixtures for continuous distributions, characterized by a mixture of distributions where each distribution has a shape similar to the base distribution and disjoint domains. This new concept is used to create generalized asymmetric versions of the Laplace and normal distributions, whic…
Enhances reinforcement learning with partial state information.
problem Improving learning under partial observability with limited privileged signals.
method Introduced informed asymmetric actor-critic framework that uses arbitrary state-dependent privileged signals.
result Unbiased policy gradient estimates with arbitrary privileged signals.