The paper calculates minimum Dehn colors for knots using symmetric local biquandle cocycles.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Langevin dynamics (LD) has been proven to be a powerful technique for optimizing a non-convex objective as an efficient algorithm to find local minima while eventually visiting a global minimum on longer time-scales. LD is based on the first-order Langevin diffusion which is reversible in time. We study two variants th…
Hill-ADAM optimizes loss landscapes by exploring state space deterministically.
Proposes a method to solve deep neural networks' local minimum problem.
In this paper, we theoretically prove that adding one special neuron per output unit eliminates all suboptimal local minima of any deep neural network, for multi-class classification, binary classification, and regression with an arbitrary loss function, under practical assumptions. At every local minimum of any deep n…
New proof shows how to identify DAGs with weakly increasing errors.
One of the main difficulties in analyzing neural networks is the non-convexity of the loss function which may have many bad local minima. In this paper, we study the landscape of neural networks for binary classification tasks. Under mild assumptions, we prove that after adding one special neuron with a skip connection…
Study shows LLC correlates with neural network compressibility.
We design a non-convex second-order optimization algorithm that is guaranteed to return an approximate local minimum in time which scales linearly in the underlying dimension and the number of training examples. The time complexity of our algorithm to find an approximate local minimum is even faster than that of gradie…
We study active learning (AL) based on Gaussian Processes (GPs) for efficiently enumerating all of the local minimum solutions of a black-box function. This problem is challenging due to the fact that local solutions are characterized by their zero gradient and positive-definite Hessian properties, but those derivative…
The L1 loss landscape of neural nets near local minima behaves differently, revealing exponential decay and increased vertex density.
This paper interprets critical scales in persistent homology for compact metric spaces.
In "Width complexes for knots and 3-manifolds," Jennifer Schultens defines the width complex for a knot in order to understand the different positions a knot can occupy in the 3-sphere and the isotopies between these positions. She poses several questions about these width complexes; in particular, she asks whether the…
Lower bound on minimum vertex degree for non-negative Lin-Lu-Yau curvature on graphs.
We study the problem of globally recovering a dictionary from a set of signals via -minimization. We assume that the signals are generated as i.i.d. random linear combinations of the atoms from a complete reference dictionary , where the linear combination coefficients are from…
For nonconvex optimization in machine learning, this article proves that every local minimum achieves the globally optimal value of the perturbable gradient basis model at any differentiable point. As a result, nonconvex machine learning is theoretically as supported as convex machine learning with a handcrafted basis …
A central challenge to many fields of science and engineering involves minimizing non-convex error functions over continuous, high dimensional spaces. Gradient descent or quasi-Newton methods are almost ubiquitously used to perform such minimizations, and it is often thought that a main source of difficulty for these l…
Due to the success of deep learning to solving a variety of challenging machine learning tasks, there is a rising interest in understanding loss functions for training neural networks from a theoretical aspect. Particularly, the properties of critical points and the landscape around them are of importance to determine …
We discuss the possibility of obtaining model-free bounds on volatility derivatives, given present market data in the form of a calibrated local volatility model. A counter-example to a wide-spread conjecture is given.
In this article, we study the small sphere limit of the Wang-Yau quasi-local energy defined in [18,19]. Given a point in a spacetime , we consider a canonical family of surfaces approaching along its future null cone and evaluate the limit of the Wang-Yau quasi-local energy. The evaluation relies on solving …
Unified proof of knot unknotting bounds using Ma-Qiu index.
We study the misclassification error for community detection in general heterogeneous stochastic block models (SBM) with noisy or partial label information. We establish a connection between the misclassification rate and the notion of minimum energy on the local neighborhood of the SBM. We develop an optimally weighte…
We target the problem of finding a local minimum in non-convex finite-sum minimization. Towards this goal, we first prove that the trust region method with inexact gradient and Hessian estimation can achieve a convergence rate of order as long as those differential estimations are sufficientl…
In this paper, we propose a new adaptive stochastic gradient Langevin dynamics (ASGLD) algorithmic framework and its two specialized versions, namely adaptive stochastic gradient (ASG) and adaptive gradient Langevin dynamics(AGLD), for non-convex optimization problems. All proposed algorithms can escape from saddle poi…
The goal of this paper is to measure the non-convexity of compact and smooth connected components of real algebraic plane curves. We study these curves first in a general setting and then in an asymptotic one. In particular, we consider sufficiently small levels of a real bivariate polynomial in a small enough neighbou…
Adaptor 'E' extends gradient-based optimizers to explore loss landscapes, improving generalization.
We propose a new topic modeling procedure that takes advantage of the fact that the Latent Dirichlet Allocation (LDA) log likelihood function is asymptotically equivalent to the logarithm of the volume of the topic simplex. This allows topic modeling to be reformulated as finding the probability simplex that minimizes …
MSGD outperforms SGD in overparametrized settings with faster convergence rates.
ECD algorithm speeds up non-convex optimization, offering quantum and stochastic enhancements.
We propose a sample efficient stochastic variance-reduced cubic regularization (Lite-SVRC) algorithm for finding the local minimum efficiently in nonconvex optimization. The proposed algorithm achieves a lower sample complexity of Hessian matrix computation than existing cubic regularization based methods. At the heart…
The study finds the minimum average area ratio on hyperbolic manifolds and its relation to scalar curvature.
Stochastic subgradient descent avoids critical points in definable functions.
A well-known Lemma in Riemannian geometry by Klingenberg says that if is a minimum point of the distance function to in the cut locus of , then either there is a minimal geodesic from to along which they are conjugate, or there is a geodesic loop at that smoothly goes throu…
Study on the noise in SGD minibatches near local minima.
We provide a formulation for Local Support Vector Machines (LSVMs) that generalizes previous formulations, and brings out the explicit connections to local polynomial learning used in nonparametric estimation literature. We investigate the simplest type of LSVMs called Local Linear Support Vector Machines (LLSVMs). For…
We use noncommutative localization to construct a chain complex which counts the critical points of a circle-valued Morse function on a manifold, generalizing the Novikov complex. As a consequence we obtain new topological lower bounds on the minimum number of critical points of a circle-valued Morse function within a …
A motif-based framework identifies local spillover structures in financial markets.
We consider the problem of learning a one-hidden-layer neural network with non-overlapping convolutional layer and ReLU activation, i.e., , in which both the convolutional weights and the output weights are paramete…
We study the problem of identifying the causal relationship between two discrete random variables from observational data. We recently proposed a novel framework called entropic causality that works in a very general functional model but makes the assumption that the unobserved exogenous variable has small entropy in t…
We analyze stochastic gradient descent for optimizing non-convex functions. In many cases for non-convex functions the goal is to find a reasonable local minimum, and the main concern is that gradient updates are trapped in saddle points. In this paper we identify strict saddle property for non-convex problem that allo…
GOTabPFN improves tabular model performance with compact tokenization for HDLSS data.
A continuous-path semimartingale market model with wealth processes discounted by a riskless asset is considered. The numeraire portfolio is the unique strictly positive wealth process that, when used as a benchmark to denominate all other wealth, makes all wealth processes local martingales. It is assumed that the num…
The ropelength of a knot is the quotient of its length by its thickness. We consider a family of energy functions for knots, depending on a power p, which approach ropelength as p increases. We describe a numerically computed trefoil knot which seems to be a local minimum for ropelength; there are nearby critical point…
We analyse the definition of quasi-local energy in GR based on a Hamiltonian analysis of the Einstein-Hilbert action initiated by Brown-York. The role of the constraint equations, in particular the Hamiltonian constraint on the timelike boundary, neglected in previous studies, is emphasized here. We argue that a consis…
We consider the problem of identifying the causal direction between two discrete random variables using observational data. Unlike previous work, we keep the most general functional model but make an assumption on the unobserved exogenous variable: Inspired by Occam's razor, we assume that the exogenous variable is sim…
New method estimates minimizer and minimum value of a regression function.
The paper analyzes how good initial guesses affect the amount of data needed for low-rank matrix recovery.
Solving statistical learning problems often involves nonconvex optimization. Despite the empirical success of nonconvex statistical optimization methods, their global dynamics, especially convergence to the desirable local minima, remain less well understood in theory. In this paper, we propose a new analytic paradigm …