Characterizes Kähler-hyperbolicity of bounded symmetric domains based on rank and genus.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Gradient flow of curve length on Sobolev metrics preserves convexity.
The -gradient flow shrinks circles with radius to a point.
Flow on curves in inversive geometry converges to loxodromics.
Proposes a differentiable STFT for more efficient optimization of hop length.
New results on the convexity of geodesic-length functions on Teichmüller space are presented. A formula for the Hessian of geodesic-length is presented. New bounds for the gradient and Hessian of geodesic-length are described. A relationship of geodesic-length functions to Weil-Petersson distance is described. Applicat…
Paper introduces a differentiable STFT for continuous window length optimization.
Truncated backpropagation through time (TBPTT) is a popular method for learning in recurrent neural networks (RNNs) that saves computation and memory at the cost of bias by truncating backpropagation after a fixed number of lags. In practice, choosing the optimal truncation length is difficult: TBPTT will not converge …
Study curves evolving by gradient flow of elastic energy, proving existence, smoothing, and convergence.
The paper studies critical points and flows of a -Hilbert functional on manifolds with circle actions.
We present Rotated Adaptive Tetra-iterated Quantizer (RATQ), a fixed-length quantizer for gradients in first order stochastic optimization. RATQ is easy to implement and involves only a Hadamard transform computation and adaptive uniform quantization with appropriately chosen dynamic ranges. For noisy gradients with al…
Recurrent neural networks and sequence to sequence models require a predetermined length for prediction output length. Our model addresses this by allowing the network to predict a variable length output in inference. A new loss function with a tailored gradient computation is developed that trades off prediction accur…
Study curves evolving on hypersurfaces with free boundaries, preserving length.
Transformers learn chain-of-thought reasoning for longer problems, proving length generalization.
Relative cup-length defined for non-Morse functions on manifolds.
In recent years, the mean field theory has been applied to the study of neural networks and has achieved a great deal of success. The theory has been applied to various neural network structures, including CNNs, RNNs, Residual networks, and Batch normalization. Inevitably, recent work has also covered the use of dropou…
Improved sampling efficiency for molecular systems using path gradients after Flow Matching.
Paper analyzes regret bounds for unconstrained online optimization.
The paper studies gradients of geodesic-length functions and systoles on Teichmüller spaces.
Paper calculates distances between strata in Teichmüller space, proving a constant separation.
The elastic flow of curves converges smoothly to a critical point.
Harmonic functions of two variables are exactly those that admit a conjugate, namely a function whose gradient has the same length and is everywhere orthogonal to the gradient of the original function. We show that there are also partial differential equations controlling the functions of three variables that admit a c…
We provide a numerically robust and fast method capable of exploiting the local geometry when solving large-scale stochastic optimisation problems. Our key innovation is an auxiliary variable construction coupled with an inverse Hessian approximation computed using a receding history of iterates and gradients. It is th…
Establishes a lower bound for Kähler hyperbolicity modulus in hyperconvex domains and bounded strongly pseudoconvex domains.
We present new computations of approximately length-minimizing polygons with fixed thickness. These curves model the centerlines of "tight" knotted tubes with minimal length and fixed circular cross-section. Our curves approximately minimize the ropelength (or quotient of length and thickness) for polygons in their kno…
The paper studies how curves evolve under area constraints and converges to a critical point.
Examines challenges and proposes new approaches in machine learning theory.
Boosted conformal procedure improves prediction intervals.
We derive bounds on the path length of gradient descent (GD) and gradient flow (GF) curves for various classes of smooth convex and nonconvex functions. Among other results, we prove that: (a) if the iterates are linearly convergent with factor , then is at most ; (b) under the Polyak-K…
Let be the Teichmüller space of marked genus , punctured Riemann surfaces with its bordification $\Tbar$ the {\em augmented Teichmüller space} of marked Riemann surfaces with nodes, \cite{Abdegn, Bersdeg}. Provided with the WP metric $\Tbar$ is a complete CAT(0) metric space, \cite{DW2, Wlcomp, Yam2…
The paper analyzes RLVR's training dynamics, proving convergence depends on aligning update direction with Gradient Gap.
We study the pull-back of the 2-parameter family of quotient elastic metrics introduced in Mio-Srivastava-Joshi on the space of arc-length parameterized loops. This point of view has the advantage of concentrating on the manifold of arc-length parameterized curves, which is a very natural manifold when the analysis of …
In this paper we use a gradient flow to deform closed planar curves to curves with least variation of geodesic curvature in the sense. Given a smooth initial curve we show that the solution to the flow exists for all time and, provided the length of the evolving curve remains bounded, smoothly converges to a mult…
Modern large scale machine learning applications require stochastic optimization algorithms to be implemented on distributed computational architectures. A key bottleneck is the communication overhead for exchanging information such as stochastic gradients among different workers. In this paper, to reduce the communica…
Maxout networks study gradients and propose initialization strategies.
Gravilon improves gradient descent for neural networks.
Transformers converge linearly to optimal models for Gaussian mixtures classification.
OMGD algorithm optimizes online convex optimization with switching costs and delayed gradients.
Uniqueness of nondegenerate blowups for planar networks shown.
We prove the analyticity of smooth critical points for O'Hara's knot energies , with and , subject to a fixed length constraint. This implies, together with the main result in \cite{BR13}, that bounded energy critical points of subject to a fixed length constraint ar…
This paper, the second of a series, deals with the function space of all smooth Kähler metrics in any given closed complex manifold in a fixed cohomology class. The previous result of the second author \cite{chen991} showed that the space is a path length space and it is geodesically convex in the sense that any tw…
Transformers learn to recall with non-orthogonal embeddings in realistic settings.
We investigate anomaly detection in an unsupervised framework and introduce Long Short Term Memory (LSTM) neural network based algorithms. In particular, given variable length data sequences, we first pass these sequences through our LSTM based structure and obtain fixed length sequences. We then find a decision functi…
Models such as Sequence-to-Sequence and Image-to-Sequence are widely used in real world applications. While the ability of these neural architectures to produce variable-length outputs makes them extremely effective for problems like Machine Translation and Image Captioning, it also leaves them vulnerable to failures o…
Study on dynamic curves with elastic energy and spontaneous curvature.
We highlight a pitfall when applying stochastic variational inference to general Bayesian networks. For global random variables approximated by an exponential family distribution, natural gradient steps, commonly starting from a unit length step size, are averaged to convergence. This useful insight into the scaling of…
The Minimum Description Length (MDL) principle states that the optimal model for a given data set is that which compresses it best. Due to practial limitations the model can be restricted to a class such as linear regression models, which we address in this study. As in other formulations such as the LASSO and forward …
A new variational method speeds up Bayesian phylogenetic inference.