Gradient flow in parameters equals linear interpolation in outputs.
problem Understanding and optimizing training algorithms in deep learning.
method Proving equivalence between gradient flow in parameter space and linear interpolation in output space, and deriving formulas for global minima.
result Gradient flow in parameters can be transformed into linear interpolation in outputs, leading to global minima.
MPF method improves parameter estimation in probabilistic models.
problem Difficulty in fitting probabilistic models due to intractable partition function.
method Minimum Probability Flow (MPF) method for parameter estimation.
result MPF outperforms existing techniques in convergence time and accuracy.
Efficient adjustment sets found for cost-minimized causal estimations.
problem Estimating interventional means with minimum cost in causal graphical models.
method Defined cost-adjustment sets, constructed flow networks, and used maximum flow algorithms.
result Minimum cost optimal adjustment sets exist and can be found efficiently.
We study the misclassification error for community detection in general heterogeneous stochastic block models (SBM) with noisy or partial label information. We establish a connection between the misclassification rate and the notion of minimum energy on the local neighborhood of the SBM. We develop an optimally weighte…
This study explains gradient flow dynamics in neural networks for small initialisation.
problem Understanding the training dynamics of neural networks for small initialisation.
method Analysis of gradient flow dynamics for one-hidden layer ReLU networks with orthogonal inputs.
result Gradient flow converges to zero loss and characterizes implicit bias towards minimum variation norm.
Using transfer entropy, we observed the strength and direction of information flow between stock indices. We uncovered that the biggest source of information flow is America. In contrast, the Asia/Pacific region the biggest is receives the most information. According to the minimum spanning tree, the GSPC is located at…
New insights into matrix factorization show strict saddles have bounded eigenvalues.
problem Understanding the nature of critical points in matrix factorization.
method Analyzing orbits of critical points under the general linear group and identifying canonical points.
result Minimum eigenvalue of strict saddles is not uniformly bounded below zero.
Study rigidity of Hamiltonians near a minimum in symplectic and magnetic settings.
problem Rigidity of Hamiltonians near a minimum in symplectic and magnetic settings.
method Analyzing Hamiltonian systems near a compact symplectic Morse-Bott minimum, focusing on Zoll flows and magnetic forms.
result A constant curvature quantity characterizes complex space forms among Kähler manifolds.
In this paper, we mainly study the mean curvature flow in Kähler surfaces with positive holomorphic sectional curvatures. We prove that if the ratio of the maximum and the minimum of the holomorphic sectional curvatures is less than 2, then there exists a positive constant δ depending on the ratio such that $\cosα\ge…
The study examines how shallow neural nets converge to training samples or manifold points during diffusion.
problem Understanding when and how shallow neural nets converge to training samples or manifold points during diffusion.
method Analysis of shallow ReLU neural network denoisers trained with minimal ℓ2 norm, comparing score flow and diffusion flow. result Probability flow converges to training points, sums of training points, or manifold points, depending on the diffusion time scheduler.
Deep linear networks can closely approximate interpolants without improving risk.
problem Understanding the risk bounds of deep linear networks compared to minimum ℓ2-norm solutions. method Bounding excess risk of interpolating deep linear networks trained using gradient flow.
result Deep linear networks can closely approximate or match minimum ℓ2-norm solutions in terms of risk. Gradient flow converges to a minimal convex structure.
problem Finding the minimal convex structure in hyperbolic manifolds.
method Weil-Petersson gradient vector field of renormalized volume.
result The flow converges to the structure with minimum convex core volume.
Gradient descent with geometrically adapted metrics drives L2 cost to global minimum at uniform rate.
problem Minimizing L2 cost in deep learning networks. method Adapting gradient descent to output layer metric in deep learning.
result Uniform exponential convergence to global minimum in L2 cost. The article analyzes the stability of a curve shortening flow for planar networks.
problem Stability analysis of anisotropic curve shortening flow for planar networks.
method Used Lojasiewicz-Simon gradient inequality to derive stability results.
result For initial data close to an energy minimizer, the flow exists globally and converges to a different energy minimum.
New method improves MMD estimation without convexity assumptions.
problem Lack of theoretical guarantees for MMD estimation algorithms.
method Preconditioned gradient descent (PGD) scheme for MMD optimization.
result PGD scheme converges globally under specific conditions.
In this paper, we define a certain "proportional volume property" for an unit vector field on a spherical domain in S3. We prove that the volume of these vector fields has an absolute minimum and this value is equal to the volume of the Hopf vector field. Some examples of such vector fields are given. We also study the…
In this paper, we use the distance comparison principle, first been developed by G. Huisken, to study the spatial curve shortening flow. We have got the result that if the initial curve is the helix, then the local minimum of the ratio of the extrinsic and intrinsic distance is non-decreasing. And we have proved a Gray…
We study the risk of minimum-norm interpolants of data in Reproducing Kernel Hilbert Spaces. Our upper bounds on the risk are of a multiple-descent shape for the various scalings of d=nα, α∈(0,1), for the input dimension d and sample size n. Empirical evidence supports our finding that minimum-norm interpo…
When maximum likelihood estimation is infeasible, one often turns to score matching, contrastive divergence, or minimum probability flow to obtain tractable parameter estimates. We provide a unifying perspective of these techniques as minimum Stein discrepancy estimators, and use this lens to design new diffusion kerne…
The modified J-flow with Calabi ansatz shows convergence or blow-up behavior based on topological constants.
problem Analyzing the behavior of the modified J-flow with Calabi ansatz.
method Using the Calabi symmetry and studying the singularities of the flow.
result The modified J-flow with Calabi ansatz converges to a solution away from a variety, and blows up along the variety.
MC-MCL improves MCL for nonlinear clustering.
problem Nonlinear clustering in data science.
method MC-MCL combines MCL with Minimum Curvilinearity for nonlinear distances.
result MC-MCL outperforms classical MCL and baseline clustering algorithms in nonlinear datasets.
Transformers learn linear models in-context without updates.
problem Understanding how transformers mimic linear models in-context.
method Gradient flow on linear regression tasks with random initialization.
result Transformers achieve prediction error competitive with best linear predictors.
New method improves MAP inference for CGMs on path graphs, avoiding approximation and maintaining integrality.
problem Improving MAP inference for aggregated count data in CGMs with small values.
method Formulated as a minimum cost flow problem, solved using DCA with efficient subroutines.
result Outputs higher quality solutions than conventional methods.
Many applications generate data with an intrinsic network structure such as time series data, image data or social network data. The network Lasso (nLasso) has been proposed recently as a method for joint clustering and optimization of machine learning models for networked data. The nLasso extends the Lasso from sparse…
The main theme of this paper is a relative version of the almost existence theorem for periodic orbits of autonomous Hamiltonian systems. We show that almost all low levels of a function on a geometrically bounded symplectically aspherical manifold carry contractible periodic orbits of the Hamiltonian flow, provided th…
The paper finds the shortest time to exploit arbitrage in multi-stock markets.
problem Finding the shortest time to exploit arbitrage in multi-stock markets.
method Characterizes the minimal time horizon for relative arbitrage in markets with 2 to 3 stocks and uses geometric flows for markets with 4 or more stocks.
result Explicit computation of minimal time horizon for 2 and 3 stocks markets, and characterization via geometric flows for markets with 4 or more stocks.
A model of open economics composed of producers and speculators is investigated by numerical simulations. The capital flows from the environment to the producers and from them to the speculators. The price fluctuations are suppressed by the speculators. When the aggressivity of the speculators grows, there is a transit…
Algorithm samples from Wasserstein barycenter of measures.
problem Sampling from Wasserstein barycenter of measures.
method Gradient flow of multimarginal formulation with penalization.
result Algorithm samples close to Wasserstein barycenter.
Curve shortening flow shrinks curves to points under certain conditions.
problem Understanding how curves shrink under curve shortening flow with ambient forces.
method Rescaling and curvature bounds analysis following Gage and Hamilton.
result Curves shrink to round points under certain curvature conditions.
Indices of acceptability are well suited to frame the axiomatic features of many performance measures, associated to terminal random cash flows.We extend this notion to classes of càdlàg processes modelling cash flows over a fixed investment horizon.We provide a representation result for bounded paths. We suggest an ac…
Simulation of high-speed train aerodynamics using RANS and machine learning.
problem Aerodynamic analysis of high-speed trains under turbulent flow conditions.
method RANS equations with turbulence model, machine learning (GEP, GPR, RF) for predictions.
result Random Forest (RF) provides the most accurate predictions for aerodynamic coefficients.
This paper shows neural networks can solve complex graph problems efficiently.
problem Solving exact maximum flow computation and minimum spanning tree problems.
method Introduces Max-Affine Arithmetic Programs and shows equivalence to neural networks.
result Two combinatorial optimization problems can be solved with polynomial-size neural networks.
Improved stock price prediction model using generalized order flow imbalance.
problem Improving stock price prediction models using new order flow imbalance indicators.
method Proposed a generalized order flow imbalance construction method and applied it to CSI 500 stocks.
result Generalized Stationarized Order Flow Imbalance (log-GOFI) shows significant improvement in explaining stock price changes.
Paper bounds PAC RL sample complexity in deterministic MDPs.
problem Identify ε-optimal policy with high probability.
method Proposes nearly matching upper and lower bounds on sample complexity, introduces deterministic return gap, uses graph-theoretical concepts and maximum-coverage exploration.
result First nearly matching upper and lower bounds on sample complexity for PAC RL in deterministic MDPs.
We notice that a generic nonsingular gradient field v=∇f on a compact 3-fold X with boundary canonically generates a simple spine K(f,v) of X. We study the transformations of K(f,v) that are induced by deformations of the data (f,v). We link the Matveev complexity c(X) of X with counting the …
Paper characterizes nc-rank using gradient flow on symmetric space.
problem Characterizing the noncommutative rank of matrices.
method Interprets residuals as gradients of a convex function on symmetric space and uses unbounded gradient flow.
result Noncommutative corank equals half the minimum gradient-norm of residuals.
In this paper, we mainly study the mean curvature flow in Kähler surfaces with positive holomorphic sectional curvatures. First, we prove that if the ratio λ of the maximum and the minimum of the holomorphic sectional curvatures <2, then there exists a positive constant $δ>\frac{29(λ-1)}{\sqrt{(48-24λ)^{2}+(29λ-29…
Paper explains neural collapse in neural networks using a new model.
problem Understanding neural collapse in neural networks during training.
method Introducing the unconstrained layer-peeled model (ULPM) to prove gradient flow convergence to critical points of a minimum-norm separation problem.
result Proves that all critical points are strict saddle points except the global minimizers exhibiting neural collapse.
We show that if K: P \to R is an autonomous Hamiltonian on a symplectic manifold (P,Ω) which attains 0 as a Morse-Bott nondegenerate minimum along a symplectic submanifold M, and if c_1(TP)|_M vanishes in real cohomology, then the Hamiltonian flow of K has contractible periodic orbits with bounded period on all suffici…
We define a quantisation of the J-flow over a projective complex manifold. As corollaries, we obtain new proofs of uniqueness of critical points of the J-flow and that these critical points achieve the absolute minimum of an associated energy functional. We show that the existence of a critical point of the J-flow impl…
New proof given for a functional's minimum condition.
problem Functional's minimum condition under varying h. method Variational method and maximum principle.
result Functional achieves minimum under Ding-Jost-Li-Wang condition.
NTK neural networks are robust to adversarial attacks in nonparametric regression.
problem Adversarial robustness of neural networks in nonparametric regression.
method Gradient flow with early stopping for NTK neural networks, proving robustness in Sobolev spaces.
result NTK neural networks achieve optimal adversarial robustness rates in Sobolev spaces.
Paper presents a new method to learn unnormalized models efficiently.
problem Scalability issues in score matching for flexible unnormalized models.
method Connects learning objectives to Wasserstein gradient flows for scalability.
result Demonstrates improved learning of unnormalized models on manifolds.
This paper analyzes convergence of large-scale Transformers with weight decay.
problem Understanding optimization guarantees in large-scale Transformer training.
method Construct mean-field limit, show gradient flow convergence to PDE, demonstrate global minimum consistency.
result Gradient flow reaches global minimum in large-scale Transformers with small weight decay.
Many important optimization problems, such as the minimum spanning tree and minimum-cost flow, can be solved optimally by a greedy method. In this work, we study a learning variant of these problems, where the model of the problem is unknown and has to be learned by interacting repeatedly with the environment in the ba…
Based on the approach of flow distances, the international trade flow system is studied from the perspective of multi-layer flow network. A model of multi-layer flow network is proposed for modelling and analyzing multiple types of flows in flow systems. Then, flow distances are introduced, and symmetric minimum flow d…
Gradient flow in phase retrieval escapes spurious minima with high probability.
problem Understanding gradient-based optimization in high-dimensional non-convex functions.
method Analytical and numerical study of gradient dynamics in phase retrieval.
result Gradient flow avoids spurious minima by drifting along unstable directions.
A new method for learning gradient flows from population dynamics.
problem Reconstructing population dynamics from limited data.
method Residual approach to enforce continuity equations, combining with data-fitting divergence.
result Demonstrated state-of-the-art performance across trajectory inference benchmarks.