State-of-the-art methods in convex and non-convex optimization employ higher-order derivative information, either implicitly or explicitly. We explore the limitations of higher-order optimization and prove that even for convex optimization, a polynomial dependence on the approximation guarantee and higher-order smoothn…
The paper connects higher order risk measures and stochastic dominance, showing their equivalence and integrating them with optimization.
problem Comparing and characterizing random outcomes in risk assessment.
method Exploring the equivalence between higher order risk measures and stochastic dominance, using stochastic optimization and expectiles as examples.
result Higher order risk measures and stochastic dominance are equivalent and can be used to characterize random outcomes.
We provide improved convergence rates for various \emph{non-smooth} optimization problems via higher-order accelerated methods. In the case of ℓ∞ regression, we achieves an O(ε−4/5) iteration complexity, breaking the O(ε−1) barrier so far present for previous methods. We arrive at a similar rate fo…
Lower bounds for higher-order methods in non-convex optimization.
problem Proving lower bounds for higher-order methods in smooth non-convex finite-sum optimization.
method Analyzing deterministic and randomized algorithms, proposing a new smoothness assumption.
result Proves optimal lower bounds for simulating pth-order regularized methods on the whole function.
Study optimizes zero-order strongly convex function minimization with higher order smoothness.
problem Optimizing a strongly convex function with noisy evaluations.
method Randomized approximation of projected gradient descent with smoothing kernel.
result Upper bounds and minimax lower bounds for the algorithm, showing near-optimality.
New model estimates higher-order interactions in stochastic processes using lower-dimensional projections.
problem Estimating higher-order interaction effects in stochastic processes with limited data.
method Additive Poisson Process (APP) combines information geometry and generalized additive models to model intensity functions in lower dimensions.
result The model can estimate higher-order intensity functions with sparse data.
The paper constructs denoisers that recover the Brenier map from higher-order score functions.
problem Estimating the Brenier map from noisy data.
method Constructs a hierarchy of denoisers using higher-order score functions.
result The T∞ denoiser recovers the Brenier map from the additive Gaussian model. Higher-order optimization problems naturally appear when investigating the effects of a patent with finite length, as in the pioneering work of Futagami and Iwaisako (2007). In this paper, we establish the Euler equations and transversality conditions necessary for analyzing such higher-order optimization problems. We …
New principle for optimal control with higher order differential constraints.
problem Optimal control problems with higher order differential constraints.
method Derivation of the Principle of Minimal Labour and generalization of Pontryagin Maximum Principle.
result Generalized Pontryagin Maximum Principle for higher order constraints.
In this paper, we describe a geometric setting for higher-order lagrangian problems on Lie groups. Using left-trivialization of the higher-order tangent bundle of a Lie group and an adaptation of the classical Skinner-Rusk formalism, we deduce an intrinsic framework for this type of dynamical systems. Interesting appli…
Paper proves higher-order flow matching preserves optimality in generative modeling.
problem Theoretical guarantees for higher-order flow matching in generative modeling.
method Neural network approximations with controlled depth, width, and sparsity.
result Proves worst case optimality for second-order flow matching.
Paper tackles non-convex optimization for higher moments in portfolio management.
problem Complexity of higher moments in optimization problems.
method Method of successive convex approximation.
result Solves mean-variance-skewness problem using non-convex optimization.
Improved algorithms for convex-concave min-max optimization and monotone variational inequalities.
problem Efficiently solving constrained convex-concave min-max problems and monotone variational inequalities.
method Higher-order methods achieving iteration complexities of O(1/T^{rac{p+1}{2}}) for p-th order derivatives.
result Achieved improved convergence rates for min-max and monotone variational inequalities.
In this paper, we examine higher order difference problems. Using the "squeezing" argument, we derive both Euler's condition and the transversality condition. In order to derive the two conditions, two needed assumptions are identified. A counterexample, in which the transversality condition is not satisfied without th…
In this paper, we investigate the popular deep learning optimization routine, Adam, from the perspective of statistical moments. While Adam is an adaptive lower-order moment based (of the stochastic gradient) method, we propose an extension namely, HAdam, which uses higher order moments of the stochastic gradient. Our …
Estimates hypergraphons for modeling complex interactions efficiently.
problem Modeling higher-order interactions using hypergraphons.
method Restricted class of Simple Lipschitz Hypergraphons (SLH) for efficient estimation.
result Optimal rates of convergence for SLH estimator.
New estimator stabilizes higher-order influence functions for stable statistical inference.
problem Numerical instability in estimating inverse population Gram matrix.
method Proposes a new stabilized higher-order estimator without sample splitting.
result Stabilized estimator exhibits more stable performance and similar statistical guarantees.
New estimator stabilizes higher-order influence functions for bilinear forms.
problem Stability issues in estimating bilinear forms using higher-order influence functions.
method Proposes a new stabilized higher-order estimator for a class of bilinear forms without sample splitting.
result New estimator exhibits more stable finite-sample performance compared to the empirical higher-order estimator.
New Riemannian optimization improves variance estimation in mixed models.
problem Challenges in estimating variance parameters in linear mixed models due to constraints.
method Formulated as an optimization problem on a Riemannian manifold, using Riemannian gradient and Hessian.
result Yields higher quality variance parameter estimates compared to existing methods.
HAMD optimizes cubic portfolios without quadratization, achieving better results.
problem Optimizing higher-order portfolio models with reduced distortion.
method Hybrid pipeline combining continuous Hamiltonian search, cardinality-preserving projection, and iterated local search.
result HAMD achieves significantly lower native cubic objective values than classical heuristics.
Local search heuristics for non-convex optimizations are popular in applied machine learning. However, in general it is hard to guarantee that such algorithms even converge to a local minimum, due to the existence of complicated saddle point structures in high dimensions. Many functions have degenerate saddle points su…
Sharp inequalities in unit ball with constraints on moments.
problem Establishing Sobolev trace inequalities with constraints.
method Constructing smooth test functions for higher order moments.
result Almost optimal Sobolev trace inequalities for 2nd and 4th orders.
We study the finite horizon Merton portfolio optimization problem in a general local-stochastic volatility setting. Using model coefficient expansion techniques, we derive approximations for the both the value function and the optimal investment strategy. We also analyze the `implied Sharpe ratio' and derive a series a…
Study improves BN TTA under distribution shift using higher-order asymptotics.
problem Improving BN TTA for changing data distributions.
method Integrates Edgeworth expansion and saddlepoint approximation with one-step M-estimation.
result Derives optimal weighting parameter for minimized mean-squared error.
Geometric formalism views optimization algorithms as discrete connections, revealing their algebraic curvature and flatness properties.
problem Understanding and optimizing the behavior of iterative optimization algorithms.
method Introducing a geometric and operator-theoretic formalism where optimization algorithms are encoded by coupled channels (drift and diffusion) whose algebraic curvature measures the deviation from ideal reversibility.
result Flat connections correspond to methods whose updates commute up to higher order, achieving minimal numerical dissipation and preserving stability.
Optimal first-order methods are shown to be fundamental limits in functional estimation.
problem Optimal functional estimation under weak conditions.
method Formalization of functional estimation with black-box nuisance function estimates and derivation of minimax lower bounds.
result First-order methods are optimal under weak conditions, but higher-order methods can outperform them when nuisance function structure is known.
Higher-order ODE solvers improve deep learning performance.
problem Improving deep learning performance using higher-order ODE solvers.
method Evaluation and improvement of Runge-Kutta (RK) methods for deep learning.
result Higher-order RK solvers can improve deep learning performance by incorporating key ingredients of optimizers.
New method improves DAG learning by using large coefficients for higher-order terms.
problem Recovering DAG structures from observational data is challenging due to combinatorial optimization.
method Proposes truncated matrix power iteration to approximate DAG constraints efficiently.
result Empirically outperforms previous methods by a factor of 3 or more in structural Hamming distance.
Transformers can approximate Newton's method for logistic regression.
problem Implementing higher order optimization methods in Transformers.
method Linear attention Transformers with ReLU layers approximating second order optimization algorithms.
result Transformers can implement a single step of Newton's iteration for matrix inversion.
Bayesian method detects mesoscale structures in pathway data networks.
problem Mesoscale structures in pathway data networks are hard to detect due to dependencies between interactions.
method Bayesian approach modeling optimal partitioning and higher-order dynamics.
result Method can recover both proximity-based and role-based groupings of nodes.
Extends RRR to capture nonlinear interactions in multi-response regression.
problem Complex relationships in real-world data cannot be adequately modeled by linear interactions.
method Introduces Higher Order Reduced Rank Regression (HORRR) using tensor representations and Tucker decomposition.
result HORRR can capture nonlinear interactions in multi-response regression.
Efficiently approximates higher-order derivatives for generative models.
problem Expensive computation of higher-order derivatives in generative models.
method Rewrite SM objective in terms of directional derivatives and use finite difference for efficient approximation.
result Comparable results to gradient-based methods but significantly more computationally efficient.
A quantum framework optimizes collateral allocation for derivatives.
problem Legal constraints and operational rules in collateral allocation for derivatives.
method Certified higher-order quantum framework that normalizes margin requirements and builds a bounded neighborhood of actions.
result Quantum framework improves certified sample quality compared to classical methods.
Paper improves stochastic bilevel optimization methods for highly-smooth problems.
problem Finding ε-stationary points in stochastic bilevel optimization. method Proposes F2SA-p methods using pth-order finite differences for hyper-gradient approximation. result Achieves upper complexity bound of ildeO(pε−4−p/2) for pth-order smooth problems. This paper introduces matrix product state (MPS) decomposition as a new and systematic method to compress multidimensional data represented by higher-order tensors. It solves two major bottlenecks in tensor compression: computation and compression quality. Regardless of tensor order, MPS compresses tensors to matrices …
Predicts node sequences in graphs using multi-order network models.
problem Predicting sequences of node traversals in graphs.
method Combines multiple higher-order network models into a multi-order model, fitting and selecting the optimal maximum order.
result Outperforms state-of-the-art algorithms for next-element and full sequence prediction.
Improves safety region certification for smoothed classifiers without changing smoothing scheme.
problem Certified safety regions for smoothed classifiers are often small compared to optimal.
method Generalizes certified radius calculation as nested optimization problem, uses 0th-1st order information, and designs efficient estimators.
result Certified safety regions are significantly larger than current methods, achieving significant improvements on various metrics.
We consider the minimization of submodular functions subject to ordering constraints. We show that this optimization problem can be cast as a convex optimization problem on a space of uni-dimensional measures, with ordering constraints corresponding to first-order stochastic dominance. We propose new discretization sch…
A key feature of inductive logic programming (ILP) is its ability to learn first-order programs, which are intrinsically more expressive than propositional programs. In this paper, we introduce techniques to learn higher-order programs. Specifically, we extend meta-interpretive learning (MIL) to support learning higher…
Unified framework for learning flexible probabilistic programs using DPP and PAC-Bayes bounds.
problem Learning and generalizing from complex probabilistic models.
method Unified DPP representation and PAC-Bayes bounds for stochastic programs.
result Improved performance and generalization prediction using flexible DPP model representations and learned complexity measures.
Optimized variable orderings improve autoregressive model performance.
problem Challenges in variable ordering affect autoregressive model efficiency.
method Learn graphical model structure to inform optimal variable orderings.
result Graph-informed orderings yield higher-fidelity samples.
For certain classes of knots we define geometric invariants called higher-order genera. Each of these invariants is a refinement of the slice genus of a knot. We find lower bounds for the higher-order genera in terms of certain von Neumann ρ-invariants, which we call higher-order signatures. The higher-order genera o…
A fundamental property of complex networks is the tendency for edges to cluster. The extent of the clustering is typically quantified by the clustering coefficient, which is the probability that a length-2 path is closed, i.e., induces a triangle in the network. However, higher-order cliques beyond triangles are crucia…
Optimized GAN discriminator using polyharmonic interpolation.
problem Optimizing the discriminator in GANs with higher-order gradient regularization.
method Polyharmonic interpolation and variational calculus.
result The optimal discriminator is a polyharmonic radial basis function.
Stability of capillary hypersurfaces with higher order mean curvature.
problem Stability of capillary hypersurfaces with constant higher order mean curvature.
method Generalization of classical stability theory for capillary hypersurfaces.
result Results on stability for capillary hypersurfaces with higher order mean curvature.
Unified framework for higher-order network analysis.
problem Complex structure of space of networks.
method Measure-theoretic formalism, Gromov-Wasserstein distance, co-optimal transport distance.
result Unified theoretical treatment of generalized networks.
We use the Frölicher-Nijenhuis formalism to reformulate the inverse problem of the calculus of variations for a system of differential equations of order 2k in terms of a semi-basic 1-form of order k. Within this general context, we use the homogeneity proposed by Crampin and Saunders in [14] to formulate and discuss t…
New methods improve estimation accuracy in noisy settings.
problem Estimating treatment effects in the presence of treatment noise.
method Developed new structure-agnostic cumulant estimators and practical procedures for higher-order robustness.
result Demonstrated that existing DML estimator is suboptimal for non-Gaussian treatment noise and introduced ACE procedures for improved accuracy.