New method uses neural operators for efficient function space optimization.
problem Optimization over function spaces with costly function evaluations.
method Sample-then-optimize approach with neural operator surrogates.
result Better sample efficiency and significant performance gains in experiments.
This research proves that quadratic regularized optimal transport can approximate the Laplace-Beltrami operator on smooth manifolds.
problem Approximating the Laplace-Beltrami operator using optimal transport with quadratic regularization.
method Deriving first-order optimal potentials and analyzing the convergence of discrete Laplace operators.
result The discrete Laplace operators converge to the Laplace-Beltrami operator on smooth manifolds.
OpEvo automates tensor operator optimization for better efficiency.
problem Manual optimization of tensor operators is inefficient and limited.
method OpEvo uses evolutionary computation with topology-aware mutation.
result OpEvo finds optimal configurations with less effort and variance.
Optimal lower bounds for eigenvalues of Dirac-Witten operator on certain submanifolds.
problem Estimating eigenvalues of the Dirac-Witten operator on specific submanifolds.
method Optimal lower bounds derived using intrinsic and extrinsic expressions.
result Limiting-cases of eigenvalues studied and optimal bounds obtained.
Optimizes eigenvalue bounds for submanifold Dirac operators.
problem Estimating eigenvalues of submanifold Dirac operators.
method Optimal lower bounds derived using intrinsic and extrinsic expressions.
result Optimal eigenvalue bounds established for submanifold Dirac operators.
Iterative thresholding algorithms seek to optimize a differentiable objective function over a sparsity or rank constraint by alternating between gradient steps that reduce the objective, and thresholding steps that enforce the constraint. This work examines the choice of the thresholding operator, and asks whether it i…
Automates GPU kernel optimization for diverse applications.
problem Lack of systematic evaluation for multi-scenario GPU kernel optimization.
method Introduces MSKernelBench and CUDAMaster for automated optimization.
result Significant speedups across various operators, outperforming existing tools.
Neural operators achieve fast convergence rates for solving PDEs.
problem Solving partial differential equations (PDEs) efficiently.
method Two-layer neural operators with gradient descent analysis in RKHS.
result Fast convergence rates are minimax optimal for early-stopped GD.
Smoothed top-k operator improves model training efficiency.
problem Discontinuous top-k operation makes models untrainable end-to-end.
method SOFT top-k operator approximates top-k as EOT solution.
result Improved performance in k-nearest neighbors and beam search.
Study optimal holomorphic extensions on complex manifolds with transitivity property.
problem Optimal holomorphic extensions on complex manifolds with transitivity property.
method Use Toeplitz operators and transitivity property for optimal holomorphic extensions.
result Transitivity property of optimal holomorphic extensions with small defect.
Operator calculus for population-based optimization provides a unified framework for analyzing convergence of various methods.
problem Convergence analysis of population-based optimization methods
method Introduce an operator calculus for describing composite mean-field algorithms as compositions of elementary operators acting on probability measures.
result Establish a modular Lyapunov principle for certifying exponential decay of state-space Lyapunov function and search errors.
Optimal multiscale learning of linear operators
problem Statistical and computational limits of learning bounded linear operators between Sobolev spaces
method Reformulate as an infinite-dimensional matrix regression problem with heterogeneous multiscale structure
result Establish minimax rates and construct a finite-resolution blockwise least-squares estimator attaining these rates
Near-optimal rates for multi-task learning with shared representations.
problem Approximation and statistical complexity of learning multiple operators.
method Multiple Neural Operators (MNO) architecture and comparison with DeepONet.
result Near-optimal upper and lower bounds for approximation and generalization.
The paper tackles data-driven optimal control of unknown nonlinear systems using RKHS.
problem Unknown nonlinear dynamics and stage cost functions.
method Embed state densities into RKHS, learn Markov operators, solve Hamilton-Jacobi-Bellman recursions.
result Solves a wide range of nonlinear control problems, including depth regulation.
Variational inference is an umbrella term for algorithms which cast Bayesian inference as optimization. Classically, variational inference uses the Kullback-Leibler divergence to define the optimization. Though this divergence has been widely used, the resultant posterior approximation can suffer from undesirable stati…
Optimizes learning Hilbert-Schmidt operators between Sobolev spaces.
problem Statistical limits of learning mappings between infinite-dimensional function spaces.
method Minimax optimal regularization and multilevel training.
result Multilevel kernel operator learning achieves optimal learning rate.
Study solves optimal portfolio selection using HJB equation.
problem Optimal portfolio selection problem.
method Maximal monotone operator method, Banach fixed-point theorem, Fourier transform, monotone operators technique.
result Existence and uniqueness of solution to HJB equation.
DeepCO uses deep learning for offline combinatorial optimization in warehouse operations.
problem Optimizing warehouse operation sequences in offline settings.
method DeepCO framework utilizing distribution regularized optimization for TSP.
result DeepCO reduces route length by 5.7% on average for TSP problems.
The impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interferes with the contraction properties of the Bellman operator. Surprisingly, despite these concerns, and independent of its effect on explorati…
DVA framework attributes value of predictive models to features, configurations, and interactions.
problem Lack of explanation for how predictive models influence operational decisions.
method Shapley-based cooperative game theory applied to predict-then-optimize systems.
result DVA can guide targeted interventions to align model beliefs with operational performance.
Study eigenvalues and eigenfunctions of fourth-order operators in annuli, proving optimal estimates and non-radiality.
problem Eigenvalue and eigenfunction analysis of fourth-order operators in degenerating annuli.
method Optimal estimates and non-radiality results for eigenfunctions in annuli.
result Nigh optimal estimate for the first eigenvalue and non-radiality of eigenfunctions in degenerating annuli.
A note on setting swap parameters for traders.
problem Determining optimal slippage parameters and trade size for wealth swapping.
method Theoretical solution and framework for optimal slippage parameters and trade size.
result Offers a method to solve optimal slippage parameters and trade size for wealth swapping.
Paper proposes a novel metric learning algorithm using Riemannian optimization.
problem Optimizing a smooth, convex function in Riemannian space with constraints.
method Developed a primal-dual algorithm with proximal operator for iterative optimization.
result Demonstrated the efficacy of the proposed metric learning algorithm on fund selection.
Self-ONNs adapt nodal operators during training for higher diversity and efficiency.
problem Limited network heterogeneity and high computational demand in ONNs.
method Self-organized ONNs with generative neurons that adapt nodal operators during training.
result Self-ONNs achieve utmost heterogeneity and computational efficiency.
A new stochastic primal--dual algorithm for solving a composite optimization problem is proposed. It is assumed that all the functions/operators that enter the optimization problem are given as statistical expectations. These expectations are unknown but revealed across time through i.i.d. realizations. The proposed al…
NEON uses neural networks to optimize functions in infinite-dimensional spaces.
problem Optimizing composite functions in function spaces.
method NEON (Neural Epistemic Operator Networks) for sequential decision-making.
result NEON achieves state-of-the-art performance with fewer parameters.
New optimizers control network width scaling, improving stability and transfer across different model sizes.
problem Designing stable optimizers for networks of varying widths.
method Interpreting optimizers as steepest descent under mean-normalized operator norms, enabling layerwise composability and width-independent bounds.
result New optimizers like row normalization and column normalization provide stable learning-rate transfer across different model widths.
Optimal scaling found to depend on operator norm across large models and datasets.
problem Lack of unifying principle for optimal hyperparameter scaling across models and datasets.
method Discovered that optimal scaling is conditioned on the operator norm of the output layer.
result The optimal learning rate/batch size pair (η∗,B∗) consistently has the same operator norm value. ICON-OCnet solves optimal execution problems with neural networks and few examples.
problem Optimal order execution in markets with unknown price impact.
method Transformer-based neural network architecture (ICON-OCnet) that learns price impact from few examples and applies it to optimal execution strategies.
result ICON-OCnet accurately infers price impact models and retrieves optimal execution strategies for various propagator kernels.
Generative operators solve many convex problems with minimal parameters.
problem Worst-case parameter bounds limit the practical use of neural operators.
method Developed generative equilibrium operators (GEOs) using realizable finite-dimensional layers.
result GEOs can uniformly approximate solutions to convex optimization problems with logarithmic growth in parameters.
Paper proposes a new optimization framework for learning eigenfunctions of operators.
problem Computing eigenvalue decomposition of high-dimensional operators.
method Operator SVD with Neural Networks via Nested Low-Rank Approximation.
result Proposed method efficiently learns top-L singular values and functions in the correct order.
Proposes differentiable and sparse top-k operators for neural networks.
problem Discontinuity of top-k operator makes it unsuitable for end-to-end training with backpropagation.
method Formulates top-k as a linear program over permutahedron, introduces p-norm regularization, and uses isotonic optimization.
result Successfully applied to neural network pruning, fine-tuning, and routing.
New method solves blind inverse problems by optimizing both operator and image parameters.
problem Solving blind inverse problems with known forward operator.
method Parallel reverse diffusion guided by gradients from intermediate stages.
result State-of-the-art performance on blind deblurring and imaging through turbulence.
We consider a new family of operators for reinforcement learning with the goal of alleviating the negative effects and becoming more robust to approximation or estimation errors. Various theoretical results are established, which include showing on a sample path basis that our family of operators preserve optimality an…
Proposes glocal hypergradient estimation for hyperparameter optimization.
problem Combining reliability and efficiency in hyperparameter optimization.
method Uses Koopman operator theory to approximate global hypergradients from local ones.
result Achieves both reliability and efficiency in hyperparameter optimization.
Optimal transport for functional data using Hilbert-Schmidt operators.
problem Optimal transport for distributions on function spaces with partially represented stochastic maps.
method Regularization technique to restrict transport maps to Hilbert-Schmidt operators, developing an efficient algorithm.
result Existence, uniqueness, and consistency of the Hilbert-Schmidt operator estimate for the transport map.
Model for open, decentralized network with task load balancing.
problem Complex computational tasks in open, decentralized networks.
method Incentive-based load balancing using economic mechanisms.
result Optimized resource allocation and enhanced system resilience.
Paper proposes SMO for solving bilevel optimization problems efficiently.
problem Solving bilevel optimization problems with nonsmooth convex lower-level and nonconvex upper-level objectives.
method Sequential minimax optimization (SMO) method using modified augmented Lagrangian and penalty schemes.
result Improves operation complexity for finding ε-KKT solutions. The optimization of electric machines at multiple operating points is crucial for applications that require frequent changes on speeds and loads, such as the electric vehicles, to strive for the machine optimal performance across the entire driving cycle. However, the number of objectives that would need to be optimize…
Decision-calibrated prediction sets improve power system operations by reducing unnecessary costs.
problem Balancing operating costs and reliability in power systems with renewable uncertainty.
method Learn conditional prediction sets as sub-level sets of norm-based score functions, calibrate uncertainty sets based on reliability of downstream decisions.
result Decision-calibrated sets lead to more efficient operations with smaller uncertainty sets and lower costs compared to standard coverage-based calibration.
Machine learning pipeline potentially consists of several stages of operations like data preprocessing, feature engineering and machine learning model training. Each operation has a set of hyper-parameters, which can become irrelevant for the pipeline when the operation is not selected. This gives rise to a hierarchica…
Solves isoperimetric problem for curl operator on compact 3-manifolds.
problem Isoperimetric problem for the curl operator on compact 3-manifolds.
method Analyzes eigenvalues and optimal domains of the curl operator.
result Optimal lower bounds for eigenvalues of curl operator are always attained.
A new method reformulates Optimal Transport Conditional Flow Matching using proximal operators.
problem Optimal Transport Conditional Flow Matching (OT-CFM) for generating models.
method Reformulate OT-CFM using proximal operators and extended Brenier potential.
result OT-CFM dynamics are terminally normally hyperbolic for manifold-supported targets.
Optimizes deep learning models for ocean dynamics using Fourier neural operators.
problem Efficiently training deep learning models for ocean dynamics with optimal hyperparameters.
method Multiobjective hyperparameter optimization with DeepHyper for Fourier neural operators.
result Optimal hyperparameters significantly improved model performance in ocean dynamics forecasting.
DOODL learns shared spectral dynamics across related dynamical systems.
problem Learning independent dynamical operators for each system limits discovery of shared structure.
method DOODL learns a dictionary of characteristic spectral dynamics on a manifold of related systems.
result DOODL achieves errors one to two orders of magnitude lower than independent operator estimation methods.
A new metric compares dynamical systems using operator eigenvalues.
problem Comparing and interpolating nonlinear dynamical systems from trajectory data.
method Representing systems as distributions of operator eigenvalues and projectors, defining a spectral-Grassmann Wasserstein metric.
result The proposed metric outperforms standard operator-based distances in machine learning applications.
Explaining neural network computation in terms of probabilistic/fuzzy logical operations has attracted much attention due to its simplicity and high interpretability. Different choices of logical operators such as AND, OR and XOR give rise to another dimension for network optimization, and in this paper, we study the o…
This study improves estimation of the first principal component in multivariate functional data.
problem Estimating the first principal component of multivariate random processes.
method Defined covariance functions and operators, introduced LASSO optimization, and established minimax lower bounds.
result The method provides an optimal variance in the minimax sense for estimating eigenelements.