We consider the problem of approximate Bayesian inference in log-supermodular models. These models encompass regular pairwise MRFs with binary variables, but allow to capture high-order interactions, which are intractable for existing approximate inference techniques such as belief propagation, mean field, and variants…
This work improves sampling of graph signals with universal bounds and greedy methods.
problem Sampling graph signals is hard due to irregularity and noise.
method Derives universal performance bounds and near-optimal guarantees for greedy sampling.
result Explicit bounds on approximate supermodularity show greedy search can be optimized with worst-case guarantees.
Paper proves supermodularity of AG-SSL objective and proposes a greedy sampling algorithm.
problem Improving semi-supervised learning with limited labeled data.
method Proves supermodularity of AG-SSL objective under Stieltjes regularization and proposes a greedy sampling algorithm.
result Proposed method achieves superior classification accuracy compared to state-of-the-art methods.
New method improves neural network generalization and robustness.
problem Improving neural network generalization and robustness to input perturbations.
method Co-Mixup with supermodular diversity, optimizing data saliency and diversity.
result Achieves state-of-the-art generalization, calibration, and localization results.
We consider log-supermodular models on binary variables, which are probabilistic models with negative log-densities which are submodular. These models provide probabilistic interpretations of common combinatorial optimization tasks such as image segmentation. In this paper, we focus primarily on parameter estimation in…
Paper reduces dimensionality for robust option pricing in 2-asset markets.
problem Robust option pricing in multi-asset markets with sub- or supermodular payoffs.
method Investigates the geometry of VMOT solutions, proving dimension reduction for 2 assets and developing a Sinkhorn algorithm.
result Dimension reduction to single-factor structure for 2-asset markets, significantly reducing computational time and improving accuracy.
Study shows how to better estimate credit provisions and economic capital.
problem Estimating credit provisions and economic capital accurately.
method Using supermodularity ordering properties and elliptically distributed latent factors.
result Convex risk measures of credit losses are nondecreasing w.r.t. various covariances.
We consider the problem of stochastic comparison of general Garch-like processes, for different parameters and different distributions of the innovations. We identify several stochastic orders that are propagated from the innovations to the Garch process itself, and discuss their interpretations. We focus on the convex…
Deep learning models' architectures, including depth and width, are key factors influencing models' performance, such as test accuracy and computation time. This paper solves two problems: given computation time budget, choose an architecture to maximize accuracy, and given accuracy requirement, choose an architecture …
Empirical risk minimization frequently employs convex surrogates to underlying discrete loss functions in order to achieve computational tractability during optimization. However, classical convex surrogates can only tightly bound modular loss functions, sub-modular functions or supermodular functions separately while …
The paper is motivated by a problem concerning the monotonicity of insurance premiums with respect to their loading parameter: the larger the parameter, the larger the insurance premium is expected to be. This property, usually called loading monotonicity, is satisfied by premiums that appear in the literature. The inc…
In this note we establish some appropriate conditions for stochastic equality of two random variables/vectors which are ordered with respect to convex ordering or with respect to supermodular ordering. Multivariate extensions of this result are also considered.
Efficiently learns perturb-and-map models using weighted log-likelihood.
problem Structured output prediction with weighted Hamming losses.
method Generalizes perturb-and-MAP framework, uses dynamic graph cuts for MAP inference, and double stochastic gradient descent for efficient learning.
result Shows efficiency in learning log-supermodular models with weak supervision.
The paper derives upper hedging prices for multivariate contingent claims using game-theoretic probability and submodularity.
problem Deriving upper hedging prices for complex financial contracts.
method Game-theoretic approach, optimization over simplexes, Lovász extension, Black-Scholes-Barenblatt equations.
result Upper and lower hedging prices can be calculated efficiently for submodular or supermodular payoff functions.
Most prior work on active learning of classifiers has focused on sequentially selecting one unlabeled example at a time to be labeled in order to reduce the overall labeling effort. In many scenarios, however, it is desirable to label an entire batch of examples at once, for example, when labels can be acquired in para…
Recurrent neural network (RNN)'s architecture is a key factor influencing its performance. We propose algorithms to optimize hidden sizes under running time constraint. We convert the discrete optimization into a subset selection problem. By novel transformations, the objective function becomes submodular and constrain…
Learning with non-modular losses is an important problem when sets of predictions are made simultaneously. The main tools for constructing convex surrogate loss functions for set prediction are margin rescaling and slack rescaling. In this work, we show that these strategies lead to tight convex surrogates iff the unde…
In this paper, we present an algorithm for minimizing the difference between two submodular functions using a variational framework which is based on (an extension of) the concave-convex procedure [17]. Because several commonly used metrics in machine learning, like mutual information and conditional mutual information…
Unified framework for hard affine SDP constraints in vRKHSs.
problem Incorporating shape constraints into predictive models for rich function classes.
method Unified convex optimization framework using second-order cone tightening.
result Unified and modular approach for handling multiple shape constraints.
Model shows AI adoption amplifies financial market risk through prediction, herding, and cognitive dependency.
problem Systemic risk in financial markets due to AI adoption.
method Developed a unified model within an extended rational expectations framework, incorporating endogenous adoption, performative prediction, algorithmic herding, and cognitive dependency.
result Systemic risk multiplier grows superlinearly with AI penetration, implying tail-loss amplification of 18-54%.
Paper connects neural network score approximation to reverse diffusion model distribution approximation.
problem Quantifying the relationship between neural network score approximation and the distribution generated by reverse diffusion models.
method Combines Hornik's universal approximation theorem, Girsanov's theorem, and data processing inequality.
result Neural network score approximation guarantees distribution approximation in reverse diffusion models.
The study provides conditions for approximating Riemannian manifolds with polyhedral metrics.
problem Approximating Riemannian manifolds with polyhedral metrics.
method Conditions on curvature tensors for Lipschitz and local polyhedral approximations.
result Conditions are sufficient for local polyhedral approximations, conjectured to be sufficient for global approximations.
Paper proposes a new adaptive multiscale value function approximation for reinforcement learning.
problem Value function approximation in reinforcement learning with varying complexity.
method Adaptive multiscale approximation using multiresolution analysis and tree approximation.
result Convergence rate of the multiscale approximation is independent of basis function regularity.
Optimal function approximation with Relu neural networks achieves minimal error.
problem Finding the minimal error in approximating convex functions with Relu networks.
method Established necessary and sufficient conditions for optimal approximations, presented neural network architectures, and proposed an algorithm for convergence.
result Proved the convergence of the proposed algorithm and validated it with experimental results.
Paper proposes MCMA architecture for neural approximate computing with higher invocation rate and energy savings.
problem Limited invocation rate of neural approximators leading to suboptimal energy efficiency.
method Introduces MCMA architecture with a multiclass classifier and multiple approximators, sharing hardware resources and efficiently swapping approximators.
result Significantly higher invocation rate and energy savings compared to existing methods.
Geometric Gaussian approximations capture any distribution.
problem Approximating complex probability distributions.
method Geometric Gaussian approximations through diffeomorphisms or exponential maps.
result Geometric Gaussian approximations are universal, capturing any distribution.
Deep learning networks are approximated using dynamical systems theory.
problem Understanding the approximation capabilities of deep learning networks.
method Modeling deep residual networks as continuous-time dynamical systems and using approximation theories in Lp. result Established general sufficient conditions for universal approximation of deep residual networks.
Method approximates Riemannian barycenter on manifolds.
problem Computing the exact Riemannian barycenter is computationally expensive.
method Uses under- and over-approximations of Riemannian distance to compute an approximate barycenter.
result Approximation method is more efficient than exact methods and steepest descent.
Efficiently reduces tensor ranks using mean-field approximation.
problem Low-rank approximation of non-negative tensors.
method Mean-field approximation of tensor rank reduction.
result Our algorithm achieves faster and competitive tensor rank reduction.
Study approximates unknown function levels with queries.
problem Approximating unknown function levels through sequential queries.
method Introduce Bisect and Approximate algorithms to reduce to local function approximation.
result Rate-optimal sample complexity guarantees for H{ö}lder functions.
We study sparse approximate solutions to convex optimization problems. It is known that in many engineering applications researchers are interested in an approximate solution of an optimization problem as a linear combination of elements from a given system of elements. There is an increasing interest in building such …
Softmax attention approximates complex functions and subsumes many known universal approximators.
problem Universal approximation of continuous sequence-to-sequence functions.
method Interpolation-based analysis of attention's internal mechanism, showing its ability to approximate ReLU functions.
result Softmax attention is a universal approximator for continuous sequence-to-sequence functions.
Deviation inequalities for stochastic approximation methods.
problem Establishing bounds on the deviation of stochastic approximation methods.
method Martingale approximation method for separately Lipschitz functions.
result Established various deviation inequalities for stochastic approximation by averaging and minimization.
Improved matrix approximation using randomized algorithms.
problem Finding better approximations of given matrices.
method Randomized algorithms to compute (HT) as an improved approximation. result Computed (HT) provides a better approximation than given F∗. AXNet combines two neural networks into one for efficient approximate computing.
problem Efficient approximate computing for error-resilient applications.
method End-to-end trainable AXNet architecture that fuses approximator and predictor.
result Significant improvement in invocation rate and reduction in training time.
Approximate symmetries of geodesic equations on 2-spheres are studied. These are the symmetries of the perturbed geodesic equations which represent approximate path of a particle rather than exact path. After giving the exact symmetries of the geodesic equations, two different approaches to study the approximate symmet…
Novel method uses MCMC to improve approximation networks.
problem Approximating complex, intractable distributions.
method Amortized MCMC with iterative refinement of approximation network.
result Improved quality of deep generative model training.
We are concerned with an approximation problem for a symmetric positive semidefinite matrix due to motivation from a class of nonlinear machine learning methods. We discuss an approximation approach that we call {matrix ridge approximation}. In particular, we define the matrix ridge approximation as an incomplete matri…
New algorithms minimize non-zero entries in low-rank approximations.
problem Minimizing non-zero entries in low-rank approximations of matrices.
method Approximation algorithms for minimizing ℓ0-norm of rank-k matrices. result First provable guarantees for ℓ0-Low Rank Approximation for k>1. Transformers use ReLUs to approximate softmax efficiently.
problem Analyzing resource usage in softmax transformer models.
method Translating ReLU approximation results to softmax attention mechanisms.
result Economic resource bounds for softmax attention mechanisms.
Approximating complex curves with simple parametric curves is widely used in CAGD, CG, and CNC. This paper presents an algorithm to compute a certified approximation to a given parametric space curve with cubic B-spline curves. By certified, we mean that the approximation can approximate the given curve to any given pr…
Recently, variational approximations such as the mean field approximation have received much interest. We extend the standard mean field method by using an approximating distribution that factorises into cluster potentials. This includes undirected graphs, directed acyclic graphs and junction trees. We derive generaliz…
Adaptive approximations improve variational inference for complex models.
problem Efficiently approximate marginal distributions and partition functions in complex probabilistic models.
method Two classes of adaptive approximations that include Bethe, tree-reweighted, and convex free energies.
result Proposed approximations automatically adapt to a given model and outperform existing methods.
Non-negative L1-approximating polynomials for Gaussian distributions are proven for certain classes of sets.
problem Existence of non-negative L1-approximating polynomials for Gaussian distributions. method Proving the existence of degree-k non-negative polynomials that approximate indicator functions of sets with Gaussian surface area in L1-norm. result Proves the existence of non-negative L1-approximating polynomials for certain classes of sets with Gaussian surface area. Variational boosting refines posterior approximations through iterative optimization.
problem Approximating intractable distributions with rich approximations.
method Iteratively solves optimization problems to refine variational approximations.
result Posterior inferences using variational boosting are more accurate and efficient.
Paper analyzes normal approximation for two-timescale stochastic algorithms, revealing interaction between fast and slow timescales.
problem Non-asymptotic bounds for accuracy of normal approximation in linear two-timescale stochastic approximation algorithms.
method Established bounds for normal approximation in terms of convex distance, focusing on last iterate and Polyak-Ruppert averaging.
result Normal approximation rate for the last iterate improves with increased timescale separation, while it decreases in the averaged setting.
One-pass algorithm finds small subset for ℓp subspace approximation with additive error.
problem Finding a small subset of data points for ℓp subspace approximation. method One-pass subset selection with additive approximation guarantee for p∈[1,∞). result First one-pass algorithm with additive error for ℓp subspace approximation. Paper introduces new approximations for lognormal sums, matching comonotonicity and moments.
problem Approximating sums of lognormal random variables accurately.
method Introduces new approximations based on weighted distribution theory, emphasizing comonotonicity and moment matching.
result Approximations perform better than classical methods, especially in the right tail of the distribution.