Paper explores subdifferential chain rules for matrix factorization and related machine learning models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Screening rules help identify active sets in optimization problems.
A new screening rule improves SLOPE efficiency for high-dimensional data.
Counterexamples show failure of uniform laws of large numbers for subdifferentials.
This work establishes uniform convergence of subdifferentials in stochastic optimization.
Study on tensor nuclear norm's decomposability and subdifferential.
We show that derivations of the differential structure of a subcartesian space satisfy the chain rule and have maximal integral curves.
The paper tackles finding stationary points in stochastic convex optimization problems.
We investigate the notion of H-subdifferential and H-normal map of a function on the Heisenberg group, based on its sub-Riemannian structure. In particular, a characterization of the convexity of a function is given via the nonemptiness of the H-subdifferential at every point.
The paper defines subdifferentials on Hadamard manifolds and identifies conditions for Fenchel conjugate equality.
In many healthcare settings, intuitive decision rules for risk stratification can help effective hospital resource allocation. This paper introduces a novel variant of decision tree algorithms that produces a chain of decisions, not a general tree. Our algorithm, -Carving Decision Chain (ACDC), sequentially carves o…
Although for neural networks with locally Lipschitz continuous activation functions the classical derivative exists almost everywhere, the standard chain rule is in general not applicable. We will consider a way of introducing a derivative for neural networks that admits a chain rule, which is both rigorous and easy to…
We derive generalization and excess risk bounds for neural nets using a family of complexity measures based on a multilevel relative entropy. The bounds are obtained by introducing the notion of generated hierarchical coverings of neural nets and by using the technique of chaining mutual information introduced in Asadi…
We prove that every function satisfies that the image of the set of critical points at which the function has Taylor expansions of order and non-empty subdifferentials of order is a Lebesgue-null set. As a by-product of our proof, for the proximal subdifferential $\partial_{…
DFR reduces the computational cost of sparse-group lasso and adaptive sparse-group lasso.
This paper presents a new methodology to compute first-order Greeks for barrier options under the framework of path-dependent payoff functions with European, Lookback, or Asian type and with time-dependent trigger levels. In particular, we develop chain rules for Wiener path integrals between two curves that arise in t…
CT compares two distributions using Bayes' theorem and chain rule.
New framework for studying eigenvalue functionals of metrics.
Mirror flows converge to a limiting flow with a convex potential.
New algorithms optimize spectral risk measures, improving interpolation between average and worst-case performance.
Method measures weight similarity in neural networks using normalization and statistical inference.
Characterizes infinite harmonic maps using 1-currents.
It is proved that the members of the Riccati hierarchy, the so-called Riccati chain equations, can be considered as particular cases of projective Riccati equations, which greatly simplifies the study of the Riccati hierarchy. This also allows us to characterize Riccati chain equations geometrically in terms of the pro…
The Widrow-Hoff rule simplifies language data simulation.
Random Intersection Chains selects important interactions from categorical features.
A conformal procedure improves CoT reasoning by aggregating reasoning paths and calibrating abstention rules.
Paper finds optimal selling rule for pairs trading with stock constraints.
We establish some perturbed minimization principles, and we develop a theory of subdifferential calculus, for functions defined on Riemannian manifolds. Then we apply these results to show existence and uniqueness of viscosity solutions to Hamilton-Jacobi equations defined on Riemannian manifolds.
Recently, based on the idea of randomizing space theory, random convex analysis has been being developed in order to deal with the corresponding problems in random environments such as analysis of conditional convex risk measures and the related variational problems and optimization problems. Random convex analysis is …
A new stopping rule based on E-values helps efficiently use sampling in Bayesian Deep Ensembles.
Study proves convergence of subgradients for optimal transport-based objectives.
This paper is concerned with an optimal stock selling rule under a Markov chain model. The objective is to find an optimal stopping time to sell the stock so as to maximize an expected return. Solutions to the associated variational inequalities are obtained. Closed-form solutions are given in terms of a set of thresho…
We propose a new class of convex penalty functions, called \emph{variational Gram functions} (VGFs), that can promote pairwise relations, such as orthogonality, among a set of vectors in a vector space. These functions can serve as regularizers in convex optimization problems arising from hierarchical classification, m…
Training activation quantized neural networks involves minimizing a piecewise constant function whose gradient vanishes almost everywhere, which is undesirable for the standard back-propagation or chain rule. An empirical way around this issue is to use a straight-through estimator (STE) (Bengio et al., 2013) in the ba…
We aim to predict and explain service failures in supply-chain networks, more precisely among last-mile pickup and delivery services to customers. We analyze a dataset of 500,000 services using (1) supervised classification with Random Forests, and (2) Association Rules. Our classifier reaches an average sensitivity of…
The concept of subdifferentiability is studied in the context of Finsler manifolds (modeled on a Banach space with a Lipschitz bump function). A class of Hamilton-Jacobi equations defined on Finsler manifolds is studied and several results related to the existence and uniqueness of viscosity solutions…
Backpropagation and the chain rule of derivatives have been prominent; however, the total derivative rule has not enjoyed the same amount of attention. In this work we show how the total derivative rule leads to an intuitive visual framework for creating gradient estimators on graphical models. In particular, previous …
SGD avoids critical points on weakly convex functions.
A one-to-one correspondence is drawn between law invariant risk measures and divergences, which we define as functionals of pairs of probability measures on arbitrary standard Borel spaces satisfying a few natural properties. Divergences include many classical information divergence measures, such as relative entropy a…
We propose a new statistical model for computational linguistics. Rather than trying to estimate directly the probability distribution of a random sentence of the language, we define a Markov chain on finite sets of sentences with many finite recurrent communicating classes and define our language model as the invarian…
In this paper we study integer multiplicity rectifiable currents carried by the subgradient (subdifferential) graphs of semi-convex functions on a -dimensional convex domain, and show a weak continuity theorem with respect to pointwise convergence for such currents. As an application, the -Hessian measures are ca…
Given a real-valued function defined on the Heisenberg group, we provide a definition of abstract convexity and Fenchel transform that takes into account the sub-Riemannian structure of the group. In our main result, we prove that, likewise the Euclidean case, a convex function can be characterized via its iterated Fen…
We study an extension of the classic stochastic multi-armed bandit problem which involves multiple plays and Markovian rewards in the rested bandits setting. In order to tackle this problem we consider an adaptive allocation rule which at each stage combines the information from the sample means of all the arms, with t…
New examples of sub-Riemannian structures satisfying Minimizing Sard conjecture found.
Generalized matrix-fractional (GMF) functions are a class of matrix support functions introduced by Burke and Hoheisel as a tool for unifying a range of seemingly divergent matrix optimization problems associated with inverse problems, regularization and learning. In this paper we dramatically simplify the support func…
In this paper we propose a novel approach for learning from data using rule based fuzzy inference systems where the model parameters are estimated using Bayesian inference and Markov Chain Monte Carlo (MCMC) techniques. We show the applicability of the method for regression and classification tasks using synthetic data…
New framework improves stochastic optimization for variational inference.
We show how risk measures originally defined in a model free framework in terms of acceptance sets and reference assets imply a meaningful underlying probability structure. Hereafter we construct a maximal domain of definition of the risk measure respecting the underlying ambiguity profile. We particularly emphasise li…