Paper proves autodiff systems are correct for non-differentiable functions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithms improve inference in non-differentiable models.
We generalize stochastic smoothing for gradient estimation of non-differentiable functions.
The study connects Hilbert entropy to non-differentiability points of limit sets in flag spaces.
We present a new algorithm for stochastic variational inference that targets at models with non-differentiable densities. One of the key challenges in stochastic variational inference is to come up with a low-variance estimator of the gradient of a variational objective. We tackle the challenge by generalizing the repa…
Study shows AD for neural nets with machine-representable numbers can be incorrect.
Unified approach for sampling non-differentiable and heavy-tailed targets.
Study improves understanding of non-differentiable penalties in high-dimensional settings.
Study identifies conditions for proxy adjustment in confounded binary treatment outcomes.
Differentiable pipeline replaces non-differentiable CAE components for shape optimization.
This paper is an attempt at understanding the quantum-like dynamics of financial markets in terms of non-differentiable price-time continuum having fractal properties. The main steps of this development are the statistical scaling, the non-differentiability hypothesis, and the equations of motion entailed by this hypot…
We show that an analogue of the Ball-Box Theorem for step 2, completely non-integrable bundles from smooth sub-Riemannian geometry hold true for a class of non-differentiable tangent subbundles that satisfy a geometric condition. In the final section of the paper we give examples of such bundles and an application to d…
We study the dynamics of a particle in a space that is non-differentiable. Non-smooth geometrical objects have an inherently probabilistic nature and, consequently, introduce stochasticity in the motion of a body that lives in their realm. We use the mathematical concept of fiber bundle to characterize the multivalued …
Paper generalizes Hardy-Rogers maps for market equilibrium analysis in duopoly markets.
Proposes a method for inference in high-dimensional classification with non-differentiable surrogate losses.
In this work we show that Evolution Strategies (ES) are a viable method for learning non-differentiable parameters of large supervised models. ES are black-box optimization algorithms that estimate distributions of model parameters; however they have only been used for relatively small problems so far. We show that it …
We show that many machine learning goals, such as improved fairness metrics, can be expressed as constraints on the model's predictions, which we call rate constraints. We study the problem of training non-convex models subject to these rate constraints (or any non-convex and non-differentiable constraints). In the non…
Complex computer simulators are increasingly used across fields of science as generative models tying parameters of an underlying theory to experimental observations. Inference in this setup is often difficult, as simulators rarely admit a tractable density or likelihood function. We introduce Adversarial Variational O…
We study a nonparametric contextual bandit problem where the expected reward functions belong to a Hölder class with smoothness parameter . We show how this interpolates between two extremes that were previously studied in isolation: non-differentiable bandits (), where rate-optimal regret is achieved by run…
Several tasks in machine learning are evaluated using non-differentiable metrics such as mean average precision or Spearman correlation. However, their non-differentiability prevents from using them as objective functions in a learning framework. Surrogate and relaxation methods exist but tend to be specific to a given…
This paper extends geometric study of neural networks to non-differentiable layers and random walks.
Study calculates slope gaps on polygon surfaces, finding non-unimodal distributions.
Bayesian optimization uses triangulation candidates for better performance.
This paper addresses the scalability challenge of architecture search by formulating the task in a differentiable manner. Unlike conventional approaches of applying evolution or reinforcement learning over a discrete and non-differentiable search space, our method is based on the continuous relaxation of the architectu…
Extends batch active learning to non-differentiable models.
Modern neural network training relies on piece-wise (sub-)differentiable functions in order to use backpropagation to update model parameters. In this work, we introduce a novel method to allow simple non-differentiable functions at intermediary layers of deep neural networks. We do so by training with a differentiable…
The paper improves ALO for -regularized models.
New method fine-tunes discrete diffusion models for RLHF tasks.
We propose an algorithm to calculate the exact solution for utility optimization problems on finite state spaces under a class of non-differentiable preferences. We prove that optimal strategies must lie on a discrete grid in the plane, and this allows us to reduce the dimension of the problem and define a very efficie…
Difference of convex (DC) functions cover a broad family of non-convex and possibly non-smooth and non-differentiable functions, and have wide applications in machine learning and statistics. Although deterministic algorithms for DC functions have been extensively studied, stochastic optimization that is more suitable …
VaR-CPO optimizes VaR-constrained RL problems with conservative policy updates.
Configuring deep Spiking Neural Networks (SNNs) is an exciting research avenue for low power spike event based computation. However, the spike generation function is non-differentiable and therefore not directly compatible with the standard error backpropagation algorithm. In this paper, we introduce a new general back…
We develop a new Low-level, First-order Probabilistic Programming Language (LF-PPL) suited for models containing a mix of continuous, discrete, and/or piecewise-continuous variables. The key success of this language and its compilation scheme is in its ability to automatically distinguish parameters the density functio…
Parametric quantile regressions are a useful tool for creating probabilistic energy forecasts. Nonetheless, since classical quantile regressions are trained using a non-differentiable cost function, their creation using complex data mining techniques (e.g., artificial neural networks) may be complicated. This article p…
We discuss a general technique that can be used to form a differentiable bound on the optima of non-differentiable or discrete objective functions. We form a unified description of these methods and consider under which circumstances the bound is concave. In particular we consider two concrete applications of the metho…
A soft presentation of hyperbolic spaces, free of differential apparatus, is offered. Fifth Euclid's postulate in such spaces is overthrown and, among other things, it is proved that spheres (equipped with great-circle distances) and hyperbolic and Euclidean spaces are the only locally compact geodesic (i.e., convex) m…
The classical Tait-Kneser theorem states that the osculating circles of a smooth plane curve, free from curvature extrema, are pairwise disjoint. We prove a number of analogs of this theorem, e.g., for ovals of osculating cubics, osculating polynomials and trigonometric polynomials; in each case, we will obtain a non-d…
Current deep learning models are mostly build upon neural networks, i.e., multiple layers of parameterized differentiable nonlinear modules that can be trained by backpropagation. In this paper, we explore the possibility of building deep models based on non-differentiable modules. We conjecture that the mystery behind…
A new method for discrete data normalizing flows using latent transformations.
New algorithm solves non-convex, non-differentiable min-max games.
SANE improves exploration of noisy, multimodal functions by finding multiple optima.
Performing inference over simulators is generally intractable as their runtime means we cannot compute a marginal likelihood. We develop a likelihood-free inference method to infer parameters for a cardiac simulator, which replicates electrical flow through the heart to the body surface. We improve the fit of a state-o…
Deep weight factorization improves neural network training through smooth optimization of sparse penalties.
Demon aligns diffusion models without retraining or backpropagation.
TREX explains tree ensembles by identifying key training examples.
We propose a penalized likelihood method that simultaneously fits the multinomial logistic regression model and combines subsets of the response categories. The penalty is non differentiable when pairs of columns in the optimization variable are equal. This encourages pairwise equality of these columns in the estimator…
This paper proposes a new randomized strategy for adaptive MCMC using Bayesian optimization. This approach applies to non-differentiable objective functions and trades off exploration and exploitation to reduce the number of potentially costly objective function evaluations. We demonstrate the strategy in the complex s…
In recent years, constrained optimization has become increasingly relevant to the machine learning community, with applications including Neyman-Pearson classification, robust optimization, and fair machine learning. A natural approach to constrained optimization is to optimize the Lagrangian, but this is not guarantee…