New method fine-tunes discrete diffusion models for RLHF tasks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Demon aligns diffusion models without retraining or backpropagation.
We study a nonparametric contextual bandit problem where the expected reward functions belong to a Hölder class with smoothness parameter . We show how this interpolates between two extremes that were previously studied in isolation: non-differentiable bandits (), where rate-optimal regret is achieved by run…
New method optimizes diffusion models without fine-tuning, integrating soft value functions.
In structured output prediction tasks, labeling ground-truth training output is often expensive. However, for many tasks, even when the true output is unknown, we can evaluate predictions using a scalar reward function, which may be easily assembled from human knowledge or non-differentiable pipelines. But searching th…
Task-specific scores are often used to optimize for and evaluate the performance of conditional text generation systems. However, such scores are non-differentiable and cannot be used in the standard supervised learning paradigm. Hence, policy gradient methods are used since the gradient can be computed without requiri…
We consider the task of Inverse Reinforcement Learning in Contextual Markov Decision Processes (MDPs). In this setting, contexts, which define the reward and transition kernel, are sampled from a distribution. In addition, although the reward is a function of the context, it is not provided to the agent. Instead, the a…
Attackers can significantly reduce team rewards in cooperative multi-agent reinforcement learning.
Paper proves autodiff systems are correct for non-differentiable functions.
New algorithms improve inference in non-differentiable models.
New method steers protein design towards desired properties.
We generalize stochastic smoothing for gradient estimation of non-differentiable functions.
The study connects Hilbert entropy to non-differentiability points of limit sets in flag spaces.
We present a new algorithm for stochastic variational inference that targets at models with non-differentiable densities. One of the key challenges in stochastic variational inference is to come up with a low-variance estimator of the gradient of a variational objective. We tackle the challenge by generalizing the repa…
Study shows AD for neural nets with machine-representable numbers can be incorrect.
Unified approach for sampling non-differentiable and heavy-tailed targets.
Study improves understanding of non-differentiable penalties in high-dimensional settings.
Study identifies conditions for proxy adjustment in confounded binary treatment outcomes.
Differentiable pipeline replaces non-differentiable CAE components for shape optimization.
This paper is an attempt at understanding the quantum-like dynamics of financial markets in terms of non-differentiable price-time continuum having fractal properties. The main steps of this development are the statistical scaling, the non-differentiability hypothesis, and the equations of motion entailed by this hypot…
We show that an analogue of the Ball-Box Theorem for step 2, completely non-integrable bundles from smooth sub-Riemannian geometry hold true for a class of non-differentiable tangent subbundles that satisfy a geometric condition. In the final section of the paper we give examples of such bundles and an application to d…
We study the dynamics of a particle in a space that is non-differentiable. Non-smooth geometrical objects have an inherently probabilistic nature and, consequently, introduce stochasticity in the motion of a body that lives in their realm. We use the mathematical concept of fiber bundle to characterize the multivalued …
A new method for optimizing models with categorical variables using diffusion.
Paper generalizes Hardy-Rogers maps for market equilibrium analysis in duopoly markets.
Proposes a method for inference in high-dimensional classification with non-differentiable surrogate losses.
In this work we show that Evolution Strategies (ES) are a viable method for learning non-differentiable parameters of large supervised models. ES are black-box optimization algorithms that estimate distributions of model parameters; however they have only been used for relatively small problems so far. We show that it …
We show that many machine learning goals, such as improved fairness metrics, can be expressed as constraints on the model's predictions, which we call rate constraints. We study the problem of training non-convex models subject to these rate constraints (or any non-convex and non-differentiable constraints). In the non…
Complex computer simulators are increasingly used across fields of science as generative models tying parameters of an underlying theory to experimental observations. Inference in this setup is often difficult, as simulators rarely admit a tractable density or likelihood function. We introduce Adversarial Variational O…
New hashing method improves document retrieval precision.
Generating novel graph structures that optimize given objectives while obeying some given underlying rules is fundamental for chemistry, biology and social science research. This is especially important in the task of molecular graph generation, whose goal is to discover novel molecules with desired properties such as …
Several tasks in machine learning are evaluated using non-differentiable metrics such as mean average precision or Spearman correlation. However, their non-differentiability prevents from using them as objective functions in a learning framework. Surrogate and relaxation methods exist but tend to be specific to a given…
This paper extends geometric study of neural networks to non-differentiable layers and random walks.
Study calculates slope gaps on polygon surfaces, finding non-unimodal distributions.
Bayesian optimization uses triangulation candidates for better performance.
This paper addresses the scalability challenge of architecture search by formulating the task in a differentiable manner. Unlike conventional approaches of applying evolution or reinforcement learning over a discrete and non-differentiable search space, our method is based on the continuous relaxation of the architectu…
Extends batch active learning to non-differentiable models.
Modern neural network training relies on piece-wise (sub-)differentiable functions in order to use backpropagation to update model parameters. In this work, we introduce a novel method to allow simple non-differentiable functions at intermediary layers of deep neural networks. We do so by training with a differentiable…
The paper improves ALO for -regularized models.
CADO optimizes heatmap-based solvers for cost minimization, overcoming performance limitations.
We propose an algorithm to calculate the exact solution for utility optimization problems on finite state spaces under a class of non-differentiable preferences. We prove that optimal strategies must lie on a discrete grid in the plane, and this allows us to reduce the dimension of the problem and define a very efficie…
We consider the learning of algorithmic tasks by mere observation of input-output pairs. Rather than studying this as a black-box discrete regression problem with no assumption whatsoever on the input-output mapping, we concentrate on tasks that are amenable to the principle of divide and conquer, and study what are it…
Program synthesis has emerged as a successful approach to the image parsing task. Most prior works rely on a two-step scheme involving supervised pretraining of a Seq2Seq model with synthetic programs followed by reinforcement learning (RL) for fine-tuning with real reference images. Fully unsupervised approaches promi…
We propose Stochastic Neural Architecture Search (SNAS), an economical end-to-end solution to Neural Architecture Search (NAS) that trains neural operation parameters and architecture distribution parameters in same round of back-propagation, while maintaining the completeness and differentiability of the NAS pipeline.…
Difference of convex (DC) functions cover a broad family of non-convex and possibly non-smooth and non-differentiable functions, and have wide applications in machine learning and statistics. Although deterministic algorithms for DC functions have been extensively studied, stochastic optimization that is more suitable …
VaR-CPO optimizes VaR-constrained RL problems with conservative policy updates.
Configuring deep Spiking Neural Networks (SNNs) is an exciting research avenue for low power spike event based computation. However, the spike generation function is non-differentiable and therefore not directly compatible with the standard error backpropagation algorithm. In this paper, we introduce a new general back…
We develop a new Low-level, First-order Probabilistic Programming Language (LF-PPL) suited for models containing a mix of continuous, discrete, and/or piecewise-continuous variables. The key success of this language and its compilation scheme is in its ability to automatically distinguish parameters the density functio…
Parametric quantile regressions are a useful tool for creating probabilistic energy forecasts. Nonetheless, since classical quantile regressions are trained using a non-differentiable cost function, their creation using complex data mining techniques (e.g., artificial neural networks) may be complicated. This article p…