SoDeep learns approximations of ranking metrics for deep learning tasks.
problem Non-differentiable metrics in machine learning tasks.
method Sorting deep (SoDeep) net trained to approximate sorting of scores.
result Competitive results on Cross-modal text-image retrieval, multi-label image classification, and visual memorability ranking tasks.
The study connects Hilbert entropy to non-differentiability points of limit sets in flag spaces.
problem Understanding non-differentiability points in limit sets of convex projective structures.
method Introduces hyperplane conicality for θ-Anosov representations and uses it to prove properties of boundary maps. result Hilbert entropy is linked to the Hausdorff dimension of non-differentiability points in flag spaces.
We show that many machine learning goals, such as improved fairness metrics, can be expressed as constraints on the model's predictions, which we call rate constraints. We study the problem of training non-convex models subject to these rate constraints (or any non-convex and non-differentiable constraints). In the non…
Differentiable pipeline replaces non-differentiable CAE components for shape optimization.
problem Gradient-based optimization is limited by non-differentiable components in CAE workflows.
method Surrogate models replace non-differentiable pipeline components, enabling gradient-based optimization.
result Gradient-based shape optimization possible without differentiable solvers.
Paper proves autodiff systems are correct for non-differentiable functions.
problem Correctness of autodiff systems for non-differentiable functions in deep learning.
method Investigation of PAP functions and introduction of intensional derivatives.
result Intensional derivatives always exist and coincide with standard derivatives for almost all inputs.
This paper extends geometric study of neural networks to non-differentiable layers and random walks.
problem Understanding the geometric properties of neural networks, especially those with non-differentiable activation functions.
method Singular Riemannian geometry approach to convolutional, residual, and recursive neural networks.
result Illustrated geometric findings with numerical experiments on image classification and thermodynamic problems.
New algorithms improve inference in non-differentiable models.
problem Inference and learning in latent variable models with non-differentiable densities.
method Proximal interacting particle Langevin algorithms (PIPLA).
result Nonasymptotic bounds and effectiveness demonstrated in various models.
Study particle dynamics in non-differentiable fractal spaces.
problem Understanding motion in non-smooth, probabilistic geometries.
method Use fiber bundle theory to characterize multivalued geodesic trajectories.
result Developed a hybrid theory combining surface and stochastic process theories.
A soft presentation of hyperbolic spaces, free of differential apparatus, is offered. Fifth Euclid's postulate in such spaces is overthrown and, among other things, it is proved that spheres (equipped with great-circle distances) and hyperbolic and Euclidean spaces are the only locally compact geodesic (i.e., convex) m…
We generalize stochastic smoothing for gradient estimation of non-differentiable functions.
problem Gradient estimation for non-differentiable functions.
method Developed a general framework for relaxation and gradient estimation of non-differentiable black-box functions using stochastic smoothing with reduced assumptions.
result Empirically validated the effectiveness of variance reduction strategies for various non-differentiable tasks.
The paper improves ALO for ℓ1-regularized models.
problem Estimating out-of-sample error for ℓ1-regularized models. method Developed a novel theory for ℓ1-regularized problems, bounding ALO error. result For ℓ1-regularized problems, ALO error goes to zero as p goes to infinity. ES for non-differentiable parameters scales to large models.
problem Learning non-differentiable parameters in large models.
method Hybrid approach combining ES for non-differentiable and gradient-based methods for differentiable parameters.
result Hybrid approach is competitive and allows training sparse models from the start.
We present a new algorithm for stochastic variational inference that targets at models with non-differentiable densities. One of the key challenges in stochastic variational inference is to come up with a low-variance estimator of the gradient of a variational objective. We tackle the challenge by generalizing the repa…
Study shows AD for neural nets with machine-representable numbers can be incorrect.
problem Correctness of AD for neural nets with machine-representable numbers.
method Analyzed two sets of parameters: incorrect and non-differentiable. Proved bounds and conditions for AD correctness.
result AD can be incorrect for machine-representable numbers, but provides a Clarke subderivative on non-differentiable set.
Unified approach for sampling non-differentiable and heavy-tailed targets.
problem Sampling non-differentiable and heavy-tailed distributions using Langevin algorithms.
method Anchored Langevin dynamics, which modifies the Langevin diffusion with a smooth reference potential and multiplicative scaling.
result Non-asymptotic guarantees in the 2-Wasserstein distance to the target distribution.
Study improves understanding of non-differentiable penalties in high-dimensional settings.
problem Theoretical understanding of non-differentiable penalties like generalized LASSO and nuclear norm in high-dimensional settings.
method Proportional high-dimensional regime analysis with finite sample upper bounds on expected squared error.
result LO provides accurate estimation of out-of-sample risk in high-dimensional settings.
Study identifies conditions for proxy adjustment in confounded binary treatment outcomes.
problem Average causal effect estimation with a non-differentially mismeasured binary confounder.
method Identifies conditions for proxy adjustment in the presence of a non-differentially mismeasured binary confounder.
result Adjusting for a non-differentially mismeasured binary proxy can improve estimation of the average causal effect.
Develops LF-PPL for non-differentiable models with automatic boundary checks.
problem Handling non-differentiable models in probabilistic programming.
method Introduces LF-PPL with automatic boundary checks and a formalism ensuring measure zero discontinuities.
result Demonstrates efficient inference for non-differentiable models using DHMC.
This paper is an attempt at understanding the quantum-like dynamics of financial markets in terms of non-differentiable price-time continuum having fractal properties. The main steps of this development are the statistical scaling, the non-differentiability hypothesis, and the equations of motion entailed by this hypot…
We show that an analogue of the Ball-Box Theorem for step 2, completely non-integrable bundles from smooth sub-Riemannian geometry hold true for a class of non-differentiable tangent subbundles that satisfy a geometric condition. In the final section of the paper we give examples of such bundles and an application to d…
Smooth Contextual Bandits bridge two previously studied extremes of non-differentiable and parametric-response bandits.
problem Nonparametric contextual bandits with Hölder smoothness.
method Developed a novel algorithm that optimally balances between non-differentiable and parametric-response bandits.
result Proved the algorithm achieves rate-optimal regret for all smoothness settings.
Paper generalizes Hardy-Rogers maps for market equilibrium analysis in duopoly markets.
problem Existence and uniqueness of market equilibrium in duopoly markets with non-differentiable, nonlinear response functions.
method Coupled fixed points approach for generalized Hardy-Rogers maps.
result Enriched understanding of market equilibrium in duopoly markets with non-differentiable response functions.
Proposes a method for inference in high-dimensional classification with non-differentiable surrogate losses.
problem Lack of inference procedures for identifying driving factors in high-dimensional classification with non-differentiable surrogate losses.
method Kernel-smoothed decorrelated score and cross-fitted version for hypothesis tests and interval estimators.
result Valid and superior inference methods for high-dimensional classification with non-differentiable surrogate losses.
Paper proposes a method to minimize non-differentiable loss functions.
problem Minimizing non-differentiable and non-decomposable loss functions.
method Learn smooth relaxations of true losses through surrogate neural networks, then optimize jointly with the prediction model.
result Empirical results show the efficiency of learning surrogate losses.
We optimize rank-based metrics using blackbox differentiation.
problem Challenges in directly optimizing rank-based metrics due to their non-differentiable and non-decomposable nature.
method Efficient, theoretically sound, and general method for differentiating rank-based metrics with mini-batch gradient descent.
result Competitive performance on standard image retrieval datasets and improved performance on object detectors.
Complex computer simulators are increasingly used across fields of science as generative models tying parameters of an underlying theory to experimental observations. Inference in this setup is often difficult, as simulators rarely admit a tractable density or likelihood function. We introduce Adversarial Variational O…
Connectionist temporal classification (CTC) is widely used for maximum likelihood learning in end-to-end speech recognition models. However, there is usually a disparity between the negative maximum likelihood and the performance metric used in speech recognition, e.g., word error rate (WER). This results in a mismatch…
In sequence generation task, many works use policy gradient for model optimization to tackle the intractable backpropagation issue when maximizing the non-differentiable evaluation metrics or fooling the discriminator in adversarial learning. In this paper, we replace policy gradient with proximal policy optimization (…
Study calculates slope gaps on polygon surfaces, finding non-unimodal distributions.
problem Understanding the distribution of slope gaps on polygon surfaces.
method Explicit computation of slope gap distributions for 2n-gons, providing bounds on non-differentiability points.
result Slope gap distributions are not always unimodal, answering a question by Athreya.
ID-ExpO fine-tunes neural networks for more faithful explanations.
problem Improving the faithfulness of explanations for complex machine learning models.
method Differentiable insertion/deletion metric-aware regularizers for optimization.
result Fine-tuned predictors produce more faithful explanations.
Bayesian optimization uses triangulation candidates for better performance.
problem Non-convex and multi-modal optimization challenges in Bayesian optimization.
method Proposes using Delaunay triangulation candidates for discrete search over continuous optimization.
result Triangulation candidates outperform numerically optimized and random alternatives.
This paper addresses the scalability challenge of architecture search by formulating the task in a differentiable manner. Unlike conventional approaches of applying evolution or reinforcement learning over a discrete and non-differentiable search space, our method is based on the continuous relaxation of the architectu…
Extends batch active learning to non-differentiable models.
problem Efficiently training machine learning models on large, initially unlabelled datasets.
method Black-box batch active learning for regression tasks that relies solely on model predictions.
result Achieves strong performance on regression datasets compared to white-box approaches for deep learning models.
New machine learning method uses algorithmic complexity for non-differentiable spaces.
problem Machine learning on non-differentiable spaces.
method Introduces complexity theory in machine learning, using algorithmic complexity for regression and classification.
result More generalizable and resilient to random attacks compared to traditional methods.
New method fine-tunes discrete diffusion models for RLHF tasks.
problem Fine-tuning discrete diffusion models with policy gradient methods is challenging.
method Proposed SEPO algorithm for efficient fine-tuning over non-differentiable rewards.
result Numerical experiments show scalability and efficiency of SEPO.
We propose an algorithm to calculate the exact solution for utility optimization problems on finite state spaces under a class of non-differentiable preferences. We prove that optimal strategies must lie on a discrete grid in the plane, and this allows us to reduce the dimension of the problem and define a very efficie…
Improves discrete latent representations using differentiable approximation bridges.
problem Improving discrete latent representations in neural networks.
method Training with a differentiable approximation bridge (DAB) neural network.
result Improves state-of-the-art performance in various domains.
Proposes a cross entropy loss for better ranking algorithms.
problem Improving the theoretical understanding and performance of ranking algorithms.
method Introduces a cross entropy-based loss function that is a convex bound on NDCG and consistent with NDCG.
result Empirically, the proposed method outperforms existing algorithms in quality and robustness.
Difference of convex (DC) functions cover a broad family of non-convex and possibly non-smooth and non-differentiable functions, and have wide applications in machine learning and statistics. Although deterministic algorithms for DC functions have been extensively studied, stochastic optimization that is more suitable …
In this paper, we propose a deep learning approach to tackle the automatic summarization tasks by incorporating topic information into the convolutional sequence-to-sequence (ConvS2S) model and using self-critical sequence training (SCST) for optimization. Through jointly attending to topics and word-level alignment, o…
VaR-CPO optimizes VaR-constrained RL problems with conservative policy updates.
problem Optimizing VaR-constrained reinforcement learning problems.
method Combines Cantelli's inequality and trust-region framework for efficient and conservative optimization.
result Achieves zero constraint violations during training in feasible environments.
Configuring deep Spiking Neural Networks (SNNs) is an exciting research avenue for low power spike event based computation. However, the spike generation function is non-differentiable and therefore not directly compatible with the standard error backpropagation algorithm. In this paper, we introduce a new general back…
This work combines GANs and A3C for high-resolution image compression.
problem Image compression for high-resolution images without loss of quality.
method Hybrid approach using GANs and A3C for end-to-end learning.
result Improves PSNR for high-resolution images through end-to-end learning.
New method optimizes collaborative filtering for better ranking metrics.
problem Improving recommendation quality metrics like top-N ranking.
method Actor-critic reinforcement learning to directly optimize ranking metrics.
result The method outperforms state-of-the-art baselines on real-world datasets.
Optimize black-box simulators with local generative models.
problem Optimizing non-differentiable, stochastic simulators with intractable likelihoods.
method Differentiable local surrogate models based on deep generative models.
result Local surrogates enable gradient-based optimization, faster than baseline methods.
We discuss a general technique that can be used to form a differentiable bound on the optima of non-differentiable or discrete objective functions. We form a unified description of these methods and consider under which circumstances the bound is concave. In particular we consider two concrete applications of the metho…
Current deep learning models are mostly build upon neural networks, i.e., multiple layers of parameterized differentiable nonlinear modules that can be trained by backpropagation. In this paper, we explore the possibility of building deep models based on non-differentiable modules. We conjecture that the mystery behind…
The classical Tait-Kneser theorem states that the osculating circles of a smooth plane curve, free from curvature extrema, are pairwise disjoint. We prove a number of analogs of this theorem, e.g., for ovals of osculating cubics, osculating polynomials and trigonometric polynomials; in each case, we will obtain a non-d…