Paper proposes a method to minimize non-differentiable loss functions.
problem Minimizing non-differentiable and non-decomposable loss functions.
method Learn smooth relaxations of true losses through surrogate neural networks, then optimize jointly with the prediction model.
result Empirical results show the efficiency of learning surrogate losses.
Proposes a method for inference in high-dimensional classification with non-differentiable surrogate losses.
problem Lack of inference procedures for identifying driving factors in high-dimensional classification with non-differentiable surrogate losses.
method Kernel-smoothed decorrelated score and cross-fitted version for hypothesis tests and interval estimators.
result Valid and superior inference methods for high-dimensional classification with non-differentiable surrogate losses.
Paper proves autodiff systems are correct for non-differentiable functions.
problem Correctness of autodiff systems for non-differentiable functions in deep learning.
method Investigation of PAP functions and introduction of intensional derivatives.
result Intensional derivatives always exist and coincide with standard derivatives for almost all inputs.
SoDeep learns approximations of ranking metrics for deep learning tasks.
problem Non-differentiable metrics in machine learning tasks.
method Sorting deep (SoDeep) net trained to approximate sorting of scores.
result Competitive results on Cross-modal text-image retrieval, multi-label image classification, and visual memorability ranking tasks.
Paper proposes FONE for efficient distributed estimation and inference.
problem Efficient distributed estimation and inference for non-differentiable convex losses.
method Proposes a multi-round distributed estimation procedure using a First-Order Newton-type Estimator (FONE).
result FONE efficiently estimates Σ−1w for non-differentiable losses, facilitating inference. Study reveals properties of local minima in ReLU networks.
problem Understanding the loss landscape of neural networks.
method Theoretical analysis of one-hidden-layer ReLU networks.
result All differentiable local minima are global within certain regions.
New algorithms improve inference in non-differentiable models.
problem Inference and learning in latent variable models with non-differentiable densities.
method Proximal interacting particle Langevin algorithms (PIPLA).
result Nonasymptotic bounds and effectiveness demonstrated in various models.
Study particle dynamics in non-differentiable fractal spaces.
problem Understanding motion in non-smooth, probabilistic geometries.
method Use fiber bundle theory to characterize multivalued geodesic trajectories.
result Developed a hybrid theory combining surface and stochastic process theories.
New machine learning method uses algorithmic complexity for non-differentiable spaces.
problem Machine learning on non-differentiable spaces.
method Introduces complexity theory in machine learning, using algorithmic complexity for regression and classification.
result More generalizable and resilient to random attacks compared to traditional methods.
We generalize stochastic smoothing for gradient estimation of non-differentiable functions.
problem Gradient estimation for non-differentiable functions.
method Developed a general framework for relaxation and gradient estimation of non-differentiable black-box functions using stochastic smoothing with reduced assumptions.
result Empirically validated the effectiveness of variance reduction strategies for various non-differentiable tasks.
Reparameterization of variational auto-encoders with continuous random variables is an effective method for reducing the variance of their gradient estimates. In the discrete case, one can perform reparametrization using the Gumbel-Max trick, but the resulting objective relies on an argmax operation and is non-dif…
The study connects Hilbert entropy to non-differentiability points of limit sets in flag spaces.
problem Understanding non-differentiability points in limit sets of convex projective structures.
method Introduces hyperplane conicality for θ-Anosov representations and uses it to prove properties of boundary maps. result Hilbert entropy is linked to the Hausdorff dimension of non-differentiability points in flag spaces.
Consider the following class of learning schemes: \begin{equation} \label{eq:main-problem1} \hat{\boldsymbolβ} := \underset{\boldsymbolβ \in \mathcal{C}}{\arg\min} \;\sum_{j=1}^n \ell(\boldsymbol{x}_j^\top\boldsymbolβ; y_j) + λR(\boldsymbolβ), \qquad \qquad \qquad (1) \end{equation} where $\boldsymbol{x}_i \in \mathbb{…
ES for non-differentiable parameters scales to large models.
problem Learning non-differentiable parameters in large models.
method Hybrid approach combining ES for non-differentiable and gradient-based methods for differentiable parameters.
result Hybrid approach is competitive and allows training sparse models from the start.
Proposes a cross entropy loss for better ranking algorithms.
problem Improving the theoretical understanding and performance of ranking algorithms.
method Introduces a cross entropy-based loss function that is a convex bound on NDCG and consistent with NDCG.
result Empirically, the proposed method outperforms existing algorithms in quality and robustness.
We present a new algorithm for stochastic variational inference that targets at models with non-differentiable densities. One of the key challenges in stochastic variational inference is to come up with a low-variance estimator of the gradient of a variational objective. We tackle the challenge by generalizing the repa…
Introduces a differentiable approximation to the zero-one loss.
problem Incompatibility of zero-one loss with gradient-based optimization.
method Smooth projection onto hypersimplex through constrained optimization.
result Achieves significant improvements in generalization under large-batch training.
Paper proposes a new black-box attack approach to minimize visual distortion.
problem Constructing adversarial examples that minimize visual distortion in a black-box threat model.
method Learning the noise distribution of adversarial examples to approximate the gradient of a non-differentiable loss function.
result The proposed attack results in much lower visual distortion compared to state-of-the-art black-box attacks.
Linear-Core Surrogates combine fast optimization and statistical efficiency in classification and structured prediction.
problem The trade-off between smoothness and margin-based losses in classification and structured prediction.
method Linear-Core (LC) Surrogates, a family of convex loss functions that stitch a linear core to a smooth tail.
result LC Surrogates achieve fast linear consistency rates while maintaining differentiability and strict H-consistency bounds. Study shows AD for neural nets with machine-representable numbers can be incorrect.
problem Correctness of AD for neural nets with machine-representable numbers.
method Analyzed two sets of parameters: incorrect and non-differentiable. Proved bounds and conditions for AD correctness.
result AD can be incorrect for machine-representable numbers, but provides a Clarke subderivative on non-differentiable set.
Unified approach for sampling non-differentiable and heavy-tailed targets.
problem Sampling non-differentiable and heavy-tailed distributions using Langevin algorithms.
method Anchored Langevin dynamics, which modifies the Langevin diffusion with a smooth reference potential and multiplicative scaling.
result Non-asymptotic guarantees in the 2-Wasserstein distance to the target distribution.
Study improves understanding of non-differentiable penalties in high-dimensional settings.
problem Theoretical understanding of non-differentiable penalties like generalized LASSO and nuclear norm in high-dimensional settings.
method Proportional high-dimensional regime analysis with finite sample upper bounds on expected squared error.
result LO provides accurate estimation of out-of-sample risk in high-dimensional settings.
Study identifies conditions for proxy adjustment in confounded binary treatment outcomes.
problem Average causal effect estimation with a non-differentially mismeasured binary confounder.
method Identifies conditions for proxy adjustment in the presence of a non-differentially mismeasured binary confounder.
result Adjusting for a non-differentially mismeasured binary proxy can improve estimation of the average causal effect.
Develops LF-PPL for non-differentiable models with automatic boundary checks.
problem Handling non-differentiable models in probabilistic programming.
method Introduces LF-PPL with automatic boundary checks and a formalism ensuring measure zero discontinuities.
result Demonstrates efficient inference for non-differentiable models using DHMC.
Differentiable pipeline replaces non-differentiable CAE components for shape optimization.
problem Gradient-based optimization is limited by non-differentiable components in CAE workflows.
method Surrogate models replace non-differentiable pipeline components, enabling gradient-based optimization.
result Gradient-based shape optimization possible without differentiable solvers.
Study proposes a differentiable surrogate loss function for optimizing Fβ score in binary classification with imbalanced data.
problem Non-differentiability of Fβ score makes it unsuitable for optimization by gradient-based learning. method Investigated relationship between Fβ score and loss functions, proposed a differentiable surrogate loss function. result Gradient paths of the proposed surrogate Fβ loss function approximate the gradient paths of the Fβ score. This paper is an attempt at understanding the quantum-like dynamics of financial markets in terms of non-differentiable price-time continuum having fractal properties. The main steps of this development are the statistical scaling, the non-differentiability hypothesis, and the equations of motion entailed by this hypot…
We show that an analogue of the Ball-Box Theorem for step 2, completely non-integrable bundles from smooth sub-Riemannian geometry hold true for a class of non-differentiable tangent subbundles that satisfy a geometric condition. In the final section of the paper we give examples of such bundles and an application to d…
Smooth Contextual Bandits bridge two previously studied extremes of non-differentiable and parametric-response bandits.
problem Nonparametric contextual bandits with Hölder smoothness.
method Developed a novel algorithm that optimally balances between non-differentiable and parametric-response bandits.
result Proved the algorithm achieves rate-optimal regret for all smoothness settings.
Bayes-consistent disagreement discrepancy loss improves model robustness.
problem Distribution shift in real-world neural network deployment.
method Introducing a novel disagreement loss that is Bayes consistent.
result Proves existing surrogates for disagreement discrepancy are not Bayes consistent.
Despite being the standard loss function to train multi-class neural networks, the log-softmax has two potential limitations. First, it involves computations that scale linearly with the number of output classes, which can restrict the size of problems we are able to tackle with current hardware. Second, it remains unc…
Paper generalizes Hardy-Rogers maps for market equilibrium analysis in duopoly markets.
problem Existence and uniqueness of market equilibrium in duopoly markets with non-differentiable, nonlinear response functions.
method Coupled fixed points approach for generalized Hardy-Rogers maps.
result Enriched understanding of market equilibrium in duopoly markets with non-differentiable response functions.
TOPNet integrates task-based evaluation into machine learning models.
problem Non-differentiable task-based evaluation criteria in real-world applications.
method Task-Oriented Prediction Network (TOPNet) with learnable surrogate loss function.
result TOPNet significantly outperforms traditional and heuristic models in financial prediction tasks.
We show that many machine learning goals, such as improved fairness metrics, can be expressed as constraints on the model's predictions, which we call rate constraints. We study the problem of training non-convex models subject to these rate constraints (or any non-convex and non-differentiable constraints). In the non…
Neural painters learn to generate brushstrokes from a non-deterministic painting program.
problem Training an agent to generate realistic brushstrokes from a non-differentiable painting program.
method A differentiable neural painter model trained on brushstrokes, optimizing for human-like strokes and intrinsic style transfer.
result Direct optimization of brushstrokes can visualize ImageNet categories and generate ideal paintings.
Differentiable PF via entropy-regularized OT for better inference.
problem Non-differentiability of traditional PF resampling methods.
method Entropy-regularized optimal transport for differentiable resampling.
result Convergent differentiable PF method with improved gradient estimates.
New hashing method improves document retrieval precision.
problem Efficiently retrieving similar documents from large text databases.
method Pairwise supervised hashing with Bernoulli VAE and unbiased gradient estimator.
result Superior performance compared to existing methods.
The Support Vector Machine (SVM) has been used in a wide variety of classification problems. The original SVM uses the hinge loss function, which is non-differentiable and makes the problem difficult to solve in particular for regularized SVMs, such as with ℓ1-regularization. This paper considers the Huberized SV…
Complex computer simulators are increasingly used across fields of science as generative models tying parameters of an underlying theory to experimental observations. Inference in this setup is often difficult, as simulators rarely admit a tractable density or likelihood function. We introduce Adversarial Variational O…
Paper relaxes stability and generalization assumptions for SGD.
problem Stability and generalization for SGD under restrictive assumptions.
method Introduces on-average model stability and develops novel bounds.
result First-ever-known fast bounds in low-noise setting using stability approach.
Study binary activated deep neural networks using PAC-Bayesian theory.
problem Generalization bounds for binary activated deep neural networks.
method Developed an end-to-end framework and provided PAC-Bayesian generalization bounds.
result Nonvacuous PAC-Bayesian generalization bounds for binary activated deep neural networks.
UNAS combines DNAS and RL for efficient architecture search.
problem Discovering high accuracy or low latency neural architectures.
method Unified framework combining differentiable and reinforcement learning approaches.
result UNAS achieves state-of-the-art accuracy on CIFAR-10, CIFAR-100, and ImageNet datasets.
Introduces robust and decomposable AP for image retrieval.
problem Challenges in training deep neural networks with AP.
method Differentiable rank approximation and loss function design.
result ROADMAP outperforms AP approximation methods and deep models.
This paper extends geometric study of neural networks to non-differentiable layers and random walks.
problem Understanding the geometric properties of neural networks, especially those with non-differentiable activation functions.
method Singular Riemannian geometry approach to convolutional, residual, and recursive neural networks.
result Illustrated geometric findings with numerical experiments on image classification and thermodynamic problems.
New stochastic algorithms solve DC functions and non-convex problems efficiently.
problem Solving non-convex, non-smooth, and non-differentiable functions efficiently.
method Proposed new stochastic optimization algorithms for DC functions and non-convex problems.
result First non-asymptotic convergence for non-convex optimization with general non-convex non-differentiable regularizers.
Deep unfolding accelerates MCMC-based COP solvers.
problem Optimizing combinatorial problems with MCMC and gradient descent.
method Combines MCMC and gradient descent, trains step sizes, uses variance estimation for non-differentiable MCMC.
result Significantly accelerates convergence speed for COPs.
Differentiable resampling improves particle filter performance.
problem Non-differentiability of traditional resampling in particle filters.
method Introduced a neural network resampler (particle transformer) trained with a likelihood-based loss function.
result Learned resampling outperforms traditional methods on synthetic and real-world tasks.
Study calculates slope gaps on polygon surfaces, finding non-unimodal distributions.
problem Understanding the distribution of slope gaps on polygon surfaces.
method Explicit computation of slope gap distributions for 2n-gons, providing bounds on non-differentiability points.
result Slope gap distributions are not always unimodal, answering a question by Athreya.