Directly estimates Fisher score for likelihood maximization.
problem Intractable likelihood functions with model simulations.
method Gradient-based optimization using local score matching and linear parameterization.
result Efficient approximation of Fisher score improves likelihood maximization.
Extends dimension reduction to data-driven settings without gradients.
problem Gradient-based dimension reduction limitations in data-driven settings.
method Score ratio matching framework, tailored parameterization, regularization, eigenvalue deflation.
result Outperforms standard score-matching for problems with low-dimensional structure.
New maximum score estimators using ReLU functions and deep neural networks.
problem Estimating parameters in models with sign restrictions.
method ReLU-based maximum score criterion and DNN architecture.
result RMS estimator achieves n−s/(2s+1) convergence rate and asymptotic normality. Study proposes a differentiable surrogate loss function for optimizing Fβ score in binary classification with imbalanced data.
problem Non-differentiability of Fβ score makes it unsuitable for optimization by gradient-based learning. method Investigated relationship between Fβ score and loss functions, proposed a differentiable surrogate loss function. result Gradient paths of the proposed surrogate Fβ loss function approximate the gradient paths of the Fβ score. New gradient-based method learns causal structures from data.
problem Challenging task in causal structure learning.
method Graph Autoencoder framework for gradient-based optimization.
result Significantly outperforms other gradient-based methods on large causal graphs.
Gradient Weighted Superpixels improve CNN interpretability without sacrificing speed.
problem Efficiency vs. interpretability trade-off in CNNs, especially for large input volumes.
method Gradient-based pixel scoring techniques applied to superpixels.
result Superpixels approximate LIME in a fraction of the time, improving interpretability.
A new method for discrete data normalizing flows using latent transformations.
problem Challenges in parameterizing bijective transformations for discrete data.
method Predict a distribution over latent transformations to make the marginal likelihood differentiable.
result Discrete-data normalizing flows can be trained using gradient-based learning with unbiased score function estimation.
ScoreStop uses gradient tests to stop gradient boosting early.
problem Overfitting in gradient boosted decision trees.
method ScoreStop uses a functional score test based on gradients to stop boosting.
result ScoreStop is competitive with loss-based early stopping methods.
Docking is an important tool in computational drug discovery that aims to predict the binding pose of a ligand to a target protein through a combination of pose scoring and optimization. A scoring function that is differentiable with respect to atom positions can be used for both scoring and gradient-based optimization…
Anomaly detection scores from VAE gradients improve tumor detection.
problem Improving anomaly detection in medical imaging.
method Using Variational Autoencoders to approximate anomaly ratings.
result Variance Autoencoder gradient-based ratings outperform other methods in tumor detection.
Gradient filters track moving parameters under noisy data and misspecification.
problem Tracking multidimensional time-varying parameters under noisy observations and model misspecification.
method Gradient-based filters update parameters using the gradient of a postulated objective function, evaluated at either the predicted or updated parameters.
result Novel sufficient conditions for exponential stability of the filtered parameter path, and finite-sample and asymptotic mean squared error bounds.
Score function estimators improve k-subset sampling efficiency.
problem Efficiently sampling k-subsets in machine learning tasks. method Revisit score function estimators, using discrete Fourier transform and control variates.
result Efficient and unbiased gradient estimates for k-subset sampling. Optimize black-box simulators with local generative models.
problem Optimizing non-differentiable, stochastic simulators with intractable likelihoods.
method Differentiable local surrogate models based on deep generative models.
result Local surrogates enable gradient-based optimization, faster than baseline methods.
EigenVI uses orthogonal function expansions for efficient variational inference.
problem Efficiently approximate complex distributions in variational inference.
method EigenVI constructs variational approximations using orthogonal function expansions, minimizing Fisher divergence.
result EigenVI provides more accurate approximations than existing methods for Gaussian BBVI.
A technique scales symbolic methods with gradients for neural model explanation.
problem Limited scalability of symbolic methods for large neural networks.
method Combines gradient-based methods with symbolic techniques using Integrated Gradients to focus on a subset of neurons.
result Produces sparser and higher saliency regions compared to gradient-based methods alone.
New robustness attacks improve evaluation of neural networks.
problem Difficulty in evaluating robustness of neural networks.
method Gradient-based adversarial attacks that are more reliable and efficient.
result Developed attacks are more reliable and efficient than existing methods.
Score matching offers efficient estimation for certain distributions.
problem Estimating probability distributions with intractable constants.
method Score matching as an alternative to maximum likelihood.
result Score matching is computationally and statistically efficient for certain distributions.
New method combines gradient optimization with constraint-based techniques for causal discovery.
problem Causal discovery from observational data, especially with small sample sizes.
method Differentiable d-separation scores using percolation theory and soft logic for gradient-based optimization of conditional independence constraints. result Empirical evaluations show robust performance in low-sample regimes, surpassing traditional methods.
Exact posterior score estimation for solving linear inverse problems
problem Solving linear inverse problems
method Derive the exact posterior score and use it as a denoising training objective
result EPS outperforms training-free and training-based baselines on various metrics
New method learns DAGs from data using neural networks.
problem Learning directed acyclic graphs from observational data.
method Adapting a continuous constrained optimization formulation to neural networks.
result New method outperforms existing continuous methods on most tasks.
A new method combines federated learning and logistic regression for better credit scoring.
problem Improving credit scoring models while protecting data privacy.
method Projected gradient-based vertical federated learning (FL-LRBC) for logistic regression.
result Significant improvement in AUC and KS statistics due to data enrichment.
Paper presents a novel gradient-based method for training models and hyperparameters simultaneously.
problem Achieving generalization in machine learning models.
method A novel gradient-based framework that trains parameters and hyperparameters simultaneously.
result Significantly smaller runtime compared to benchmark methods for equivalent prediction scores.
Kernel-Gradient Drifting improves generative modeling for non-Euclidean data.
problem Challenges in generative modeling for non-Euclidean data.
method Replaces Euclidean displacement with kernel-induced directions, exposing score-based structure.
result Kernel-gradient drifting enables state-of-the-art one-step generation for non-Euclidean data.
In training speech recognition systems, labeling audio clips can be expensive, and not all data is equally valuable. Active learning aims to label only the most informative samples to reduce cost. For speech recognition, confidence scores and other likelihood-based active learning methods have been shown to be effectiv…
New method learns CTBN structures from incomplete data.
problem Learning CTBN structures from incomplete data is computationally infeasible.
method Gradient-based optimization of mixture weights combined with variational method.
result Scalable structure learning of CTBNs from incomplete data.
Method identifies root causes of anomalies in causal processes.
problem Identifying root causes of anomalies in causal processes.
method Noisy functional causal model, Bayesian learning, gradient-based attribution.
result Proposes efficient method to compute anomaly attribution scores.
Transformer-based method for causal discovery with prior knowledge integration.
problem Complex nonlinear dependencies and spurious correlations in time series data.
method Multi-layer Transformer forecaster with gradient-based causal structure extraction and attention masking for prior knowledge integration.
result Significant improvement in causal discovery and causal lag estimation compared to state-of-the-art methods.
Efficiently approximates higher-order derivatives for generative models.
problem Expensive computation of higher-order derivatives in generative models.
method Rewrite SM objective in terms of directional derivatives and use finite difference for efficient approximation.
result Comparable results to gradient-based methods but significantly more computationally efficient.
New objective reduces bias and variance in reinforcement learning derivatives.
problem Estimating derivatives in reinforcement learning with unknown dynamics.
method Derives an objective function compatible with any advantage estimators, allowing trade-off between bias and variance.
result Demonstrates effectiveness in both theoretical and practical settings.
Method generates visual explanations for similarity models without classification.
problem Lack of visual explanations for similarity models trained without classification loss.
method Gradient-based visual attention using learned feature embeddings.
result Attention maps improve model performance and can be used as constraints.
LOGAN optimizes GAN training by improving adversarial dynamics.
problem Training GANs is challenging due to delicate adversarial dynamics and potential divergence.
method Integrates natural gradient-based latent optimisation into CS-GAN.
result Significant improvement in GAN training performance, achieving state-of-the-art results.
Many machine learning algorithms are vulnerable to almost imperceptible perturbations of their inputs. So far it was unclear how much risk adversarial perturbations carry for the safety of real-world machine learning applications because most methods used to generate such perturbations rely either on detailed model inf…
Gradient descent on DDPM objective learns Gaussian mixtures efficiently.
problem Learning Gaussian mixtures using gradient-based methods.
method Gradient descent on DDPM objective, connecting to EM and spectral methods.
result Gradient descent can efficiently recover Gaussian mixture parameters under certain conditions.
WaveGrad generates high-fidelity audio using gradient estimation.
problem Generating high-fidelity audio efficiently.
method Conditional model using score matching and diffusion models, iteratively refining a Gaussian white noise signal.
result WaveGrad can generate high-fidelity audio samples using as few as six iterations.
A new method detects changes in machine learning models over time.
problem Automatic monitoring of machine learning models trained on evolving data.
method Score-based statistical hypothesis test for change detection.
result The method can detect changes in any number of model components.
Paper introduces a new anomaly detection framework combining density estimation and deep learning.
problem Detecting anomalies in data with varying dimensions.
method Two versions: shallow approach using adaptive Fourier features and density matrices; deep approach using autoencoder.
result Both methods achieve comparable or superior performance compared to state-of-the-art methods.
Differentiable structure learning addresses DAGs with multiple global minimizers.
problem Identify the true DAG from global minimizers of acyclicity-constrained optimization problems.
method Carefully regularize the likelihood to identify the sparsest model in the Markov equivalence class.
result Regularization of the likelihood defines a score that identifies the sparsest model in general models and likelihoods.
Explaining the output of a deep network remains a challenge. In the case of an image classifier, one type of explanation is to identify pixels that strongly influence the final decision. A starting point for this strategy is the gradient of the class score function with respect to the input image. This gradient can be …
New method attacks GNNs with limited node access, increasing misclassification rate.
problem Attacking GNNs with limited node access and limited attack nodes.
method Generalized gradient-based attacks using importance scores derived from random walks.
result Proposed greedy procedure significantly increases misclassification rate.
LAWN normalizes logits to improve deep network adaptability and generalization.
problem Large logits and weights lead to overfitting in deep networks.
method Logit Attenuating Weight Normalization (LAWN) constrains weight norms in the final sub-network.
result LAWN improves generalization and adaptability of deep networks.
Kernel SVGD improves high-dimensional inference with noise adaptation.
problem Challenges in high-dimensional inference with SVGD.
method Noise Conditional Kernel SVGD (NCK-SVGD) with entropic regularization.
result NCK-SVGD produces samples comparable to GANs and SGLD on computer vision benchmarks.
Improved text generation using transferable rewards from related tasks.
problem Non-differentiable task-specific scores limit the use of policy gradient methods in text generation.
method Transferable Reward Learner that uses model-based rewards for sentence-level and phrase-level similarity.
result Improved performance on semantic evaluation measures in image captioning tasks.
PDNAS optimizes GNN architectures for diverse datasets.
problem Inadequate adaptability and combinatorial search space in GNNs.
method Dual architecture search (micro- and macro-architectures) with gradient-based optimization.
result PDNAS finds deeper GNNs with better performance on diverse datasets.
Review of gradient-based algorithms in statistical inference problems.
problem Understanding the dynamics of gradient-based algorithms in statistical inference.
method Insights from physics of glassy systems.
result Quantitative and qualitative understanding of algorithm performance.
Gradient-based methods improve understanding of deep learning survival models.
problem Limited interpretability of deep learning survival models hinders their adoption.
method Gradient-based explanation methods tailored to survival neural networks.
result Gradient-based methods capture feature effects and temporal dynamics.
Extract class-specific subnetworks from neural models for better understanding and improved explanations.
problem Understanding and explaining the complex behavior of deep neural networks.
method For each semantic class, extract a class-specific subnetwork with a compressed structure that maintains comparable performance.
result Extracted subnetworks improve explanation saliency and adversarial example detection.
MPM-ParVI uses particle sampling for variational inference.
problem Variational inference for complex probabilistic models.
method Material Point Method (MPM) for particle-based simulation.
result Deterministic sampling and inference for intractable densities.
Introduces new gradient-based methods for machine learning problems.
problem New challenges in machine learning due to decision-making and multi-agent problems.
method Gradient-based optimization and variational inequalities.
result Shifts focus from pattern recognition to decision-making and multi-agent problems.