Characterizes functions in Carnot groups of step 2.
problem Understanding intrinsic Lipschitz functions in Carnot groups.
method Characterization via intrinsic distributional gradients.
result Characterization of locally intrinsic Lipschitz functions in Carnot groups of step 2.
The paper analyzes the intrinsic exploration terms in policy-gradient algorithms.
problem Exploration in policy-gradient algorithms and its impact on policy optimization.
method Numerical optimization criteria and stochastic gradient analysis.
result Exploration techniques improve policy optimization by smoothing the learning objective and modifying gradient estimates.
Innovative method solves nonconvex optimization on manifolds.
problem Nonconvex optimization problems on Riemannian manifolds.
method Intrinsic Riemannian proximal gradient method.
result Converges for nonconvex or nonembedded problems.
CNNs trained by gradient descent can learn intrinsic image rank robustly to background noises.
problem Understanding the intrinsic dimension of data in over-parameterized CNNs.
method Theoretical analysis and experiments on synthetic and real datasets.
result CNNs trained by gradient descent can learn the intrinsic dimension of clean images robustly to background noises.
New method for efficient sketching of gradients and Hessians.
problem Memory constraints in training machine learning models.
method A novel framework for scalable gradient and HVP sketching tailored for modern hardware.
result Theoretical guarantees and practical applications in training data attribution and Hessian spectrum analysis.
In many sequential decision making tasks, it is challenging to design reward functions that help an RL agent efficiently learn behavior that is considered good by the agent designer. A number of different formulations of the reward-design problem, or close variants thereof, have been proposed in the literature. In this…
Reinforcement learning for embodied agents is a challenging problem. The accumulated reward to be optimized is often a very rugged function, and gradient methods are impaired by many local optimizers. We demonstrate, in an experimental setting, that incorporating an intrinsic reward can smoothen the optimization landsc…
A new method solves convex optimization problems on manifolds efficiently.
problem Optimization on Hadamard manifolds with convex objectives.
method Intrinsic Riemannian proximal gradient method.
result Sublinear and linear convergence rates for convex and strongly convex problems, respectively.
The paper explores properties of projections and gradient methods in hyperbolic space forms.
problem Optimization problems in hyperbolic space forms.
method Intrinsic κ-projection and gradient projection methods.
result Every accumulation point of the sequence generated by the gradient projection method is a stationary point.
In this paper we provide a characterization of intrinsic Lipschitz graphs in the sub-Riemannian Heisenberg groups in terms of their distributional gradients. Moreover, we prove the equivalence of different notions of continuous weak solutions to the equation φ_y+ [φ^{2}/2]_t=w, where w is a bounded function depending o…
Paper analyzes dataset distillation for efficient encoding of task-relevant information.
problem Efficiently encoding task-relevant information from gradient-based learning of non-linear tasks.
method Theoretical analysis of dataset distillation applied to two-layer neural networks with gradient-based training.
result Low-dimensional structure of the problem is efficiently encoded into distilled data, reproducing a model with high generalization ability.
Conformal Autoencoders infer intrinsic dimensionality and impose invariance.
problem Detecting intrinsic dimensionality and imposing invariance in nonlinear manifold data.
method Imposing orthogonality conditions on latent variables to infer intrinsic dimensionality and build coordinate invariance.
result The method can infer intrinsic dimensionality and build coordinate invariance on submanifolds.
This work uses model uncertainty for efficient exploration in sparse reward environments.
problem Challenging exploration in sparse reward reinforcement learning environments.
method Implicit generative modeling approach to estimate Bayesian uncertainty of the agent's belief of the environment dynamics.
result Our implicit generative model consistently outperforms competing approaches in data efficiency for exploration.
A new method for SVGD reduces variance in high dimensions.
problem High-dimensional variance in SVGD.
method Grassmann Stein Variational Gradient Descent (GSVGD) projects onto arbitrary subspaces and uses coupled Grassmann-valued diffusion.
result GSVGD explores high-dimensional problems with intrinsic low-dimensional structure efficiently.
Improved sampling for diffusion models and log-concave distributions.
problem Efficient sampling for diffusion models and log-concave distributions.
method Algorithms for sampling with δ-error in polylog(1/δ) steps using accurate score estimates. result Exponential improvement in complexity over previous results.
This paper presents the Homeo-Heterostatic Value Gradients (HHVG) algorithm as a formal account on the constructive interplay between boredom and curiosity which gives rise to effective exploration and superior forward model learning. We envisaged actions as instrumental in agent's own epistemic disclosure. This motiva…
Proposes intrinsic methods to detect overfitting in models.
problem Detecting overfitting in models without relying on external test sets.
method Counterfactual Simulation (CFS) to analyze model flow through training data.
result CFS can separate models with different levels of overfit using only their logic circuit representations.
A new flow on null manifolds yields gradient estimates.
problem Understanding geometric properties of globally null manifolds.
method Introducing a degenerate Ricci-type flow in a Riemannian leaf of the manifold.
result Proved several new gradient estimates for the flow.
Gradient descent recovers low-rank matrices from corrupted measurements with double over-parameterization.
problem Robust recovery of low-rank matrices from grossly corrupted measurements.
method Gradient descent with discrepant learning rates for double over-parameterized models.
result Gradient descent with discrepant learning rates provably recovers the underlying matrix without prior knowledge on rank or sparsity.
We propose a conjugate gradient type optimization technique for the computation of the Karcher mean on the set of complex linear subspaces of fixed dimension, modeled by the so-called Grassmannian. The identification of the Grassmannian with Hermitian projection matrices allows an accessible introduction of the geometr…
Proposes ridge regression on Riemannian manifolds for time-series prediction.
problem Time-series prediction on Riemannian manifolds.
method Combines Riemannian least-squares fitting via Bézier curves, empirical covariance on manifolds, and Mahalanobis distance regularization.
result Significant error reduction in synthetic spherical experiments and hurricane forecasting.
FRA-Attack improves adversarial transferability for closed-source MLLMs by aligning visual focus across models.
problem Improving adversarial transferability for closed-source MLLMs, especially with high accuracy.
method Unified frequency-domain regularization approach: high-pass DCT objective for feature alignment and Frequency-domain Gradient Regularization (FGR) for gradient optimization.
result FRA-Attack achieves superior cross-model transferability, especially on GPT-5.4, Claude-Opus-4.6, and Gemini-3-flash.
Two new exploration methods for multi-agent systems improve team performance.
problem Exploration in transition-dependent multi-agent settings.
method EITI and EDTI, using mutual information and VoI to encourage coordinated exploration.
result Significant improvement in multi-agent performance through coordinated exploration.
Derivative formulas on measure spaces of Riemannian manifolds are characterized.
problem Characterizing derivatives in measure spaces on Riemannian manifolds.
method Introducing and characterizing derivatives in measure spaces for functions on the space of finite measures over a Riemannian manifold.
result Derivatives in measure spaces for functions on Riemannian manifolds are linked and calculated.
A new method for Bayesian inference in high dimensions using projected Stein variational gradient descent.
problem Bayesian inference challenges in high-dimensional data.
method Adapting Stein variational gradient descent to exploit intrinsic low dimensionality of data.
result pSVGD is more accurate and efficient than SVGD, especially in high-dimensional settings.
New score matching method estimates local intrinsic dimension efficiently.
problem Quantifying the local intrinsic dimension of complex data.
method Denoising score matching loss and equivalent implicit score matching loss.
result Denoising score matching loss is a highly competitive and scalable LID estimator.
Three training regimes found for scale-invariant neural networks on the sphere.
problem Training scale-invariant neural networks on the sphere with varying effective learning rate.
method Investigated three regimes of training: convergence, chaotic equilibrium, and divergence.
result Discovered three distinct training regimes with unique characteristics.
The paper classifies flows of SU(2)-structures on 4-manifolds.
problem Classifying flows of SU(2)-structures on 4-manifolds.
method Adapting a representation-theoretic method from Bryant for G2 geometry. result Explicit expressions for Ricci and self-dual Weyl curvature in terms of intrinsic torsion.
A new method for Bayesian inference tackles high-dimensional problems.
problem Bayesian inference in high-dimensional settings with kernel density estimation issues.
method Projected Wasserstein gradient descent (pWGD) method to overcome curse of dimensionality.
result pWGD method effectively addresses high-dimensional Bayesian inference problems.
Kernel-Gradient Drifting improves generative modeling for non-Euclidean data.
problem Challenges in generative modeling for non-Euclidean data.
method Replaces Euclidean displacement with kernel-induced directions, exposing score-based structure.
result Kernel-gradient drifting enables state-of-the-art one-step generation for non-Euclidean data.
FGBoost boosts gradient boosting for complex data.
problem Gradient boosting struggles with non-Euclidean data.
method Introduces FGBoost for geodesic metric spaces.
result FGBoost performs well on complex data.
A novel approach to computing barycenters on graph-supported probability measures.
problem Computing weighted averages of measures on graphs.
method Dynamic optimal transport formulation on the simplex, gradient descent on the probability simplex.
result Intrinsic gradient descent provides a coherent framework for synthesizing and analyzing measures on graphs.
This study uses continuous-time analysis to understand how momentum affects the optimisation of diagonal linear networks.
problem The effect of momentum on the optimisation trajectory of gradient descent.
method Leveraging a continuous-time approach to analyze momentum gradient descent with step size γ and momentum parameter β.
result Small values of λ help recover sparse solutions in overparametrised regression settings.
This research enhances ML models using gradient information from neural networks.
problem Improving the accuracy of machine learning models.
method Leveraging gradients extracted from neural networks to improve model performance.
result Gradient information can effectively enhance machine learning models with existing datasets.
Introduces intrinsic Hopf-Lax semigroup linking to intrinsic slope.
problem Understanding intrinsic Hopf-Lax semigroup and its relation to intrinsic slope.
method Introduces and proves the link between intrinsic Hopf-Lax semigroup and intrinsic slope.
result Intrinsic Hopf-Lax semigroup is a subsolution of Hamilton-Jacobi type equality.
On a sub-Riemannian manifold we define two type of Laplacians. The \emph{macroscopic Laplacian} Δω, as the divergence of the horizontal gradient, once a volume ω is fixed, and the \emph{microscopic Laplacian}, as the operator associated with a sequence of geodesic random walks. We consider a general class of rando…
Advances geometric structure flows, proving short-time existence and uniqueness for various flows.
problem Analyzing flows of geometric structures, focusing on non-isometric flows and specific subgroups.
method Developed algebra and compared two flows: negative gradient and Ricci-harmonic. Proved existence and uniqueness for Ricci-harmonic flow.
result Proved short-time existence and uniqueness for Ricci-harmonic flow for arbitrary lower-order torsion-quadratic terms.
Variational Bayesian neural nets combine the flexibility of deep learning with Bayesian uncertainty estimation. Unfortunately, there is a tradeoff between cheap but simple variational families (e.g.~fully factorized) or expensive and complicated inference procedures. We show that natural gradient ascent with adaptive w…
We develop Riemannian Stein Variational Gradient Descent (RSVGD), a Bayesian inference method that generalizes Stein Variational Gradient Descent (SVGD) to Riemann manifold. The benefits are two-folds: (i) for inference tasks in Euclidean spaces, RSVGD has the advantage over SVGD of utilizing information geometry, and …
In recent years, neural networks have demonstrated outstanding effectiveness in a large amount of applications.However, recent works have shown that neural networks are susceptible to adversarial examples, indicating possible flaws intrinsic to the network structures. To address this problem and improve the robustness …
Neural networks provide a rich class of high-dimensional, non-convex optimization problems. Despite their non-convexity, gradient-descent methods often successfully optimize these models. This has motivated a recent spur in research attempting to characterize properties of their loss surface that may explain such succe…
Adapts IG for better feature attributions and robustness.
problem Reliability concerns in feature attributions for deep learning models.
method Adaptation of path-based feature attribution to Riemannian geometry of data manifolds.
result IG along geodesics generates more intuitive and robust explanations.
New method optimizes 3D training data generation for deep networks.
problem Challenges in generating realistic 3D training data for deep networks.
method Hybrid gradient optimization of design decisions in graphics-based generation pipelines.
result Our approach outperforms prior methods in computational efficiency and performance.
Active subspaces on Riemannian manifolds generalize Euclidean principles.
problem Understanding how scalar-valued quantities change over Riemannian manifolds.
method Generalization of active subspaces from Euclidean to Riemannian spaces using parallel transport.
result The method provides a new way to study scalar-valued quantities on manifolds, differing from extrinsic approaches.
SDPG algorithm improves sample efficiency and reward in DRL for continuous action spaces.
problem Improving sample efficiency and reward in distributional reinforcement learning for continuous action spaces.
method SDPG algorithm models return distribution using samples via reparameterization technique.
result SDPG shows better sample efficiency and higher reward in OpenAI Gym environments.
Paper explores low-precision SGLD for neural networks, reducing costs without sacrificing performance.
problem Infeasibility of low-precision sampling in large-scale scenarios.
method Developed low-precision SGLD with quantization function and full-precision gradient accumulators.
result Low-precision SGLD achieves comparable performance to full-precision SGLD with only 8 bits.
Corrected CBOW performs similarly to Skip-gram.
problem CBOW embeddings underperform Skip-gram embeddings in word2vec.
method Fixed a bug in CBOW gradient update to improve performance.
result Corrected CBOW embeddings are competitive with Skip-gram on various tasks.
We give a geometrically intrinsic construction of a global time function for relatively compact diamond-shaped regions in arbitrary spacetimes. In the case of Minkowski spacetime, the flow of diffeomorphisms associated to a suitably normalized gradient of this time function becomes the conformal isotropy subgroup of th…