A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
An unsupervised learning algorithm to cluster hyperspectral image (HSI) data is proposed that exploits spatially-regularized random walks. Markov diffusions are defined on the space of HSI spectra with transitions constrained to near spatial neighbors. The explicit incorporation of spatial regularity into the diffusion…
In-network distributed estimation of sparse parameter vectors via diffusion LMS strategies has been studied and investigated in recent years. In all the existing works, some convex regularization approach has been used at each node of the network in order to achieve an overall network performance superior to that of th…
We propose regularization strategies for learning discriminative models that are robust to in-class variations of the input data. We use the Wasserstein-2 geometry to capture semantically meaningful neighborhoods in the space of images, and define a corresponding input-dependent additive noise data augmentation model. …
Gradient guidance improves diffusion models for optimizing specific objectives.
problem Improving diffusion models for specific optimization tasks.
method Established a mathematical framework for gradient-guided diffusion, linking it to optimization theory. Developed a modified gradient guidance method and iteratively fine-tuned version.
result Gradient-guided diffusion models are essentially solutions to regularized optimization problems, preserving latent structure.
An active learning algorithm for the classification of high-dimensional images is proposed in which spatially-regularized nonlinear diffusion geometry is used to characterize cluster cores. The proposed method samples from estimated cluster cores in order to generate a small but potent set of training labels which prop…
This paper examines how Higher-Order Langevin Dynamics reduces memorization in diffusion models.
problem Memorization of training samples in diffusion models, violating copyright and privacy.
method Introduces Higher-Order Langevin Dynamics (HOLD) to regularize diffusion model trajectories.
result The dynamics of the data variable in HOLD are governed by a low-pass-filtered version of the learned score function, with smoothness increasing with model order.
The existence of complete Radner equilibria is established in an economy which parameters are driven by a diffusion process. Our results complement those in the literature. In particular, we work under essentially minimal regularity conditions and treat time-inhomogeneous case.
We establish Schauder a priori estimates and regularity for solutions to a class of boundary-degenerate elliptic linear second-order partial differential equations. Furthermore, given a smooth source function, we prove regularity of solutions up to the portion of the boundary where the operator is degenerate. Degenerat…
Diffusion models can generalize well even with coarse scores, thanks to the manifold hypothesis.
problem Understanding why diffusion models generate novel samples with coarse scores.
method Exploring the manifold hypothesis to explain diffusion model behavior.
result Diffusion models trained with coarse scores can achieve near-parametric rates of generalization, faster than estimating the full data distribution.
We study the problem of identifying the source of a diffusion spreading over a regular tree. When the degree of each node is at least three, we show that it is possible to construct confidence sets for the diffusion source with size independent of the number of infected nodes. Our estimators are motivated by analogous …
We consider a Markov process X, which is the solution of a stochastic differential equation driven by a Lévy process Z and an independent Wiener process W. Under some regularity conditions, including non-degeneracy of the diffusive and jump components of the process as well as smoothness of the Lévy density of $Z…
The purpose of this work is to develop and study a distributed strategy for Pareto optimization of an aggregate cost consisting of regularized risks. Each risk is modeled as the expectation of some loss function with unknown probability distribution while the regularizers are assumed deterministic, but are not required…
We present an adaptive regularization algorithm that can be effectively applied to the optimization problem in deep learning framework. Our regularization algorithm aims to take into account the fitness of data to the current state of model in the determination of regularity to achieve better generalization. The degree…
Recently, Mahoney and Orecchia demonstrated that popular diffusion-based procedures to compute a quick \emph{approximation} to the first nontrivial eigenvector of a data graph Laplacian \emph{exactly} solve certain regularized Semi-Definite Programs (SDPs). In this paper, we extend that result by providing a statistica…
The value function of an optimal stopping problem for jump diffusions is known to be a generalized solution of a variational inequality. Assuming that the diffusion component of the process is nondegenerate and a mild assumption on the singularity of the Lévy measure, this paper shows that the value function of this op…
We report statistical regularities of the opening and closing auctions of French equities, focusing on the diffusive properties of the indicative auction price. Two mechanisms are at play as the auction end time nears: the typical price change magnitude decreases, favoring underdiffusion, while the rate of these events…
In this paper we study the stochastic area swept by a regular time-homogeneous diffusion till a stopping time. This unifies some recent literature in this area. Through stochastic time change we establish a link between the stochastic area and the stopping time of another associated time-homogeneous diffusion. Then we …
There are several (mathematical) reasons why Dupire's formula fails in the non-diffusion setting. And yet, in practice, ad-hoc preconditioning of the option data works reasonably well. In this note we attempt to explain why. In particular, we propose a regularization procedure of the option data so that Dupire's local …
We introduce the {\it diffusion K-means} clustering method on Riemannian submanifolds, which maximizes the within-cluster connectedness based on the diffusion distance. The diffusion K-means constructs a random walk on the similarity graph with vertices as data points randomly sampled on the manifolds and edges as …