This paper proposes a method to approximate non-Gaussian likelihoods in Gaussian Processes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Estimates change-points and graph structures in a time-varying Ising model.
Bayesian method detects change points and clusters in piece-wise constant signals.
Neural models can realize decision trees with parameter sharing and improved performance.
XOFM explains attribute effects in ordinal regression using piece-wise linear functions.
RUMBoost combines RUMs and deep learning for better choice modelling.
PyChEst detects changes in non-stationary time series without distributional assumptions.
We provide larger step-size restrictions for which gradient descent based algorithms (almost surely) avoid strict saddle points. In particular, consider a twice differentiable (non-convex) objective function whose gradient has Lipschitz constant L and whose Hessian is well-behaved. We prove that the probability of init…
This paper extends depth separation results to piece-wise oscillatory functions.
Study tackles inverse problems on low-dimensional manifolds, proving stability and proposing a reconstruction algorithm.
The fused lasso is analyzed for high-dimensional piecewise-constant regression coefficients.
Mode connectivity is a surprising phenomenon in the loss landscape of deep nets. Optima -- at least those discovered by gradient-based optimization -- turn out to be connected by simple paths on which the loss function is almost constant. Often, these paths can be chosen to be piece-wise linear, with as few as two segm…
Optimal order execution strategies for brokers under reference benchmarks.
In this paper, we investigate a transition from an elastica to a piece-wised elastica whose connected point defines the hinge angle ; we refer the piece-wised elastica -elastica or -elastica. The transition appears in the bending beam experiment; we compress elastic beams gradually and then suddenly du…
Efficiently infers switching nonlinear systems with collapsed amortized variational inference.
L*ReLU improves deep learning for fine-grained image classification.
We consider billiard ball motion in a convex domain of the Euclidean plane bounded by a piece-wise smooth curve influenced by the constant magnetic field. We show that if there exists a polynomial in velocities integral of the magnetic billiard flow then every smooth piece of the boundary must be algebraic and eith…
In this survey article, we review the relation between heat kernels and path integrals. In particular, we review recent results on the approximation of the Wiener measure on compact manifold by measures on (finite-dimensional) spaces of piece-wise geodesics.
AdaPID optimizes diffusion-based samplers by dynamically adjusting schedules.
DAMI uses interpretable regions to select informative samples for deep learning models.
CTR prediction in real-world business is a difficult machine learning problem with large scale nonlinear sparse data. In this paper, we introduce an industrial strength solution with model named Large Scale Piece-wise Linear Model (LS-PLM). We formulate the learning problem with and regularizers, leadin…
In this paper we analyze the asymptotic properties of l1 penalized maximum likelihood estimation of signals with piece-wise constant mean values and/or variances. The focus is on segmentation of a non-stationary time series with respect to changes in these model parameters. This change point detection and estimation pr…
We study the problem of estimating a temporally varying coefficient and varying structure (VCVS) graphical model underlying nonstationary time series data, such as social states of interacting individuals or microarray expression profiles of gene networks, as opposed to i.i.d. data from an invariant model widely consid…
This work simplifies adversarial attacks using neural networks, reducing computation and improving training convergence.
We consider the detection of activations over graphs under Gaussian noise, where signals are piece-wise constant over the graph. Despite the wide applicability of such a detection algorithm, there has been little success in the development of computationally feasible methods with proveable theoretical guarantees for ge…
The paper proves Sard's theorem for polynomial maps in infinite dimensions.
Considering Wirtinger's inequality for piece-wise equipartite functions we find a discrete version of this classical inequality. The main tool we use is the theorem of classification of isometries. Our approach provides a new elementary proof of Wirtinger's inequality that also allows to study the case of equality. Mor…
Assessing world-wide financial integration constitutes a recurrent challenge in macroeconometrics, often addressed by visual inspections searching for data patterns. Econophysics literature enables us to build complementary, data-driven measures of financial integration using graphs. The present contribution investigat…
Method to create rational Seifert surfaces for knots in Lens space.
It is shown that most of the well-known basic results for Sobolev-Slobodeckii and Bessel potential spaces, known to hold on bounded smooth domains in , continue to be valid on a wide class of Riemannian manifolds with singularities and boundary, provided suitable weights, which reflect the nature of the s…
Let S be a triangulated 2-sphere with fixed triangulation T. We apply the methods of thin position from knot theory to obtain a simple version of the three geodesics theorem for the 2-sphere [5]. In general these three geodesics may be unstable, corresponding, for example, to the three equators of an ellipsoid. Using a…
Given two points on a soup can or conical cup with lid, we find and classify all paths of minimal length connecting them. When the number of minimal paths is finite, there are at most four on a can and three on a cup. At worst, minimal paths are piece-wise smooth with three components, each of which is a classical geod…
Regularization is typically understood as improving generalization by altering the landscape of local extrema to which the model eventually converges. Deep neural networks (DNNs), however, challenge this view: We show that removing regularization after an initial transient period has little effect on generalization, ev…
This paper accelerates TV regularization algorithms by unrolling proximal gradient descent.
A new algorithm finds optimal solutions for constrained decision processes.
We investigate the functional determinant of the laplacian on piece-wise flat two-dimensional surfaces, with conical singularities in the interior and/or corners on the boundary. Our results extend earlier investigations of the determinants on smooth surfaces with smooth boundaries. The differences to the smooth case a…
In this survey, we present and compare different approaches to estimate Mutual Information (MI) from data to analyse general dependencies between variables of interest in a system. We demonstrate the performance difference of MI versus correlation analysis, which is only optimal in case of linear dependencies. First, w…
Tree ensemble kernels improve Bayesian optimization for mixed features and constraints.
The study examines generalization bounds for regression and classification tasks on adaptive input domains.
Method finds differential equations for integrable billiard tables.
New algorithm handles both decaying and non-decaying bandit problems.
In many applications we seek to maximize an expectation with respect to a distribution over discrete variables. Estimating gradients of such objectives with respect to the distribution parameters is a challenging problem. We analyze existing solutions including finite-difference (FD) estimators and continuous relaxatio…
We propose a strategy for approximating Pareto optimal sets based on the global analysis framework proposed by Smale (Dynamical systems, New York, 1973, pp. 531-544). The method highlights and exploits the underlying manifold structure of the Pareto sets, approximating Pareto optima by means of simplicial complexes. Th…
Sorting an array is a fundamental routine in machine learning, one that is used to compute rank-based statistics, cumulative distribution functions (CDFs), quantiles, or to select closest neighbors and labels. The sorting function is however piece-wise constant (the sorting permutation of a vector does not change if th…
Most of machine learning approaches have stemmed from the application of minimizing the mean squared distance principle, based on the computationally efficient quadratic optimization methods. However, when faced with high-dimensional and noisy data, the quadratic error functionals demonstrated many weaknesses including…
New methods for calculating curvature in graph theory.
The functional determinant of an elliptic operator with positive, discrete spectrum may be defined as , where , the zeta function, is the sum analytically continued to around the origin. In this paper is calculated for the Laplace operator with Dirichlet boundary…
New algorithm for active bipartite ranking with continuous distributions.