Paper presents ABGD for efficient piecewise linear regression in high dimensions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Max-affine regression method converges linearly using GD and SGD.
New AMP algorithm estimates signals and latent variables in mixed regression models.
Paper presents an efficient algorithm for estimating Lipschitz functions from noisy data.
Single-head attention approximates any function under various norms.
New method uses DC functions for piecewise linear regression.
Paper proposes Sp-GD for sparse max-affine regression with theoretical guarantees.
We build a rigorous bridge between deep networks (DNs) and approximation theory via spline functions and operators. Our key result is that a large class of DNs can be written as a composition of max-affine spline operators (MASOs), which provide a powerful portal through which to view and analyze their inner workings. …
Max-affine regression refers to a model where the unknown regression function is modeled as a maximum of unknown affine functions for a fixed . This generalizes linear regression and (real) phase retrieval, and is closely related to convex regression. Working within a non-asymptotic framework, we study th…
This paper shows neural networks can solve complex graph problems efficiently.
Spectrahedral regression fits convex functions via a non-convex optimization problem.
Tropical Geometry and Mathematical Morphology share the same max-plus and min-plus semiring arithmetic and matrix algebra. In this chapter we summarize some of their main ideas and common (geometric and algebraic) structure, generalize and extend both of them using weighted lattices and a max- algebra with an ar…
We study the geometry of deep (neural) networks (DNs) with piecewise affine and convex nonlinearities. The layers of such DNs have been shown to be {\em max-affine spline operators} (MASOs) that partition their input space and apply a region-dependent affine mapping to their input to produce their output. We demonstrat…
Generative networks are analyzed using spline operators to understand their properties and limitations.
Nonlinearity is crucial to the performance of a deep (neural) network (DN). To date there has been little progress understanding the menagerie of available nonlinearities, but recently progress has been made on understanding the rôle played by piecewise affine and convex nonlinearities like the ReLU and absolute value …
We prove that functions defined on a lattice in a finite dimensional torus with bounded finite differences can be smoothly extended to the whole torus, and relate the bounds on the extension's derivatives with bounds on the original function's finite differences.
The homological and homotopical Dehn functions are different ways of measuring the difficulty of filling a closed curve inside a group or a space. The homological Dehn function measures fillings of cycles by chains, while the homotopical Dehn function measures fillings of curves by disks. Since the two definitions invo…
Deep networks adapt to function regularity and data distribution.
Paper examines power consumption in neural networks using various activation functions.
Scalarizing functions have been widely used to convert a multiobjective optimization problem into a single objective optimization problem. However, their use in solving (computationally) expensive multi- and many-objective optimization problems in Bayesian multiobjective optimization is scarce. Scalarizing functions ca…
Temporal difference learning explained through gradient splitting, improving convergence times.
Proves functions can have two indices differing by two.
Different optimizer choices lead to different financial model predictions.
Deep Neural Networks have been shown to be beneficial for a variety of tasks, in particular allowing for end-to-end learning and reducing the requirement for manual design decisions. However, still many parameters have to be chosen in advance, also raising the need to optimize them. One important, but often ignored sys…
EPIC quantifies reward differences without policy optimization.
We consider the problem of estimating the difference between two functional undirected graphical models with shared structures. In many applications, data are naturally regarded as high-dimensional random function vectors rather than multivariate scalars. For example, electroencephalography (EEG) data are more appropri…
In the article the necessary and sufficient conditions for a representation of Lipschitz function of two variables as a difference of two convex functions are formulated. An algorithm of this representation is given. The outcome of this algorithm is a sequence of pairs of convex functions that converge uniformly to a p…
We show that the travel time difference functions, measured on the boundary, determine a compact Riemannian manifold with smooth boundary up to Riemannian isometry, if boundary satisfies a certain visibility condition. This corresponds with the inverse microseismicity problem. The novelty of our paper is a new type of …
Regularizers change the geometric properties of loss functions in neural networks.
New filling functions for groups with coefficients show different asymptotic behavior.
New decompositions misattribute differences between populations, even when outcomes are identical.
Automatically discovers effective activation functions for deep learning.
FuDGE estimates differences between functional graphs in high-dimensional settings.
The approximation power of general feedforward neural networks with piecewise linear activation functions is investigated. First, lower bounds on the size of a network are established in terms of the approximation error and network depth and width. These bounds improve upon state-of-the-art bounds for certain classes o…
We define a generalized likelihood function based on uncertainty measures and show that maximizing such a likelihood function for different measures induces different types of classifiers. In the probabilistic framework, we obtain classifiers that optimize the cross-entropy function. In the possibilistic framework, we …
Temporal difference learning and Residual Gradient methods are the most widely used temporal difference based learning algorithms; however, it has been shown that none of their objective functions is optimal w.r.t approximating the true value function . Two novel algorithms are proposed to approximate the true value…
Unified framework for PE and TD methods in continuous time and space.
One of the main obstacles to broad application of reinforcement learning methods is the parameter sensitivity of our core learning algorithms. In many large-scale applications, online computation and function approximation represent key strategies in scaling up reinforcement learning algorithms. In this setting, we hav…
Quantum dilogarithm function proven from a linear difference equation.
We propose unsupervised representation learning and feature extraction from dendrograms. The commonly used Minimax distance measures correspond to building a dendrogram with single linkage criterion, with defining specific forms of a level function and a distance function over that. Therefore, we extend this method to …
We consider active, semi-supervised learning in an offline transductive setting. We show that a previously proposed error bound for active learning on undirected weighted graphs can be generalized by replacing graph cut with an arbitrary symmetric submodular function. Arbitrary non-symmetric submodular functions can be…
We apply supervised deep neural networks (DNNs) for pricing and calibration of both vanilla and exotic options under both diffusion and pure jump processes with and without stochastic volatility. We train our neural network models under different number of layers, neurons per layer, and various different activation fun…
The paper uses distance correlation for brain connectivity and a novel multi-task learning model for age prediction.
PBVFs generalize across policies using learned value functions.
Study examines maximal domains of radial harmonic functions across different curvature types.
Schizophrenia, a mental disorder that is characterized by abnormal social behavior and failure to distinguish one's own thoughts and ideas from reality, has been associated with structural abnormalities in the architecture of functional brain networks. Using various methods from network analysis, we examine the effect …
The paper analyzes the efficiency of gradient estimation methods in noisy function evaluations.
Manifolds uniquely identified by boundary distance differences.