Global minima found for multidimensional scaling with penalties.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper, we propose the idea of radial scaling in frequency domain and activation functions with compact support to produce a multi-scale DNN (MscaleDNN), which will have the multi-scale capability in approximating high frequency and high dimensional functions and speeding up the solution of high dimensional PDEs…
We establish decoupled functional CLTs for two-time-scale stochastic approximation.
A main goal of regression is to derive statistical conclusions on the conditional distribution of the output variable Y given the input values x. Two of the most important characteristics of a single distribution are location and scale. Support vector machines (SVMs) are well established to estimate location functions …
In theoretical analysis of deep learning, discovering which features of deep learning lead to good performance is an important task. In this paper, using the framework for analyzing the generalization error developed in Suzuki (2018), we derive a fast learning rate for deep neural networks with more general activation …
Introduces new performance measures using scaled utility functions.
We study the concept of coarse disjointness and large scale -to- functions. As a byproduct, we obtain an Ostrand-type characterization of asymptotic dimension for coarse structures. It is shown that properties like finite asymptotic dimension, coarse finitism, large scale weak paracompactness, ect. are all invari…
In this paper, we address the problem of measuring and analysing sensation, the subjective magnitude of one's experience. We do this in the context of the method of triads: the sensation of the stimulus is evaluated via relative judgments of the form: "Is stimulus S_i more similar to stimulus S_j or to stimulus S_k?". …
The paper explores neural scaling laws for deep operator networks, offering a theoretical foundation.
Q()-Learning improves Q-Learning by separating action-value functions into different time scales.
The optimal dividend problem by De Finetti (1957) has been recently generalized to the spectrally negative Lévy model where the implementation of optimal strategies draws upon the computation of scale functions and their derivatives. This paper proposes a phase-type fitting approximation of the optimal strategy. We con…
Study analyzes price response and spread impact in foreign exchange markets.
Paper finds sparse representation of functions using inverse scale space flow.
The most widely used activation functions in current deep feed-forward neural networks are rectified linear units (ReLU), and many alternatives have been successfully applied, as well. However, none of the alternatives have managed to consistently outperform the rest and there is no unified theory connecting properties…
Much recent work has concerned sparse approximations to speed up the Gaussian process regression from the unfavorable O(n3) scaling in computational time to O(nm2). Thus far, work has concentrated on models with one covariance function. However, in many practical situations additive models with multiple covariance func…
Several multiscale methods account for sub-grid scale features using coarse scale basis functions. For example, in the Multiscale Finite Volume method the coarse scale basis functions are obtained by solving a set of local problems over dual-grid cells. We introduce a data-driven approach for the estimation of these co…
This work explores variably scaled kernels to improve non-stationary Gaussian processes.
We determine when an arithmetic subgroup of a reductive group defined over a global function field is of type FP_\infty by comparing its large-scale geometry to the large-scale geometry of lattices in real semisimple Lie groups.
New method constructs potential functions for Kähler-Einstein metrics.
The height function of various surfaces decomposes into finite sums of scaled and translated versions of itself.
Locally adaptive clustering for tree delineation.
We prove that any proper, geodesic metric space whose Dehn function grows asymptotically like the Euclidean one has asymptotic cones which are non-positively curved in the sense of Alexandrov, thus are . This is new already in the setting of Riemannian manifolds and establishes in particular the borderlin…
Share price returns on different time scales can be well modelled by a superstatistical dynamics. Here we provide an investigation which type of superstatistics is most suitable to properly describe share price dynamics on various time scales. It is shown that while chi-square superstatistics works well on a time scale…
In this paper we combine two important extensions of ordinary least squares regression: regularization and optimal scaling. Optimal scaling (sometimes also called optimal scoring) has originally been developed for categorical data, and the process finds quantifications for the categories that are optimal for the regres…
Exact 1-Wasserstein distance between location-scale distributions derived, with privacy effects studied.
A homogeneous nilpotent Lie group has a scaling automorphism determined by a grading of its Lie algebra. Many proofs of upper bounds for the Dehn function of such a group depend on being able to fill curves with discs compatible with this grading; the area of such discs changes predictably under the scaling automorphis…
New algorithm reduces regret bounds for Bayesian optimization with unknown hyperparameters.
The dependency of the generalization error of neural networks on model and dataset size is of critical importance both in practice and for understanding the theory of neural networks. Nevertheless, the functional form of this dependency remains elusive. In this work, we present a functional form which approximates well…
New proof for global rigidity of vertex scaling on polyhedral surfaces.
Recently, self-normalizing neural networks (SNNs) have been proposed with the intention to avoid batch or weight normalization. The key step in SNNs is to properly scale the exponential linear unit (referred to as SELU) to inherently incorporate normalization based on central limit theory. SELU is a monotonically incre…
A simple model explains inference scaling in neural models.
Study rates of convergence for approximate solutions to linear ill-posed problems in Hilbert scales.
We discover scaling laws for kernel regression loss under various learning rate schedules.
Model shows feature learning can improve neural scaling laws for hard tasks.
We consider strictly stationary heavy tailed time series whose finite-dimensional exponent measures are concentrated on axes, and hence their extremal properties cannot be tackled using classical multivariate regular variation that is suitable for time series with extremal dependence. We recover relevant information ab…
Designing a covariance function that represents the underlying correlation is a crucial step in modeling complex natural systems, such as climate models. Geospatial datasets at a global scale usually suffer from non-stationarity and non-uniformly smooth spatial boundaries. A Gaussian process regression using a non-stat…
The study examines correlations of logarithms of integers at different scalings.
We construct a new map from a convex function to a distribution on its domain, with the property that this distribution is a multi-scale exploration of the function. We use this map to solve a decade-old open problem in adversarial bandit convex optimization by showing that the minimax regret for this problem is $\tild…
Differentially private log-location-scale regression models improve privacy in statistical analysis.
Inverse depth scaling found in LLMs due to similar layers averaging error.
Deep ResNets exhibit distinct scaling properties with depth, challenging neural ODE models.
We investigate properties that intuitively ought to be satisfied by graph clustering quality functions, that is, functions that assign a score to a clustering of a graph. Graph clustering, also known as network community detection, is often performed by optimizing such a function. Two axioms tailored for graph clusteri…
Pruned neural networks' error scales predictably with architecture and task.
We are motivated by large scale submodular optimization problems, where standard algorithms that treat the submodular functions in the \emph{value oracle model} do not scale. In this paper, we present a model called the \emph{precomputational complexity model}, along with a unifying memoization based framework, which l…
The scaling properties of oil price fluctuations are described as a non-stationary stochastic process realized by a time series of finite length. An original model is used to extract the scaling exponent of the fluctuation functions within a non-stationary process formulation. It is shown that, when returns are measure…
Estimates box dimension of fractal interpolation surfaces using oscillation vectors.
We introduce a framework to study the effective objectives at different time scales of financial market microstructure. The financial market can be regarded as a complex adaptive system, where purposeful agents collectively and simultaneously create and perceive their environment as they interact with it. It has been s…
QS-BO optimizes functions using only rank-based feedback.