The paper analyzes fixed step-size SA schemes on Riemannian manifolds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Deep neural networks can approximate complex functions through repeated compositions of a fixed-size ReLU network.
New activation functions achieve arbitrary-accuracy Sobolev approximation by fixed-size neural networks.
Determinantal Point Processes (DPPs) are popular models for point processes with repulsion. They appear in numerous contexts, from physics to graph theory, and display appealing theoretical properties. On the more practical side of things, since DPPs tend to select sets of points that are some distance apart (repulsion…
Formula found for minimum ARI between clusterings of fixed sizes.
Methods for learning feature representations for Offline Handwritten Signature Verification have been successfully proposed in recent literature, using Deep Convolutional Neural Networks to learn representations from signature pixels. Such methods reported large performance improvements compared to handcrafted feature …
Ensembles of random-feature models can't outperform a single large model.
New neural network models learn symmetric functions of varying input sizes.
In this paper, we develop a new aligned vertex convolutional network model to learn multi-scale local-level vertex features for graph classification. Our idea is to transform the graphs of arbitrary sizes into fixed-sized aligned vertex grid structures, and define a new vertex convolution operation by adopting a set of…
In this paper, we present our method of using fixed-size ordinally forgetting encoding (FOFE) to solve the word sense disambiguation (WSD) problem. FOFE enables us to encode variable-length sequence of words into a theoretically unique fixed-size representation that can be fed into a feed forward neural network (FFNN),…
In this paper, we investigate the significance of choosing an appropriate tessellation strategy for a spatio-temporal taxi demand-supply modeling framework. Our study compares (i) the variable-sized polygon based Voronoi tessellation, and (ii) the fixed-sized grid based Geohash tessellation, using taxi demand-supply GP…
The softmax content-based attention mechanism has proven to be very beneficial in many applications of recurrent neural networks. Nevertheless it suffers from two major computational limitations. First, its computations for an attention lookup scale linearly in the size of the attended sequence. Second, it does not enc…
New TD algorithms stabilize RL tasks by reformulating updates into fixed point equations.
Transforms any test into anytime-valid with sample savings.
An investor with constant relative risk aversion trades a safe and several risky assets with constant investment opportunities. For a small fixed transaction cost, levied on each trade regardless of its size, we explicitly determine the leading-order corrections to the frictionless value function and optimal policy.
The paper explores properties of continuous actions on manifolds, proving bounds on subgroup size and fixed points.
Adaptive batch sizes improve active learning efficiency and flexibility.
New method improves model accuracy in Byzantine-robust distributed learning by optimizing batch size.
In Statistical Learning, the Vapnik-Chervonenkis (VC) dimension is an important combinatorial property of classifiers. To our knowledge, no theoretical results yet exist for the VC dimension of edited nearest-neighbour (1NN) classifiers with reference set of fixed size. Related theoretical results are scattered in the …
Resolving a conjecture of Abbe, Bandeira and Hall, the authors have recently shown that the semidefinite programming (SDP) relaxation of the maximum likelihood estimator achieves the sharp threshold for exactly recovering the community structure under the binary stochastic block model of two equal-sized clusters. The s…
A market fix serves as a benchmark for foreign exchange (FX) execution, and is employed by many institutional investors to establish an exact reference at which execution takes place. The currently most popular FX fix is the World Market Reuters (WM/R) 4pm fix. Execution at the WM/R 4pm fix is a service offered by FX b…
Accurate taxi demand-supply forecasting is a challenging application of ITS (Intelligent Transportation Systems), due to the complex spatial and temporal patterns. We investigate the impact of different spatial partitioning techniques on the prediction performance of an LSTM (Long Short-Term Memory) network, in the con…
Modeling functional data, this study uncovers the size-and-shape of functions under noisy observations.
Training deep neural networks with Stochastic Gradient Descent, or its variants, requires careful choice of both learning rate and batch size. While smaller batch sizes generally converge in fewer training epochs, larger batch sizes offer more parallelism and hence better computational efficiency. We have developed a n…
Minimal submanifolds in matrix spaces proven for specific ranks.
Empirical study on SGD hyperparameters and adversarial robustness.
A government has to finance a risk for its population. It shares the charges among the population with a fixed scale based on economic criteria. Various organisms have to collect and to redistribute fairly the subsidies. Under these conditions, when the size of the organisms is varied, the distribution's laws of the cr…
Classifies knots by lattice size, finding unknot ratios and crossing numbers.
This manuscript shows that AdaBoost and its immediate variants can produce approximate maximum margin classifiers simply by scaling step size choices with a fixed small constant. In this way, when the unscaled step size is an optimal choice, these results provide guarantees for Friedman's empirically successful "shrink…
We explore the impact of learning paradigms on training deep neural networks for the Travelling Salesman Problem. We design controlled experiments to train supervised learning (SL) and reinforcement learning (RL) models on fixed graph sizes up to 100 nodes, and evaluate them on variable sized graphs up to 500 nodes. Be…
Implicit Q-learning and SARSA adjust step-sizes automatically, improving stability and performance.
We present two instances, L-GAE and L-VGAE, of the variational graph auto-encoding family (VGAE) based on separating feature propagation operations from graph convolution layers typically found in graph learning methods to a single linear matrix computation made prior to input in standard auto-encoder architectures. Th…
For a fixed marked surface , we show that the problem of deciding whether or not a mapping class is reducible lies in . As usual this immediately gives an exponential time algorithm to decide whether or not a mapping class is reducible. To do this we use an (ideal) triangulation to obtain a coordinate s…
Stochastic Gradient Descent (SGD) is a central tool in machine learning. We prove that SGD converges to zero loss, even with a fixed (non-vanishing) learning rate - in the special case of homogeneous linear classifiers with smooth monotone loss functions, optimized on linearly separable data. Previous works assumed eit…
There is significant recent interest to parallelize deep learning algorithms in order to handle the enormous growth in data and model sizes. While most advances focus on model parallelization and engaging multiple computing agents via using a central parameter server, aspect of data parallelization along with decentral…
Improved batched SH algorithm maintains original performance.
Study on CEF discount in Bangladesh, finds size and maturity impact, turnover negative.
Scaling laws for neural language models reveal optimal model size and compute allocation.
We present a mathematical analysis of a non-convex energy landscape for robust subspace recovery. We prove that an underlying subspace is the only stationary point and local minimizer in a specified neighborhood under a deterministic condition on a dataset. If the deterministic condition is satisfied, we further show t…
Improved computational complexity in statistical models using second-order information.
Convolutional neural networks (CNNs) are commonly trained using a fixed spatial image size predetermined for a given model. Although trained on images of aspecific size, it is well established that CNNs can be used to evaluate a wide range of image sizes at test time, by adjusting the size of intermediate feature maps.…
When can reliable inference be drawn in the "Big Data" context? This paper presents a framework for answering this fundamental question in the context of correlation mining, with implications for general large scale inference. In large scale data applications like genomics, connectomics, and eco-informatics the dataset…
Polyak step size GD reaches final radius of convergence after log iterations.
We discuss our recent work [4] in which gravitational radiation was studied by evaluating the Wang-Yau quasi-local mass of surfaces of fixed size at the infinity of both axial and polar perturbations of the Schwarzschild spacetime, à la Chandrasekhar [1].
We provide a detailed study on the implicit bias of gradient descent when optimizing loss functions with strictly monotone tails, such as the logistic loss, over separable datasets. We look at two basic questions: (a) what are the conditions on the tail of the loss function under which gradient descent converges in the…
Proposes continuous convolution layers for flexible feature map resizing.
A new memory system handles non-stationary environments by self-sizing and retaining memories.
SGDm with fixed step-size diverges under covariate shift, similar to a parametric oscillator.