Study predicts risk of true-lumen narrowing after ATAAD surgery using CT data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study of infinitely deep but narrow neural networks using NTK theory.
We show that deep narrow Boltzmann machines are universal approximators of probability distributions on the activities of their visible units, provided they have sufficiently many hidden layers, each containing the same number of units as the visible layer. We show that, within certain parameter domains, deep Boltzmann…
New approach finds minimum width for deep, narrow MLPs.
We investigate the macroeconomic consequences of narrow banking in the context of stock-flow consistent models. We begin with an extension of the Goodwin-Keen model incorporating time deposits, government bills, cash, and central bank reserves to the base model with loans and demand deposits and use it to describe a fr…
Study on ion travel time on curved surfaces.
We prove the equidistribution of (weighted) periodic orbits of the geodesic ow on noncompact negatively curved manifolds toward equilibrium states in the narrow topology, i.e. in the dual of bounded continuous functions. We deduce an exact asymptotic counting for periodic orbits (weighted or not), which was previously …
Emergent misalignment is influenced by training dynamics, model priors, and data.
Paper calculates topological complexity of robot movement in narrow aisles.
Wide networks are often believed to have a nice optimization landscape, but what rigorous results can we prove? To understand the benefit of width, it is important to identify the difference between wide and narrow networks. In this work, we prove that from narrow to wide networks, there is a phase transition from havi…
We show that for neural network functions that have width less or equal to the input dimension all connected components of decision regions are unbounded. The result holds for continuous and strictly monotonic activation functions as well as for the ReLU activation function. This complements recent results on approxima…
Recent theoretical work has demonstrated that deep neural networks have superior performance over shallow networks, but their training is more difficult, e.g., they suffer from the vanishing gradient problem. This problem can be typically resolved by the rectified linear unit (ReLU) activation. However, here we show th…
Large learning rates improve generalization, but optimal ranges are narrower than previously thought.
AI system narrows human decision options for better outcomes.
Embedding principle explains loss landscape of deep neural networks.
Improved bounds on neural network expressivity.
Search is a prominent channel for discovering products on an e-commerce platform. Ranking products retrieved from search becomes crucial to address customer's need and optimize for business metrics. While learning to Rank (LETOR) models have been extensively studied and have demonstrated efficacy in the context of web …
We study Weil-Petersson (WP) geodesics with narrow end invariant and develop techniques to control length-functions and twist parameters along them and prescribe their itinerary in the moduli space of Riemann surfaces. This class of geodesics is rich enough to provide for examples of closed WP geodesics in the thin par…
As a testament to their success, the theory of random forests has long been outpaced by their application in practice. In this paper, we take a step towards narrowing this gap by providing a consistency result for online random forests.
We study weak solutions to degenerate quasilinear elliptic equations, involving first order terms, in unbounded tubular domains. In particular we show that, under suitable hypotheses, the weak comparison principle holds if the domain is narrow enough.
Study proves deep narrow RNNs can approximate any function, with minimum width independent of data length.
Bayesian Quadrature improves ensembling for neural networks with dispersed likelihood peaks.
This paper focuses on a traditional relation extraction task in the context of limited annotated data and a narrow knowledge domain. We explore this task with a clinical corpus consisting of 200 breast cancer follow-up treatment letters in which 16 distinct types of relations are annotated. We experiment with an approa…
Model predicts bid and ask price dynamics with spread-dependent intensities.
New method narrows prediction intervals for individual treatment effects.
Analyzes Lévy flights on manifolds for finding small targets.
We generalize recent theoretical work on the minimal number of layers of narrow deep belief networks that can approximate any probability distribution on the states of their visible units arbitrarily well. We relax the setting of binary units (Sutskever and Hinton, 2008; Le Roux and Bengio, 2008, 2010; Montúfar and Ay,…
Random shuffle method boosts HF dataset size 10-21 times.
New method predicts aphasia severity with narrower uncertainty intervals.
ST-BCP narrows the coverage gap in BCP by transforming nonconformity scores.
GN algorithm solves batched bandit for nondegenerate functions near-optimally.
Paper finds wide minima are better for generalization and proposes a new learning rate schedule.
Copulas outperform marginal models in multivariate risk forecasting, reducing model risk by narrowing down the set of models.
The classical Universal Approximation Theorem holds for neural networks of arbitrary width and bounded depth. Here we consider the natural `dual' scenario for networks of bounded width and arbitrary depth. Precisely, let be the number of inputs neurons, be the number of output neurons, and let be any nonaff…
This paper describes a distributed MapReduce implementation of the minimum Redundancy Maximum Relevance algorithm, a popular feature selection method in bioinformatics and network inference problems. The proposed approach handles both tall/narrow and wide/short datasets. We further provide an open source implementation…
Improves matrix multiplication throughput for asymmetric bit-width operands.
New method calibrates probabilistic regression models without restrictive assumptions.
We show that the strong asymptotic class of Weil-Petersson (WP) geodesics with narrow end invariant and bounded annular coefficients is determined by the forward ending lamination. This generalizes the Recurrent Ending Lamination Theorem of Brock-Masur-Minsky. As an application we provide a symbolic condition for diver…
This paper makes a small step towards a non-stochastic version of superhedging duality relations in the case of one traded security with a continuous price path. Namely, we prove the coincidence of game-theoretic and measure-theoretic expectation for lower semicontinuous positive functionals. We consider a new broad de…
Deep learning uncovers patterns between knot types.
A new method reduces memory requirements for Graph Transformers by sparsely training a network.
Tilting loss functions improves machine learning performance.
In an appendix to an earlier paper (cf. arXiv:1703.00984) we showed we showed how to construct tunnels of positive scalar curvature and of arbitrarily small length and volume connecting points in a \emph{three dimensional} manifold of \emph{constant sectional curvature}. Here we generalize the construction to arbitrary…
In this paper we adopted state-of-the-art machine learning algorithms, namely: random forest (RF) and least squares boosting, to model crash data and identify the optimum model to study the impact of narrow lanes on the safety of arterial roads. Using a ten-year crash dataset in four cities in Nebraska, two machine lea…
We review recent results about the maximal values of the Kullback-Leibler information divergence from statistical models defined by neural networks, including naive Bayes models, restricted Boltzmann machines, deep belief networks, and various classes of exponential families. We illustrate approaches to compute the max…
New concept: reward hacking, where optimizing a flawed reward function can hurt performance.
English translation of "Solitony Ricciego" (Wiadomości Matematyczne 48, 2012, no. 1, pp. 1-32). Despite the general-sounding title, the text covers just a few narrow topics: Perelman's proof of the fact that compact Ricci solitons are of the gradient type, and a detailed unified description of Page's and Berard Bergery…
Randomly trained neural networks can generalize well if there's a simpler underlying teacher model.