Improved forecasting of suicide attempts using LSGPs for patients with little data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper proves all endomorphisms of framed little disk operad are automorphisms.
Maps self-duality in little disks operad to framed manifolds.
Bayesian Neural Networks show unexpected collapse of epistemic uncertainty with large models and little data.
Certain six-dimensional (1,0) supersymmetric little string theories, when compactified on , have moduli spaces of vacua given by smooth K3 surfaces. Using ideas of Gaiotto-Moore-Neitzke, we show that this provides a systematic procedure for determining the Ricci-flat metric on a smooth K3 surface in terms of BPS d…
This paper gives a partial description of the homotopy type of K, the space of long knots in 3-dimensional Euclidean space. The primary result is the construction of a homotopy equivalence between K and the free little 2-cubes object over the space of prime knots. In proving the freeness result, a close correspondence …
New RNN model forecasts unseen time series with little training data.
Conditional density estimation generalizes regression by modeling a full density f(yjx) rather than only the expected value E(yjx). This is important for many tasks, including handling multi-modality and generating prediction intervals. Though fundamental and widely applicable, nonparametric conditional density estimat…
The Whitney embedding theorem gives an upper bound on the smallest embedding dimension of a manifold. If a data set lies on a manifold, a random projection into this reduced dimension will retain the manifold structure. Here we present an algorithm to find a projection that distorts the data as little as possible.
Many engineers wish to deploy modern neural networks in memory-limited settings; but the development of flexible methods for reducing memory use is in its infancy, and there is little knowledge of the resulting cost-benefit. We propose structural model distillation for memory reduction using a strategy that produces a …
Random forests are a scheme proposed by Leo Breiman in the 2000's for building a predictor ensemble with a set of decision trees that grow in randomly selected subspaces of data. Despite growing interest and practical use, there has been little exploration of the statistical properties of random forests, and little is …
We provide elementary proofs of Lemmas 7.1 and 7.4 appearing in "The Cartan-Hadamard conjecture and the Little Prince", by B. Kloeckner and G. Kuperberg. The Lemmas play an important role in the derivation of novel isoperimetric inequalities. The original proofs relied on Sage, a symbolic algebra package, to factor cer…
New algorithm speeds up causal inference for large data.
New method preserves spectral clustering performance under aggressive sparsification and quantization.
Study fast learning rates for square loss in dependent data with hypercontractivity condition.
In the past decade, tracking health trends using social media data has shown great promise, due to a powerful combination of massive adoption of social media around the world, and increasingly potent hardware and software that enables us to work with these new big data streams. At the same time, many challenging proble…
Transfer learning from natural image datasets, particularly ImageNet, using standard large models and corresponding pretrained weights has become a de-facto method for deep learning applications to medical imaging. However, there are fundamental differences in data sizes, features and task specifications between natura…
A new method reduces sample complexity for meta-learning.
Regularization is typically understood as improving generalization by altering the landscape of local extrema to which the model eventually converges. Deep neural networks (DNNs), however, challenge this view: We show that removing regularization after an initial transient period has little effect on generalization, ev…
Motivated by string topology and the arc operad, we introduce the notion of quasi-operads and consider four (quasi)-operads which are different varieties of the operad of cacti. These are cacti without local zeros (or spines) and cacti proper as well as both varieties with fixed constant size one of the constituting lo…
This note is about a little extension of Nash's embedding theorem in the case of complete manifolds.
The problem of frequent pattern mining has been studied quite extensively for various types of data, including sets, sequences, and graphs. Somewhat surprisingly, another important type of data, namely rank data, has received very little attention in data mining so far. In this paper, we therefore addresses the problem…
Many modern machine learning models are trained to achieve zero or near-zero training error in order to obtain near-optimal (but non-zero) test error. This phenomenon of strong generalization performance for "overfitted" / interpolated classifiers appears to be ubiquitous in high-dimensional data, having been observed …
The Adversarially Learned Mixture Model (AMM) is a generative model for unsupervised or semi-supervised data clustering. The AMM is the first adversarially optimized method to model the conditional dependence between inferred continuous and categorical latent variables. Experiments on the MNIST and SVHN datasets show t…
The performance of a Part-of-speech (POS) tagger is highly dependent on the domain ofthe processed text, and for many domains there is no or only very little training data available. This work addresses the problem of POS tagging noisy user-generated text using a neural network. We propose an architecture that trains a…
With the advent of large labelled datasets and high-capacity models, the performance of machine vision systems has been improving rapidly. However, the technology has still major limitations, starting from the fact that different vision problems are still solved by different models, trained from scratch or fine-tuned o…
The growing use of Machine Learning has produced significant advances in many fields. For image-based tasks, however, the use of deep learning remains challenging in small datasets. In this article, we review, evaluate and compare the current state-of-the-art techniques in training neural networks to elucidate which te…
New method for efficient inference in large datasets.
Study finds little progress in medical machine learning benchmarks over 3 years.
When working with asymptotically hyperbolic initial data sets for general relativity it is convenient to assume certain simplifying properties. We prove that the subset of initial data sets with such properties is dense in the set of physically reasonable asymptotically hyperbolic initial data sets. More specifically, …
Extends co-clustering to mixed numerical and binary data.
We consider the problem of reconstructing signals and images from periodic nonlinearities. For such problems, we design a measurement scheme that supports efficient reconstruction; moreover, our method can be adapted to extend to compressive sensing-based signal and image acquisition systems. Our techniques can be pote…
The Gaussian process (GP) is a popular way to specify dependencies between random variables in a probabilistic model. In the Bayesian framework the covariance structure can be specified using unknown hyperparameters. Integrating over these hyperparameters considers different possible explanations for the data when maki…
Generative adversarial networks (GANs) are innovative techniques for learning generative models of complex data distributions from samples. Despite remarkable recent improvements in generating realistic images, one of their major shortcomings is the fact that in practice, they tend to produce samples with little divers…
This paper introduces GEMINI, a new mutual information metric for unsupervised neural network training.
Generative Adversarial Networks (GAN) boast impressive capacity to generate realistic images. However, like much of the field of deep learning, they require an inordinate amount of data to produce results, thereby limiting their usefulness in generating novelty. In the same vein, recent advances in meta-learning have o…
Motivated by the fact that most of the information relevant to the prediction of target tokens is drawn from the source sentence , we propose truncating the target-side window used for computing self-attention by making an -gram assumption. Experiments on WMT EnDe and EnFr data sets show that the…
The leverage effect refers to the generally negative correlation between the return of an asset and the changes in its volatility. There is broad agreement in the literature that the effect should be present for theoretical reasons, and it has been consistently found in empirical work. However, a few papers have pointe…
Although operator-valued kernels have recently received increasing interest in various machine learning and functional data analysis problems such as multi-task learning or functional regression, little attention has been paid to the understanding of their associated feature spaces. In this paper, we explore the potent…
Much of the work in metalearning has focused on classifier selection, combined more recently with hyperparameter optimization, with little concern for data preprocessing. Yet, it is generally well accepted that machine learning applications require not only model building, but also data preprocessing. In other words, p…
Neuroscience is experiencing a data revolution in which many hundreds or thousands of neurons are recorded simultaneously. Currently, there is little consensus on how such data should be analyzed. Here we introduce LFADS (Latent Factor Analysis via Dynamical Systems), a method to infer latent dynamics from simultaneous…
This paper introduces GEMINI, a new metric for unsupervised neural network training that avoids the need for regularizations.
To improve the ability of VAE to disentangle in the latent space, existing works mostly focus on enforcing independence among the learned latent factors. However, the ability of these models to disentangle often decreases as the complexity of the generative factors increases. In this paper, we investigate the little-ex…
Machine learning portfolios perform well with simple imputation of missing data.
We propose a novel data-dependent structured gradient regularizer to increase the robustness of neural networks vis-a-vis adversarial perturbations. Our regularizer can be derived as a controlled approximation from first principles, leveraging the fundamental link between training with noise and regularization. It adds…
We characterize and study variable importance (VIMP) and pairwise variable associations in binary regression trees. A key component involves the node mean squared error for a quantity we refer to as a maximal subtree. The theory naturally extends from single trees to ensembles of trees and applies to methods like rando…
We study the out-of-sample properties of robust empirical optimization problems with smooth -divergence penalties and smooth concave objective functions, and develop a theory for data-driven calibration of the non-negative "robustness parameter" that controls the size of the deviations from the nominal model. Bu…
A little complement concerning the dynamics of non-metric manifolds is provided, by showing that any flow on an -bounded surface with non-zero Euler character has a fixed point.