In this paper we address the following question, given a face representation, how many identities can it resolve? In other words, what is the capacity of the face representation? A scientific basis for estimating the capacity of a given face representation will not only benefit the evaluation and comparison of differen…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Researchers analyze neural process architectures and their representational capacities.
Paper applies theorem to find optimal investment boundary in stochastic capacity expansion.
gLSTM improves graph neural networks by increasing storage capacity to prevent over-squashing.
We extend manifold capacity to nonlinear neural representations with contextual information.
The brain optimizes memory by forgetting what's predictable, improving generalization.
Memory capacity of DAM scales exponentially with feature separation, unaffected by correlations.
Semi-supervised node classification in attributed graphs, i.e., graphs with node features, involves learning to classify unlabeled nodes given a partially labeled graph. Label predictions are made by jointly modeling the node and its' neighborhood features. State-of-the-art models for node classification on such attrib…
There is some theoretical evidence that deep neural networks with multiple hidden layers have a potential for more efficient representation of multidimensional mappings than shallow networks with a single hidden layer. The question is whether it is possible to exploit this theoretical advantage for finding such represe…
Study shows how correlations between neural activity affect classification capacity.
Proposes a new graph representation method using tensor products.
Recurrent neural networks are powerful models for processing sequential data, but they are generally plagued by vanishing and exploding gradient problems. Unitary recurrent neural networks (uRNNs), which use unitary recurrence matrices, have recently been proposed as a means to avoid these issues. However, in previous …
We study a stochastic, continuous time model on a finite horizon for a firm that produces a single good. We model the production capacity as an Ito diffusion controlled by a nondecreasing process representing the cumulative investment. The firm aims to maximize its expected total net profit by choosing the optimal inve…
New analysis tightens memory capacity of Hopfield models using spherical codes.
Bézier-GAN optimizes airfoil design by reducing shape complexity.
SOC-ICNN expands neural network representational capacity by using conic optimization.
This paper focuses on the discrimination capacity of aggregation functions: these are the permutation invariant functions used by graph neural networks to combine the features of nodes. Realizing that the most powerful aggregation functions suffer from a dimensionality curse, we consider a restricted setting. In partic…
Study on risk measures using distorted Choquet integrals with random distortions.
Generative autoencoders offer a promising approach for controllable text generation by leveraging their latent sentence representations. However, current models struggle to maintain coherent latent spaces required to perform meaningful text manipulations via latent vector operations. Specifically, we demonstrate by exa…
This paper investigates how data augmentation improves linear separation of manifold data.
A grand challenge in representation learning is to learn the different explanatory factors of variation behind the high dimen- sional data. Encoder models are often determined to optimize performance on training data when the real objective is to generalize well to unseen data. Although there is enough numerical eviden…
Model for cross-border markets with limited transmission capacities.
With the advent of large labelled datasets and high-capacity models, the performance of machine vision systems has been improving rapidly. However, the technology has still major limitations, starting from the fact that different vision problems are still solved by different models, trained from scratch or fine-tuned o…
Sparse codes improve optimal control tasks with correlated inputs.
Simple object representations improve model-free RL performance.
We obtain a dual representation of the Kantorovich functional defined for functions on the Skorokhod space using quotient sets. Our representation takes the form of a Choquet capacity generated by martingale measures satisfying additional constraints to ensure compatibility with the quotient sets. These sets contain st…
Wide neural networks can degrade performance, contrary to conventional wisdom.
We present new intuitions and theoretical assessments of the emergence of disentangled representation in variational autoencoders. Taking a rate-distortion theory perspective, we show the circumstances under which representations aligned with the underlying generative factors of variation of data emerge when optimising…
Holomorphic networks on modular arithmetic show clear success or failure, no in-between.
GANs can improve image reconstruction by using intermediate layers.
Paper investigates Lambda Value-at-Risk under ambiguity and risk sharing.
Objects are represented in sensory systems by continuous manifolds due to sensitivity of neuronal responses to changes in physical features such as location, orientation, and intensity. What makes certain sensory representations better suited for invariant decoding of objects by downstream networks? We present a theory…
The problem of high-dimensional and large-scale representation of visual data is addressed from an unsupervised learning perspective. The emphasis is put on discrete representations, where the description length can be measured in bits and hence the model capacity can be controlled. The algorithmic infrastructure is de…
This paper extends financial theory to measure learnable market structure under computational constraints.
Intelligent behaviour in the real-world requires the ability to acquire new knowledge from an ongoing sequence of experiences while preserving and reusing past knowledge. We propose a novel algorithm for unsupervised representation learning from piece-wise stationary visual data: Variational Autoencoder with Shared Emb…
Sequential learning, also called lifelong learning, studies the problem of learning tasks in a sequence with access restricted to only the data of the current task. In this paper we look at a scenario with fixed model capacity, and postulate that the learning process should not be selfish, i.e. it should account for fu…
IRMAE learns compact latent spaces by minimizing rank.
We address the problem of one-to-many mappings in supervised learning, where a single instance has many different solutions of possibly equal cost. The framework of conditional variational autoencoders describes a class of methods to tackle such structured-prediction tasks by means of latent variables. We propose to in…
Representation learning has become an invaluable approach for learning from symbolic data such as text and graphs. However, while complex symbolic datasets often exhibit a latent hierarchical structure, state-of-the-art methods typically learn embeddings in Euclidean vector spaces, which do not account for this propert…
The recently developed variational autoencoders (VAEs) have proved to be an effective confluence of the rich representational power of neural networks with Bayesian methods. However, most work on VAEs use a rather simple prior over the latent variables such as standard normal distribution, thereby restricting its appli…
Characterizes continuity of monotone functionals in mixed topology.
New Feedback Transformer architecture improves model performance by exposing past representations to future.
Understanding the representational power of Restricted Boltzmann Machines (RBMs) with multiple layers is an ill-understood problem and is an area of active research. Motivated from the approach of \emph{Inherent Structure formalism} (Stillinger & Weber, 1982), extensively used in analysing Spin Glasses, we propose a no…
Proves generalization bounds for SGD using Feller processes and Hausdorff dimension.
Variational methods that rely on a recognition network to approximate the posterior of directed graphical models offer better inference and learning than previous methods. Recent advances that exploit the capacity and flexibility in this approach have expanded what kinds of models can be trained. However, as a proposal…
We address the problem of communicating domain knowledge from a user to the designer of a clustering algorithm. We propose a protocol in which the user provides a clustering of a relatively small random sample of a data set. The algorithm designer then uses that sample to come up with a data representation under which …
New analysis explains pathology of deep Gaussian processes.
We study various capacities on compact Kähler manifolds which generalize the Bedford-Taylor Monge-Ampère capacity. We then use these capacities to study the existence and the regularity of solutions of complex Monge-Ampère equations.