CDEFs reduce model complexity and uncover time correlations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We prove that if two Tambara-Yamagami categories TY(A,χ,ν) and TY(A',χ',ν') give rise to the same state sum invariants of 3-manifolds and the order of one of the groups A, A' is odd, then ν=ν' and there is a group isomorphism A\approx A' carrying χto χ'. The proof is based on an explicit computation of the state sum in…
We present trellis networks, a new architecture for sequence modeling. On the one hand, a trellis network is a temporal convolutional network with special structure, characterized by weight tying across depth and direct injection of the input into deep layers. On the other hand, we show that truncated recurrent network…
Transformers can be made to implement specific algorithms by controlling training parameters.
We introduce a Deep Boltzmann Machine model suitable for modeling and extracting latent semantic representations from a large unstructured collection of documents. We overcome the apparent difficulty of training a DBM with judicious parameter tying. This parameter tying enables an efficient pretraining algorithm and a …
This paper takes a step towards theoretical analysis of the relationship between word embeddings and context embeddings in models such as word2vec. We start from basic probabilistic assumptions on the nature of word vectors, context vectors, and text generation. These assumptions are well supported either empirically o…
Study trace systoles on surfaces, finding optimal bounds and implications.
We describe the infinitesimal moduli space of pairs where is a manifold with holonomy, and is a vector bundle on with an instanton connection. These structures arise in connection to the moduli space of heterotic string compactifications on compact and non-compact seven dimensional spaces, e.…
Deep equilibrium models converge globally without explicit computation.
Recent work (Cohen & Welling, 2016) has shown that generalizations of convolutions, based on group theory, provide powerful inductive biases for learning. In these generalizations, filters are not only translated but can also be rotated, flipped, etc. However, coming up with exact models of how to rotate a 3 x 3 filter…
Recurrent neural networks have been very successful at predicting sequences of words in tasks such as language modeling. However, all such models are based on the conventional classification framework, where the model is trained against one-hot targets, and each word is represented both as an input and as an output in …
Satellite constructions on a knot can be thought of as taking some strands of a knot and then tying in another knot. Using satellite constructions one can construct many distinct isotopy classes of knots. Pushing this further one can construct distinct concordance classes of knots which preserve some algebraic invarian…
Heterotic string compactifications on integrable structure manifolds with instanton bundles yield supersymmetric three-dimensional vacua that are of interest in physics. In this paper, we define a covariant exterior derivative and show that it is equivalent to a heterotic …
SE(3)-Transformers maintain equivariance for 3D data under rotations and translations.
Consider a feedforward neural network such that , where is a smooth function, therefore must satisfy pointwise. We prove a theorem that a network with more than one hidden layer…
We present three new inequalities tying the signature, the simplicial volume and the Euler characteristic of surface bundles over surfaces. Two of them are true for any surface bundle, while the third holds on a specific family of surface bundles, namely the ones that arise through a ramified covering. These are the ma…
We give bounds on the gap functions of the singularities of a cuspidal plane curve of arbitrary genus, generalising recent work of Borodzik and Livingston. We apply these inequalities to unicuspidal curves whose singularity has one Puiseux pair: we prove two identities tying the parameters of the singularity, the genus…
Knowledge graphs contain knowledge about the world and provide a structured representation of this knowledge. Current knowledge graphs contain only a small subset of what is true in the world. Link prediction approaches aim at predicting new links for a knowledge graph given the existing links among the entities. Tenso…
Paper finds conditions for minimal hypersurfaces in S^6 with constant scalar curvature.
We propose to study equivariance in deep neural networks through parameter symmetries. In particular, given a group that acts discretely on the input and output of a standard neural network layer , we show that is equivariant with respect to -action iff $\m…
This is a continuation of our first paper in [WY16]. There are two purposes of this paper: One is to give a proof of the main result in [WY16] without going through the argument depending on numerical effectiveness. The other one is to provide a proof of our conjecture, mentioned in [TY], where the assumption of negati…
We propose a simple and easy to implement neural network compression algorithm that achieves results competitive with more complicated state-of-the-art methods. The key idea is to modify the original optimization problem by adding K independent Gaussian priors (corresponding to the k-means objective) over the network p…
The paper extends two-step homogeneous geodesics to homogeneous Finsler spaces.
Tying knots and linking microscopic loops of polymers, macromolecules, or defect lines in complex materials is a challenging task for material scientists. We demonstrate the knotting of microscopic topological defect lines in chiral nematic liquid crystal colloids into knots and links of arbitrary complexity by using l…
We study the minimal crossing number of composite knots , where and are prime, by relating it to the minimal crossing number of spatial graphs, in particular the -theta curve that results from tying of the edges of the planar embedding of the $2n…
Recently, the deep learning community has given growing attention to neural architectures engineered to learn problems in relational domains. Convolutional Neural Networks employ parameter sharing over the image domain, tying the weights of neural connections on a grid topology and thus enforcing the learning of a numb…
Optimal score function estimation via empirical risk minimization
Knowledge graphs are graphical representations of large databases of facts, which typically suffer from incompleteness. Inferring missing relations (links) between entities (nodes) is the task of link prediction. A recent state-of-the-art approach to link prediction, ConvE, implements a convolutional neural network to …
Pruned neural networks' error scales predictably with architecture and task.
Powerful generative models, particularly in Natural Language Modelling, are commonly trained by maximizing a variational lower bound on the data log likelihood. These models often suffer from poor use of their latent variable, with ad-hoc annealing factors used to encourage retention of information in the latent variab…
Complex computer simulators are increasingly used across fields of science as generative models tying parameters of an underlying theory to experimental observations. Inference in this setup is often difficult, as simulators rarely admit a tractable density or likelihood function. We introduce Adversarial Variational O…
New approach ties loss curvature to model performance in deep learning.
In this paper we consider backward stochastic differential equations with time-delayed generators of a moving average type. The classical framework with linear generators depending on is extended and we investigate linear generators depending on . We…
We study geodesics of the form , $X,Y\in \fr{g}=\operatorname{Lie}(G)$, in homogeneous spaces , where is the natural projection. These curves naturally generalise homogeneous geodesics, that is orbits of one-parameter subgroups of (i.e. , $X\in …
Some machine learning applications require continual learning - where data comes in a sequence of datasets, each is used for training and then permanently discarded. From a Bayesian perspective, continual learning seems straightforward: Given the model posterior one would simply use this as the prior for the next task.…
ARBITER learns SPX-VIX term structures without arbitrage constraints.
This note revisits recent results regarding the geometry and moduli of solutions of the heterotic string on manifolds with a structure. In particular, such heterotic systems can be rephrased in terms of a differential acting on a complex , where ${\cal Q}=T^*Y\…
Solvents can induce helical knots in simulated biopolymer tubes.
Bayesian bandits misspecification affects UX optimization, revealing new models.
GS-BSE improves label shift estimation by smoothing priors on a graph.
New method improves dictionary recovery from over-realized models.
We study geodesics in generalized Wallach spaces which are expressed as orbits of products of three exponential terms. These are homogeneous spaces whose isotropy representation decomposes into a direct sum of three submodules , satisfying the relations $[\fr…
The paper offers a framework to analyze machine learning problems using concentration of measure.
Classifies instantons on ALF multi-Taub-NUT spaces and ties them to bow solutions.
Reinforcement learning (RL) for robotics is challenging due to the difficulty in hand-engineering a dense cost function, which can lead to unintended behavior, and dynamical uncertainty, which makes exploration and constraint satisfaction challenging. We address these issues with a new model-based reinforcement learnin…
Let be a compact, oriented 3-manifold with a contact form and a metric . Suppose that is a principal bundle with structure group such that is the principal SO(3) bundle of orthonormal frames for . A unitary connection on the Hermitian line bundle $…
Investors optimize their portfolios within a Wasserstein ball to match a benchmark's risk profile.
CNPs improve function approximation by contrastive learning.