Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

60120179239 · Jun 202019922001200920172026
48 results for weight tying

We prove that if two Tambara-Yamagami categories TY(A,χ,ν) and TY(A',χ',ν') give rise to the same state sum invariants of 3-manifolds and the order of one of the groups A, A' is odd, then ν=ν' and there is a group isomorphism A\approx A' carrying χto χ'. The proof is based on an explicit computation of the state sum in…

2010-09-09abs ↗pdf ↗

We present trellis networks, a new architecture for sequence modeling. On the one hand, a trellis network is a temporal convolutional network with special structure, characterized by weight tying across depth and direct injection of the input into deep layers. On the other hand, we show that truncated recurrent network…

2018-10-15abs ↗pdf ↗

Transformers can be made to implement specific algorithms by controlling training parameters.

problem Understanding when and how weight-tied looped transformers implement specific algorithms.
method Controlled experiments on group word problems, analyzing training contracts and convergence times.
result The speed and halting rule of weight-tied looped transformers are determined by the training contract and the number of loops.

We introduce a Deep Boltzmann Machine model suitable for modeling and extracting latent semantic representations from a large unstructured collection of documents. We overcome the apparent difficulty of training a DBM with judicious parameter tying. This parameter tying enables an efficient pretraining algorithm and a …

2013-09-26abs ↗pdf ↗

This paper takes a step towards theoretical analysis of the relationship between word embeddings and context embeddings in models such as word2vec. We start from basic probabilistic assumptions on the nature of word vectors, context vectors, and text generation. These assumptions are well supported either empirically o…

2019-02-26abs ↗pdf ↗

Study trace systoles on surfaces, finding optimal bounds and implications.

problem Optimal systolic inequalities on hyperbolic manifolds and non-Fuchsian representations.
method Defined trace systole, used Markoff maps correspondence, computed bounds.
result Explicit optimal bounds for one-holed torus, four-holed sphere, and non-orientable surface of genus 3.

We describe the infinitesimal moduli space of pairs (Y,V)(Y, V) where YY is a manifold with G2G_2 holonomy, and VV is a vector bundle on YY with an instanton connection. These structures arise in connection to the moduli space of heterotic string compactifications on compact and non-compact seven dimensional spaces, e.…

2016-07-12abs ↗pdf ↗

Recent work (Cohen & Welling, 2016) has shown that generalizations of convolutions, based on group theory, provide powerful inductive biases for learning. In these generalizations, filters are not only translated but can also be rotated, flipped, etc. However, coming up with exact models of how to rotate a 3 x 3 filter…

2019-05-12abs ↗pdf ↗

Satellite constructions on a knot can be thought of as taking some strands of a knot and then tying in another knot. Using satellite constructions one can construct many distinct isotopy classes of knots. Pushing this further one can construct distinct concordance classes of knots which preserve some algebraic invarian…

2016-01-11abs ↗pdf ↗

Heterotic string compactifications on integrable G2G_2 structure manifolds YY with instanton bundles (V,A),(TY,θ~)(V,A), (TY,\tildeθ) yield supersymmetric three-dimensional vacua that are of interest in physics. In this paper, we define a covariant exterior derivative D\cal D and show that it is equivalent to a heterotic G2G_2

2017-04-27abs ↗pdf ↗

SE(3)-Transformers maintain equivariance for 3D data under rotations and translations.

problem Ensuring stable and predictable performance in 3D data under transformations.
method Introducing a self-attention module that is equivariant under continuous 3D roto-translations.
result The SE(3)-Transformer outperforms non-equivariant and non-attention models on real-world datasets.

Consider a feedforward neural network ψ:RdRdψ: \mathbb{R}^d\rightarrow \mathbb{R}^d such that ψfψ\approx \nabla f, where f:RdRf:\mathbb{R}^d \rightarrow \mathbb{R} is a smooth function, therefore ψψ must satisfy jψi=iψj\partial_j ψ_i = \partial_i ψ_j pointwise. We prove a theorem that a ψψ network with more than one hidden layer…

2019-10-28abs ↗pdf ↗

We give bounds on the gap functions of the singularities of a cuspidal plane curve of arbitrary genus, generalising recent work of Borodzik and Livingston. We apply these inequalities to unicuspidal curves whose singularity has one Puiseux pair: we prove two identities tying the parameters of the singularity, the genus…

2014-09-11abs ↗pdf ↗

Knowledge graphs contain knowledge about the world and provide a structured representation of this knowledge. Current knowledge graphs contain only a small subset of what is true in the world. Link prediction approaches aim at predicting new links for a knowledge graph given the existing links among the entities. Tenso…

2018-02-13abs ↗pdf ↗

Paper finds conditions for minimal hypersurfaces in S^6 with constant scalar curvature.

problem Finding conditions for minimal hypersurfaces in S^6 with constant scalar curvature.
method Assumptions on principal curvatures for isoparametric hypersurfaces.
result Rigidity result: Hypersurfaces with exactly two distinct principal curvatures are Clifford tori.

We propose to study equivariance in deep neural networks through parameter symmetries. In particular, given a group G\mathcal{G} that acts discretely on the input and output of a standard neural network layer φW:MNφ_{W}: \Re^{M} \to \Re^{N}, we show that φWφ_{W} is equivariant with respect to G\mathcal{G}-action iff $\m…

2017-02-27abs ↗pdf ↗

The paper extends two-step homogeneous geodesics to homogeneous Finsler spaces.

problem Extending two-step homogeneous geodesics to Finsler spaces.
method Providing sufficient conditions for (α,β)(α,β) spaces and decomposable cubic spaces to have two-step Finsler geodesic orbit spaces.
result Presented examples of two-step Finsler geodesic orbit spaces.

Tying knots and linking microscopic loops of polymers, macromolecules, or defect lines in complex materials is a challenging task for material scientists. We demonstrate the knotting of microscopic topological defect lines in chiral nematic liquid crystal colloids into knots and links of arbitrary complexity by using l…

2011-07-08abs ↗pdf ↗

We study the minimal crossing number c(K1#K2)c(K_{1}\# K_{2}) of composite knots K1#K2K_{1}\# K_{2}, where K1K_1 and K2K_2 are prime, by relating it to the minimal crossing number of spatial graphs, in particular the 2n2n-theta curve θK1,K2nθ_{K_{1},K_{2}}^n that results from tying nn of the edges of the planar embedding of the $2n…

2017-09-15abs ↗pdf ↗

Recently, the deep learning community has given growing attention to neural architectures engineered to learn problems in relational domains. Convolutional Neural Networks employ parameter sharing over the image domain, tying the weights of neural connections on a grid topology and thus enforcing the learning of a numb…

2019-01-23abs ↗pdf ↗

Knowledge graphs are graphical representations of large databases of facts, which typically suffer from incompleteness. Inferring missing relations (links) between entities (nodes) is the task of link prediction. A recent state-of-the-art approach to link prediction, ConvE, implements a convolutional neural network to …

2018-08-21abs ↗pdf ↗

Pruned neural networks' error scales predictably with architecture and task.

problem Understanding the predictability of pruning across different scales and architectures.
method Functionally approximated the error of pruned networks, showing it is predictable in terms of invariant tying width, depth, and pruning level.
result The error of pruned networks follows a scaling law with interpretable coefficients that depend on architecture and task.

Powerful generative models, particularly in Natural Language Modelling, are commonly trained by maximizing a variational lower bound on the data log likelihood. These models often suffer from poor use of their latent variable, with ad-hoc annealing factors used to encourage retention of information in the latent variab…

2018-06-12abs ↗pdf ↗

Complex computer simulators are increasingly used across fields of science as generative models tying parameters of an underlying theory to experimental observations. Inference in this setup is often difficult, as simulators rarely admit a tractable density or likelihood function. We introduce Adversarial Variational O…

2017-07-22abs ↗pdf ↗

In this paper we consider backward stochastic differential equations with time-delayed generators of a moving average type. The classical framework with linear generators depending on (Y(t),Z(t))(Y(t),Z(t)) is extended and we investigate linear generators depending on (1t0tY(s)ds,1t0tZ(s)ds)(\frac{1}{t}\int_0^tY(s)ds, \frac{1}{t}\int_0^tZ(s)ds). We…

2010-08-22abs ↗pdf ↗

We study geodesics of the form γ(t)=π(exp(tX)exp(tY))γ(t)=π(\exp(tX)\exp(tY)), $X,Y\in \fr{g}=\operatorname{Lie}(G)$, in homogeneous spaces G/KG/K, where π:GG/Kπ:G\rightarrow G/K is the natural projection. These curves naturally generalise homogeneous geodesics, that is orbits of one-parameter subgroups of GG (i.e. γ(t)=π(exp(tX))γ(t)=π(\exp (tX)), $X\in …

2016-11-14abs ↗pdf ↗

Some machine learning applications require continual learning - where data comes in a sequence of datasets, each is used for training and then permanently discarded. From a Bayesian perspective, continual learning seems straightforward: Given the model posterior one would simply use this as the prior for the next task.…

2019-02-18abs ↗pdf ↗

This note revisits recent results regarding the geometry and moduli of solutions of the heterotic string on manifolds YY with a G2G_2 structure. In particular, such heterotic G2G_2 systems can be rephrased in terms of a differential Dˇ\check {\cal D} acting on a complex Ωˇ(Y,Q)\checkΩ^*(Y , {\cal Q}), where ${\cal Q}=T^*Y\…

2017-09-20abs ↗pdf ↗

Bayesian bandits misspecification affects UX optimization, revealing new models.

problem Misspecification of value models in Bayesian bandits impacts UX optimization.
method Formulated UXO as a restless, sleeping bandit with unobserved confounders and optional stopping. Provided model extensions to address misspecifications.
result Common misspecifications lead to sub-optimal rewards, demonstrating overdispersion's effects on bandit performance.

GS-B3^3SE improves label shift estimation by smoothing priors on a graph.

problem Label shift adaptation when source and target distributions share conditional but not marginal probabilities.
method Graph-Smoothed Bayesian Black-Box Shift Estimator (GS-B3^3SE) places Laplacian-Gaussian priors on log-priors and confusion-matrix columns tied by a label-similarity graph.
result GS-B3^3SE produces a tractable posterior with HMC or Newton-CG schemes, proving identifiability, contraction, and robustness.

We study geodesics in generalized Wallach spaces which are expressed as orbits of products of three exponential terms. These are homogeneous spaces M=G/KM=G/K whose isotropy representation decomposes into a direct sum of three submodules m=m1m2m3\frak{m}=\frak{m}_1\oplus\frak{m}_2\oplus\frak{m}_3, satisfying the relations $[\fr…

2015-03-14abs ↗pdf ↗

The paper offers a framework to analyze machine learning problems using concentration of measure.

problem Analyzing machine learning algorithms defined by implicit equations.
method Develops a concentration of measure framework to solve convex problems and implicit formulations.
result Provides precise estimations for the first moments of the solution, describing the behavior and performance of machine learning classifiers.

Let YY be a compact, oriented 3-manifold with a contact form aa and a metric ds2ds^2. Suppose that FYF\to Y is a principal bundle with structure group U(2)=SU(2)×±1S1U(2) = SU(2)\times_{\pm1}S^1 such that F/S1F/S^1 is the principal SO(3) bundle of orthonormal frames for TYTY. A unitary connection A0A_0 on the Hermitian line bundle $…

2013-07-17abs ↗pdf ↗

Investors optimize their portfolios within a Wasserstein ball to match a benchmark's risk profile.

problem Optimizing portfolio performance while maintaining risk proximity to a benchmark.
method Optimal dynamic strategy selection based on minimizing distortion risk measures within a Wasserstein ball.
result An optimal dynamic strategy exists and can be calculated through isotonic projections.

CNPs improve function approximation by contrastive learning.

problem Learning from non-i.i.d function instantiations in high-dimensional, noisy spaces.
method CNPs with TCL and FCL contrastive branches for better function approximation.
result CNPs outperform other variants in function distribution reconstruction and parameter identification.