Multi-task learning shares information between related tasks, sometimes reducing the number of parameters required. State-of-the-art results across multiple natural language understanding tasks in the GLUE benchmark have previously used transfer from a single large task: unsupervised pre-training with BERT, where a sep…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Fine-tuning large pre-trained models is an effective transfer mechanism in NLP. However, in the presence of many downstream tasks, fine-tuning is parameter inefficient: an entire new model is required for every task. As an alternative, we propose transfer with adapter modules. Adapter modules yield a compact and extens…
Downscaled models outperform larger ones on GLUE tasks.
TextHide secures language understanding tasks by adding minimal encryption.
Pea-KD improves BERT student models by 4.4% on average in GLUE tasks.
This is an expository paper which aims to give a simple proof of the existence of Ricci-flat metrics on certain K3 surfaces, as an illustration of general "glueing" techniques.
We formulate a more conceptual interpretation of the Cappell-Lee-Miller glueing/splitting theorem using the new language of asymptotic maps and asymptotic exactness. Additionally, we present an asymptotic description of the Mayer-Vietoris sequence naturally associated to the Cech cohomology of the sheaf of local soluti…
The paper develops glueing theory for topological spaces and applies it to compactifications.
We discuss some additivity properties of the simplicial volume for manifolds with boundary: we give proofs of additivity for glueing amenable boundary components and of superadditivity for glueing amenable submanifolds of the boundary, and we discuss doubling of 3-manifolds.
Time delay estimation (TDE) is a critical and challenging step in all ultrasound elastography methods. A growing number of TDE techniques require an approximate but robust and fast method to initialize solving for TDE. Herein, we present a fast method for calculating an approximate TDE between two radio frequency (RF) …
We introduce HUBERT which combines the structured-representational power of Tensor-Product Representations (TPRs) and BERT, a pre-trained bidirectional Transformer language model. We show that there is shared structure between different NLP datasets that HUBERT, but not BERT, is able to learn and leverage. We validate …
One of the questions that arises when designing models that learn to solve multiple tasks simultaneously is how much of the available training budget should be devoted to each individual task. We refer to any formalized approach to addressing this problem (learned or otherwise) as a task selection policy. In this work …
New method glues Scherk surfaces into minimal surfaces, limiting possible outcomes.
A compact 4-dimensional manifold is a non-singular graph-manifold if it can be obtained by the glueing T^2-bundles over compact surfaces (with boundary) of negative Euler characteristics. If none of glueing diffeomorphisms respect the bundle structures, the graph-structure is called reduced. We prove that any homotopy …
We present a method to desingularize a compact G_2 manifold with isolated conical singularities by cutting out a neighbourhood of each singular point and glueing in an asymptotically conical G_2 manifold. Controlling the error on the overlap glueing region enables us to use a result of Joyce to conclude that the result…
BERT fine-tuning is unstable due to optimization issues, not forgetting or dataset size.
We show that two smooth nearby Riemannian metrics can be glued interpolating their scalar curvature. The resulting smooth metric is the same as the starting ones outside the gluing region and has scalar curvature interpolating between the original ones. One can then glue metrics while maintaining inequalities satisfied…
This paper improves neural network compression by using robust low-rank approximations.
We describe a glueing construction for a certain self-dual reduction of the Yang-Mills equations in dimension 8.
BERT-based architectures currently give state-of-the-art performance on many NLP tasks, but little is known about the exact mechanisms that contribute to its success. In the current work, we focus on the interpretation of self-attention, which is one of the fundamental underlying components of BERT. Using a subset of G…
A new method compresses NLP networks by using multiple subspaces instead of a single one.
We extend the definition of analytic and Reidemeister torsion from closed compact Riemannian manifolds to compact Riemannian manifolds with boundary , given a flat bundle $\Cal F$ of $\Cal A$-Hilbert modules of finite type and a decomposition of the boundary …
Multi-task learning (MTL) has achieved success over a wide range of problems, where the goal is to improve the performance of a primary task using a set of relevant auxiliary tasks. However, when the usefulness of the auxiliary tasks w.r.t. the primary task is not known a priori, the success of MTL models depends on th…
In this note we discuss the problem of resolving conically singular cscK varieties to construct smooth cscK manifolds, showing a glueing result for (some) crepant resolutions of cscK varieties with discrete automorphism groups.
CERT improves language understanding by contrastively learning sentence-level semantics.
In natural language processing, it has been observed recently that generalization could be greatly improved by finetuning a large-scale language model pretrained on a large unlabeled corpus. Despite its recent success and wide adoption, finetuning a large pretrained language model on a downstream task is prone to degen…
We start with a disk with vertices along its boundary where pairs of vertices are connected with strips with certain restrictions. This forms a {\it pairing}. To relate two pairings, we define an operator called a cut-and-glue operation. We show that this operation does not change an invariant of pairings know…
The purpose of this thesis is to define a "local" version of Ozsváth and Szabó's Heegaard Floer homology for links in the 3-dimensional sphere, i.e. a Heegaard Floer homology for tangles in the closed 3-ball. After studying basic properties of $\operatorname…
MixKD improves large-scale language model compression and generalization.
Recently, pre-trained language representation flourishes as the mainstay of the natural language understanding community, e.g., BERT. These pre-trained language representations can create state-of-the-art results on a wide range of downstream tasks. Along with continuous significant performance improvement, the size an…
We describe a glueing construction for the Yang-Mills equations in dimension . Our method is based on a construction of approximate solutions, and a detailed analysis of the linearized operator near an approximate solution.
I give a formula for computing the number of regular -coverings of closed orientable Seifert 3-manifolds, for a given finite group . The number is computed using a 3d TQFT with finite gauge group, through a cut-and-glue process.
We study a variant of the bandit problem where side information in the form of bounds on the mean of each arm is provided. We prove that these translate to tighter estimates of subgaussian factors and develop novel algorithms that exploit these estimates. In the linear setting, we present the Restricted-set OFUL (R-OFU…
As machine learning is applied more widely, data scientists often struggle to find or create end-to-end machine learning systems for specific tasks. The proliferation of libraries and frameworks and the complexity of the tasks have led to the emergence of "pipeline jungles" - brittle, ad hoc ML systems. To address thes…
We cut a hyperbolic surface of finite area along some analytic simple closed curves, and glue in cylinders of varying moduli. We prove that as the moduli of the glued cylinders go to infinity, the Fenchel-Nielsen twist coordinates for the resulting surface around those cylinders converge.
Study proposes memory-efficient backpropagation for linear layers in neural networks.
Attention-based models have shown significant improvement over traditional algorithms in several NLP tasks. The Transformer, for instance, is an illustrative example that generates abstract representations of tokens inputted to an encoder based on their relationships to all tokens in a sequence. Recent studies have sho…
In the symplectic category there is a `connect sum' operation that glues symplectic manifolds by identifying neighborhoods of embedded codimension two submanifolds. This paper establishes a formula for the Gromov-Witten invariants of a symplectic sum Z=X#Y in terms of the relative GW invariants of X and Y. Several appl…
Weight Squeezing transfers knowledge from large models to smaller ones, improving performance and speed.
Improved NLP performance with fewer parameters and less data using conditional multi-task learning.
We establish a canonical gluing procedure for Seiberg-Witten monopoles on the two pieces of a closed, oriented 4-manifold X which is split along a 3-dimensional closed, oriented submanifold. We only assume that the (unperturbed) character variety is Kuranishi-smooth and the limiting maps are transversal -- then we will…
John Conway created pairs of domains that sound the same for a special kind of music.
PosCal training improves classification models by calibrating posterior probabilities.
New Morse theory techniques glue nontransverse flowlines.
With a 4-ended tangle , we associate a Heegaard Floer invariant , the peculiar module of . Based on Zarev's bordered sutured Heegaard Floer theory, we prove a glueing formula for this invariant which recovers link Floer homology . Moreover, we classify…
We present BART, a denoising autoencoder for pretraining sequence-to-sequence models. BART is trained by (1) corrupting text with an arbitrary noising function, and (2) learning a model to reconstruct the original text. It uses a standard Tranformer-based neural machine translation architecture which, despite its simpl…
We establish a glueing theorem for the Ginzburg-Landau equations in dimension . To this end, we consider a nondegenerate minimal submanifold of codimension 2, and construct a one-parameter family of solutions to the Ginzburg-Landau equations such that the energy density concentrates near this submanifold. The pr…
This paper shows that many hyperbolic manifolds obtained by glueing arithmetic pieces embed into higher-dimensional hyperbolic manifolds as codimension-one totally geodesic submanifolds. As a consequence, many Gromov--Pyatetski-Shapiro and Agol--Belolipetsky--Thomson non-arithmetic manifolds embed geodesically. Moreove…