Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

17355269 · May 202619922001200920172026
48 results for Turk's head knot

The minimum number of colors is a challenging knot invariant since, by definition, its calculation requires taking the minimum over infinitely many minima. In this article we estimate and in some cases calculate the minimum number of colors for the Turk's head knots on three strands.

2010-02-25abs ↗pdf ↗

Fox's trapezoidal conjecture for four-strand Turk's head knots is proven.

problem Proving log-concavity of the coefficient sequence of Dn(z)D_n(z) for four-strand Turk's head knots.
method Four-block smoothing theorem for products of reciprocal quartics.
result The coefficient sequence of Dn(z)D_n(z) is log-concave.

Study knots with large SL2(C)\mathrm{SL}_2(\mathbb{C}) character varieties.

problem Knots with high-dimensional character varieties.
method Two diagrammatic constructions: split link diagrams and rational tangle replacements; braids and orientation-reversing involutions.
result Conjecture that all Turk's head knots Th(p,q)Th(p,q) with pp and qq odd are X\mathcal{X}-large.

We compute the Kauffman bracket polynomial of the three-lead Turk's head, the chain sinnet and the figure-eight chain shadow diagrams. Each of these knots can in fact be constructed by repeatedly concatenating the same 3-tangle, respectively, then taking the closure. The bracket is then evaluated by expressing the stat…

2018-07-13abs ↗pdf ↗

The m,n Turk's Head Knot, THK(m,n), is an "alternating (m,n) torus knot." We prove the Harary-Kauffman conjecture for all THK(m,n) except for the case where m \geq 5 is odd and n \geq 3 is relatively prime to m. We also give evidence in support of the conjecture in that case. Our proof rests on the observation that non…

2008-11-01abs ↗pdf ↗

Research classifies knots based on sliceness and amphichirality.

problem Classifying odd-stranded Turk's head knots based on sliceness and amphichirality.
method Constructing commuting pairs of ambient involutions and analyzing the equivariant Fox-Milnor square condition.
result Established a sharp parity dichotomy for equivariant rational sliceness and Klein amphichirality of odd-stranded Turk's head knots.

We derive a factorization of the Alexander polynomial of the 4-strand Turk's head knot using hypergeometric representations.

problem Deriving a factorization of the Alexander polynomial of the 4-strand Turk's head knot
method Using the reduced Burau representation and multivariable resultant elimination over reciprocal constraints
result Deriving a factorization of the Alexander polynomial in terms of Chebyshev polynomials

In this paper, we prove a formula for the 2-head of the colored Jones polynomial for an infinite family of pretzel knots. Following Hall, the proof utilizes skein-theoretic techniques and a careful examination of higher order stability properties for coefficients of the colored Jones polynomial.

2019-02-19abs ↗pdf ↗

We show that the head and tail functions of the colored Jones polynomial of adequate links are the product of head and tail functions of the colored Jones polynomial of alternating links that can be read-off an adequate diagram of the link. We apply this to strengthen a theorem of Kalfagianni, Futer and Purcell on the …

2013-10-16abs ↗pdf ↗

We investigate the coefficients of the highest and lowest terms (also called the head and the tail) of the colored Jones polynomial and show that they stabilize for alternating links and for adequate links. To do this we apply techniques from skein theory.

2011-12-16abs ↗pdf ↗

The colored Jones polynomial is a series of one variable Laurent polynomials J(K,n) associated with a knot K in 3-space. We will show that for an alternating knot K the absolute values of the first and the last three leading coefficients of J(K,n) are independent of n when n is sufficiently large. Computation of sample…

2006-04-10abs ↗pdf ↗

Building machine translation (MT) test sets is a relatively expensive task. As MT becomes increasingly desired for more and more language pairs and more and more domains, it becomes necessary to build test sets for each case. In this paper, we investigate using Amazon's Mechanical Turk (MTurk) to make MT test sets chea…

2014-10-20abs ↗pdf ↗

Multi-head attention outperforms single-head in in-context linear regression tasks.

problem Comparing performance of transformer with single-/multi-head attention in in-context learning.
method Theoretical analysis of performance of transformers with different attention mechanisms in linear regression tasks.
result Multi-head attention with a substantial embedding dimension outperforms single-head attention in in-context linear regression tasks.

New theory shows how multi-head attention reduces variance and decorrelates outputs.

problem Understanding and optimizing multi-head attention in neural networks.
method Developed a statistical theory linking multi-head attention to ensemble Nadaraya-Watson estimators.
result MHA variance reduction depends on head decorrelation, not just head count.

Geometric theory of projection heads in self-supervised learning.

problem Dimensional collapse and information invariance trade-off in projection heads.
method Geometric modeling of projection heads as Riemannian metrics, analyzing Hessian eigenvalues, and tracking optimization geometry.
result Smooth nonlinear heads induce negative curvature, preventing collapse; linear and ReLU heads cannot.

New approach improves multi-head attention by making heads less similar.

problem Multi-head attention can lead to similar features, reducing model expressiveness.
method Proposes a non-parametric approach using Bayesian techniques to make heads repel each other.
result Improves feature diversity, leading to better representations and performance.

Deep forecasting models show output heads significantly improve performance on fat-tailed financial returns.

problem Improving deep learning models for forecasting fat-tailed financial returns.
method Comparison of backbone architectures and output heads (point, Gaussian, Gaussian mixture) on S&P 500 monthly log-returns.
result Switching from point to Gaussian heads improves CRPS by about 1.3 percent, and from Gaussian to mixture adds another 2.4 percent.

We study a certain skein element in the relative Kauffman bracket skein module of the disk with some marked points, and expand this element in terms linearly independent elements of this module. This expansion is used to compute and study the head and the tail of the colored Jones polynomial and in particular we give a…

2012-12-10abs ↗pdf ↗

Multi-head attention mechanism is capable of learning various representations from sequential data while paying attention to different subsequences, e.g., word-pieces or syllables in a spoken word. From the subsequences, it retrieves richer information than a single-head attention which only summarizes the whole sequen…

2019-10-10abs ↗pdf ↗

Attention mechanisms have become ubiquitous in NLP. Recent architectures, notably the Transformer, learn powerful context-aware word representations through layered, multi-headed attention. The multiple heads learn diverse types of word relationships. However, with standard softmax attention, all attention heads are de…

2019-08-30abs ↗pdf ↗

This work proposes a collaborative multi-head attention layer to reduce model size without sacrificing accuracy.

problem Over-parameterization in transformer models trained with large datasets.
method Proposes a collaborative multi-head attention layer that shares key/query projections.
result Reduction in model size by 4 for same accuracy and speed.

Study on multi-head softmax attention dynamics for in-context learning.

problem Understanding and optimizing multi-head softmax attention models for multi-task linear regression.
method Gradient flow analysis and spectral mapping technique.
result Gradient flow converges to optimal multi-head softmax attention model, with task allocation emerging during training.

Transformer-MGK replaces redundant heads with Gaussian key mixtures, improving efficiency and performance.

problem Redundant attention heads in transformers degrade performance and efficiency.
method Transformer-MGK replaces redundant heads with a mixture of Gaussian keys.
result Transformer-MGK accelerates training and inference, reduces parameters and FLOPs, and achieves comparable or better accuracy.

Minimalistic model captures head direction system properties.

problem Representing head direction system in a high-dimensional space.
method A minimalistic representation model of the rotation group U(1), including fully connected and convolutional versions.
result Emergence of Gaussian-like tuning profiles and 2D circle geometry in both model versions.

Develops a mean-field theory for multi-head self-attention under cross-entropy training.

problem Mean-field analysis of multi-head self-attention under cross-entropy training.
method Mean-field theory for a simplified single-layer causal multi-head self-attention model.
result Proves a static finite-head approximation bound for the optimal risk.

Investigates the benefits of multi-head attention in Transformers, deriving convergence and generalization guarantees.

problem Underexplored dynamics of multi-head attention in Transformer training and generalization.
method Derives convergence and generalization guarantees for gradient-descent training of a multi-head self-attention model.
result Establishes conditions for initialization that ensure multi-head attention's realizability.

New deep learning model generates accurate personalized human head models for electromagnetic dosimetry.

problem Challenges in generating accurate human head models for personalized electromagnetic dosimetry.
method Proposed ForkNet architecture for segmentation of whole human head structures using deep learning.
result Generated head models exhibit strong matching with manual segmentation results.

Investigates optimal parameter allocation in Transformers for efficiency and expressivity.

problem Balancing expressivity and efficiency in Transformer model parameters.
method Mathematical analysis and theoretical characterization of attention heads and head dimensions.
result Later layers can operate more efficiently with reduced parameters due to saturation of softmax activations.

We demystify attention patterns in multi-head softmax models for linear data.

problem Understanding the training dynamics and emergent patterns in multi-head softmax attention models.
method Extensive empirical experiments and rigorous theoretical analysis.
result Multi-head softmax attention models approximate a debiased gradient descent predictor, outperforming single-head attention and achieving near-Bayesian optimality.

Improved robot navigation using multi-head attention for natural language instructions.

problem Improving robot navigation in unfamiliar environments.
method Proposes a multi-head attention mechanism blending layer in a neural network model.
result Significant performance gains in translating instructions for unseen environments.

Study shows the number of attention heads affects transformer performance.

problem Understanding how the number of attention heads impacts transformer performance.
method Introduced a generalized DD-retrieval task, established upper and lower bounds on parameter complexity, and validated with experiments.
result Transformers with many heads can efficiently approximate functions, while few heads require a large number of parameters.

New method removes contrastive loss by adding a prediction head, revealing learning mechanisms.

problem Understanding why neural networks learn competitive representations despite trivial optima.
method Empirical and theoretical analysis of a trainable, identity-initialized prediction head.
result The trainable prediction head enables learning all features, preventing dimensional collapse.

As the Portable Document Format (PDF) file format increases in popularity, research in analysing its structure for text extraction and analysis is necessary. Detecting headings can be a crucial component of classifying and extracting meaningful data. This research involves training a supervised learning model to detect…

2018-08-31abs ↗pdf ↗

Single-head attention approximates any function under various norms.

problem Universal approximation of functions using attention mechanisms.
method Interpreting attention as partitioning and summing linear transformations.
result Single-head attention can approximate any continuous function under LL_\infty-norm and Lebesgue integrable functions under LpL_p-norm.