MetNet forecasts precipitation up to 8 hours with high spatial and temporal resolution.
problem Precise weather forecasting for long lead times.
method Neural network architecture using axial self-attention for global context aggregation.
result MetNet outperforms Numerical Weather Prediction at forecasts of up to 8 hours.
Improved speech enhancement with MNTFA using time-frequency attention.
problem Speech enhancement with limited model size and memory.
method Designing MNTFA with self-attention modules for long sequences and joint training.
result MNTFA achieves better performance with fewer parameters than DPCRN.
3D Axial-Attention improves lung nodule classification accuracy.
problem Limited 3D attention in existing methods.
method Proposes 3D Axial-Attention network with 3D positional encoding.
result 3D Axial-Attention achieves state-of-the-art performance.
For singular corank 1 surfaces in R3 we introduce a distinguished normal vector called the axial vector. Using this vector and the curvature parabola we define a new type of curvature called the axial curvature, which generalizes the singular curvature for frontal type singularities. We then study contact pr…
The article improves Beckner's inequality for axially symmetric functions on the n-dimensional sphere.
problem Improving Beckner's inequality for axially symmetric functions on Sn. method Uniqueness and existence results for Q-curvature type equations with a Paneitz operator on Sn for axially symmetric functions. result Improved Beckner's inequality for axially symmetric functions on Sn. New method constructs axial vector fields and defines quasi-local spin-angular momentum.
problem Constructing axial vector fields on Riemannian two-spheres.
method Using centre-of-mass unit sphere reference systems and Lie-propagated unit sphere reference systems.
result Constructive definition of quasi-local spin-angular momentum and balance relations.
Improved Beckner's inequality for axially symmetric functions on S^4.
problem Proving axially symmetric solutions to a constant Q-curvature type equation must be constant.
method Analyzing constant Q-curvature type equations on S^4, using Pohozaev-type identities and bifurcation methods.
result Improved Beckner's inequality for axially symmetric functions on S^4.
New surfaces described that are symmetric and solve a specific equation.
problem Understanding symmetric shapes of membranes.
method Characterized and described axially symmetric Helfrich spheres using the reduced membrane equation.
result These surfaces are symmetric and belong to a specific family.
Finite index solutions to Bernoulli problem are always axially symmetric.
problem Entire solutions to the Bernoulli free boundary problem with finite Morse index in 3D.
method Proof of axial symmetry for finite index solutions.
result Finite index solutions to the Bernoulli problem in 3D are axially symmetric.
Sharp inequality proven for symmetric functions on a 4D sphere.
problem Proving a sharp Beckner's inequality for axially symmetric functions on S4. method Utilized pointwise properties of Gegenbauer polynomials.
result Sharp Beckner's inequality established for axially symmetric functions on S4. Smooth minimizers found for Willmore energy surfaces.
problem Finding minimizers for Willmore energy surfaces.
method Existence and smoothness established through axially symmetric surfaces with prescribed isoperimetric ratio.
result Existence and smoothness of minimizers proven.
Axial-LOB predicts stock prices from LOB data using attention layers.
problem Predicting stock price from LOB data with long-range dependencies.
method Axial-LOB uses gated position-sensitive axial attention layers to incorporate global interactions.
result Axial-LOB achieves state-of-the-art performance in stock price prediction.
In this paper are studied the simplest patterns of axial curvature lines (along which the normal curvature vector is at a vertex of the ellipse of curvature) near a critical point of a surface mapped into R4. These critical points, where the rank of the mapping drops from 2 to 1, occur isolated in generic one parameter…
We study the long time existence theory for a non local flow associated to a free boundary problem for a trapped non liquid drop. The drop has free boundary components on two horizontal plates and its free energy is anisotropic and axially symmetric. For axially symmetric initial surfaces with sufficiently large volume…
For any n>1 we give an explicit example of an n-axially symmetric Cartesian current in B^3 x S^2 with non-trivial vertical part and non-constant graph part minimizing the relaxed Dirichlet energy among the n-axially symmetric Cartesian currents with the same boundary. This stands in sharp contrast with a results of Har…
Defines axial curvatures for corank 1 singular manifolds in higher dimensions.
problem Characterizing singular n-manifolds in Rn+k with corank 1 singular points. method Using curvature locus and second fundamental form, defining up to l(n−1) axial curvatures. result Umbilic curvatures are absolute values of our axial curvatures.
We study mean curvature flow of smooth, axially symmetric surfaces in R3 with Neumann boundary data. We show that all singularities at the first singular time must be of type I.
Based on the Hamiltonian dimensional reduction of 3+1 axially symmetric, Ricci-flat Lorentzian spacetimes to a 2+1 Einstein-wave map system with the (negatively curved) hyperbolic 2-plane target, we construct a positive-definite, (spacetime) gauge-invariant energy functional for linear axially symmetric perturbatio…
New minimal hypersurfaces found via transformations.
problem Finding new axially symmetric minimal hypersurfaces in 4D Minkowski space.
method Combining scaling symmetries and a non-obvious symmetry (analogous to Bianchi's transformation) to generate new hypersurfaces.
result Infinitely many axially symmetric minimal hypersurfaces can be generated from any given one.
We investigate the formation of singularities for surfaces evolving by volume preserving mean curvature flow. For axially symmetric flows - surfaces of revolution - in R3 with Neumann boundary conditions, we prove that the first developing singularity is of Type I. The result is obtained without any additio…
Derives a Hamiltonian model for 3D axially symmetric magnetohydrodynamics.
problem Modeling of 3D axially symmetric magnetohydrodynamics.
method Hamiltonian formulation and matrix discretization.
result First discrete model for 3D magnetohydrodynamics compatible with underlying Lie-Poisson structure.
GSA-Nets apply group equivariance to self-attention for vision tasks.
problem Improving self-attention networks for vision tasks.
method Define group-equivariant positional encodings.
result GSA-Nets outperform non-equivariant self-attention networks on vision benchmarks.
This work proves Kerr black holes are dynamically stable under certain perturbations.
problem Dynamical stability of Kerr black holes under axially symmetric perturbations.
method Dimensional reduction to 2+1 Einstein-wave map system, construction of positive-definite energy functional, proving boundary terms vanish.
result Strictly conserved positive energy for axially symmetric linear perturbations of Kerr black holes.
Researchers found a way to measure energy in black hole perturbations.
problem Lack of positive-definite and conserved energy in black hole stability.
method Dimensional reduction and construction of a positive-definite energy functional.
result Conserved Hamiltonian energy for axially symmetric perturbations of Kerr black holes.
Self-attention prefers sparse functions of input sequences, reducing sample complexity.
problem Understanding the inductive biases of self-attention in modeling long-range dependencies.
method Theoretical analysis and synthetic experiments to probe sample complexity of learning sparse functions with Transformers.
result Bounded-norm Transformer networks can represent sparse functions of the input sequence with logarithmic sample complexity.
For Riemannian metrics of constant positive curvature on a punctured sphere with conic singularities at the punctures and co-axial monodromy of the developing map, possible angles at the singularities are completely described. This completes the recent result of Mondello and Panov. The related problem of describing pos…
Self-attention models benefit equally from width and depth, but beyond a certain point, depth becomes less efficient.
problem Understanding the optimal balance between depth and width in self-attention models.
method Theoretical predictions and empirical ablations on networks of varying depths and widths.
result An optimal width of 30K is recommended for a 1-Trillion parameter network, marking a significant width for self-attention models.
Study algebraic invariants from lightning self-attention models.
problem Understanding polynomial coefficients of self-attention mechanisms.
method Identify algebraic invariants using polynomial coefficients and coordinate geometry.
result Found linear and nonlinear families of algebraic invariants.
Random forests with attention and self-attention improve regression performance.
problem Improving regression model performance on various datasets.
method Proposes new models using attention and self-attention mechanisms to solve regression problems.
result The models improve model performance on many datasets.
Paper investigates Lipschitz constants of self-attention modules in neural networks.
problem Lipschitz constants of self-attention modules in neural networks.
method Proved standard dot-product self-attention is not Lipschitz for unbounded input domain. Proposed L2 self-attention that is Lipschitz. Derived upper bound on L2 self-attention's Lipschitz constant.
result Proved standard self-attention is not Lipschitz for unbounded input domain and proposed an alternative L2 self-attention that is Lipschitz.
We study the provenance of singularity formation under mean curvature flow and volume preserving mean curvature flow in an axially symmetric setting. We prove that if the mean curvature is uniformly bounded on any finite time interval, then no singularities can develop during that time under both mean curvature flow an…
The key to a Transformer model is the self-attention mechanism, which allows the model to analyze an entire sequence in a computationally efficient manner. Recent work has suggested the possibility that general attention mechanisms used by RNNs could be replaced by active-memory mechanisms. In this work, we evaluate wh…
Self attention mechanisms have become a key building block in many state-of-the-art language understanding models. In this paper, we show that the self attention operator can be formulated in terms of 1x1 convolution operations. Following this observation, we propose several novel operators: First, we introduce a 2D ve…
The Einstein/Maxwell equations reduce in the stationary and axially symmetric case to a harmonic map with prescribed singularities phi: R^3Σ-> H^2_C, where Sigma is a subset of the axis of symmetry, and H^2_C is the complex hyperbolic plane. Motivated by this problem, we prove the existence and uniqueness of harmonic m…
Kernel PCA explains self-attention mechanisms in deep learning models.
problem Understanding and explaining self-attention mechanisms in deep learning models.
method Deriving self-attention from kernel principal component analysis (kernel PCA).
result RPC-Attention, a robust attention mechanism, outperforms softmax attention in various tasks.
Study reveals self-attention's role in learning and generalizing interactions.
problem Understanding self-attention's theoretical role in neural architectures.
method Interacting entities analysis, including multi-agent RL and genetic sequences.
result Self-attention efficiently represents, learns, and generalizes pairwise interactions.
Linformer reduces transformer complexity to linear, improving efficiency.
problem High cost of training and deploying large transformer models for long sequences.
method Approximates self-attention with low-rank matrix, proposing Linformer with O(n) complexity. result Linformer performs similarly to standard transformers but is more memory- and time-efficient.
Here are described the axiumbilic points that appear in generic one parameter families of surfaces immersed in R4. At these points the ellipse of curvature of the immersion, Little, Garcia - Sotomayor has equal axes. A review is made on the basic preliminaries on axial curvature lines and the associated axiumbilic poin…
Recent studies identified that sequential Recommendation is improved by the attention mechanism. By following this development, we propose Relation-Aware Kernelized Self-Attention (RKSA) adopting a self-attention mechanism of the Transformer with augmentation of a probabilistic model. The original self-attention of Tra…
Skeinformer accelerates self-attention for long sequences with linear complexity.
problem Efficiency of Transformer models in processing long sequences.
method Matrix sketching and column sampling to reduce quadratic complexity to linear.
result Skeinformer outperforms alternatives with smaller time/space footprint.
New theory shows how membranes can break symmetry.
problem Understanding symmetry breaking in membranes with boundaries.
method Applied bifurcation theory and reduced membrane equation.
result Existence of symmetry breaking bifurcation in membrane solutions.
We discuss the existence of Killing tensors for certain (physically motivated) stationary and axially symmetric vacuum space-times. We show nonexistence of a nontrivial Killing tensor for a Tomimatsu-Sato metric (up to valence 7), for a C-metric (up to valence 9) and for a Zipoy-Voorhees metric (up to valence 11). The …
We study the convergence of an axially symmetric hypersurface evolving by volume preserving mean curvature flow. Assuming the surface is not pinching off along the axis at any time during the flow, and without any additional conditions, as for example on the curvature, we prove that it converges to a hemisphere, when t…
We make a detailed study of the moduli space of winding number two (k=2) axially symmetric vortices (or equivalently, of co-axial composite of two fundamental vortices), occurring in U(2) gauge theory with two flavors in the Higgs phase, recently discussed by Hashimoto-Tong (hep-th/0506022) and Auzzi-Shifman-Yung (hep-…
Paper connects MoE and self-attention, proposing active-attention.
problem Improving efficiency and performance of self-attention mechanisms.
method Established connection between MoE and self-attention, analyzed quadratic gating functions, proposed active-attention mechanism.
result Active-attention outperforms standard self-attention in various tasks.
Gradient descent converges geometrically to optimal self-attention parameters.
problem Training softmax self-attention layers for linear regression.
method Structure-aware gradient descent with preconditioner and regularizer.
result Gradient descent converges geometrically to global minima.
Transformers' self-attention mechanism is mapped to a generalized Potts model.
problem Uncertainty in what type of data distribution self-attention can efficiently learn.
method Decouple word positions and embeddings, then show self-attention learns a generalized Potts model.
result Training self-attention is equivalent to solving the inverse Potts problem.
Model detects electricity theft with high accuracy.
problem Detecting electricity theft on imbalanced datasets.
method Multi-head self-attention mechanism with dilated convolutions and binary mask.
result Achieved AUC of 0.926, improving previous work by 17%.