MetNet forecasts precipitation up to 8 hours with high spatial and temporal resolution.
problem Precise weather forecasting for long lead times.
method Neural network architecture using axial self-attention for global context aggregation.
result MetNet outperforms Numerical Weather Prediction at forecasts of up to 8 hours.
Improved speech enhancement with MNTFA using time-frequency attention.
problem Speech enhancement with limited model size and memory.
method Designing MNTFA with self-attention modules for long sequences and joint training.
result MNTFA achieves better performance with fewer parameters than DPCRN.
3D Axial-Attention improves lung nodule classification accuracy.
problem Limited 3D attention in existing methods.
method Proposes 3D Axial-Attention network with 3D positional encoding.
result 3D Axial-Attention achieves state-of-the-art performance.
New curvature defined for corank 1 singular surfaces in 3D.
problem Defining curvature for singular surfaces.
method Introducing axial vector and curvature parabola to define axial curvature.
result Relates axial curvature to Gaussian curvature of a blow-up for certain singularities.
The article improves Beckner's inequality for axially symmetric functions on the n-dimensional sphere.
problem Improving Beckner's inequality for axially symmetric functions on Sn. method Uniqueness and existence results for Q-curvature type equations with a Paneitz operator on Sn for axially symmetric functions. result Improved Beckner's inequality for axially symmetric functions on Sn. New method constructs axial vector fields and defines quasi-local spin-angular momentum.
problem Constructing axial vector fields on Riemannian two-spheres.
method Using centre-of-mass unit sphere reference systems and Lie-propagated unit sphere reference systems.
result Constructive definition of quasi-local spin-angular momentum and balance relations.
Improved Beckner's inequality for axially symmetric functions on S^4.
problem Proving axially symmetric solutions to a constant Q-curvature type equation must be constant.
method Analyzing constant Q-curvature type equations on S^4, using Pohozaev-type identities and bifurcation methods.
result Improved Beckner's inequality for axially symmetric functions on S^4.
New surfaces described that are symmetric and solve a specific equation.
problem Understanding symmetric shapes of membranes.
method Characterized and described axially symmetric Helfrich spheres using the reduced membrane equation.
result These surfaces are symmetric and belong to a specific family.
Finite index solutions to Bernoulli problem are always axially symmetric.
problem Entire solutions to the Bernoulli free boundary problem with finite Morse index in 3D.
method Proof of axial symmetry for finite index solutions.
result Finite index solutions to the Bernoulli problem in 3D are axially symmetric.
Sharp inequality proven for symmetric functions on a 4D sphere.
problem Proving a sharp Beckner's inequality for axially symmetric functions on S4. method Utilized pointwise properties of Gegenbauer polynomials.
result Sharp Beckner's inequality established for axially symmetric functions on S4. Smooth minimizers found for Willmore energy surfaces.
problem Finding minimizers for Willmore energy surfaces.
method Existence and smoothness established through axially symmetric surfaces with prescribed isoperimetric ratio.
result Existence and smoothness of minimizers proven.
Axial-LOB predicts stock prices from LOB data using attention layers.
problem Predicting stock price from LOB data with long-range dependencies.
method Axial-LOB uses gated position-sensitive axial attention layers to incorporate global interactions.
result Axial-LOB achieves state-of-the-art performance in stock price prediction.
In this paper are studied the simplest patterns of axial curvature lines (along which the normal curvature vector is at a vertex of the ellipse of curvature) near a critical point of a surface mapped into R4. These critical points, where the rank of the mapping drops from 2 to 1, occur isolated in generic one parameter…
We study the long time existence theory for a non local flow associated to a free boundary problem for a trapped non liquid drop. The drop has free boundary components on two horizontal plates and its free energy is anisotropic and axially symmetric. For axially symmetric initial surfaces with sufficiently large volume…
For any n>1 we give an explicit example of an n-axially symmetric Cartesian current in B^3 x S^2 with non-trivial vertical part and non-constant graph part minimizing the relaxed Dirichlet energy among the n-axially symmetric Cartesian currents with the same boundary. This stands in sharp contrast with a results of Har…
Defines axial curvatures for corank 1 singular manifolds in higher dimensions.
problem Characterizing singular n-manifolds in Rn+k with corank 1 singular points. method Using curvature locus and second fundamental form, defining up to l(n−1) axial curvatures. result Umbilic curvatures are absolute values of our axial curvatures.
We study mean curvature flow of smooth, axially symmetric surfaces in R3 with Neumann boundary data. We show that all singularities at the first singular time must be of type I.
Based on the Hamiltonian dimensional reduction of 3+1 axially symmetric, Ricci-flat Lorentzian spacetimes to a 2+1 Einstein-wave map system with the (negatively curved) hyperbolic 2-plane target, we construct a positive-definite, (spacetime) gauge-invariant energy functional for linear axially symmetric perturbatio…
New minimal hypersurfaces found via transformations.
problem Finding new axially symmetric minimal hypersurfaces in 4D Minkowski space.
method Combining scaling symmetries and a non-obvious symmetry (analogous to Bianchi's transformation) to generate new hypersurfaces.
result Infinitely many axially symmetric minimal hypersurfaces can be generated from any given one.
We investigate the formation of singularities for surfaces evolving by volume preserving mean curvature flow. For axially symmetric flows - surfaces of revolution - in R3 with Neumann boundary conditions, we prove that the first developing singularity is of Type I. The result is obtained without any additio…
Derives a Hamiltonian model for 3D axially symmetric magnetohydrodynamics.
problem Modeling of 3D axially symmetric magnetohydrodynamics.
method Hamiltonian formulation and matrix discretization.
result First discrete model for 3D magnetohydrodynamics compatible with underlying Lie-Poisson structure.
GSA-Nets apply group equivariance to self-attention for vision tasks.
problem Improving self-attention networks for vision tasks.
method Define group-equivariant positional encodings.
result GSA-Nets outperform non-equivariant self-attention networks on vision benchmarks.
This work proves Kerr black holes are dynamically stable under certain perturbations.
problem Dynamical stability of Kerr black holes under axially symmetric perturbations.
method Dimensional reduction to 2+1 Einstein-wave map system, construction of positive-definite energy functional, proving boundary terms vanish.
result Strictly conserved positive energy for axially symmetric linear perturbations of Kerr black holes.
Paper proposes multiscale self-attentive convolutions for vision and language.
problem Improving language and vision understanding models using self-attention.
method Developed 1D and 2D Self Attentive Convolutions (SAC), multiscale SAC (MSAC).
result MSAC enhances model performance for vision and language tasks.
Active-memory mechanisms can replace self-attention in Transformers, but optimal results often require both.
problem Replacing self-attention with active-memory mechanisms in Transformers.
method Evaluation of various active-memory mechanisms in a Transformer model.
result Active-memory mechanisms can achieve comparable results to self-attention for language modeling, but optimal results are often achieved by combining both mechanisms.
Researchers found a way to measure energy in black hole perturbations.
problem Lack of positive-definite and conserved energy in black hole stability.
method Dimensional reduction and construction of a positive-definite energy functional.
result Conserved Hamiltonian energy for axially symmetric perturbations of Kerr black holes.
Self-attention prefers sparse functions of input sequences, reducing sample complexity.
problem Understanding the inductive biases of self-attention in modeling long-range dependencies.
method Theoretical analysis and synthetic experiments to probe sample complexity of learning sparse functions with Transformers.
result Bounded-norm Transformer networks can represent sparse functions of the input sequence with logarithmic sample complexity.
For Riemannian metrics of constant positive curvature on a punctured sphere with conic singularities at the punctures and co-axial monodromy of the developing map, possible angles at the singularities are completely described. This completes the recent result of Mondello and Panov. The related problem of describing pos…
Self-attention models benefit equally from width and depth, but beyond a certain point, depth becomes less efficient.
problem Understanding the optimal balance between depth and width in self-attention models.
method Theoretical predictions and empirical ablations on networks of varying depths and widths.
result An optimal width of 30K is recommended for a 1-Trillion parameter network, marking a significant width for self-attention models.
Improves sequential recommendation with relation-aware self-attention.
problem Improving accuracy in sequential recommendation.
method Integrates Transformer's self-attention mechanism with a probabilistic model of recommendation context.
result Significant improvements over recent baseline models.
Study algebraic invariants from lightning self-attention models.
problem Understanding polynomial coefficients of self-attention mechanisms.
method Identify algebraic invariants using polynomial coefficients and coordinate geometry.
result Found linear and nonlinear families of algebraic invariants.
Random forests with attention and self-attention improve regression performance.
problem Improving regression model performance on various datasets.
method Proposes new models using attention and self-attention mechanisms to solve regression problems.
result The models improve model performance on many datasets.
Paper investigates Lipschitz constants of self-attention modules in neural networks.
problem Lipschitz constants of self-attention modules in neural networks.
method Proved standard dot-product self-attention is not Lipschitz for unbounded input domain. Proposed L2 self-attention that is Lipschitz. Derived upper bound on L2 self-attention's Lipschitz constant.
result Proved standard self-attention is not Lipschitz for unbounded input domain and proposed an alternative L2 self-attention that is Lipschitz.
We study the provenance of singularity formation under mean curvature flow and volume preserving mean curvature flow in an axially symmetric setting. We prove that if the mean curvature is uniformly bounded on any finite time interval, then no singularities can develop during that time under both mean curvature flow an…
The Einstein/Maxwell equations reduce in the stationary and axially symmetric case to a harmonic map with prescribed singularities phi: R^3Σ-> H^2_C, where Sigma is a subset of the axis of symmetry, and H^2_C is the complex hyperbolic plane. Motivated by this problem, we prove the existence and uniqueness of harmonic m…
Kernel PCA explains self-attention mechanisms in deep learning models.
problem Understanding and explaining self-attention mechanisms in deep learning models.
method Deriving self-attention from kernel principal component analysis (kernel PCA).
result RPC-Attention, a robust attention mechanism, outperforms softmax attention in various tasks.
Proposes a faster Transformer decoding method by truncating target-side self-attention windows.
problem Efficiency in Transformer decoding with minimal BLEU score loss.
method N-gram assumption to truncate target-side self-attention windows.
result N-gram masked self-attention model maintains BLEU score for N values from 4 to 8. Study reveals self-attention's role in learning and generalizing interactions.
problem Understanding self-attention's theoretical role in neural architectures.
method Interacting entities analysis, including multi-agent RL and genetic sequences.
result Self-attention efficiently represents, learns, and generalizes pairwise interactions.
Linformer reduces transformer complexity to linear, improving efficiency.
problem High cost of training and deploying large transformer models for long sequences.
method Approximates self-attention with low-rank matrix, proposing Linformer with O(n) complexity. result Linformer performs similarly to standard transformers but is more memory- and time-efficient.
Here are described the axiumbilic points that appear in generic one parameter families of surfaces immersed in R4. At these points the ellipse of curvature of the immersion, Little, Garcia - Sotomayor has equal axes. A review is made on the basic preliminaries on axial curvature lines and the associated axiumbilic poin…
Skeinformer accelerates self-attention for long sequences with linear complexity.
problem Efficiency of Transformer models in processing long sequences.
method Matrix sketching and column sampling to reduce quadratic complexity to linear.
result Skeinformer outperforms alternatives with smaller time/space footprint.
New theory shows how membranes can break symmetry.
problem Understanding symmetry breaking in membranes with boundaries.
method Applied bifurcation theory and reduced membrane equation.
result Existence of symmetry breaking bifurcation in membrane solutions.
We discuss the existence of Killing tensors for certain (physically motivated) stationary and axially symmetric vacuum space-times. We show nonexistence of a nontrivial Killing tensor for a Tomimatsu-Sato metric (up to valence 7), for a C-metric (up to valence 9) and for a Zipoy-Voorhees metric (up to valence 11). The …
We study the convergence of an axially symmetric hypersurface evolving by volume preserving mean curvature flow. Assuming the surface is not pinching off along the axis at any time during the flow, and without any additional conditions, as for example on the curvature, we prove that it converges to a hemisphere, when t…
We make a detailed study of the moduli space of winding number two (k=2) axially symmetric vortices (or equivalently, of co-axial composite of two fundamental vortices), occurring in U(2) gauge theory with two flavors in the Higgs phase, recently discussed by Hashimoto-Tong (hep-th/0506022) and Auzzi-Shifman-Yung (hep-…
Paper connects MoE and self-attention, proposing active-attention.
problem Improving efficiency and performance of self-attention mechanisms.
method Established connection between MoE and self-attention, analyzed quadratic gating functions, proposed active-attention mechanism.
result Active-attention outperforms standard self-attention in various tasks.
Gradient descent converges geometrically to optimal self-attention parameters.
problem Training softmax self-attention layers for linear regression.
method Structure-aware gradient descent with preconditioner and regularizer.
result Gradient descent converges geometrically to global minima.
New method embeds time span into self-attention for better temporal pattern recognition.
problem Capturing temporal patterns in event sequences without recurrent networks.
method Functional time representation learning with Bochner's and Mercer's Theorems.
result Proposed methods outperform baseline models in various continuous-time event sequence prediction tasks.