Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

6481,2961,9442,592 · Jun 202019922001200920172026
48 results for softmax partition of unity

Paper studies Transformer learning theory for Euclidean and Riemannian domains.

problem Understanding and optimizing Transformer networks for regression tasks.
method Constructive approximation framework using softmax partition of unity and attention mechanism.
result Transformer can achieve uniform ε-approximation error with minimal parameters.

POUnets combine partitions of unity and monomials for efficient deep learning.

problem Efficiently approximating functions with deep neural networks in high dimensions.
method Integrates partitions of unity and monomials into neural network architecture.
result POUnets achieve hp-convergence for smooth functions and outperform MLPs for discontinuous functions.

Optimizes Lipschitz estimates for partitions of unity and characterizes spaces with Assouad-Nagata dimension.

problem Understanding the properties of partitions of unity and their Lipschitz bounds.
method Analyzes the standard partition of unity and its p\ell^p-generalizations, using the approximate midpoint property and Lebesgue number.
result Optimal Lipschitz bounds for partitions of unity and characterizes metric spaces with Assouad-Nagata dimension.

We consider the Witten-Reshetikhin-Turaev invariants or Chern-Simons partition function at or around roots of unity q=e2πi1Kq=e^{2πi \frac{1}{K}} with rational level K=rsK=\frac{r}{s} where rr and ss are coprime integers. From the exact expression for the G=SU(2)G=SU(2) Witten-Reshetikhin-Turaev invariants of Seifert manifolds at…

2019-06-28abs ↗pdf ↗

This paper is devoted to dualization of paracompactness to the coarse category via the concept of RR-disjointness. Property A of G.Yu can be seen as a coarse variant of amenability via partitions of unity and leads to a dualization of paracompactness via partitions of unity. On the other hand, finite decomposition com…

2013-07-15abs ↗pdf ↗

Enhances POU-Nets with probabilistic noise model for efficient spatial data clustering.

problem Improving the efficiency and accuracy of deep learning models for spatial data.
method Integrates Gaussian noise model into POU-Nets to enable gradient-based optimization and hierarchical refinement.
result Achieves sharp spatial partitions and higher-order polynomial approximation without regularizers.

AMORE uses neural operators to efficiently predict multiple thermochemical states in stiff chemical kinetics.

problem Efficiently integrating stiff chemical kinetics systems to reduce computational cost.
method Developed AMORE, a framework of adaptive multi-output operator network with two adaptive loss functions.
result Demonstrated improved accuracy and efficiency in predicting thermochemical states from initial conditions.

A new operator based on t-distributions improves NN classifiers' robustness to out-of-distribution samples.

problem NN classifiers assign extreme probabilities to out-of-distribution samples, leading to unreliable predictions.
method Derive a novel operator using t-distributions to model uncertainty more accurately.
result Classifiers using the new operator are more robust to out-of-distribution samples.

A-manifolds and A-bundles are manifolds and vector bundles modelled on a projective finitely generated module over a topological algebra A. In this paper we investigate the conditions under which an A-bundle is provided with an A-valued hermitian structure and a compatible connection, in case A is a commutative complet…

1998-10-15abs ↗pdf ↗

Gromov \cite{Gr1_1} and Dranishnikov \cite{Dr1_1} introduced asymptotic and coarse dimensions of proper metric spaces via quite different ways. We define coarse and asymptotic dimension of all metric spaces in a unified manner and we investigate relationships between them generalizing results of Dranishnikov \cite{Dr…

2005-06-27abs ↗pdf ↗

For non-compact manifolds with boundary we prove that bounded geometry defined by coordinate-free curvature bounds is equivalent to bounded geometry defined using bounds on the metric tensor in geodesic coordinates. We produce a nice atlas with subordinate partition of unity on manifolds with boundary of bounded geomet…

2000-01-19abs ↗pdf ↗

We study the Chern-Simons partition function of orthogonal quantum group invariants, and propose a new orthogonal Labastida-Mariño-Ooguri-Vafa conjecture as well as degree conjecture for free energy associated to the orthogonal Chern-Simons partition function. We prove the degree conjecture and some interesting cases o…

2010-07-09abs ↗pdf ↗

Resurgent analysis reveals full partition function for 3-manifold invariants.

problem Analyzing resurgence in 3-manifold invariants for SL(2,C)SL(2, \mathbb{C}).
method Resurgent analysis applied to infinite families of Seifert manifolds and torus knot complements.
result The contribution from abelian flat connections contains information of all non-abelian flat connections, indicating a full partition function.

We show a Whitney Approximation Theorem for a continuous map from a manifold to a smooth CW complex. This enables us to show that a topological CW complex is homotopy equivalent to a smooth CW complex in a category of topological spaces. It is also shown that, for any open covering of a smooth CW complex, there exists …

2020-01-09abs ↗pdf ↗

New insights into the top-K sparse softmax gating function for deep learning.

problem Understanding the theoretical effects of the top-K sparse softmax gating function on density and parameter estimations.
method Using a Gaussian mixture of experts, novel loss functions, and theoretical analysis.
result The convergence rates of density and parameter estimations are parametric under certain conditions, but slow under over-specified models.

We formally prove the connection between k-means clustering and the predictions of neural networks based on the softmax activation layer. In existing work, this connection has been analyzed empirically, but it has never before been mathematically derived. The softmax function partitions the transformed input space into…

2020-01-07abs ↗pdf ↗

Continuing the study of bounded geometry for Riemannian foliations, begun by Sanguiao, we introduce a chart-free definition of this concept. Our main theorem states that it is equivalent to a condition involving certain normal foliation charts. For this type of charts, it is also shown that the derivatives of the chang…

2013-08-02abs ↗pdf ↗

Given an open cover of a paracompact topological space X, there are two natural ways to construct a map from the cohomology of the nerve of the cover to the cohomology of X. One of them is based on a partition of unity, and is more topological in nature, while the other one relies on the Mayer-Vietoris double complex, …

2019-12-16abs ↗pdf ↗

K-Means and RBF networks are shown to be equivalent under certain conditions.

problem Discrete clustering vs. continuous optimization in machine learning.
method Established variational and gradient-based equivalence between K-Means and RBF networks.
result Gradient-based updates of RBF centers recover K-Means centroid update rule.

Differential chains are a proper subspace of de Rham currents given as an inductive limit of Banach spaces endowed with a geometrically defined strong topology. Boundary is a continuous operator, as are operators that dualize to Hodge star, Lie derivative, pullback and interior product. Partitions of unity exist in thi…

2012-10-16abs ↗pdf ↗

Recent research in coarse geometry revealed similarities between certain concepts of analysis, large scale geometry, and topology. Property A of G.Yu is the coarse analog of amenability for groups and its generalization (exact spaces) was later strengthened to be the large scale analog of paracompact spaces using parti…

2012-08-13abs ↗pdf ↗

We propose a new model for pricing Quanto CDS and risky bonds. The model operates with four stochastic factors, namely: hazard rate, foreign exchange rate, domestic interest rate, and foreign interest rate, and also allows for jumps-at-default in the FX and foreign interest rates. Corresponding systems of PDEs are deri…

2017-11-20abs ↗pdf ↗

Unified framework for SGMoE resolves estimation and selection issues.

problem Non-identifiability, coupled differential relations, and tight coupling in softmax-Gated models.
method Unified statistical framework with Voronoi-type loss functions and dendrograms of mixing measures.
result Consistent selection of the number of experts without model sweeps, optimal parameter rates under overfitting.

Softmax is an output activation function for modeling categorical probability distributions in many applications of deep learning. However, a recent study revealed that softmax can be a bottleneck of representational capacity of neural networks in language modeling (the softmax bottleneck). In this paper, we propose an…

2018-05-28abs ↗pdf ↗

The computational cost of training with softmax cross entropy loss grows linearly with the number of classes. For the settings where a large number of classes are involved, a common method to speed up training is to sample a subset of classes and utilize an estimate of the loss gradient based on these classes, known as…

2019-07-24abs ↗pdf ↗

The paper studies Nijenhuis operators with a unity and their connection to F-manifolds.

problem Understanding Nijenhuis operators and their relationship to F-manifolds.
method Established a Splitting Theorem for Nijenhuis operators with a unity and proved their equivalence to F-manifolds.
result The class of regular F-manifolds coincides with the class of Nijenhuis manifolds with a cyclic unity.

The paper extends ternary algebra concepts using cube roots of unity.

problem Extending algebraic structures from binary to ternary multiplication.
method Introducing ternary associator, commutator, and Lie algebra at cube roots of unity.
result Derived an identity for ternary commutator based on GA(1,5)GA(1,5).

For an arbitrary positive integer n, we construct infinitely many one-cusped hyperbolic 3-manifolds where each manifold's A-polynomial detects every n-th root of unity. This answers a question of Cooper, Culler, Gillet, Long, and Shalen as to which roots of unity arise in this manner.

2004-11-09abs ↗pdf ↗

Revises logistic-softmax likelihood for Bayesian meta-learning in few-shot classification.

problem Inherent uncertainty in logistic-softmax leads to suboptimal performance in meta-learning.
method Redesigns logistic-softmax likelihood with a temperature parameter for better control of prior confidence.
result Achieves well-calibrated uncertainty estimates and comparable/superior performance on benchmark datasets.

Paper introduces Balanced Meta-Softmax for better long-tailed visual recognition.

problem Long-tailed distribution mismatch between training and testing data.
method Balanced Meta-Softmax, an unbiased extension of Softmax, using a Meta Sampler.
result Balanced Meta-Softmax outperforms state-of-the-art solutions on visual recognition and instance segmentation.