Improved Transformer performance by addressing 'explaining away' effect.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study finds discrepancies in model explanations for clean vs. adversarial inputs.
FCDD explains deep anomaly detection by mapping anomalies away and providing heatmap explanations.
Principal component analysis (PCA) is a mainstay of modern data analysis - a black box that is widely used but (sometimes) poorly understood. The goal of this paper is to dispel the magic behind this black box. This manuscript focuses on building a solid intuition for how and why principal component analysis works. Thi…
We introduce a factor analysis model that summarizes the dependencies between observed variable groups, instead of dependencies between individual variables as standard factor analysis does. A group may correspond to one view of the same set of objects, one of many data sets tied by co-occurrence, or a set of alternati…
Classifiers used in the wild, in particular for safety-critical systems, should not only have good generalization properties but also should know when they don't know, in particular make low confidence predictions far away from the training data. We show that ReLU type neural networks which yield a piecewise linear cla…
We investigate the problem of estimating a given real symmetric signal matrix from a noisy observation matrix in the limit of large dimension. We consider the case where the noisy measurement comes either from an arbitrary additive or multiplicative rotational invariant perturbati…
We propose a simple probabilistic model to explain the spatial structure of the rent distribution of housing market in city of Sapporo. Here we modify the mathematical model proposed by Gauvin et. al. Especially, we consider the competition between two distances, namely, the distance between house and center, and the d…
A new ODE model explains gradient descent dynamics near edge of stability.
An important preprocessing step in most data analysis pipelines aims to extract a small set of sources that explain most of the data. Currently used algorithms for blind source separation (BSS), however, often fail to extract the desired sources and need extensive cross-validation. In contrast, their rarely used probab…
Market inefficiencies arise from density-dependent returns in a noisy environment.
Throwing away data can improve worst-group error in imbalanced datasets.
Explains classic quantitative strategies and their workings.
We propose a novel classification model for weak signal data, building upon a recent model for Bayesian multi-view learning, Group Factor Analysis (GFA). Instead of assuming all data to come from a single GFA model, we allow latent clusters, each having a different GFA model and producing a different class distribution…
The study examines 4D steady gradient Ricci solitons with nonnegative curvature away from a compact set.
We consider complete spacelike hypersurfaces with constant mean curvature in the open region of de Sitter space known as the steady state space. We prove that if the hypersurface is bounded away from the infinity of the ambient space, then the mean curvature must be H=1. Moreover, in the 2-dimensional case we obtain th…
We study the blowup behavior at infinity of the normalized Kahler-Ricci flow on a Fano manifold which does not admit Kahler-Einstein metrics. We prove an estimate for the Kahler potential away from a multiplier ideal subscheme, which implies that the volume forms along the flow converge to zero locally uniformly away f…
Markov's theorem classifies the worst irrational numbers with respect to rational approximation and the indefinite binary quadratic forms whose values for integer arguments stay farthest away from zero. The main purpose of this paper is to present a new proof of Markov's theorem using hyperbolic geometry. The main ingr…
New theory explains how equivariant self-supervised learning improves feature extraction.
Smooths metrics with nonnegative scalar curvature near singular sets.
A closed Riemannian manifold is said to have cross blocking if whenever distinct points p and q are at distance less than the diameter, all light rays from p can be shaded away from q with at most two point shades. Similarly, a closed Riemannian manifold is said to have sphere blocking if for each point p, all the ligh…
We show the optimal regularity of geodesics in nef and big cohomology class on Kähler manifolds away from the non-Kähler locus, assuming sufficiently regular initial data. As a special case, we prove the regularity of geodesics of Kähler metrics on compact Kähler varieties away from the singular loc…
We use a counting argument and surgery theory to show that if is a sufficiently general algebraic hypersurface in , then any local diffeomorphism of simply connected manifolds which is a -sheeted cover away from has degree or (however all degrees are poss…
In high-dimensional data, structured noise caused by observed and unobserved factors affecting multiple target variables simultaneously, imposes a serious challenge for modeling, by masking the often weak signal. Therefore, (1) explaining away the structured noise in multiple-output regression is of paramount importanc…
Unified framework for efficient Frank-Wolfe optimization of Dominant Set Clustering.
Improved Frank-Wolfe algorithm for polytopes converges linearly with dimension dependence on optimal face.
We provide a simple convergence proof for Adam and Adagrad.
We show that every knot is one crossing change away from a knot of arbitrarily high bridge number and arbitrarily high bridge distance.
Private anchors affect how information is communicated and can improve or distort transmission.
Recently, there has been a renewed interest in the machine learning community for variants of a sparse greedy approximation procedure for concave optimization known as {the Frank-Wolfe (FW) method}. In particular, this procedure has been successfully applied to train large-scale instances of non-linear Support Vector M…
Smooth solutions found for a specific type of Yamabe problem.
Stochastic gradient descent approximates Gaussian process posteriors efficiently.
In this article, we study the higher-order regularity of the Kähler-Ricci flow on compact Kähler manifolds with semi-ample canonical line bundle. We proved, using a parabolic analogue of Hein-Tosatti's work on collapsing Calabi-Yau metrics, that when the generic fibers of the Iitaka fibration are biholomorphic to each …
Gravitational waves are predicted by the general theory of relativity. In [6] D. Christodoulou showed that gravitational waves have a nonlinear memory. We proved in [3] that the electromagnetic field contributes at highest order to the nonlinear memory effect of gravitational waves. In the present paper, we study this …
We study "how far away" a finite index subgroup G of SL(2,Z) is from being a congruence group. For this we define its deficiency of being a congruence group. We show that the index of the image of G in SL(2,Z/nZ) is biggest, if n is the general Wohlfahrt level. We furthermore show that the Veech groups of origamis (or …
We establish new obstruction results to the existence of Riemannian metrics on tori satisfying mixed bounds on both their sectional and Ricci curvatures. More precisely, from Lohkamp's theorem, every torus of dimension at least three admits Riemannian metrics with negative Ricci curvature. We show that the sectional cu…
In this note we explore a connection between finite covers of surfaces and the Teichmüller polynomial of a fibered face of a hyperbolic 3--manifold. We consider the action of a homological pseudo-Anosov homeomorphism on the homology groups of a class of finite abelian covers of a surface . Eigenspaces of t…
Several works have aimed to explain why overparameterized neural networks generalize well when trained by Stochastic Gradient Descent (SGD). The consensus explanation that has emerged credits the randomized nature of SGD for the bias of the training process towards low-complexity models and, thus, for implicit regulari…
We prove that, generically, magnetic geodesics on surfaces will turn away from points with lightlike tangent planes, and we motivate our result with numerical solutions for closed magnetic geodesics.
The study improves norms of spectral projectors on specific surfaces.
In this paper we prove several results on the geometry of surfaces immersed in with small or bounded norm of . For instance, we prove that if the norm of and the norm of , , are sufficiently small, then such a surface is graphical away from its boundary. We also prove …
We prove that the displacement energy of a stable coisotropic submanifold is bounded away from zero if the ambient symplectic manifold is closed, rational and satisfies a mild topological condition.
Deep neural networks have dramatically achieved great success on a variety of challenging tasks. However, most successful DNNs have an extremely complex structure, leading to extensive research on model compression.As a significant area of progress in model compression, traditional gradual pruning approaches involve an…
We prove the existence of Cannon-Thurston maps for Kleinian groups corresponding to pared manifolds whose boundary is incompressible away from cusps. We also describe the structure of these maps in terms of ending laminations.
We prove a criterion of convergence in the augmented Teichmueller space that can be phrased in terms of convergence of the hyperbolic metrics or of quasiconformal convergence away from the nodes.
Negative curvature manifolds have vanishing bounded volume class if and only if Cheeger constant is positive.
In this paper we prove convergence and compactness results for Ricci flows with bounded scalar curvature and entropy. More specifically, we show that Ricci flows with bounded scalar curvature converge smoothly away from a singular set of codimension . We also establish a general form of the Hamilton-Tian Conjec…
The study compares different game-theoretic attribution methods and finds that interventional Shapley values yield less consistent results than Aumann-Shapley due to path symmetry.