Paper refines cross-lingual word embeddings using Manhattan norm.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Any Sasakian structure can be closely mimicked by embeddings into weighted spheres.
We develop embeddings for nonlinear subspaces preserving vector norms.
After appropriate normalizations an embedded disk whose second fundamental form has large norm contains a multi-valued graph, provided the L^P norm of the mean curvature is sufficiently small. This generalizes to non-minimal surfaces a well known result of Colding and Minicozzi.
New insights show embedding lengths correlate with semantic properties.
In this paper, we propose a method for estimating the Sobolev type embedding constant on a domain with minimally smooth boundary. We estimate the embedding constant by constructing an extension operator and computing its operator norm. We also present some examples of estimating the embedding constant for certain domai…
We present a lower bound for a fragmentation norm and construct a bi-Lipschitz embedding with respect to the fragmentation norm on the group of Hamiltonian diffeomorphisms of a symplectic manifold . As an application, we provide an answer to Brandenbursk…
We prove that the embedding of the quaternionic hyperbolic disc into quaternionic hyperbolic -space is tight and thereby obtain the value of the Gromov norm of the quaternionic Kähler class.
Study improves Poincaré-Sobolev inequalities for differential forms.
Kernel interpolation is inconsistent for norms with smoothness above a constant.
A new clustering method improves recovery guarantees by re-embedding data.
In this paper we prove several results on the geometry of surfaces immersed in with small or bounded norm of . For instance, we prove that if the norm of and the norm of , , are sufficiently small, then such a surface is graphical away from its boundary. We also prove …
Clustering is an effective technique in data mining to group a set of objects in terms of some attributes. Among various clustering approaches, the family of K-Means algorithms gains popularity due to simplicity and efficiency. However, most of existing K-Means based clustering algorithms cannot deal with outliers well…
Estimates for the norm of the second fundamental form, , play a crucial role in studying the geometry of surfaces. In fact, when is bounded the surface cannot bend too sharply. In this paper we prove that for an embedded geodesic disk with bounded norm of , is bounded at interior points, pro…
Adversarial attacks aim to confound machine learning systems, while remaining virtually imperceptible to humans. Attacks on image classification systems are typically gauged in terms of -norm distortions in the pixel feature space. We perform a behavioral study, demonstrating that the pixel -norm for any $0\le p …
Optimal subspace embedding with near-optimal sparsity for high-dimensional data.
We consider the smooth inverse mean curvature flow of strictly convex hypersurfaces with boundary embedded in which are perpendicular to the unit sphere from the inside. We prove that the flow hypersurfaces converge to the embedding of a flat disk in the norm of
Learning rates for least-squares regression are typically expressed in terms of -norms. In this paper we extend these rates to norms stronger than the -norm without requiring the regression function to be contained in the hypothesis space. In the special case of Sobolev reproducing kernel Hilbert spaces used …
We propose and evaluate new techniques for compressing and speeding up dense matrix multiplications as found in the fully connected and recurrent layers of neural networks for embedded large vocabulary continuous speech recognition (LVCSR). For compression, we introduce and study a trace norm regularization technique f…
Survey of Thurston norm properties and connections to 3-manifold invariants.
The paper studies how norms of random vectors are preserved by random projections.
In this paper we prove that an embedded and simply connected constant mean curvature surface with curvature large at a point contains a multi-valued graph around that point on the scale of , where is the norm squared of the second fundamental form. This generalizes Colding and Minicozzi's result for mini…
Every element in the first cohomology group of a 3--manifold is dual to embedded surfaces. The Thurston norm measures the minimal `complexity' of such surfaces. For instance the Thurston norm of a knot complement determines the genus of the knot in the 3--sphere. We show that the degrees of twisted Alexander polynomial…
For graphs generated from stochastic blockmodels, adjacency spectral embedding is asymptotically consistent. Further, adjacency spectral embedding composed with universally consistent classifiers is universally consistent to achieve the Bayes error. However when the graph contains private or sensitive information, trea…
Let be an -component link () with pairwise nonzero linking numbers in a rational homology -sphere . Assume the link complement has nondegenerate Thurston norm. In this paper, we study when a Thurston norm-minimizing surface properly embedded in remains norm-minimizing after…
We give a direct proof for the asymptotic faithfulness of the quantum representations of the mapping class groups using peak sections in Kodaira embedding. We give also estimates on the norm of the parallell transport of the projective connection on the Verlinde bundle. The faithfulness has been proved earlier …
Improved eigenvalue bounds for minimal hypersurfaces in spheres.
Estimates spectral projections restricted to uniformly embedded submanifolds.
We prove, using the subspace embedding guarantee in a black box way, that one can achieve the spectral norm guarantee for approximate matrix multiplication with a dimensionality-reducing map having rows. Here is the maximum stable rank, i.e. squared ratio of Frobenius and op…
For supervised and unsupervised learning, positive definite kernels allow to use large and potentially infinite dimensional feature spaces with a computational cost that only depends on the number of observations. This is usually done through the penalization of predictor functions by Euclidean or Hilbertian norms. In …
We consider closed immersed hypersurfaces evolving by surface diffusion flow, and perform an analysis based on local and global integral estimates. First we show that a properly immersed stationary (ΔH \equiv 0) hypersurface in \R^3 or \R^4 with restricted growth of the curvature at infinity and small total tracefree c…
MCE reduces embedding instability in nonlinear dimensionality reduction.
In this paper, we study the problem of approximately computing the product of two real matrices. In particular, we analyze a dimensionality-reduction-based approximation algorithm due to Sarlos [1], introducing the notion of nuclear rank as the ratio of the nuclear norm over the spectral norm. The presented bound has i…
Proposes DP-MERF for privacy-preserving synthetic data generation.
Improved molecular property prediction using WL embedding in GNNs.
New learning rates for embeddings in RKHSs, even when the target is not Hilbert-Schmidt.
The extraction of clusters from a dataset which includes multiple clusters and a significant background component is a non-trivial task of practical importance. In image analysis this manifests for example in anomaly detection and target detection. The traditional spectral clustering algorithm, which relies on the lead…
Low-rank matrix recovery has found many applications in science and engineering such as machine learning, signal processing, collaborative filtering, system identification, and Euclidean embedding. But the low-rank matrix recovery problem is an NP hard problem and thus challenging. A commonly used heuristic approach is…
Algorithms compute length spectra of torus graphs efficiently.
We study biinvariant word metrics on groups. We provide an efficient algorithm for computing the biinvariant word norm on a finitely generated free group and we construct an isometric embedding of a locally compact tree into the biinvariant Cayley graph of a nonabelian free group. We investigate the geometry of cyclic …
New flows represent Thurston norm ball faces, differing by veering mutations.
Most of existing manifold learning methods rely on Mean Squared Error (MSE) or norm. However, for the problem of image quality assessment, these are not promising measure. In this paper, we introduce the concept of an image structure manifold which captures image structure features and discriminates image dist…
In this note we prove that for each positive integer there exists a bi-Lipschitz embedding , where is equipped with the entropy metric. In particular, the same result holds when the entropy metric is substituted with the autonomous metric.
Let K be an algebraically closed field of characteristic zero, endowed with a complete nonarchimedean norm. Let X be a K-rigid analytic variety and Σa semianalytic subset of X. Then the closure of Σin X with respect to the canonical topology is again semianalytic. The proof uses Embedded Resolution of Singularities.
The study determines -Thurston norms in Sol manifolds and embeds non-orientable surfaces.
TensorSketch is an oblivious linear sketch introduced in Pagh'13 and later used in Pham, Pagh'13 in the context of SVMs for polynomial kernels. It was shown in Avron, Nguyen, Woodruff'14 that TensorSketch provides a subspace embedding, and therefore can be used for canonical correlation analysis, low rank approximation…
Using sparse-inducing norms to learn robust models has received increasing attention from many fields for its attractive properties. Projection-based methods have been widely applied to learning tasks constrained by such norms. As a key building block of these methods, an efficient operator for Euclidean projection ont…
Existing ordinal embedding methods usually follow a two-stage routine: outlier detection is first employed to pick out the inconsistent comparisons; then an embedding is learned from the clean data. However, learning in a multi-stage manner is well-known to suffer from sub-optimal solutions. In this paper, we propose a…