Sparse random projection (RP) is a popular tool for dimensionality reduction that shows promising performance with low computational complexity. However, in the existing sparse RP matrices, the positions of non-zero entries are usually randomly selected. Although they adopt uniform sampling with replacement, due to lar…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithm optimizes positions of CountSketch non-zero entries for better data compression.
We consider the problem of selecting non-zero entries of a matrix in order to produce a sparse sketch of it, , that minimizes . For large matrices, such that (for example, representing observations over attributes) we give sampling distributions that exhibit four importa…
This work proposes a method to learn sparse representations that are more efficient for large-scale data retrieval.
Improved matrix completion for non-uniformly sampled data.
We show that complete uniform visibility manifolds of finite volume with sectional curvature have positive simplicial volumes. This implies that their minimal volumes are non-zero.
In this paper, we consider the streaming memory-limited matrix completion problem when the observed entries are noisy versions of a small random fraction of the original entries. We are interested in scenarios where the matrix size is very large so the matrix is very hard to store and manipulate. Here, columns of the o…
Wedge Sampling improves tensor completion with nearly-linear sample complexity.
Screening is the problem of finding a superset of the set of non-zero entries in an unknown p-dimensional vector β* given n noisy observations. Naturally, we want this superset to be as small as possible. We propose a novel framework for screening, which we refer to as Multiple Grouping (MuG), that groups variables, pe…
Matrix factorization (MF) has been widely used to discover the low-rank structure and to predict the missing entries of data matrix. In many real-world learning systems, the data matrix can be very high-dimensional but sparse. This poses an imbalanced learning problem, since the scale of missing entries is usually much…
Matrix completion is a classical problem in data science wherein one attempts to reconstruct a low-rank matrix while only observing some subset of the entries. Previous authors have phrased this problem as a nuclear norm minimization problem. Almost all previous work assumes no explicit structure of the matrix and uses…
Study on materials with disclinations, limiting their size.
In this paper, we consider matrix completion from non-uniformly sampled entries including fully observed and partially observed columns. Specifically, we assume that a small number of columns are randomly selected and fully observed, and each remaining column is partially observed with uniform sampling. To recover the …
We consider the problem of estimating the support of a vector based on observations contaminated by noise. A significant body of work has studied behavior of -relaxations when applied to measurement matrices drawn from standard dense ensembles (e.g., Gaussian, Bernoulli). In this paper,…
In this note, we establish a relationship between fractional Dehn twist coefficients of Riemann surface automorphisms and modular invariants of holomorphic families of algebraic curves. Specially, we give a characterization of pseudo-periodic maps with nontrivial fractional Dehn twist coefficients. We also obtain some …
Large-scale regression problems where both the number of variables, , and the number of observations, , may be large and in the order of millions or more, are becoming increasingly more common. Typically the data are sparse: only a fraction of a percent of the entries in the design matrix are non-zero. Neverthele…
We derive high-order compact finite difference schemes for option pricing in stochastic volatility models on non-uniform grids. The schemes are fourth-order accurate in space and second-order accurate in time for vanishing correlation. In our numerical study we obtain high-order numerical convergence also for non-zero …
This work discusses the problem of sparse signal recovery when there is correlation among the values of non-zero entries. We examine intra-vector correlation in the context of the block sparse model and inter-vector correlation in the context of the multiple measurement vector model, as well as their combination. Algor…
Uniform approximations for RHTs improve kernel approximation and distance estimation.
A low rank matrix X has been contaminated by uniformly distributed noise, missing values, outliers and corrupt entries. Reconstruction of X from the singular values and singular vectors of the contaminated matrix Y is a key problem in machine learning, computer vision and data science. In this paper we show that common…
Improved 2-bit covariance estimator with reduced operator norm error and no tuning needed.
This paper improves matrix completion by leveraging element importance and non-uniform sampling.
New systolic inequality for 3D contact forms on Seifert bundles.
We present high-order compact schemes for a linear second-order parabolic partial differential equation (PDE) with mixed second-order derivative terms in two spatial dimensions. The schemes are applied to option pricing PDE for a family of stochastic volatility models. We use a non-uniform grid with more grid-points ar…
We present a system and a set of techniques for learning linear predictors with convex losses on terascale datasets, with trillions of features, {The number of features here refers to the number of non-zero entries in the data matrix.} billions of training examples and millions of parameters in an hour using a cluster …
New method corrects bias in missing data for matrix completion.
Detecting a planted submatrix in random matrices with non-asymptotic methods.
Active seriation recovers item order from noisy pairwise similarity measurements.
We consider a high dimensional linear regression problem where the goal is to efficiently recover an unknown vector from noisy linear observations , for known and unknown . Unlike most of the literature on this model we make no spa…
The task of reconstructing a matrix given a sample of observedentries is known as the matrix completion problem. It arises ina wide range of problems, including recommender systems, collaborativefiltering, dimensionality reduction, image processing, quantum physics or multi-class classificationto name a few. Most works…
We present a technique for significantly speeding up Alternating Least Squares (ALS) and Gradient Descent (GD), two widely used algorithms for tensor factorization. By exploiting properties of the Khatri-Rao product, we show how to efficiently address a computationally challenging sub-step of both algorithms. Our algor…
Low-rank matrix completion is an important problem with extensive real-world applications. When observations are uniformly sampled from the underlying matrix entries, existing methods all require the matrix to be incoherent. This paper provides the first working method for coherent matrix completion under the standard …
A new method for matrix completion with model-free weights.
The Burau representation of braid group B4 is shown to be faithful almost everywhere.
This paper studies the matrix completion problem under arbitrary sampling schemes. We propose a new estimator incorporating both max-norm and nuclear-norm regularization, based on which we can conduct efficient low-rank matrix recovery using a random subset of entries observed with additive noise under general non-unif…
Optimal subspace embedding with near-optimal sparsity for high-dimensional data.
Uniform proof of -injectivity for certain maps in low dimensions.
Example surfaces with curvature bounds but no Laplacian eigenvalue lower bound.
The support recovery problem consists of determining a sparse subset of variables that is relevant in generating a set of observations. In this paper, we study the support recovery problem in the phase retrieval model consisting of noisy phaseless measurements, which arises in a diverse range of settings such as optica…
We build upon probabilistic models for Boolean Matrix and Boolean Tensor factorisation that have recently been shown to solve these problems with unprecedented accuracy and to enable posterior inference to scale to Billions of observation. Here, we lift the restriction of a pre-specified number of latent dimensions by …
Mixture models and topic models generate each observation from a single cluster, but standard variational posteriors for each observation assign positive probability to all possible clusters. This requires dense storage and runtime costs that scale with the total number of clusters, even though typically only a few clu…
We introduce a variant of (sparse) PCA in which the set of feasible support sets is determined by a graph. In particular, we consider the following setting: given a directed acyclic graph on vertices corresponding to variables, the non-zero entries of the extracted principal component must coincide with vertice…
Low-rank matrix completion (LRMC) problems arise in a wide variety of applications. Previous theory mainly provides conditions for completion under missing-at-random samplings. This paper studies deterministic conditions for completion. An incomplete matrix is finitely rank- completable if there are at …
We consider the high-dimensional sparse linear regression problem of accurately estimating a sparse vector using a small number of linear measurements that are contaminated by noise. It is well known that the standard cadre of computationally tractable sparse regression algorithms---such as the Lasso, Orthogonal Matchi…
The task of estimating a matrix given a sample of observed entries is known as the \emph{matrix completion problem}. Most works on matrix completion have focused on recovering an unknown real-valued low-rank matrix from a random sample of its entries. Here, we investigate the case of highly quantized observations when …
The split version of the Freudenthal-Tits magic square stems from Lie theory and constructs a Lie algebra starting from two split composition algebras [3, 17, 18]. The geometries appearing in the second row are Severi-Brauer varieties [20]. We provide an easy uniform axiomatization of these geometries and related ones,…
Structure discovery in graphical models is the determination of the topology of a graph that encodes conditional independence properties of the joint distribution of all variables in the model. For some class of probability distributions, an edge between two variables is present if and only if the corresponding entry i…
The paper proves conditions for non-uniform expansion in partially hyperbolic systems.