This paper introduces a new data-driven methodology for estimating sparse covariance matrices of the random coefficients in logit mixture models. Researchers typically specify covariance matrices in logit mixture models under one of two extreme assumptions: either an unrestricted full covariance matrix (allowing correl…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Nonnegative Matrix Factorization (NMF) has been a popular representation method for pattern classification problem. It tries to decompose a nonnegative matrix of data samples as the product of a nonnegative basic matrix and a nonnegative coefficient matrix, and the coefficient matrix is used as the new representation. …
The fused lasso is analyzed for high-dimensional piecewise-constant regression coefficients.
Improved portfolio optimization using Kendall-like correlation coefficients.
Wilson lines generate positive Laurent polynomials in decorated triangulations.
The paper improves matrix completion with auxiliary covariates using LS estimation.
We introduce a method to predict which correlation matrix coefficients are likely to change their signs in the future in the high-dimensional regime, i.e. when the number of features is larger than the number of samples per feature. The stability of correlation signs, two-by-two relationships, is found to depend on thr…
We categorify the coefficients of the Burau representation matrix using elementary geometrical methods. We show the faithfulness of this categorification in the sense that it detects the trivial braid.
We propose a novel algorithm for efficiently computing a sparse directed adjacency matrix from a group of time series following a causal graph process. Our solution is scalable for both dense and sparse graphs and automatically selects the LASSO coefficient to obtain an appropriate number of edges in the adjacency matr…
New method improves DAG learning by using large coefficients for higher-order terms.
The paper extends Pearson correlation to multi-variables, useful for noise measurement and feature selection.
We consider the problem of predicting several response variables using the same set of explanatory variables. This setting naturally induces a group structure over the coefficient matrix, in which every explanatory variable corresponds to a set of related coefficients. Most of the existing methods that utilize this gro…
Activity coefficients, which are a measure of the non-ideality of liquid mixtures, are a key property in chemical engineering with relevance to modeling chemical and phase equilibria as well as transport processes. Although experimental data on thousands of binary mixtures are available, prediction methods are needed t…
Paper finds a lower bound for estimating low-rank matrices in logistic regression.
We present a formula for the trace of any symmetric power of a matrix (with coefficients in a field) in terms of the ordinary powers of the matrix, an arbitrarily chosen linear function which vanishes on the identity matrix, and polynomial functions defined recursively.
Classical scalar-response regression methods treat covariates as a vector and estimate a corresponding vector of regression coefficients. In medical applications, however, regressors are often in a form of multi-dimensional arrays. For example, one may be interested in using MRI imaging to identify which brain regions …
Anosov groups study matrix coefficients and orbit counting in symmetric spaces.
We propose dimension reduction methods for sparse, high-dimensional multivariate response regression models. Both the number of responses and that of the predictors may exceed the sample size. Sometimes viewed as complementary, predictor selection and rank reduction are the most popular strategies for obtaining lower-d…
We compare some methods recently used in the literature to detect the existence of a certain degree of common behavior of stock returns belonging to the same economic sector. Specifically, we discuss methods based on random matrix theory and hierarchical clustering techniques. We apply these methods to a portfolio of s…
CAST improves spectral clustering for multi-scale data by integrating reachability similarity.
The paper constructs Goeritz matrices from Dehn colorings.
Motivated by the Bagging Partial Least Squares (PLS) and Principal Component Analysis (PCA) algorithms, we propose a Principal Model Analysis (PMA) method in this paper. In the proposed PMA algorithm, the PCA and the PLS are combined. In the method, multiple PLS models are trained on sub-training sets, derived from the…
In this paper, we propose an online algorithm to compute matrix factorizations. Proposed algorithm updates the dictionary matrix and associated coefficients using a single observation at each time. The algorithm performs low-rank updates to dictionary matrix. We derive the algorithm by defining a simple objective funct…
Subspace recovery from corrupted and missing data is crucial for various applications in signal processing and information theory. To complete missing values and detect column corruptions, existing robust Matrix Completion (MC) methods mostly concentrate on recovering a low-rank matrix from few corrupted coefficients w…
Develops a new multivariate regression model for complex outcomes.
Multivariate regression model is a natural generalization of the classical univari- ate regression model for fitting multiple responses. In this paper, we propose a high- dimensional multivariate conditional regression model for constructing sparse estimates of the multivariate regression coefficient matrix that accoun…
Study examines NFT market dynamics using correlation and noise analysis.
Derives adjoint formulas for matrix operations and applies them to specific cases.
We present two formulas for Chern classes of the tensor product of two vector bundles. In the first formula we consider a matrix containing Chern classes of the first bundle and we take a polynomial of this matrix with Chern classes of the second bundle as coefficients. The determinant of this expression equals the Che…
A representative model in integrative analysis of two high-dimensional correlated datasets is to decompose each data matrix into a low-rank common matrix generated by latent factors shared across datasets, a low-rank distinctive matrix corresponding to each dataset, and an additive noise matrix. Existing decomposition …
We consider the problem of multivariate regression in a setting where the relevant predictors could be shared among different responses. We propose an algorithm which decomposes the coefficient matrix into the product of a long matrix and a wide matrix, with an elastic net penalty on the former and an penalty …
This paper concerns cluster algebras with principal coefficients A(S,M) associated to bordered surfaces (S,M), and is a companion to a concurrent work of the authors with Schiffler [MSW2]. Given any (generalized) arc or loop in the surface -- with or without self-intersections -- we associate an element of (the fractio…
We consider the dictionary learning problem, where the aim is to model the given data as a linear combination of a few columns of a matrix known as a dictionary, where the sparse weights forming the linear combination are known as coefficients. Since the dictionary and coefficients, parameterizing the linear model are …
Novel approach ensures stability of compact schemes for variable PDEs.
This talk is a report on joint work with A. Vaintrob [arXiv:math.CO/0109104 and math.GT/0111102]. It is organised as follows. We begin by recalling how the classical Matrix-Tree Theorem relates two different expressions for the lowest degree coefficient of the Alexander-Conway polynomial of a link. We then state our fo…
Ridge leverage scores provide a balance between low-rank approximation and regularization, and are ubiquitous in randomized linear algebra and machine learning. Deterministic algorithms are also of interest in the moderately big data regime, because deterministic algorithms provide interpretability to the practitioner …
SCOPE estimator improves covariance and precision matrix estimation.
Meta-learning improves predictions with generalized ridge regression in high-dimensional settings.
For the problems of low-rank matrix completion, the efficiency of the widely-used nuclear norm technique may be challenged under many circumstances, especially when certain basis coefficients are fixed, for example, the low-rank correlation matrix completion in various fields such as the financial market and the low-ra…
We consider a fundamental algorithmic question in spectral graph theory: Compute a spectral sparsifier of random-walk matrix-polynomial where is the adjacency matrix of a weighted, undirected graph, is the diagonal matrix of weighted degrees, and are nonn…
Linear regression models depend directly on the design matrix and its properties. Techniques that efficiently estimate model coefficients by partitioning rows of the design matrix are increasingly popular for large-scale problems because they fit well with modern parallel computing architectures. We propose a simple me…
This paper describes the connection between scattering matrices on conformally compact asymptotically Einstein manifolds and conformally invariant objects on their boundaries at infinity. The conformally invariant powers of the Laplacian arise as residues of the scattering matrix and Branson's Q-curvature in even dimen…
We study the cross-correlation matrix of inventory variations of the most active individual and institutional investors in an emerging market to understand the dynamics of inventory variations. We find that the distribution of cross-correlation coefficient has a power-law form in the bulk followed by …
The performance of Orthogonal Matching Pursuit (OMP) for variable selection is analyzed for random designs. When contrasted with the deterministic case, since the performance is here measured after averaging over the distribution of the design matrix, one can have far less stringent sparsity constraints on the coeffici…
We introduce the probabilistic sequential matrix factorization (PSMF) method for factorizing time-varying and non-stationary datasets consisting of high-dimensional time-series. In particular, we consider nonlinear Gaussian state-space models where sequential approximate inference results in the factorization of a data…
We use the explicit relation between genus filtrated -loop means of the Gaussian matrix model and terms of the genus expansion of the Kontsevich--Penner matrix model (KPMM), which is the generating function for volumes of discretized (open) moduli spaces (discrete volumes), to express Gaussian means…
Existing nonnegative matrix factorization methods focus on learning global structure of the data to construct basis and coefficient matrices, which ignores the local structure that commonly exists among data. In this paper, we propose a new type of nonnegative matrix factorization method, which learns local similarity …
The multiplicative update (MU) algorithm has been extensively used to estimate the basis and coefficient matrices in nonnegative matrix factorization (NMF) problems under a wide range of divergences and regularizers. However, theoretical convergence guarantees have only been derived for a few special divergences withou…