Solves portfolio optimization with cardinality constraints using column generation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper introduces MU for NMF with -divergences and disjoint constraints.
Paper tackles BNSL with IP, improving quality of solutions.
New method solves large-scale linear programming problems for sparse signal reconstruction.
We study the problem of instance segmentation in biological images with crowded and compact cells. We formulate this task as an integer program where variables correspond to cells and constraints enforce that cells do not overlap. To solve this integer program, we propose a column generation formulation where the prici…
New algorithms solve L1-regularized SVMs and related LPs, outperforming existing methods.
The paper tackles online resource allocation with uncertain coefficients and chance constraints.
We consider the problem of performing matrix completion with side information on row-by-row and column-by-column similarities. We build upon recent proposals for matrix estimation with smoothness constraints with respect to row and column graphs. We present a novel iterative procedure for directly minimizing an informa…
Paper presents an efficient algorithm for learning minimax risk classifiers with large-scale data.
NMF with specific constraints is equivalent to LDA.
New rule-based method for classification with scalability, interpretability, and fairness.
Develops a non-parametric Dirichlet process method for probabilistic biclustering.
A common problem in large-scale data analysis is to approximate a matrix using a combination of specifically sampled rows and columns, known as CUR decomposition. Unfortunately, in many real-world environments, the ability to sample specific individual rows or columns of the matrix is limited by either system constrain…
Optimized portfolio turnover strategies enhance wealth and reduce costs.
Nonnegative matrix factorization (NMF) has been shown to be identifiable under the separability assumption, under which all the columns(or rows) of the input data matrix belong to the convex cone generated by only a few of these columns(or rows) [1]. In real applications, however, such separability assumption is hard t…
RECol generates error columns to improve outlier detection.
Transposable data represents interactions among two sets of entities, and are typically represented as a matrix containing the known interaction values. Additional side information may consist of feature vectors specific to entities corresponding to the rows and/or columns of such a matrix. Further information may also…
Paper introduces a novel matrix-wise sparse MNNLS formulation and algorithm.
The column group is a subgroup of the symmetric group on the elements of a finite blackboard birack generated by the column permutations in the birack matrix. We use subgroups of the column group associated to birack homomorphisms to define an enhancement of the integral birack counting invariant and give examples whic…
The paper tackles one-for-many counterfactual explanations using column generation.
It has been recently shown that numerical semiparametric bounds on the expected payoff of fi- nancial or actuarial instruments can be computed using semidefinite programming. However, this approach has practical limitations. Here we use column generation, a classical optimization technique, to address these limitations…
New algorithms optimize matrix manifolds, converging faster than existing methods.
This paper defines a generalized column subset selection problem which is concerned with the selection of a few columns from a source matrix A that best approximate the span of a target matrix B. The paper then proposes a fast greedy algorithm for solving this problem and draws connections to different problems that ca…
New method recovers matrix column space with active sampling for better results.
TGAN models tabular data using GAN for realistic synthetic data.
Matrix completion, i.e., the exact and provable recovery of a low-rank matrix from a small subset of its elements, is currently only known to be possible if the matrix satisfies a restrictive structural constraint---known as {\em incoherence}---on its row and column spaces. In these cases, the subset of elements is sam…
A collaborative convex framework for factoring a data matrix into a non-negative product , with a sparse coefficient matrix , is proposed. We restrict the columns of the dictionary matrix to coincide with certain columns of the data matrix , thereby guaranteeing a physically meaningful dictionary and …
We consider the problem of matrix column subset selection, which selects a subset of columns from an input matrix such that the input can be well approximated by the span of the selected columns. Column subset selection has been applied to numerous real-world data applications such as population genetics summarization,…
This paper studies quandles with one non-trivial column and their properties.
In this paper, we present GASG21 (Grassmannian Adaptive Stochastic Gradient for norm minimization), an adaptive stochastic gradient algorithm to robustly recover the low-rank subspace from a large matrix. In the presence of column outliers, we reformulate the batch mode matrix norm minimization with…
In this paper, we consider matrix completion from non-uniformly sampled entries including fully observed and partially observed columns. Specifically, we assume that a small number of columns are randomly selected and fully observed, and each remaining column is partially observed with uniform sampling. To recover the …
Double autoencoder improves missing value imputation in recommender systems.
End-to-end pipeline for data-driven decision making in mixed-integer optimization.
Neuromorphic column performs online unsupervised clustering.
Unified methods for fast column selection in various applications.
Paper tackles low-rank matrix recovery with column -norm regularization.
We propose a general framework for reduced-rank modeling of matrix-valued data. By applying a generalized nuclear norm penalty we can directly model low-dimensional latent variables associated with rows and columns. Our framework flexibly incorporates row and column features, smoothing kernels, and other sources of sid…
In this note we describe how some objects from generalized geometry appear in the qualitative analysis and numerical simulation of mechanical systems. In particular we discuss double vector bundles and Dirac structures. It turns out that those objects can be naturally associated to systems with constraints -- we recall…
Representation learning is typically applied to only one mode of a data matrix, either its rows or columns. Yet in many applications, there is an underlying geometry to both the rows and the columns. We propose utilizing this coupled structure to perform co-manifold learning: uncovering the underlying geometry of both …
This paper explores the use of Column Generation (CG) techniques in constructing univariate binary decision trees for classification tasks. We propose a novel Integer Linear Programming (ILP) formulation, based on root-to-leaf paths in decision trees. The model is solved via a Column Generation based heuristic. To spee…
We propose and study a row-and-column affine measurement scheme for low-rank matrix recovery. Each measurement is a linear combination of elements in one row or one column of a matrix . This setting arises naturally in applications from different domains. However, current algorithms developed for standard matrix rec…
This paper considers the problem of completing a matrix with many missing entries under the assumption that the columns of the matrix belong to a union of multiple low-rank subspaces. This generalizes the standard low-rank matrix completion problem to situations in which the matrix rank can be quite high or even full r…
Proposes a new matrix factorization model for interval-valued matrices.
The paper identifies redundant columns in matrices for feature selection and clustering.
We consider the problem of recovering a complete (i.e., square and invertible) matrix , from with , provided is sufficiently sparse. This recovery problem is central to theoretical understanding of dictionary learnin…
The paper proposes a method to optimize rule-based models for better accuracy and interpretability.
Paper tackles fair low-rank approximation and column subset selection.
Sherlock uses deep learning to accurately detect data types from column headers.