DeepTMR reorders matrices without prior knowledge of structural patterns.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
AutoLL uses neural networks to automatically reorder graph nodes for linear layouts.
PRRO generates synthetic tabular data that improves SL performance and class distribution.
New algorithm optimizes matrix reordering for noisy disordered matrices.
We explore under what conditions one can obtain a nontrivial knot, given a collection of vectors. First, we show how to get a crossing from any 3 vectors equal in magnitude, by arbitrarily picking 2 vectors and identifying the sufficient and necessary criteria for picking a third vector that will guarantee a crossi…
Maximal extractable value in CFMMs can degrade or improve routing quality, with reordering MEV showing logarithmic impact.
This work introduces uncertainty principles to mitigate Maximal Extractable Value in blockchain systems.
Ethereum block builders can earn up to $14M/month by reordering transactions, harming users.
Study examines trading costs on Uniswap, finding adversarial slippage is significant for large trades and certain assets.
Efficiently learns and transports posterior densities for real-time inference.
Designing deep learning models for highly-constrained hardware would allow imbuing many edge devices with intelligence. Microcontrollers (MCUs) are an attractive platform for building smart devices due to their low cost, wide availability, and modest power usage. However, they lack the computational resources to run ne…
Unified optimization framework for matrix seriation.
New method recovers matrix column space with active sampling for better results.
Paper improves uncertainty estimation in LLM-as-a-judge systems.
We consider the problem of matrix column subset selection, which selects a subset of columns from an input matrix such that the input can be well approximated by the span of the selected columns. Column subset selection has been applied to numerous real-world data applications such as population genetics summarization,…
This paper studies quandles with one non-trivial column and their properties.
cGAP visualizes high-dimensional categorical data with interpretable geometric structure.
cGAP visualizes high-dimensional categorical data with interpretable geometric structure.
The column group is a subgroup of the symmetric group on the elements of a finite blackboard birack generated by the column permutations in the birack matrix. We use subgroups of the column group associated to birack homomorphisms to define an enhancement of the integral birack counting invariant and give examples whic…
Hyperfitting improves LLM generation quality by enhancing diversity, contrary to simple temperature scaling.
In this paper, we consider matrix completion from non-uniformly sampled entries including fully observed and partially observed columns. Specifically, we assume that a small number of columns are randomly selected and fully observed, and each remaining column is partially observed with uniform sampling. To recover the …
Double autoencoder improves missing value imputation in recommender systems.
Neuromorphic column performs online unsupervised clustering.
Paper tackles low-rank matrix recovery with column -norm regularization.
RECol generates error columns to improve outlier detection.
Representation learning is typically applied to only one mode of a data matrix, either its rows or columns. Yet in many applications, there is an underlying geometry to both the rows and the columns. We propose utilizing this coupled structure to perform co-manifold learning: uncovering the underlying geometry of both …
The paper identifies redundant columns in matrices for feature selection and clustering.
We consider a column of a rotating stationary surface in Euclidean space. We obtain a value in such way that if the length of column satisfies , then the surface is instable. This extends, in some sense, previous results due to Plateau and Rayleigh for columns of surfaces with constant mean curvature…
Paper tackles fair low-rank approximation and column subset selection.
This paper considers the problem of matrix completion when some number of the columns are completely and arbitrarily corrupted, potentially by a malicious adversary. It is well-known that standard algorithms for matrix completion can return arbitrarily poor results, if even a single column is corrupted. One direct appl…
We consider analysis of relational data (a matrix), in which the rows correspond to subjects (e.g., people) and the columns correspond to attributes. The elements of the matrix may be a mix of real and categorical. Each subject and attribute is characterized by a latent binary feature vector, and an inferred matrix map…
This paper defines a generalized column subset selection problem which is concerned with the selection of a few columns from a source matrix A that best approximate the span of a target matrix B. The paper then proposes a fast greedy algorithm for solving this problem and draws connections to different problems that ca…
We consider the problem of performing matrix completion with side information on row-by-row and column-by-column similarities. We build upon recent proposals for matrix estimation with smoothness constraints with respect to row and column graphs. We present a novel iterative procedure for directly minimizing an informa…
Correctly detecting the semantic type of data columns is crucial for data science tasks such as automated data cleaning, schema matching, and data discovery. Existing data preparation and analysis systems rely on dictionary lookups and regular expression matching to detect semantic types. However, these matching-based …
It has been recently shown that numerical semiparametric bounds on the expected payoff of fi- nancial or actuarial instruments can be computed using semidefinite programming. However, this approach has practical limitations. Here we use column generation, a classical optimization technique, to address these limitations…
The paper tackles one-for-many counterfactual explanations using column generation.
A common problem in large-scale data analysis is to approximate a matrix using a combination of specifically sampled rows and columns, known as CUR decomposition. Unfortunately, in many real-world environments, the ability to sample specific individual rows or columns of the matrix is limited by either system constrain…
We study the problem of instance segmentation in biological images with crowded and compact cells. We formulate this task as an integer program where variables correspond to cells and constraints enforce that cells do not overlap. To solve this integer program, we propose a column generation formulation where the prici…
Unified methods for fast column selection in various applications.
In standard clustering problems, data points are represented by vectors, and by stacking them together, one forms a data matrix with row or column cluster structure. In this paper, we consider a class of binary matrices, arising in many applications, which exhibit both row and column cluster structure, and our goal is …
We study a data model in which the data matrix D can be expressed as D = L + S + C, where L is a low rank matrix, S an element-wise sparse matrix and C a matrix whose non-zero columns are outlying data points. To date, robust PCA algorithms have solely considered models with either S or C, but not both. As such, existi…
Proposes a new matrix factorization model for interval-valued matrices.
We study the column subset selection problem with respect to the entrywise -norm loss. It is known that in the worst case, to obtain a good rank- approximation to a matrix, one needs an arbitrarily large number of columns to obtain a -approximation to the best entrywise -norm low ra…
We propose and study a row-and-column affine measurement scheme for low-rank matrix recovery. Each measurement is a linear combination of elements in one row or one column of a matrix . This setting arises naturally in applications from different domains. However, current algorithms developed for standard matrix rec…
We develop an efficient algorithm for low-rank approximation with improved approximation guarantees.
This paper considers the problem of completing a matrix with many missing entries under the assumption that the columns of the matrix belong to a union of multiple low-rank subspaces. This generalizes the standard low-rank matrix completion problem to situations in which the matrix rank can be quite high or even full r…
The problem of biclustering consists of the simultaneous clustering of rows and columns of a matrix such that each of the submatrices induced by a pair of row and column clusters is as uniform as possible. In this paper we approximate the optimal biclustering by applying one-way clustering algorithms independently on t…
Paper tackles BNSL with IP, improving quality of solutions.