Develops a method to complete binary matrices using all types of observed entries.
problem Completing binary matrices from partial observations.
method Combines risks from Davenport et al. (2014) and Hsieh et al. (2015) to use all types of entries.
result Improves matrix completion performance by using all types of entries.
An algorithm finds the maximum entry of a stochastic low-rank matrix from noisy observations.
problem Finding the maximum entry of a stochastic low-rank matrix from sequential observations.
method LowRankElim algorithm, which is a statistical approach to find the maximum entry of a non-negative matrix.
result An upper bound on the regret of $O((K + L) \poly(d) Δ^{-1} \log n)$, where K and L are the number of rows and columns, d is the rank of the matrix, and Δ is the minimum gap. Efficient algorithm for mixed membership models with large p.
problem Learning mixed membership models with high p and k.
method Efficient algorithm reducing tensor decomposition to sub-tensor factorization.
result Provable guarantees and competitive empirical results.
The completion of low rank matrices from few entries is a task with many practical applications. We consider here two aspects of this problem: detectability, i.e. the ability to estimate the rank r reliably from the fewest possible random entries, and performance in achieving small reconstruction error. We propose a …
Nonnegative matrix factorization (NMF) factorizes a non-negative matrix into product of two non-negative matrices, namely a signal matrix and a mixing matrix. NMF suffers from the scale and ordering ambiguities. Often, the source signals can be monotonous in nature. For example, in source separation problem, the source…
A fast matrix factorization method for sparse data with non-uniform missing data weights.
problem Sparse and imbalanced data in real-world learning systems.
method Non-uniform weighting of missing data, efficient learning method with truncated SVD and eALS.
result Improved performance in downstream applications compared to uniform weighting.
This work generalizes transformer attention to capture higher-order correlations efficiently.
problem Detecting triple-wise connections that were impossible for transformers.
method Developed a generalized attention scheme using Kronecker computation, showing near-linear time algorithms for bounded entries.
result A near-linear time algorithm for generalized attention computation in the bounded-entry setting.
New solutions found for elliptic systems with mixed couplings.
problem Existence of fully nontrivial solutions to elliptic systems with mixed couplings.
method Study of fully nontrivial solutions to the system with mixed couplings in a bounded or unbounded domain.
result New existence and multiplicity results of fully nontrivial solutions.
The paper presents presentations for Euclidean Picard modular groups.
problem Finding presentations for arithmetic non-cocompact lattices in Euclidean domains.
method Applying Macbeath's theorem to a Γ-invariant covering by horoballs.
result Presentations for Picard modular groups with d=2,11 are obtained.
Study of a generalized geometric Brownian motion with varying entry and exit rates.
problem Understanding the long-run behavior of economic systems with growth, volatility, entry, and exit.
method Generalized geometric Brownian motion framework with varying entry and exit rates, analyzing moments and survival probability.
result Optimal exit rate minimizes mean first-passage time, influencing system outcome.
New research shows larger language models improve data processing for diverse entries.
problem Optimizing data processing for tables with diverse string entries.
method Analytical tasks on tables with varying language model sizes and a fuzzy join benchmark.
result Larger language models improve data processing for diverse entries, but fine-tuning is necessary.
New algorithm tackles non-convex matrix completion in semi-random settings.
problem Matrix completion in semi-random environments with varying observation probabilities.
method Proposes a pre-processing step to re-weight semi-random input, followed by a nearly-linear time algorithm.
result Recovering ground-truth matrix using non-convex local minima after pre-processing.
We propose a general framework for reconstructing and denoising single entries of incomplete and noisy entries. We describe: effective algorithms for deciding if and entry can be reconstructed and, if so, for reconstructing and denoising it; and a priori bounds on the error of each entry, individually. In the noiseless…
For any matrix A in R^(m x n) of rank ρ, we present a probability distribution over the entries of A (the element-wise leverage scores of equation (2)) that reveals the most influential entries in the matrix. From a theoretical perspective, we prove that sampling at most s = O ((m + n) ρ^2 ln (m + n)) entries of the ma…
Improved matrix completion for non-uniformly sampled data.
problem Estimating unobserved entries in a matrix with varying sampling probabilities.
method Developed entry-specific bounds for low-rank matrix completion under structured non-uniform sampling.
result Error bounds for each entry match minimax lower bounds under certain conditions.
Solved a specific case of Salter's question on Burau representation.
problem Under what conditions are matrices in the image of the Burau representation of B3. method Algorithmically constructed a counterexample to Salter's specific question.
result The central quotient of the Burau image group is not the central quotient of a certain subgroup of the unitary group.
New NMF method recovers archetypes without separability condition.
problem Non-negative matrix factorization (NMF) for non-separable data.
method Optimizes convex envelope of archetypes and data points, with regularization.
result Estimator is robust and finds good solutions for real and synthetic data.
Study on optimal bubble riding with price-dependent entry times in a mean field game model.
problem Optimal bubble riding with price-dependent entry times.
method Mean field game of controls with common noise and random entry time, existence result obtained through discretization and limit analysis.
result Existence of equilibrium in the mean field game model.
Counterexamples show Salter's question on Burau image is negative for n=4.
problem Conditions for a matrix to be in the Burau image of B4. method Analyzing the central quotient and using counterexamples.
result The central quotient of the Burau image group does not coincide with the central quotient of a specific subgroup of the unitary group for n=4. Research shows how deepfakes can be used to manipulate accounting systems.
problem The vulnerability of CAATs to adversarial attacks.
method Developed a thread model to camouflage anomalies, used adversarial autoencoder neural networks to learn latent factors, demonstrated misuse of model to generate misleading entries.
result Adversarial autoencoder neural networks can learn and manipulate accounting data to deceive CAATs.
We present a novel algebraic combinatorial view on low-rank matrix completion based on studying relations between a few entries with tools from algebraic geometry and matroid theory. The intrinsic locality of the approach allows for the treatment of single entries in a closed theoretical and practical framework. More s…
Paper improves tensor completion by reducing sample entries needed.
problem Reducing the number of required sample entries for tensor completion.
method Utilizes multi-rank and unitary transformation in tensor singular value decomposition.
result Provides a bound on the number of required sample entries for tensor completion.
BGRL learns graph representations without costly negative examples.
problem Efficient representation learning on large graphs without labels.
method Bootstrapped Graph Latents (BGRL) learns by predicting augmentations.
result BGRL achieves state-of-the-art performance with 2-10x memory savings.
Reducing barriers to entry in large-scale ML markets, study shows multi-objective learning can lower data requirements.
problem Barriers to entry in emerging markets for large-scale machine learning models.
method Defined a multi-objective high-dimensional regression framework to study reputational damage and data requirements.
result The number of data points needed for a new company to enter the market can be significantly smaller than the incumbent company's dataset size.
A new tensor completion method handles missing data with missing not at random entries.
problem Handling missing data in tensors where the probability of observation depends on other entries.
method Estimate propensities using convex relaxation, then use higher-order SVD with inverse propensities weights.
result Finite-sample error bounds on the completed tensor are provided.
Groups of matrices with integer-like entries are studied.
problem Characterizing groups of matrices with algebraic integer entries.
method Analyzing traces and subgroups of matrices in number fields.
result Irreducible or completely reducible subgroups with algebraic integer traces are numerical.
Study on market entry timing in stock liquidation with trading constraints.
problem Optimal timing of market entry and exit in portfolio liquidation with trading restrictions.
method Mean-field game approach to model N-player and mean-field games of optimal portfolio liquidation. result Existence of unique equilibrium in both mean-field and N-player games. Estimates low-rank distributional matrices from incomplete samples.
problem Matrix completion for distributional entries with limited observed data.
method Kernel mean embeddings, Tucker rank, functional unfolding operators.
result Effective estimator for distributional matrix completion established.
Scalable and robust TR decomposition for large-scale data with missing entries and outliers.
problem Handling large-scale tensor data with missing entries and outliers.
method Auto-weighted steepest descent method for missing entries and outliers identification, FGMC and RStS strategies.
result Outperforms existing TR decomposition methods in the presence of outliers and runs faster than robust tensor completion algorithms.
In this paper, we consider the streaming memory-limited matrix completion problem when the observed entries are noisy versions of a small random fraction of the original entries. We are interested in scenarios where the matrix size is very large so the matrix is very hard to store and manipulate. Here, columns of the o…
We give an algorithm for completing an order-m symmetric low-rank tensor from its multilinear entries in time roughly proportional to the number of tensor entries. We apply our tensor completion algorithm to the problem of learning mixtures of product distributions over the hypercube, obtaining new algorithmic result…
The covariance matrix of a p-dimensional random variable is a fundamental quantity in data analysis. Given n i.i.d. observations, it is typically estimated by the sample covariance matrix, at a computational cost of O(np2) operations. When n,p are large, this computation may be prohibitively slow. Moreover, …
New unitary RNN architecture using complex Cayley transform outperforms existing methods.
problem Vanishing or exploding gradient problem in RNNs.
method Developed a unitary RNN architecture based on a complex scaled Cayley transform.
result scuRNN achieves comparable or better results than existing unitary RNNs.
Paper analyzes SSC for data with missing entries, improving performance.
problem Theoretical analysis of SSC with missing data entries.
method Analyzes theoretical guarantees for SSC with incomplete data, projecting zero-filled data onto observation pattern.
result Improves performance of SSC with incomplete data by projecting zero-filled data onto observation pattern.
The paper completes matrices from non-uniformly sampled entries, especially when columns are randomly selected and fully observed.
problem Matrix completion from non-uniformly sampled entries, including fully and partially observed columns.
method First, recover the column space from fully observed columns. Then, for each partially observed column, find a vector in the recovered column space with the observed entries. For low-rank matrices, recover them from Ω(rnlnn) entries. result The algorithm can exactly recover a low-rank matrix from merely Ω(rnlnn) entries. A graph connects Specht and web bases; matrix is unipotent with vanishing entries.
problem Comparing two bases of irreducible representations of the symmetric group.
method Graph theory and combinatorial analysis to describe relations between bases and prove properties of the transition matrix.
result The transition matrix between Specht and web bases is unipotent with additional vanishing entries.
A new method for streaming PCA provides confidence intervals for eigenvector entries.
problem Uncertainty quantification for individual entries in streaming PCA.
method Oja's algorithm, Bernstein-type concentration bound, Central Limit Theorem, subsampling algorithm.
result Sharp concentration bound and Central Limit Theorem for streaming PCA entries.
PACE-GGM uses Gaussian mechanism for private covariance estimation.
problem Private estimation of covariance matrices in high dimensions.
method Data-adaptive selection of entries, Gaussian mechanism, maximum-entropy reconstruction.
result Consistent improvements in estimation error compared to Gaussian mechanism and baselines.
We study low rank matrix and tensor completion and propose novel algorithms that employ adaptive sampling schemes to obtain strong performance guarantees. Our algorithms exploit adaptivity to identify entries that are highly informative for learning the column space of the matrix (tensor) and consequently, our results …
Let M be a random (alpha n) x n matrix of rank r<<n, and assume that a uniformly random subset E of its entries is observed. We describe an efficient algorithm that reconstructs M from |E| = O(rn) observed entries with relative root mean square error RMSE <= C(rn/|E|)^0.5 . Further, if r=O(1), M can be reconstructed ex…
New spectral methods improve matrix estimation in RL with low-rank structure.
problem Estimating matrices with low-rank structure in reinforcement learning.
method Spectral-based matrix estimation approaches.
result Spectral methods efficiently recover singular subspaces and minimize entry-wise error.
TATD predicts missing entries in time-evolving tensors by exploiting temporal dependency and sparsity.
problem Predict missing entries in time-evolving tensors with temporal dependency and sparsity issues.
method TATD (Time-Aware Tensor Decomposition) integrates temporal dependency and time-varying sparsity through a smoothing regularization with Gaussian kernel and alternating optimization.
result TATD achieves state-of-the-art accuracy for decomposing temporal tensors.
A new method detects and compacts saturated entries in antisparse coding.
problem Efficiently solving antisparse coding problems with ℓ∞-norm penalties. method Safe squeezing methodology to detect and compact saturated entries, reducing problem dimensionality.
result The method accelerates the computation of antisparse representation by detecting and compacting saturated entries.
We study algebraic properties of matrices whose rows are mutual neighbours, and are also neigbours of 0 ("neighbour" in the sense of a certain nilpotency condition). The intended application is in synthetic differential geometry. For a square matrix of this kind, the product of the diagonal entries equals the determina…
LEMN improves memory networks for learning from streaming data.
problem Efficiently processing large data streams in memory-augmented neural networks.
method LEMN uses a RNN-based retention agent to learn and replace less important memory entries based on their importance and historical context.
result LEMN achieves significant improvements over existing methods in learning from streaming data.
New algorithm optimizes positions of CountSketch non-zero entries for better data compression.
problem Optimizing positions of CountSketch non-zero entries for better data compression.
method Learning algorithm that optimizes both values and positions of CountSketch non-zero entries.
result Improves accuracy for low rank approximation and other problems like k-means clustering.
We consider the problem of reconstructing a low rank matrix from a subset of its entries and analyze two variants of the so-called Alternating Minimization algorithm, which has been proposed in the past. We establish that when the underlying matrix has rank r=1, has positive bounded entries, and the graph $\mathcal{G…
Study heavy-tailed weights' impact on neural network's spectral distribution.
problem Analyzing spectral distribution of conjugate kernel matrices with heavy-tailed weights.
method Computed limiting eigenvalue distribution through moments, considering heavy-tailed distributions and nonlinear activation functions.
result Heavy-tailed weights induce strong correlations, leading to fundamentally different spectral behavior.