Gradient methods converge better for alternating updates in bilinear zero-sum games.
problem Understanding the dynamics of gradient algorithms for bilinear zero-sum games.
method Systematic analysis of popular gradient updates for simultaneous and alternating versions of bilinear zero-sum games.
result Alternating updates converge better than simultaneous ones, with optimal parameter setup and rates.
Paper analyzes game dynamics with negative momentum for improved stability and convergence.
problem Complexity and instability in game dynamics, especially in adversarial settings.
method Analyzed gradient-based methods with negative momentum on simple games and adversarial problems.
result Alternating gradient updates with negative momentum achieve convergence in difficult adversarial problems.
LD-SGD improves communication in decentralized SGD.
problem Efficiently combining local updates and decentralized communication.
method Proposes LD-SGD integrating local updates and decentralized SGD, with a convergence analysis.
result LD-SGD converges to a critical point for non-convex objectives with non-identically distributed data.
Alt-GDA outperforms Sim-GDA in minimax games with near-optimal local convergence.
problem Minimax optimization convergence rate comparison
method Alternating Gradient Descent-Ascent (Alt-GDA) vs. Simultaneous Gradient Descent-Ascent (Sim-GDA)
result Alt-GDA achieves near-optimal local convergence rate for strongly convex-strongly concave problems, while Sim-GDA converges slower.
A new algorithm balances fairness in clustering to avoid discrimination.
problem Clustering data can unfairly discriminate against different demographic groups.
method Designing a stochastic alternating balance fair k-means algorithm (SAfairKM) that alternates between k-means updates and group swap updates.
result The algorithm efficiently constructs well-spread and high-quality Pareto fronts on synthetic and real datasets.
A dual-learner strategy tracks concept drift in nonstationary data streams.
problem Learning from nonstationary data with abrupt or gradual changes.
method Alternating learners framework with long- and short-memory models.
result Effective tracking and prediction of concept drift in streaming data.
Efficiently completes low-rank matrices with nearly linear time complexity.
problem Completing low-rank matrices from a few observed entries.
method Robust alternating minimization framework with approximate updates.
result Achieves nearly linear time complexity in matrix completion.
Paper proposes an algorithm to recover non-negative matrix factorization with mild conditions.
problem Understanding and guaranteeing recovery of non-negative matrix factorization.
method Alternates between updating features and decoding weights using ReLU.
result Proves recovery of ground-truth under mild conditions, including linear independence of features.
Model analyzes cooccurrence data for recommender systems and item relevance.
problem High-dimensional cooccurrence data from online platforms.
method Shared parameter Alternating Tweedie (SA-Tweedie) model with Fisher scoring and learning rate adjustment.
result SA-Tweedie model outperforms other methods in optimizing parameters.
Develops a new method for matrix factorization problems.
problem Solving a general matrix factorization model with potential function.
method Non-monotone alternating updating method based on a potential function.
result The method can outperform existing methods for specific applications.
We propose a new stochastic dual coordinate ascent technique that can be applied to a wide range of regularized learning problems. Our method is based on Alternating Direction Multiplier Method (ADMM) to deal with complex regularization functions such as structured regularizations. Although the original ADMM is a batch…
The full causal ladder of spacetimes is constructed, and their updated main properties are developed. Old concepts and alternative definitions of each level of the ladder are revisited, with emphasis in minimum hypotheses. The implications of the recently solved ``folk questions on smoothability'', and alternative prop…
New method improves matrix factorization speed and accuracy.
problem Matrix factorization optimization problems suffer from biased solutions and lack of convergence guarantees.
method Proposes a novel Bregman distance for matrix factorization, enabling non-alternating schemes with convergence proof.
result Convergence to a stationary point proved for matrix factorization problems.
AuON is a linear-time optimizer that improves upon Muon's performance without approximate orthogonal matrices.
problem High memory and computational costs of orthogonal momentum updates.
method AuON uses normalized nonlinear scaling and a 'emergency brake' to handle exploding attention logits.
result AuON achieves strong performance without approximate orthogonal matrices, preserving structural alignment and reconditioning.
In this paper, we provide local and global convergence guarantees for recovering CP (Candecomp/Parafac) tensor decomposition. The main step of the proposed algorithm is a simple alternating rank-1 update which is the alternating version of the tensor power iteration adapted for asymmetric tensors. Local convergence g…
Study how learning impacts decision-making in project expansion.
problem Optimal timing of expansion decisions based on uncertain project profitability.
method Bayesian updating model with two irreversible alternatives (exit or expansion).
result Time-to-decision is not monotonic with information arrival rate.
We propose two new alternating direction methods to solve "fully" nonsmooth constrained convex problems. Our algorithms have the best known worst-case iteration-complexity guarantee under mild assumptions for both the objective residual and feasibility gap. Through theoretical analysis, we show how to update all the al…
MPI-FAUN tackles NMF for big data, offering scalable parallel algorithms.
problem Efficient parallel algorithms for NMF on big data.
method MPI-based framework for NMF, solving alternating NLS subproblems.
result Significant performance improvements over baseline implementations.
Many applications in signal processing benefit from the sparsity of signals in a certain transform domain or dictionary. Synthesis sparsifying dictionaries that are directly adapted to data have been popular in applications such as image denoising, inpainting, and medical image reconstruction. In this work, we focus in…
Federated learning is vulnerable to model poisoning attacks by a single malicious agent.
problem Vulnerability of federated learning to model poisoning attacks by a single non-colluding agent.
method Exploration of model poisoning attacks, including boosting, alternating minimization, and parameter estimation.
result Even a constrained adversary can successfully carry out model poisoning attacks while maintaining stealth.
New CT image reconstruction method reduces X-ray dose while improving image quality.
problem Reducing X-ray dose in CT while maintaining image quality.
method Combines PWLS with learned sparsifying transform using alternating optimization and relaxed OS-LALM.
result Proposed method improves image quality for low dose levels compared to existing methods.
New algorithms improve NMF for extracting patterns from time series data.
problem Extracting short-lived temporal motifs from high-dimensional time series data.
method Extended HALS and ANLS algorithms for CNMF model.
result Improved performance on large-scale data compared to multiplicative updates.
Drop-Muon updates only some layers, speeding up training.
problem Conventional deep learning optimizers update all layers at once, which can be inefficient.
method Drop-Muon updates only a subset of layers per step, with randomized schedules.
result Drop-Muon achieves up to 1.4x faster training time with similar accuracy.
A new policy improvement method using CEM for Actor-Critic.
problem Improving policy efficiency and robustness in reinforcement learning.
method Greedy Actor-Critic (Greedy AC) using Conditional Cross-Entropy Method (CCEM).
result Greedy AC outperforms Soft Actor-Critic and is less sensitive to entropy regularization.
A new resampling strategy, Importance Resampling, improves sample efficiency and reduces variance in off-policy prediction.
problem High variance updates in importance sampling for off-policy prediction.
method Importance Resampling (IR) resamples experience from a replay buffer and applies standard on-policy updates, avoiding importance sampling ratios.
result Importance Resampling (IR) shows improved sample efficiency and lower variance updates compared to other methods.
A new method for training CNNs that decouples layer updates.
problem Update locking inefficiency in neural network training.
method Decoupled Greedy Learning (DGL) that relaxes joint training objective.
result DGL leads to better generalization than sequential greedy optimization.
Proposes time-smoothed gradients for more stable online forecasting.
problem Stability and efficiency in online forecasting with SGD.
method Introduces time-smoothed gradients within SGD update rules.
result Time-smoothed gradients yield more stable results than existing methods.
New method for LVEBMs using saddle-point optimization and Langevin updates.
problem Expressive generative modeling of latent variables with hidden structure.
method Reformulate LVEBM training as a saddle problem, using Langevin updates and gradient flows.
result Proves existence and convergence of the algorithm under standard assumptions, with improved ELBO bounds.
Develops a model to learn shared and idiosyncratic patterns in point processes.
problem Learning shared and unique patterns in point processes from diverse observations.
method Developed a parametric point process model with alternating optimization for learning shared structure and idiosyncratic effects.
result The method yields explainable point process models that perform well compared to existing methods.
This paper proposes an alternating back-propagation algorithm for learning the generator network model. The model is a non-linear generalization of factor analysis. In this model, the mapping from the continuous latent factors to the observed signal is parametrized by a convolutional neural network. The alternating bac…
Holdout set improves risk score accuracy without biasing predictions.
problem Updating risk scores can lead to biased estimates when directly applied.
method Use a holdout set of non-intervention population to update risk scores.
result Optimal holdout size reduces adverse outcomes to optimal level.
Interpolates between SPG and NeuRD with Capped Implicit Exploration.
problem Combining SPG and NeuRD for better performance in non-stationary environments.
method Introduces Capped Implicit Exploration (CIX) to interpolate between SPG and NeuRD.
result NeuRD-CIX performs well more consistently than NeuRD while retaining NeuRD's advantages.
We propose a general algorithmic framework for constrained matrix and tensor factorization, which is widely used in signal processing and machine learning. The new framework is a hybrid between alternating optimization (AO) and the alternating direction method of multipliers (ADMM): each matrix factor is updated in tur…
Proposes RNSE for clustering with adaptive similarity matrix learning.
problem Sub-optimal results due to mismatch between stages in Spectral Clustering.
method End-to-end single-stage learning with adaptive similarity matrix and non-negative constraints.
result Superior clustering performance on synthetic and real-world datasets.
Proposes a hierarchical deep generative model for natural images.
problem Analyzing piecewise smooth signals like natural images.
method Hierarchical deep generative model with alternating minimization algorithm.
result Demonstrates the model's representation capabilities and classification performance.
dYdX updates liquidity provider incentives to enhance trading efficiency.
problem Incentivizing liquidity providers to maintain efficient market structures.
method Analyzed various metrics (makerVolume, depths, spreads) and used historical trades to update the LP Incentives Programme.
result Updated the LP Incentives Programme to encourage more active and efficient liquidity.
Bayesian uncertainty quantification is flawed, according to new research.
problem Flawed interpretation of Bayesian uncertainty quantification.
method Discussion of Bayesian updating and optimization-based perspective, proposing measures of quality.
result Bayesian uncertainty quantification is not coherent with optimization-based perspective.
Efficiently optimizes boolean functions using multilinear polynomials and exponential weight updates.
problem Optimizing boolean functions over the boolean hypercube with high computational cost.
method Proposes a computationally efficient algorithm using multilinear polynomials and exponential weight updates.
result Improves computational time up to several orders of magnitude compared to state-of-the-art algorithms.
Bio-inspired neural networks use predictive coding for efficient weight updates.
problem Training artificial neural networks efficiently and biologically plausibly.
method Predictive Coding (PC) updates weights locally using only local information.
result PC provides theoretical advantages like automatic gradient scaling.
We develop methods for parameter estimation in settings with large-scale data sets, where traditional methods are no longer tenable. Our methods rely on stochastic approximations, which are computationally efficient as they maintain one iterate as a parameter estimate, and successively update that iterate based on a si…
New TD method stabilizes average-reward learning.
problem Stability issues in average-reward TD learning.
method Implicit fixed point update for average-reward TD(λ). result Improved numerical stability and broader step-size range.
Pion optimizes LLMs by preserving weight matrix singular values.
problem Training large language models (LLMs) with standard optimizers leads to unstable weight matrices.
method Pion uses orthogonal transformations to update weight matrices, preserving their singular values.
result Pion offers a stable alternative to standard optimizers for LLM pretraining and finetuning.
New algorithm speeds up NMF with β-divergence.
problem Efficiently factorize nonnegative matrices with β-divergence. method Joint majorization-minimization with multiplicative updates.
result Significant reduction in computation time for NMF.
Paper tackles risk-sensitive decision-making under uncertainty.
problem Risk-sensitive decision-making problem under uncertainty.
method Formulated as a stochastic control problem, delineated necessary optimality conditions.
result Illustrative examples from optimal betting and inventory management support the theory.
A networked learning method for correlated data outperforms federated learning in precision.
problem Estimating models from correlated data distributed across a network.
method Local linear model estimation with network regularization and information exchange.
result The weighted ensemble average estimate converges faster and more precisely than federated learning.
Dual Policy Iteration combines fast and slow policies for better reinforcement learning performance.
problem Improving reinforcement learning algorithms for practical applications.
method Alternates between a fast, reactive policy and a slow, non-reactive policy, optimizing both under each other's supervision.
result Demonstrates improved performance on various continuous control Markov Decision Processes.
Develops DP-SCD for stochastic coordinate descent, making it differentially private.
problem Privacy leak in auxiliary information during stochastic coordinate descent training.
method Develops DP-SCD, leveraging independent noise addition and decoupling/parallelizing coordinate updates.
result Demonstrates competitive performance against DP-SGD with less tuning.
Algorithm recovers factors of rank-1 matrices from noisy measurements.
problem Estimating factors of a rank-1 matrix from nonlinearly transformed and noisy measurements.
method Alternating minimization with random initialization and analysis of empirical error recursion.
result Algorithm converges geometrically fast from random initialization, with sharp guarantees.