This work develops scalable model selection methods with fast update and selection.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
SFPO optimizes LLM reasoning by repositioning before updating, improving stability and efficiency.
Efficiently handles contextual bandits with diffusion models.
Paper introduces a new gradient statistic to improve deep learning convergence.
We describe -MLE, a fast and efficient local search algorithm for learning finite statistical mixtures of exponential families such as Gaussian mixture models. Mixture models are traditionally learned using the expectation-maximization (EM) soft clustering technique that monotonically increases the incomplete (expec…
Gradient-based meta-learning has proven to be highly effective at learning model initializations, representations, and update rules that allow fast adaptation from a few samples. The core idea behind these approaches is to use fast adaptation and generalization -- two second-order metrics -- as training signals on a me…
New algorithms for fast online decision making using neural networks and martingale posteriors.
New framework explains fast transfer of hyperparameters across model scales.
We propose a new algorithm for solving the graph-fused lasso (GFL), a method for parameter estimation that operates under the assumption that the signal tends to be locally constant over a predefined graph structure. Our key insight is to decompose the graph into a set of trails which can then each be solved efficientl…
Dictionary learning is the task of determining a data-dependent transform that yields a sparse representation of some observed data. The dictionary learning problem is non-convex, and usually solved via computationally complex iterative algorithms. Furthermore, the resulting transforms obtained generally lack structure…
Paper analyzes ensemble Kalman updates for effective dimension and localization.
In-Place TTT enhances LLMs with dynamic parameter updates at inference time.
Hierarchical pretraining with slow-fast ODEs
A new method improves EEG classification across subjects efficiently.
Identifying recurring patterns in high-dimensional time series data is an important problem in many scientific domains. A popular model to achieve this is convolutive nonnegative matrix factorization (CNMF), which extends classic nonnegative matrix factorization (NMF) to extract short-lived temporal motifs from a long …
Online learning algorithms have impressive convergence properties when it comes to risk minimization and convex games on very large problems. However, they are inherently sequential in their design which prevents them from taking advantage of modern multi-core architectures. In this paper we prove that online learning …
Generative adversarial network improves geosteering in fluvial reservoirs.
Fast feature selection for SHM using canonical correlation.
This paper introduces a novel theoretically sound approach for the celebrated CMA-ES algorithm. Assuming the parameters of the multi variate normal distribution for the minimum follow a conjugate prior distribution, we derive their optimal update at each iteration step. Not only provides this Bayesian framework a justi…
In this paper we provide a new analysis of the SEM algorithm. Unlike previous work, we focus on the analysis of a single run of the algorithm. First, we discuss the algorithm for general mixture distributions. Second, we consider Gaussian mixture models and show that with high probability the update equations of the EM…
In this paper we develop a Bayesian procedure for estimating multivariate stochastic volatility (MSV) using state space models. A multiplicative model based on inverted Wishart and multivariate singular beta distributions is proposed for the evolution of the volatility, and a flexible sequential volatility updating is …
FSNet improves online time series forecasting by balancing fast adaptation and old knowledge.
Adaptive sparse GP model for non-stationary data.
Gradient-EM Bayesian meta-learning accelerates adaptation with reduced computation and improved robustness.
We establish decoupled functional CLTs for two-time-scale stochastic approximation.
User and item features of side information are crucial for accurate recommendation. However, the large number of feature dimensions, e.g., usually larger than 10^7, results in expensive storage and computational cost. This prohibits fast recommendation especially on mobile applications where the computational resource …
We investigate numerically efficient approximations of eigenspaces associated to symmetric and general matrices. The eigenspaces are factored into a fixed number of fundamental components that can be efficiently manipulated (we consider extended orthogonal Givens or scaling and shear transformations). The number of the…
New algorithm reduces best-in-class regret in contextual bandits.
We present a new online boosting algorithm for adapting the weights of a boosted classifier, which yields a closer approximation to Freund and Schapire's AdaBoost algorithm than previous online boosting algorithms. We also contribute a new way of deriving the online algorithm that ties together previous online boosting…
We investigate finite-time decoupled convergence in nonlinear two-time-scale stochastic approximation.
Paper develops a fast Bayesian method to predict toxic trades in financial transactions.
New approach to meta-learning with variational Bayes for unlabeled data.
Aims to describe neural network training dynamics using two-time-scale models.
FLOP algorithm speeds up causal structure learning for linear models.
This paper speeds up OCSSVM training using SMO.
Study on rich regime training in deep learning, finding active parameters in bottom layers.
We consider the problem of fast time-series data clustering. Building on previous work modeling the correlation-based Hamiltonian of spin variables we present an updated fast non-expensive Agglomerative Likelihood Clustering algorithm (ALC). The method replaces the optimized genetic algorithm based approach (f-SPC) wit…
Partial differential equations (PDEs) are widely used across the physical and computational sciences. Decades of research and engineering went into designing fast iterative solution methods. Existing solvers are general purpose, but may be sub-optimal for specific classes of problems. In contrast to existing hand-craft…
Optimizes financial auditor schedules to reduce time and costs.
Natural-gradient methods enable fast and simple algorithms for variational inference, but due to computational difficulties, their use is mostly limited to \emph{minimal} exponential-family (EF) approximations. In this paper, we extend their application to estimate \emph{structured} approximations such as mixtures of E…
Living review of ML for particle physics, updated frequently.
A fast method for Lasso and Logistic Lasso problems.
FedNNNN improves FL by adjusting model update vector norms.
Forward stagewise regression follows a very simple strategy for constructing a sequence of sparse regression estimates: it starts with all coefficients equal to zero, and iteratively updates the coefficient (by a small amount ) of the variable that achieves the maximal absolute inner product with the current residua…
Learning to infer Bayesian posterior from a few-shot dataset is an important step towards robust meta-learning due to the model uncertainty inherent in the problem. In this paper, we propose a novel Bayesian model-agnostic meta-learning method. The proposed method combines scalable gradient-based meta-learning with non…
A new method speeds up sampling in diffusion models.
Efficient algorithm for mobile health provides timely physical activity suggestions.
A novel decentralized algorithm improves minimax optimization in federated learning.