Proposes a new method for kernel density estimation using stagewise minimization and a simple dictionary.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Forward stagewise regression follows a very simple strategy for constructing a sequence of sparse regression estimates: it starts with all coefficients equal to zero, and iteratively updates the coefficient (by a small amount ) of the variable that achieves the maximal absolute inner product with the current residua…
Proposes MM-DUST for efficient generalized lasso solution paths.
Although stochastic gradient descent (SGD) method and its variants (e.g., stochastic momentum methods, AdaGrad) are the choice of algorithms for solving non-convex problems (especially deep learning), there still remain big gaps between the theory and the practice with many questions unresolved. For example, there is s…
Stagewise training strategy is widely used for learning neural networks, which runs a stochastic algorithm (e.g., SGD) starting with a relatively large step size (aka learning rate) and geometrically decreasing the step size after a number of iterations. It has been observed that the stagewise SGD has much faster conve…
Stagewise boosting improves gradient boosting for distributional regression.
The performance of EM in learning mixtures of product distributions often depends on the initialization. This can be problematic in crowdsourcing and other applications, e.g. when a small number of 'experts' are diluted by a large number of noisy, unreliable participants. We develop a new EM algorithm that is driven by…
We present a method of variable selection for the sparse generalized additive model. The method doesn't assume any specific functional form, and can select from a large number of candidates. It takes the form of incremental forward stagewise regression. Given no functional form is assumed, we devised an approach termed…
Boosting methods are highly popular and effective supervised learning methods which combine weak learners into a single accurate model with good statistical performance. In this paper, we analyze two well-known boosting methods, AdaBoost and Incremental Forward Stagewise Regression (FS), by establishing t…
A machine learning approach to record fusion with high accuracy.
We propose a sparse and low-rank tensor regression model to relate a univariate outcome to a feature tensor, in which each unit-rank tensor from the CP decomposition of the coefficient tensor is assumed to be sparse. This structure is both parsimonious and highly interpretable, as it implies that the outcome is related…
Divide-and-conquer method speeds sparse factorization for large matrices.
This paper presents a natural extension of stagewise ranking to the the case of infinitely many items. We introduce the infinite generalized Mallows model (IGM), describe its properties and give procedures to estimate it from data. For estimation of multimodal distributions we introduce the Exponential-Blurring-Mean-Sh…
Least Angle Regression is a promising technique for variable selection applications, offering a nice alternative to stepwise regression. It provides an explanation for the similar behavior of LASSO (-penalized regression) and forward stagewise regression, and provides a fast implementation of both. The idea has…
Distributed stochastic gradient descent~(DSGD) has been widely used for optimizing large-scale machine learning models, including both convex and non-convex models. With the rapid growth of model size, huge communication cost has been the bottleneck of traditional DSGD. Recently, many communication compression methods …
A new method SEBS optimizes SGD batch size for better performance.
This paper addresses the general problem of modelling and learning rank data with ties. We propose a probabilistic generative model, that models the process as permutations over partitions. This results in super-exponential combinatorial state space with unknown numbers of partitions and unknown ordering among them. We…
This work provides simple algorithms for multi-class (and multi-label) prediction in settings where both the number of examples n and the data dimension d are relatively large. These robust and parameter free algorithms are essentially iterative least-squares updates and very versatile both in theory and in practice. O…
In this paper we analyze boosting algorithms in linear regression from a new perspective: that of modern first-order methods in convex optimization. We show that classic boosting algorithms in linear regression, namely the incremental forward stagewise algorithm (FS) and least squares boosting (LS-Boost($…
Paper tackles robust optimization under uncertainty using nested distance.
The paper decomposes probabilistic scores into reliability, uncertainty, and information loss.
Lookahead, also known as non-myopic, Bayesian optimization (BO) aims to find optimal sampling policies through solving a dynamic program (DP) that maximizes a long-term reward over a rolling horizon. Though promising, lookahead BO faces the risk of error propagation through its increased dependence on a possibly mis-sp…
TSL learns separable models to avoid signal cancellation and off-support extrapolation.
Enforcing safety is a key aspect of many problems pertaining to sequential decision making under uncertainty, which require the decisions made at every step to be both informative of the optimal decision and also safe. For example, we value both efficacy and comfort in medical therapy, and efficiency and safety in robo…
CASP selects reliable policies for two-stage recommender systems by considering both value and support.
Efficiently performs robust and sparse kernel regression.
REGAIN learns optimal auxiliary directions for forecast reconciliation.
New method selects variables for GP regression using sparse projection.
Estimates density ratio for two-sample comparison using tree models.
STL-SGD accelerates Local SGD by gradually increasing communication periods.
Algorithm infers sampling distribution from i.i.d. samples without supervision.
Local graph clustering methods aim to find small clusters in very large graphs. These methods take as input a graph and a seed node, and they return as output a good cluster in a running time that depends on the size of the output cluster but that is independent of the size of the input graph. In this paper, we adopt a…
Minimal networks minimize length and mass in certain configurations.
The study finds conditions for area-minimizing cones over submanifolds.
The paper studies deformations of singular minimal hypersurfaces in dimensions 7 and above.
Minimal surfaces in 3-sphere created by reflections from polygons, with new examples based on pentagons.
Study on minimal surfaces in a 3D space with 2m-norm.
Some elementary considerations are presented concerning Catenoids and their stability, separable minimal hypersurfaces, minimal surfaces obtainable by rotating shapes, determinantal varieties, minimal tori in S3, the minimality in Rnk of the ordered set of k orthogonal equal-length n-vectors, and U(1)-invariant minimal…
Study counts minimal surfaces in curved 3D spaces, finding hyperbolic space minimizes area.
Proves unique continuation for area minimizing currents.
New inequality helps map stability in minimal surfaces.
Minimal submanifolds in spheres can be produced via Clifford type minimal products, and their Morse indices and nullities are calculated.
Round balls minimize liquid drop model volumes ≤ 1.
Minimal generating sets of Reidemeister moves identified and classified.
The paper creates symmetrical discrete minimal nets using Schwarz reflection.
In this paper, we prove that every conformal minimal immersion of a compact bordered Riemann surface into a minimally convex domain can be approximated, uniformly on compacts in , by proper complete conformal minimal immersions . We also obtain a …
We prove that there are no minimal hypersurfaces properly immersed in any region of the Euclidean space bounded by unstable minimal cones. We also prove the analogous result for -minimal hypersurfaces.
New minimal surfaces grow area very quickly.