In this paper, we propose a new volume-preserving flow and show that it performs similarly to the linear general normalizing flow. The idea is to enrich a linear Inverse Autoregressive Flow by introducing multiple lower-triangular matrices with ones on the diagonal and combining them using a convex combination. In the …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
It is well known that Expected Shortfall (also called Average Value-at-Risk) is a convex risk measure, i. e. Expected Shortfall of a convex linear combination of arbitrary risk positions is not greater than a convex linear combination with the same weights of Expected Shortfalls of the same risk positions. In this shor…
Combination theorems for convex projective geometry subgroups.
Under appropriate cooperation protocols and parameter choices, fully decentralized solutions for stochastic optimization have been shown to match the performance of centralized solutions and result in linear speedup (in the number of agents) relative to non-cooperative approaches in the strongly-convex setting. More re…
We consider learning a convex combination of basis models, and present some new theoretical and empirical results that demonstrate the effectiveness of a greedy approach. Theoretically, we first consider whether we can use linear, instead of convex, combinations, and obtain generalization results similar to existing on…
Sliced inverse regression is a popular tool for sufficient dimension reduction, which replaces covariates with a minimal set of their linear combinations without loss of information on the conditional distribution of the response given the covariates. The estimated linear combinations include all covariates, making res…
It is generally believed that ensemble approaches, which combine multiple algorithms or models, can outperform any single algorithm at machine learning tasks, such as prediction. In this paper, we propose Bayesian convex and linear aggregation approaches motivated by regression applications. We show that the proposed a…
Nonnegative matrix factorization (NMF) is a widely used linear dimensionality reduction technique for nonnegative data. NMF requires that each data point is approximated by a convex combination of basis elements. Archetypal analysis (AA), also referred to as convex NMF, is a well-known NMF variant imposing that the bas…
Develops a Riemannian archetypal analysis for interpretable non-linear data.
Optimal domain adaptation model using Fisher's Linear Discriminant.
Dual explanation method using convex hulls and example-based vectors.
Paper introduces -DER for regression tasks using morphological operators and convex-concave procedure.
Defines diversification as a binary relationship between financial portfolios.
Study iterative regularization for linear models with convex bias, improving robust sparse recovery.
Predict covariance from features using convex optimization.
Region-specific linear models are widely used in practical applications because of their non-linear but highly interpretable model representations. One of the key challenges in their use is non-convexity in simultaneous optimization of regions and region-specific models. This paper proposes novel convex region-specific…
ZeroS improves Transformers by adding negative weights, matching or beating softmax attention.
New method handles robust and adaptive control of linear systems with non-convex costs.
The paper proves a new inequality for 3-manifolds with noncompact boundaries.
Generalizing Courant's nodal domain theorem, the "Extended Courant property" is the statement that a linear combination of the first eigenfunctions has at most nodal domains. In a previous paper (Documenta Mathematica, 2018, Vol. 23, pp. 1561--1585), we gave simple counterexamples to this property, including co…
The minimization of convex objectives coming from linear supervised learning problems, such as penalized generalized linear models, can be formulated as finite sums of convex functions. For such problems, a large set of stochastic first-order solvers based on the idea of variance reduction are available and combine bot…
Drago optimizes DRO problems with faster convergence.
Efficiently representing real world data in a succinct and parsimonious manner is of central importance in many fields. We present a generalized greedy pursuit framework, allowing us to efficiently solve structured matrix factorization problems, where the factors are allowed to be from arbitrary sets of structured vect…
Hadwiger's Theorem states that Euclidean-invariant convex-continuous valuations of definable sets are linear combinations of intrinsic volumes. We lift this result from sets to data distributions over sets, specifically, to definable real-valued functions on n-dimensional Euclidean space. This generalizes intrinsic vol…
The vector of periodic, compound returns of a typical investment portfolio is almost never a convex combination of the return vectors of the securities in the portfolio. As a result the ex post version of Harry Markowitz's "standard mean-variance portfolio selection model" does not apply to compound return data. We pro…
Matching pursuit algorithms are an important class of algorithms in signal processing and machine learning. We present a blended matching pursuit algorithm, combining coordinate descent-like steps with stronger gradient descent steps, for minimizing a smooth convex function over a linear space spanned by a set of atoms…
Proposes a new loss function for robust learning.
We introduce a new class of lower bounds on the log partition function of a Markov random field which makes use of a reversed Jensen's inequality. In particular, our method approximates the intractable distribution using a linear combination of spanning trees with negative weights. This technique is a lower-bound count…
We address the problem of solving convex optimization problems with many convex constraints in a distributed setting. Our approach is based on an extension of the alternating direction method of multipliers (ADMM) that recently gained a lot of attention in the Big Data context. Although it has been invented decades ago…
We consider a discriminative learning (regression) problem, whereby the regression function is a convex combination of k linear classifiers. Existing approaches are based on the EM algorithm, or similar techniques, without provable guarantees. We develop a simple method based on spectral techniques and a `mirroring' tr…
Paper introduces SMM for forecasting multiple time series with missing values.
Learning rate annealing helps even in convex problems, improving generalization.
Push-SAGA is a decentralized algorithm for directed graphs that converges linearly.
According to Courant's theorem, an eigenfunction as\-sociated with the -th eigenvalue has at most nodal domains. A footnote in the book of Courant and Hilbert, states that the same assertion is true for any linear combination of eigenfunctions associated with eigenvalues less than or equal to . We c…
Optimum in Convex Hulls (OCH) generalizes clinical trial results to broader populations.
Distributed learning of probabilistic models from multiple data repositories with minimum communication is increasingly important. We study a simple communication-efficient learning framework that first calculates the local maximum likelihood estimates (MLE) based on the data subsets, and then combines the local MLEs t…
Active-set algorithm improves Cox regression for shape-restricted covariates.
We develop a projected Nesterov's proximal-gradient (PNPG) approach for sparse signal reconstruction that combines adaptive step size with Nesterov's momentum acceleration. The objective function that we wish to minimize is the sum of a convex differentiable data-fidelity (negative log-likelihood (NLL)) term and a conv…
Let be a compact 3-manifold with boundary, which admits a convex co-compact hyperbolic metric. We consider the hyperbolic metrics on such that the boundary is smooth and strictly convex. We show that the induced metrics on the boundary are exactly the metrics with curvature , and that the th…
We study sparse approximate solutions to convex optimization problems. It is known that in many engineering applications researchers are interested in an approximate solution of an optimization problem as a linear combination of elements from a given system of elements. There is an increasing interest in building such …
We consider the dictionary learning problem, where the aim is to model the given data as a linear combination of a few columns of a matrix known as a dictionary, where the sparse weights forming the linear combination are known as coefficients. Since the dictionary and coefficients, parameterizing the linear model are …
Efficiently combines probabilistic predictions using kernel embeddings.
Linear Transformer Block combines MLP and linear attention for near-optimal ICL in linear regression.
Paper proposes a new model for better engine control.
Demixing problems in many areas such as hyperspectral imaging and differential optical absorption spectroscopy (DOAS) often require finding sparse nonnegative linear combinations of dictionary elements that match observed data. We show how aspects of these problems, such as misalignment of DOAS references and uncertain…
Linear-Core Surrogates combine fast optimization and statistical efficiency in classification and structured prediction.
A new method solves diagonally constrained SDPs quickly and accurately.
The generalized partially linear additive model (GPLAM) is a flexible and interpretable approach to building predictive models. It combines features in an additive manner, allowing each to have either a linear or nonlinear effect on the response. However, the choice of which features to treat as linear or nonlinear is …