We consider the problem of solving a large-scale Quadratically Constrained Quadratic Program. Such problems occur naturally in many scientific and web applications. Although there are efficient methods which tackle this problem, they are mostly not scalable. In this paper, we develop a method that transforms the quadra…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Gaussian equivalence fails for simple polynomial embeddings in quadratic scaling RF models.
Study how generalization scales with model size and data in quadratic neural networks.
Paper presents an ADMM-based approach to efficiently integrate quadratic programming layers into neural networks.
New model captures asymmetric rough volatility with Zumbach effect.
We study super--replication of European contingent claims in an illiquid market with insider information. Illiquidity is captured by quadratic transaction costs and insider information is modeled by an investor who can peek into the future. Our main result describes the scaling limit of the super--replication prices wh…
Study on SGD dynamics and scaling laws for training quadratic neural networks in high dimensions.
We propose HAMSI (Hessian Approximated Multiple Subsets Iteration), which is a provably convergent, second order incremental algorithm for solving large-scale partially separable optimization problems. The algorithm is based on a local quadratic approximation, and hence, allows incorporating curvature information to sp…
Study on neural network dynamics in high dimensions with quadratic activation.
Global minima found for multidimensional scaling with penalties.
We consider the problem of high-dimensional classification between the two groups with unequal covariance matrices. Rather than estimating the full quadratic discriminant rule, we propose to perform simultaneous variable selection and linear dimension reduction on original data, with the subsequent application of quadr…
Linear memory stores associations up to a logarithmic scale, but listwise retrieval can handle a quadratic scale.
The study sets limits on how well systems can be controlled adaptively.
A streaming algorithm estimates quadratic covariation from financial data efficiently.
Study shows certainty equivalent policy minimizes regret in continuous-time systems.
New algorithms achieve logarithmic regret in learning linear quadratic control systems.
This paper gives examples of explicit arbitrage-free term structure models with Lévy jumps via state price density approach. By generalizing quadratic Gaussian models, it is found that the probability density function of a Lévy process is a "natural" scale for the process to be the state variable of a market.
Proposes SPFB method for optimizing partition functions in stochastic learning.
We reconsider the problem of optimal trading in the presence of linear and quadratic costs, for arbitrary linear costs but in the limit where quadratic costs are small. Using matched asymptotic expansion techniques, we find that the trading speed vanishes inside a band that is narrower than in the absence of quadratic …
Paper connects MoE and self-attention, proposing active-attention.
Enhances power of covariance matrix tests for high-dimensional data.
Motivated by electricity consumption metering, we extend existing nonnegative matrix factorization (NMF) algorithms to use linear measurements as observations, instead of matrix entries. The objective is to estimate multiple time series at a fine temporal scale from temporal aggregates measured on each individual serie…
New model-free algorithm achieves similar LQR regret guarantees.
New algorithm achieves logarithmic regret for adversarial online control.
BBVI with STL converges geometrically under perfect specification, with quadratic variance bound.
We study the performance of the certainty equivalent controller on Linear Quadratic (LQ) control problems with unknown transition dynamics. We show that for both the fully and partially observed settings, the sub-optimality gap between the cost incurred by playing the certainty equivalent controller on the true system …
Improved kernel Stein discrepancy for large-scale data.
We study the problem of learning similarity functions over very large corpora using neural network embedding models. These models are typically trained using SGD with sampling of random observed and unobserved pairs, with a number of samples that grows quadratically with the corpus size, making it expensive to scale to…
We propose a novel end-to-end non-minimax algorithm for training optimal transport mappings for the quadratic cost (Wasserstein-2 distance). The algorithm uses input convex neural networks and a cycle-consistency regularization to approximate Wasserstein-2 distance. In contrast to popular entropic and quadratic regular…
Based on the new type of random walk process called the Potentials of Unbalanced Complex Kinetics (PUCK) model, we theoretically show that the price diffusion in large scales is amplified 2/(2 + b) times, where b is the coefficient of quadratic term of the potential. In short time scales the price diffusion depends on …
This paper concerns integral varifolds of arbitrary dimension in an open subset of Euclidean space satisfying integrability conditions on their first variation. Firstly, the study of pointwise power decay rates almost everywhere of the quadratic tilt-excess is completed by establishing the precise decay rate for two-di…
Paper solves Gromov-Wasserstein for point clouds efficiently.
Quadratic discriminant analysis (QDA) is a standard tool for classification due to its simplicity and flexibility. Because the number of its parameters scales quadratically with the number of the variables, QDA is not practical, however, when the dimensionality is relatively large. To address this, we propose a novel p…
Proposes sparse QSVM for better generalization and interpretability.
Study utility indifference pricing with delayed investment information in a Bachelier model.
New method distinguishes stochastic from deterministic signals using excursion counts.
Bayesian neural networks are shown to be minimax and admissible under certain conditions.
Unified framework for fast large-scale portfolio optimization.
We attempt to unveil the fine structure of volatility feedback effects in the context of general quadratic autoregressive (QARCH) models, which assume that today's volatility can be expressed as a general quadratic form of the past daily returns. The standard ARCH or GARCH framework is recovered when the quadratic kern…
New method for online inference of constrained optimization problems.
A homogeneous nilpotent Lie group has a scaling automorphism determined by a grading of its Lie algebra. Many proofs of upper bounds for the Dehn function of such a group depend on being able to fill curves with discs compatible with this grading; the area of such discs changes predictably under the scaling automorphis…
Learning DAG or Bayesian network models is an important problem in multi-variate causal inference. However, a number of challenges arises in learning large-scale DAG models including model identifiability and computational complexity since the space of directed graphs is huge. In this paper, we address these issues in …
Semidefinite programs (SDP) are important in learning and combinatorial optimization with numerous applications. In pursuit of low-rank solutions and low complexity algorithms, we consider the Burer--Monteiro factorization approach for solving SDPs. We show that all approximate local optima are global optima for the pe…
Improved loss scaling for stochastic momentum algorithms in high dimensions.
New Q-Newton's method avoids saddle points and converges quadratically.
Training of one-vs.-rest SVMs can be parallelized over the number of classes in a straight forward way. Given enough computational resources, one-vs.-rest SVMs can thus be trained on data involving a large number of classes. The same cannot be stated, however, for the so-called all-in-one SVMs, which require solving a …
A new robust and flexible classification method for non-Gaussian data.
Mamba struggles with long context lengths, but spectrum scaling improves performance.