This work interprets GELU and related activations via a first-order loss function.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Solves asset allocation for investors with utility functions and limits.
Paper discusses new stochastic algorithms for sparse signal recovery.
Unified framework for shrinkage, thresholding, and regularization in normal mean estimation and linear regression.
In this paper, we investigate a multivariate multi-response (MVMR) linear regression problem, which contains multiple linear regression models with differently distributed design matrices, and different regression and output vectors. The goal is to recover the support union of all regression vectors using -reg…
New algorithms estimate function levels with near-optimal efficiency.
Paper analyzes robust matrix completion with efficient nonconvex method and leave-one-out analysis.
Thresholded Lasso bandit minimizes regret in sparse linear bandits.
We compute the log canonical thresholds of non-negatively curved singular hermitian metrics on ample linearized line bundles on bi-equivariant group compactifications of complex reductive groups. To this end, we associate to any such metric a convex function whose asymptotic behavior determines the log canonical thresh…
Improved learning bounds for corrupted data using thresholded gradient descent.
Noise makes learning linear thresholds hard, but algorithms can still learn near-optimal thresholds.
In this paper, we study the proximal gradient algorithm with extrapolation for minimizing the sum of a Lipschitz differentiable function and a proper closed convex function. Under the error bound condition used in [19] for analyzing the convergence of the proximal gradient algorithm, we show that there exists a thresho…
Paper proposes a new activation function to reduce overfitting and large weight update issues.
Iterative thresholding algorithms seek to optimize a differentiable objective function over a sparsity or rank constraint by alternating between gradient steps that reduce the objective, and thresholding steps that enforce the constraint. This work examines the choice of the thresholding operator, and asks whether it i…
In this paper, we propose a communication- and computation-efficient algorithm to solve a convex consensus optimization problem defined over a decentralized network. A remarkable existing algorithm to solve this problem is the alternating direction method of multipliers (ADMM), in which at every iteration every node up…
Optimal algorithm for identifying best arm in stochastic linear bandits with fixed confidence.
LinearAPT optimizes decision-making under resource constraints for a linear threshold problem.
New method identifies extreme risk propagation in financial networks.
The paper tackles reward-relevance in offline RL with sparse decision dynamics.
This work efficiently learns linear threshold functions from label proportions using Gaussian distributions.
Starting from an exact relationship between news, threshold and price return distributions in the stationary state, I discuss the ability of the Ghoulmie-Cont-Nadal model of traders to produce fat-tailed price returns. Under normal conditions, this model is not able to transform Gaussian news into fat-tailed price retu…
Noise in linear networks minimizes sharpness and leads to shrinkage-thresholding.
FILTER model uses fusion penalized logistic threshold regression for high-dimensional data with unknown cut points.
Robust learning mixtures of linear regressions improve robustness.
This paper describes a fast algorithm for recovering low-rank matrices from their linear measurements contaminated with Poisson noise: the Poisson noise Maximum Likelihood Singular Value thresholding (PMLSV) algorithm. We propose a convex optimization formulation with a cost function consisting of the sum of a likeliho…
Paper develops algorithms to maximize AUC in imbalanced classification.
Learning a classifier with control on the false-positive rate plays a critical role in many machine learning applications. Existing approaches either introduce prior knowledge dependent label cost or tune parameters based on traditional classifiers, which lack consistency in methodology because they do not strictly adh…
Paper provides linear convergence guarantees for KZIHT and KZPT methods.
To estimate a sparse linear model from data with Gaussian noise, consilience from lasso and compressed sensing literatures is that thresholding estimators like lasso and the Dantzig selector have the ability in some situations to identify with high probability part of the significant covariates asymptotically, and are …
We consider the problem of learning a non-negative linear classifier with a -norm of at most , and a fixed threshold, under the hinge-loss. This problem generalizes the problem of learning a -monotone disjunction. We prove that we can learn efficiently in this setting, at a rate which is linear in both and…
We study the problem of robust linear regression with response variable corruptions. We consider the oblivious adversary model, where the adversary corrupts a fraction of the responses in complete ignorance of the data. We provide a nearly linear time estimator which consistently estimates the true regression vector, e…
Optimal algorithm for high-dimensional stochastic linear bandits with sparse parameters.
Method constructs confidence regions for linear models with arbitrary predictors.
This paper establishes minimax rates for online regression with arbitrary classes of functions and general losses. We show that below a certain threshold for the complexity of the function class, the minimax rates depend on both the curvature of the loss function and the sequential complexities of the class. Above this…
Unified formula for training dynamics of linear networks combining lazy and balanced regimes.
New method trains neural networks with threshold activation functions efficiently.
New stability thresholds detect K-stability in Fano manifolds.
A new SSL method uses instance-dependent thresholds to improve accuracy.
The study optimizes polynomial regression for learning under Gaussian distributions.
Study the geometry of matrix multiplication in deep neural networks.
We study the fundamental limits on learning latent community structure in dynamic networks. Specifically, we study dynamic stochastic block models where nodes change their community membership over time, but where edges are generated independently at each time step. In this setting (which is a special case of several e…
We consider the problem of the optimal trading strategy in the presence of linear costs, and with a strict cap on the allowed position in the market. Using Bellman's backward recursion method, we show that the optimal strategy is to switch between the maximum allowed long position and the maximum allowed short position…
Improved bounds on combining hypothesis classes for binary functions.
Lower bound proves ridgeless regression performs poorly near interpolation threshold.
Quadratic regression involves modeling the response as a (generalized) linear function of not only the features but also of quadratic terms . The inclusion of such higher-order "interaction terms" in regression often provides an easy way to increase accuracy in already-high-dimensional problem…
The execution flow drives market dynamics, validated on real data.
We consider the problem of online active learning to collect data for regression modeling. Specifically, we consider a decision maker with a limited experimentation budget who must efficiently learn an underlying linear population model. Our main contribution is a novel threshold-based algorithm for selection of most i…
Iterative shrinkage/thresholding algorithm (ISTA) is a well-studied method for finding sparse solutions to ill-posed inverse problems. In this letter, we present a data-driven scheme for learning optimal thresholding functions for ISTA. The proposed scheme is obtained by relating iterations of ISTA to layers of a simpl…