Neural model predicts procedure names in stripped binaries.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In the top-down approach to multi-name credit modeling, calculation of singe name sensitivities appears possible, at least in principle, within the so-called random thinning (RT) procedure which dissects the portfolio risk into individual contributions. We make an attempt to construct a practical RT framework that enab…
We present the first tree-based regressor whose convergence rate depends only on the intrinsic dimension of the data, namely its Assouad dimension. The regressor uses the RPtree partitioning procedure, a simple randomized variant of k-d trees.
Empirical study shows GANs overfit and drop modes when training is deterministic.
New method calibrates crypto option prices more robustly.
ARK improves knockoffs robustness to feature distribution misspecification.
Given a finite family of functions, the goal of model selection aggregation is to construct a procedure that mimics the function from this family that is the closest to an unknown regression function. More precisely, we consider a general regression model with fixed design and measure the distance between functions by …
In this article, we consider the sparse tensor singular value decomposition, which aims for dimension reduction on high-dimensional high-order data with certain sparsity structure. A method named Sparse Tensor Alternating Thresholding for Singular Value Decomposition (STAT-SVD) is proposed. The proposed procedure featu…
We consider the problem of performing matrix completion with side information on row-by-row and column-by-column similarities. We build upon recent proposals for matrix estimation with smoothness constraints with respect to row and column graphs. We present a novel iterative procedure for directly minimizing an informa…
Researchers create a method to join hyperboloidal data sets without violating the shear-free condition.
Proposes cost-sensitive feature selection for SVMs.
Random forests are powerful non-parametric regression method but are severely limited in their usage in the presence of randomly censored observations, and naively applied can exhibit poor predictive performance due to the incurred biases. Based on a local adaptive representation of random forests, we develop its regre…
New method estimates multiscaling exponents for financial risk assessment.
Framework mitigates risk non-monotonicity in high-dimensional predictions.
We solve the integration problem for generalized complex manifolds, obtaining as the natural integrating object a weakly holomorphic symplectic groupoid, which is a real symplectic groupoid with a compatible complex structure defined only on the associated stack, i.e., only up to Morita equivalence. We explain how such…
AdaVol adapts QML for real-time GARCH volatility prediction.
Method estimates noise variance in Gaussian process regression.
The paper proves deep ReLU networks can die and proposes a new initialization method to prevent it.
Develops a method to estimate quantiles in censored data using random forests.
A method to uniformly sample graph-encoded surfaces of fixed size.
Non-convex optimization is ubiquitous in machine learning. Majorization-Minimization (MM) is a powerful iterative procedure for optimizing non-convex functions that works by optimizing a sequence of bounds on the function. In MM, the bound at each iteration is required to \emph{touch} the objective function at the opti…
This paper deals with finding an -dimensional solution to a system of quadratic equations of the form for , which is also known as phase retrieval and is NP-hard in general. We put forth a novel procedure for minimizing the amplitude-based least-squares empirical los…
We present a methodology to extract the backbone of complex networks based on the weight and direction of links, as well as on nontopological properties of nodes. We show how the methodology can be applied in general to networks in which mass or energy is flowing along the links. In particular, the procedure enables us…
Extends knockoff filter for composite null hypotheses in variable selection.
The study sets limits on dihedral angles of large hyperbolic polyhedra.
Regularisation method studies Lie algebroids via foliated structures.
We study adaptive data-dependent dimensionality reduction in the context of supervised learning in general metric spaces. Our main statistical contribution is a generalization bound for Lipschitz functions in metric spaces that are doubling, or nearly doubling. On the algorithmic front, we describe an analogue of PCA f…
ISLET efficiently estimates low-rank tensors with optimal performance and speed.
We study a transformation of metric measure spaces introduced by Gigli and Mantegazza consisting in replacing the original distance with the length distance induced by the transport distance between heat kernel measures. We study the smoothing effect of this procedure in two important examples. Firstly, we show that in…
Wave propagator constructed on globally hyperbolic spacetimes.
A new method improves language model extrapolation without changing training.
Twinning splits data into fast, statistically similar sets.
We study a maturity randomization technique for approximating optimal control problems. The algorithm is based on a sequence of control problems with random terminal horizon which converges to the original one. This is a generalization of the so-called Canadization procedure suggested by Carr [Review of Financial Studi…
Efficiently estimates shrinkage coefficient for RTME using LOOCV approximation.
Paper fine-tunes LLMs using user edits, unifying preference, supervision, and reward feedback.
DNA-SE uses deep learning to solve semiparametric problems efficiently.
PASOA optimizes Bayesian design by improving SMC samplers and EIG.
The completion of tensors, or high-order arrays, attracts significant attention in recent research. Current literature on tensor completion primarily focuses on recovery from a set of uniformly randomly measured entries, and the required number of measurements to achieve recovery is not guaranteed to be optimal. In add…
Bayesian Lasso Sparse model provides sparse estimates in linear and nonlinear regression.
There has been a recent surge of interest in studying permutation-based models for ranking from pairwise comparison data. Despite being structurally richer and more robust than parametric ranking models, permutation-based models are less well understood statistically and generally lack efficient learning algorithms. In…
The equivalence problem of curves with values in a Riemannian manifold, is solved. The domain of validity of Frenet's theorem is shown to be the spaces of constant curvature. For a general Riemannian manifold new invariants must thus be added. There are two important generic classes of curves; namely, Frenet curves and…
We show that the herding procedure of Welling (2009) takes exactly the form of a standard convex optimization algorithm--namely a conditional gradient algorithm minimizing a quadratic moment discrepancy. This link enables us to invoke convergence results from convex optimization and to consider faster alternatives for …
Novel method for learning Gaussian graphical models from paired data.
In this paper, we consider the problem of "hyper-sparse aggregation". Namely, given a dictionary of functions, we look for an optimal aggregation algorithm that writes with as many zero coefficients as possible. This problem is of particular interest when…
Probabilistic graphical models are graphical representations of probability distributions. Graphical models have applications in many fields including biology, social sciences, linguistic, neuroscience. In this paper, we propose directed acyclic graphs (DAGs) learning via bootstrap aggregating. The proposed procedure i…
Quadratic discriminant analysis (QDA) is a standard tool for classification due to its simplicity and flexibility. Because the number of its parameters scales quadratically with the number of the variables, QDA is not practical, however, when the dimensionality is relatively large. To address this, we propose a novel p…
The paper analyzes two ISGD modes for statistical inference, deriving error bounds and confidence intervals.
Enhances count process modelling with Markov-modulated non-homogeneous Poisson process.