A new algorithm selects data subsets avoiding outliers and high leverage points.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Active learning aims to obtain a classifier of high accuracy by using fewer label requests in comparison to passive learning by selecting effective queries. Many active learning methods have been developed in the past two decades, which sample queries based on informativeness or representativeness of unlabeled data poi…
D2D converts CLDs into SDMs to explore leverage points under uncertainty.
New robust estimator for high-dimensional data with outliers and leverage points.
We present a simple agent-based model of a financial system composed of leveraged investors such as banks that invest in stocks and manage their risk using a Value-at-Risk constraint, based on historical observations of asset prices. The Value-at-Risk constraint implies that when perceived risk is low, leverage is high…
Modeling price formation with interacting Hawkes processes leading to stochastic volatility with leverage.
We propose a fast inference method for Bayesian nonlinear support vector machines that leverages stochastic variational inference and inducing points. Our experiments show that the proposed method is faster than competing Bayesian approaches and scales easily to millions of data points. It provides additional features …
We give the first algorithm for kernel Nyström approximation that runs in *linear time in the number of training points* and is provably accurate for all kernel matrices, without dependence on regularity or incoherence conditions. The algorithm projects the kernel onto a set of landmark points sampled by their *rid…
Upper bounds on fixed points in PWL neural networks with hyperplane analysis.
The leverage effect refers to the generally negative correlation between the return of an asset and the changes in its volatility. There is broad agreement in the literature that the effect should be present for theoretical reasons, and it has been consistently found in empirical work. However, a few papers have pointe…
Effective risk control must make a tradeoff between the microprudential risk of exogenous shocks to individual institutions and the macroprudential risks caused by their systemic interactions. We investigate a simple dynamical model for understanding this tradeoff, consisting of a bank with a leverage target and an unl…
A new robust PCA method uses Innovation Search and Leverage Scores.
Extends importance sampling to nonlinear models using adjoint operators.
t-SNE algorithm's points remain bounded under gradient flow.
Estimates volatility of volatility and leverage effect using high-frequency options data.
New method assesses individual training points' privacy risk without retraining.
New method upsamples sparse, non-uniform point clouds more accurately.
TAGM models time-varying connections between variables.
Historical daily data for eleven years of the fifty constituent stocks of the NIFTY index traded on the National Stock Exchange have been analyzed to check for the stylized facts in the Indian market. It is observed that while some stylized facts of other markets are also observed in Indian market, there are significan…
Point processes are becoming very popular in modeling asynchronous sequential data due to their sound mathematical foundation and strength in modeling a variety of real-world phenomena. Currently, they are often characterized via intensity function which limits model's expressiveness due to unrealistic assumptions on i…
We extend the Frank-Wolfe (FW) optimization algorithm to solve constrained smooth convex-concave saddle point (SP) problems. Remarkably, the method only requires access to linear minimization oracles. Leveraging recent advances in FW optimization, we provide the first proof of convergence of a FW-type saddle point solv…
Modeling bank leverage dynamics using dynamical systems and neural networks.
Detects graph topology changes from noisy signals using prior spectral information.
ASkotch solves large-scale KRR faster and better than existing methods.
Existing techniques to compress point cloud attributes leverage either geometric or video-based compression tools. We explore a radically different approach inspired by recent advances in point cloud representation learning. Point clouds can be interpreted as 2D manifolds in 3D space. Specifically, we fold a 2D grid on…
DDEQs extend DEQs to discrete measure inputs using Wasserstein gradient flows.
Current approaches for explaining machine learning models fall into two distinct classes: antecedent event influence and value attribution. The former leverages training instances to describe how much influence a training point exerts on a test point, while the latter attempts to attribute value to the features most pe…
Short-term load forecasting is a critical element of power systems energy management systems. In recent years, probabilistic load forecasting (PLF) has gained increased attention for its ability to provide uncertainty information that helps to improve the reliability and economics of system operation performances. This…
This letter presents a new spectral-clustering-based approach to the subspace clustering problem. Underpinning the proposed method is a convex program for optimal direction search, which for each data point d finds an optimal direction in the span of the data that has minimum projection on the other data points and non…
Recent years have seen increased interest in performance guarantees of gradient descent algorithms for non-convex optimization. A number of works have uncovered that gradient noise plays a critical role in the ability of gradient descent recursions to efficiently escape saddle-points and reach second-order stationary p…
Solves capillary curvature problems for specific p values.
Study reveals different types of critical points in shallow neural networks.
Proposes a novel approach for cluster-aware matching using Laplacian Optimal Transport.
We propose networked exponential families to jointly leverage the information in the topology as well as the attributes (features) of networked data points. Networked exponential families are a flexible probabilistic model for heterogeneous datasets with intrinsic network structure. These models can be learnt efficient…
Gaussian processes (GPs) are flexible non-parametric models, with a capacity that grows with the available data. However, computational constraints with standard inference procedures have limited exact GPs to problems with fewer than about ten thousand training points, necessitating approximations for larger datasets. …
Paper explores how DPP sampling can implicitly regularize kernel regression.
New method assesses prediction intervals across different operating points.
Selecting diverse and important items, called landmarks, from a large set is a problem of interest in machine learning. As a specific example, in order to deal with large training sets, kernel methods often rely on low rank matrix Nyström approximations based on the selection or sampling of landmarks. In this context, …
Given a matrix and a vector , we show how to compute an -approximate solution to the regression problem in time where …
SE(3)-Transformers maintain equivariance for 3D data under rotations and translations.
We present a novel approach to leverage large unlabeled datasets by pre-training state-of-the-art deep neural networks on randomly-labeled datasets. Specifically, we train the neural networks to memorize arbitrary labels for all the samples in a dataset and use these pre-trained networks as a starting point for regular…
New method detects and locates changes in spatio-temporal point processes.
MPMC generates low-discrepancy points using graph neural networks.
Although over 100 languages are supported by strong off-the-shelf machine translation systems, only a subset of them possess large annotated corpora for named entity recognition. Motivated by this fact, we leverage machine translation to improve annotation-projection approaches to cross-lingual named entity recognition…
In healthcare, patient risk stratification models are often learned using time-series data extracted from electronic health records. When extracting data for a clinical prediction task, several formulations exist, depending on how one chooses the time of prediction and the prediction horizon. In this paper, we show how…
Unified framework denoises data and abstains from uncertain predictions.
A registration-free framework monitors shape and color in 4D point clouds.
IGNN captures long-range graph dependencies using fixed-point equations.