Estimates the upper bound of linear regions in spheres centered at specific data points in ReLU neural networks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In label-noise learning, \textit{noise transition matrix}, denoting the probabilities that clean labels flip into noisy labels, plays a central role in building \textit{statistically consistent classifiers}. Existing theories have shown that the transition matrix can be learned by exploiting \textit{anchor points} (i.e…
Proves critical points of ADM mass correspond to specific initial data sets.
Data augmentation is a ubiquitous technique for increasing the size of labeled training sets by leveraging task-specific data transformations that preserve class labels. While it is often easy for domain experts to specify individual transformations, constructing and tuning the more sophisticated compositions typically…
This work defines observation-specific explanations for black-box models.
This paper proposes a centroid-based clustering algorithm which is capable of clustering data-points with n-features, without having to specify the number of clusters to be formed. The core logic behind the algorithm is a similarity measure, which collectively decides whether to assign an incoming data-point to a pre-e…
Captures data influence changes during training.
The paper classifies circle actions on 6D manifolds with isolated fixed points.
The paper develops methods to create private synthetic spatial point patterns.
A new method clusters rows of a matrix of point processes.
Gradient ascent method successfully removes specific data points from neural networks without retraining.
We design a general framework for answering adaptive statistical queries that focuses on providing explicit confidence intervals along with point estimates. Prior work in this area has either focused on providing tight confidence intervals for specific analyses, or providing general worst-case bounds for point estimate…
Develops a faster model selection method using influence functions.
We study the relationship between social media output and National Football League (NFL) games, using a dataset containing messages from Twitter and NFL game statistics. Specifically, we consider tweets pertaining to specific teams and games in the NFL season and use them alongside statistical game data to build predic…
Proposes a new method to unlearn from specific data points in conformal predictors.
New privacy framework tailored to specific data distributions.
Improved local multivariable regression for better inference with limited data.
We study parameter estimation and asymptotic inference for sparse nonlinear regression. More specifically, we assume the data are given by , where is nonlinear. To recover , we propose an -regularized least-squares estimator. Unlike classical linear regression, the correspondin…
High-dimensional data often lie in low-dimensional subspaces corresponding to different classes they belong to. Finding sparse representations of data points in a dictionary built using the collection of data helps to uncover low-dimensional subspaces and address problems such as clustering, classification, subset sele…
The paper is devoted to elaboration of a novel specific indicator based on the modified Holder exponents. This indicator has been used for forecasting critical points of financial time series and crashes of the USA stock market. The proposed approach is based on the hypothesis, which claims that before market critical …
Evaluates change point detection algorithms on real-world data.
Proposes first privacy-preserving method for estimating Hawkes processes.
Deep learning models can infer individual trajectories from sparse data.
Gaussian processes (GPs) are flexible non-parametric models, with a capacity that grows with the available data. However, computational constraints with standard inference procedures have limited exact GPs to problems with fewer than about ten thousand training points, necessitating approximations for larger datasets. …
Deep neural networks predict prostate motion from MR images.
The paper develops a neural network-based method for detecting change points in large-scale time-evolving data.
This paper extends the analysis of Muni Toke and Yoshida (2020) to the case of marked point processes. We consider multiple marked point processes with intensities defined by three multiplicative components, namely a common baseline intensity, a state-dependent component specific to each process, and a state-dependent …
Bayesian method detects outliers and uncertain points in data.
This paper describes a novel approach to change-point detection when the observed high-dimensional data may have missing elements. The performance of classical methods for change-point detection typically scales poorly with the dimensionality of the data, so that a large number of observations are collected after the t…
The problem of clustering noisy and incompletely observed high-dimensional data points into a union of low-dimensional subspaces and a set of outliers is considered. The number of subspaces, their dimensions, and their orientations are assumed unknown. We propose a simple low-complexity subspace clustering algorithm, w…
The fuzzy ROC extends Receiver Operating Curve (ROC) visualization to the situation where some data points, falling in an indeterminacy region, are not classified. It addresses two challenges: definition of sensitivity and specificity bounds under indeterminacy; and visual summarization of the large number of possibili…
The paper explores how to select data points for optimal learning performance.
Score matching fails for general point processes, a new estimator improves accuracy.
Improved modeling of persistence diagrams for data analysis.
We present a novel technique based on deep learning and set theory which yields exceptional classification and prediction results. Having access to a sufficiently large amount of labelled training data, our methodology is capable of predicting the labels of the test data almost always even if the training data is entir…
It is a key to construct a similarity graph in graph-oriented subspace learning and clustering. In a similarity graph, each vertex denotes a data point and the edge weight represents the similarity between two points. There are two popular schemes to construct a similarity graph, i.e., pairwise distance based scheme an…
New method helps interpret complex models by visualizing feature shifts.
Graph change-point detection method learns graph similarity from data.
We continue our computation, using a combinatorial method based on Gronthendieck's dessins d'enfant, of the number of (weak) equivalence classes of surface branched covers matching certain specific branch data. In this note we concentrate on data with the surface of genus g as source surface, the sphere as target surfa…
NDDV estimates data point value from a single stochastic trajectory.
This research tackles data deletion in linear regression with noisy SGD, finding perfect deleted points.
We propose a family of near-metrics based on local graph diffusion to capture similarity for a wide class of data sets. These quasi-metametrics, as their names suggest, dispense with one or two standard axioms of metric spaces, specifically distinguishability and symmetry, so that similarity between data points of arbi…
Repo dealers' market power affects bond prices by up to 2 percentage points.
New method estimates data influence efficiently by leveraging test samples.
In an earlier work we identified the types and numbers of static equilibrium points of solids arising from fine, equidistant -discretrizations of smooth, convex surfaces. We showed that such discretizations carry equilibrium points on two scales: the local scale corresponds to the discretization, the global scale to…
New method detects text changes under dependencies, outperforming baselines.
Fair active learning selects data points to balance model accuracy and fairness.
Study minimax estimation of stratified structure from i.i.d. samples.