Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

3256499741,298 · Jun 202019922001200920172026
48 results for data point removal

PUMA augments models to remove unique data points without performance loss.

problem Preserving model performance while removing unique training data points.
method Explicitly models data influence, reweights remaining data optimally.
result PUMA effectively removes unique data points without performance degradation.

Unified framework for removing unwanted information from machine learning models.

problem Removing undesirable features or data points from machine learning models while preserving utility.
method Information-theoretic regularization approach for data point and feature unlearning.
result Unified mathematical framework with provable guarantees for both data point and feature unlearning.

Ray Singer torsion is a numerical invariant associated with a compact Riemannian manifold equipped with a flat bundle and a Hermitian structure on this bundle. In this note we show how one can remove the dependence on the Riemannian metric and on the Hermitian structure with the help of a base point and of an Euler str…

1998-07-02abs ↗pdf ↗

A statistical framework for removing unwanted data domains in machine learning.

problem Removing unwanted data domains in machine learning while preserving desired performance.
method Modeling domains as probability distributions and using hypothesis testing to select samples to remove.
result Characterization of allowable edited data distributions and removal-preservation Pareto frontiers for various distribution families.

Paper presents a new way to estimate model changes without full model evaluation.

problem Efficiently estimating changes in model parameters and outputs due to data point removal.
method Dual representation of influence functions for linearizable models, reducing computational complexity.
result The dual representation can be an efficient alternative to original influence functions, especially for large models.

We assess cluster stability by trimming extreme points and tracking data range reduction.

problem Assessing stability of one-dimensional clusters.
method Probabilistic method using diameter-shrinkage ratio to track data range reduction.
result Our method achieves higher accuracy than classical tests in small or noisy samples.

Good data stewardship requires removal of data at the request of the data's owner. This raises the question if and how a trained machine-learning model, which implicitly stores information about its training data, should be affected by such a removal request. Is it possible to "remove" data from a machine-learning mode…

2019-11-08abs ↗pdf ↗

In machine learning, statistics, econometrics and statistical physics, cross-validation (CV) is used asa standard approach in quantifying the generalisation performance of a statistical model. A directapplication of CV in time-series leads to the loss of serial correlations, a requirement of preserving anynon-stationar…

2019-10-21abs ↗pdf ↗

We clarify the status of log-periodicity associated with speculative bubbles preceding financial crashes. In particular, we address Feigenbaum's [2001] criticism and show how it can be rebuked. Feigenbaum's main result is as follows: ``the hypothesis that the log-periodic component is present in the data cannot be reje…

2001-06-26abs ↗pdf ↗

The study calculates the index distribution of Brownian loops in various geometrical settings.

problem Calculating the distribution of the index of Brownian loops in specific geometrical settings.
method Analysis based on the geometry of Hopf and anti-de Sitter fibrations, and the relationship between winding and area forms.
result Explicit formulas and asymptotics for the distribution of the index of the Brownian loop.

Selective removal of data subsets can efficiently unlearn unwanted distributions.

problem Efficiently removing unwanted data subsets without losing important information.
method Formalized as distributional unlearning, using Kullback-Leibler divergence constraints to select a small subset of data.
result Proposed method achieves corresponding log-loss bounds and is quadratically more sample-efficient than random removal.

Bayesian inference forgetting framework removes influence of single data points.

problem Enforcement of the right to be forgotten in machine learning causes high costs for companies.
method Develops forgetting algorithms for variational and Markov chain Monte Carlo in Bayesian inference.
result Proves removal of influence of single datums on learned models with guaranteed generalizability.

Paper improves DP-ERM for binary linear classification with large-margin subsets.

problem Differentially private binary linear classification with large-margin subsets.
method Efficient (ε,δ)(\varepsilon,δ)-DP algorithm with empirical zero-one risk bound.
result Improved empirical zero-one risk bound for binary linear classification.

Gradient ascent method successfully removes specific data points from neural networks without retraining.

problem Addressing privacy and ethical concerns by removing specific data points from trained models.
method Gradient ascent approach to unlearning, leveraging the implicit bias of gradient descent towards margin maximization conditions.
result Gradient ascent method can successfully unlearn specific data points from two-layer ReLU neural networks without retraining.

Following on from ``Hyperbolic Plateau problems'' (by the same author), we provide a complete geometric description of solutions to the Plateau problem (S,φ)(S,φ) when SS is a compact Riemann surface with a finite number of points removed.

2005-06-13abs ↗pdf ↗

The importance of training robust neural network grows as 3D data is increasingly utilized in deep learning for vision tasks in robotics, drone control, and autonomous driving. One commonly used 3D data type is 3D point clouds, which describe shape information. We examine the problem of creating robust models from the …

2019-08-16abs ↗pdf ↗

Traditional text classifiers are limited to predicting over a fixed set of labels. However, in many real-world applications the label set is frequently changing. For example, in intent classification, new intents may be added over time while others are removed. We propose to address the problem of dynamic text classifi…

2019-11-04abs ↗pdf ↗

We study isolated singularities of two dimensional Yang-Mills-Higgs fields defined on a fiber bundle, where the fiber space is a compact Riemannian manifold and the structure group is a compact connected Lie group. In general the singularity can not be removed due to possibly non-vanishing limit holonomy around the sin…

2019-07-16abs ↗pdf ↗

Influence functions estimate the effect of removing a training point on a model without the need to retrain. They are based on a first-order Taylor approximation that is guaranteed to be accurate for sufficiently small changes to the model, and so are commonly used to study the effect of individual points in large data…

2019-05-30abs ↗pdf ↗

Proposes a new method to unlearn from specific data points in conformal predictors.

problem Challenges of existing unlearning methods in conformal predictors.
method Formalizes conformal unlearning, introduces practical metrics, and presents an optimization algorithm.
result Demonstrates effective removal of targeted information while preserving utility.

This paper proposes a novel Gaussian process approach to fault removal in time-series data. Fault removal does not delete the faulty signal data but, instead, massages the fault from the data. We assume that only one fault occurs at any one time and model the signal by two separate non-parametric Gaussian process model…

2015-07-02abs ↗pdf ↗

Data selection methods, such as active learning and core-set selection, are useful tools for machine learning on large datasets. However, they can be prohibitively expensive to apply in deep learning because they depend on feature representations that need to be learned. In this work, we show that we can greatly improv…

2019-06-26abs ↗pdf ↗

Let X be a compact 2-manifold with nonempty boundary dX and let f: (X, dX) --> (X, dX) be a boundary-preserving map. Denote by MF_d[f] the minimum number of fixed point among all boundary-preserving maps that are homotopic through boundary-preserving maps to f. The relative Nielsen number N_d(f) is the sum of the numbe…

2004-02-20abs ↗pdf ↗

We pursue the analogy of a framed flow category with the flow data of a Morse function. In classical Morse theory, Morse functions can sometimes be locally altered and simplified by the Morse moves. These moves include the Whitney trick which removes two oppositely framed flowlines between critical points of adjacent i…

2015-07-13abs ↗pdf ↗

Cluster analysis and outlier detection are strongly coupled tasks in data mining area. Cluster structure can be easily destroyed by few outliers; on the contrary, outliers are defined by the concept of cluster, which are recognized as the points belonging to none of the clusters. Unfortunately, most existing studies do…

2018-01-05abs ↗pdf ↗

This paper studies the regularity of constrained Willmore immersions into Rm3\R^{m\ge3} locally around both "regular" points and around branch points, where the immersive nature of the map degenerates. We develop local asymptotic expansions for the immersion, its first, and its second derivatives, given in terms of resi…

2012-11-19abs ↗pdf ↗

We demonstrate, theoretically and empirically, that adversarial robustness can significantly benefit from semisupervised learning. Theoretically, we revisit the simple Gaussian model of Schmidt et al. that shows a sample complexity gap between standard and robust classification. We prove that unlabeled data bridges thi…

2019-05-31abs ↗pdf ↗