Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

75149224298 · Jun 202019922001200920172026
48 results for forgotten examples

The paper analyzes how forgetting in LLMs is linked to simple task-upstream example associations.

problem Forgetting of upstream knowledge in fine-tuned LLMs.
method Empirical analysis of forgotten examples in NN upstream examples after MM new tasks, using low-rank matrix approximation.
result Forgetting can be predicted efficiently using matrix completion over empirical associations.

Rediscovered by a systematic search, a forgotten class of integrable surfaces is shown to disprove the Finkel-Wu conjecture. The associated integrable nonlinear partial differential equation zyy+(1/z)xx+2=0 z_{yy} + (1/z)_{xx} + 2 = 0 possesses a zero curvature representation, a third-order symmetry, and a nonlocal transformatio…

2010-02-04abs ↗pdf ↗

Proposes a new method to unlearn from specific data points in conformal predictors.

problem Challenges of existing unlearning methods in conformal predictors.
method Formalizes conformal unlearning, introduces practical metrics, and presents an optimization algorithm.
result Demonstrates effective removal of targeted information while preserving utility.

AI models forget statistics' lesson: correlation doesn't imply causation.

problem AI models often produce flawed causal models due to ignoring correlation vs causation.
method Demonstrates examples of flawed AI models and proposes rethinking core models.
result Current efforts to make AI models ethical are insufficient.

This work bridges continual learning, active learning, and open set recognition in deep neural networks.

problem Protecting previously acquired representations from catastrophic forgetting in deep neural networks.
method Surveying the literature and proposing a consolidated view to integrate open set recognition and active learning principles.
result Joint improvement in alleviating catastrophic forgetting, querying data, selecting task orders, and robust open world application.

Cycloids, hipocycloids and epicycloids have an often forgotten common property: they are homothetic to their evolutes. But what if use convex symmetric polygons as unit balls, can we define evolutes and cycloids which are genuinely discrete? Indeed, we can! We define discrete cycloids as eigenvectors of a discrete doub…

2017-02-02abs ↗pdf ↗

Bayesian inference forgetting framework removes influence of single data points.

problem Enforcement of the right to be forgotten in machine learning causes high costs for companies.
method Develops forgetting algorithms for variational and Markov chain Monte Carlo in Bayesian inference.
result Proves removal of influence of single datums on learned models with guaranteed generalizability.

The paper proposes selective forgetting for deep neural networks at a finer level than samples.

problem Selective forgetting of deep neural networks to handle outliers, poisoned data, or sensitive information.
method Formulated selective forgetting at a finer level than samples, introduced as an optimization problem on three criteria.
result Experimental results show the model can forget specific information for classification, improving accuracy in specific cases.

Intense recent discussions have focused on how to provide individuals with control over when their data can and cannot be used --- the EU's Right To Be Forgotten regulation is an example of this effort. In this paper we initiate a framework studying what to do when it is no longer permissible to deploy models derivativ…

2019-07-11abs ↗pdf ↗

Abstract sketches historical development of Lie brackets, crossed modules, and Lie-Rinehart algebras.

problem Characterizing and understanding the relationships between Lie brackets, crossed modules, and Lie-Rinehart algebras.
method Historical review and combinatorial group theory considerations.
result The mutual relationship between Lie-Rinehart algebras and Lie brackets, and the historical development of these concepts.

New method quantifies uncertainty in fine-tuned LLMs using LoRA ensembles.

problem Uncertainty in fine-tuned LLMs and how to trust their predictions.
method Posterior approximations using low-rank adaptation ensembles.
result Unexpected retention of acquired knowledge during fine-tuning in overfitting regime.

This paper presents the contemporary Fundamental Theorem of Asset Pricing as being equivalent to approaches to pricing that emerged before 1700 in the context of Virtue Ethics. This is done by considering the history of science and mathematics in the thirteenth and seventeenth century. An explanation as to why these ap…

2012-10-19abs ↗pdf ↗

This paper develops efficient federated learning and unlearning methods in Bayesian models.

problem Managing epistemic uncertainty and legal right to be forgotten in decentralized networks.
method Develops federated variational inference solutions based on decentralized local free energy minimization.
result Demonstrates efficient unlearning mechanisms in federated learning and unlearning.

Paper uses second-order differential geometry to study stochastic mechanics.

problem Stochastic differential equations and their symmetries.
method Develops second-order differential geometry to study symmetries of SDEs and constructs stochastic mechanics.
result Establishes stochastic Lagrangian and Hamiltonian mechanics and their relations with HJB equations.

DVWU framework improves model performance by considering data value heterogeneity.

problem Existing machine unlearning algorithms ignore data value heterogeneity, potentially degrading model performance.
method Data Value-Weighted Unlearning (DVWU) framework that integrates data values into the unlearning process.
result DVWU achieves superior predictive performance and robustness compared to conventional unlearning approaches.

A novel method for parallel transport and geodesics on submanifolds.

problem Understanding parallel transport and geodesics on submanifolds.
method Rolling tangent space to visualize and analyze parallel transport and geodesics.
result Conditions for parallel transport and geodesics are simplified and visualized in the tangent space.

Conditional GANs are at the forefront of natural image synthesis. The main drawback of such models is the necessity for labeled data. In this work we exploit two popular unsupervised learning techniques, adversarial training and self-supervision, and take a step towards bridging the gap between conditional and uncondit…

2018-11-27abs ↗pdf ↗

Study on deleting user data in linear regression models to maintain limited memory.

problem Deleting user data in a limited time frame for statistical models.
method Proposed FIFD-OLS and FIFD-Adaptive Ridge algorithms for low-dimensional and online settings.
result Demonstrated effectiveness of FIFD-Adaptive Ridge in maintaining statistical efficiency.

Procedure removes training data dependency from deep networks, improving generalization.

problem Removing dependency on training data in deep networks for better generalization.
method Deterministic and stochastic parts to ensure forgetting, leveraging activation and weight dynamics.
result New bound on information extraction from black-box networks, ensuring forgetting in activations.

New method removes specific training data influence from neural networks.

problem Removing specific training data influence from neural networks for privacy and regulatory reasons.
method Noisy fine-tuning on retain data to ensure provable unlearning guarantees without restrictive assumptions.
result Achieves formal unlearning guarantees and performs effectively in practice.

Algorithm removes specific training data from models efficiently in high-dimensional settings.

problem Efficiently removing specific training data from high-dimensional models without full retraining.
method Starts from original model parameters, performs Newton steps, adds isotropic Laplacian noise.
result Two Newton steps are sufficient for effective unlearning in high-dimensional problems.

Paper proposes machine unlearning method to forget user data from neural networks.

problem Memorization of user data in neural networks violates GDPR's right to be forgotten.
method Proposes Forsaken method to measure and achieve high forgetting rates without significant accuracy loss.
result Forsaken method achieves over 90% forgetting rate with less than 5% accuracy loss.

A framework for certified unlearning in decentralized federated learning.

problem Privacy-preserving machine learning in decentralized federated learning.
method Newton-style updates to quantify and correct data influence, using Fisher information matrices for scalability.
result The proposed framework ensures that the unlearned model is difficult to distinguish from a retrained model without the deleted data.

Modified PCA algorithm with continual learning preserves features of previous modes for multimode process monitoring.

problem Catastrophic forgetting of previous modes in monitoring models for successive modes.
method Modified PCA algorithm with elastic weight consolidation (EWC) to preserve features of previous modes.
result PCA-EWC algorithm effectively monitors multimode processes without performance decrease.

We consider parametric exponential families of dimension KK on the real line. We study a variant of \textit{boundary crossing probabilities} coming from the multi-armed bandit literature, in the case when the real-valued distributions form an exponential family of dimension KK. Formally, our result is a concentration…

2017-05-24abs ↗pdf ↗

Paper shows incorrectness of approximate unlearning definitions and challenges exact unlearning verification.

problem Incorrectness of approximate unlearning definitions and challenges in verifying exact unlearning.
method Analysis of machine unlearning approaches, including exact and approximate methods.
result Unlearning is only well-defined at the algorithmic level, and auditable claims are limited.

A new method for forgetting data from trained models using information theory.

problem Efficiently forgetting private or copyrighted data from trained machine learning models.
method An information-theoretic approach to zero-shot unlearning, minimizing gradient smoothing.
result Our method successfully unlearns data while maintaining model performance.

Paper proposes ManiF-SMC for effective approximate machine unlearning.

problem Limited unlearning effectiveness and potential to undermine original learning objectives.
method Reformulates approximate unlearning as pushing erased samples towards semantic neighbors in retained data, using a margin-based triplet loss.
result Achieves unlearning effectiveness comparable to state-of-the-art methods while operating purely in representation space.