Efficiently poisons offline RLHF models by flipping preference labels.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study shows flipping a small subset of labels can severely damage machine learning models.
Ridge regression shows different behaviors in binary classification with noisy labels.
New method improves LLM judge accuracy by accounting for dependencies in aggregated binary labels.
Deep Partition Aggregation defends against poisoning attacks with provable certificates.
Identifies minimal training subset to flip a prediction.
Many machine learning systems rely on data collected in the wild from untrusted sources, exposing the learning algorithms to data poisoning. Attackers can inject malicious data in the training dataset to subvert the learning process, compromising the performance of the algorithm producing errors in a targeted or an ind…
We consider the problem of learning linear classifiers when both features and labels are binary. In addition, the features are noisy, i.e., they could be flipped with an unknown probability. In Sy-De attribute noise model, where all features could be noisy together with same probability, we show that - loss ($l_{…
The paper proposes a least squares method for binary compressive sampling with low intrinsic dimension signals.
Neural networks have been criticized for their lack of easy interpretation, which undermines confidence in their use for important applications. Here, we introduce a novel technique, interpreting a trained neural network by investigating its flip points. A flip point is any point that lies on the boundary between two o…
A double pants decomposition of a 2-dimensional surface is a collection of two pants decomposition of this surface introduced in arXiv:1005.0073v2. There are two natural operations acting on double pants decompositions: flips and handle twists. It is shown in arXiv:1005.0073v2 that the groupoid generated by flips and h…
We consider the weakly supervised binary classification problem where the labels are randomly flipped with probability . Although there exist numerous algorithms for this problem, it remains theoretically unexplored how the statistical accuracies and computational efficiency of these algorithms depend on the degr…
Machine learning algorithms are known to be susceptible to data poisoning attacks, where an adversary manipulates the training data to degrade performance of the resulting classifier. In this work, we present a unifying view of randomized smoothing over arbitrary functions, and we leverage this novel characterization t…
In this paper, we use reinforcement learning to find effective decoding strategies for binary linear codes. We start by reviewing several iterative decoding algorithms that involve a decision-making process at each step, including bit-flipping (BF) decoding, residual belief propagation, and anchor decoding. We then ill…
Noisy PN learning is the problem of binary classification when training examples may be mislabeled (flipped) uniformly with noise rate rho1 for positive examples and rho0 for negative examples. We propose Rank Pruning (RP) to solve noisy PN learning and the open problem of estimating the noise rates, i.e. the fraction …
The paper solves pentagon equations using triangulations and edge transformations.
It has been a long-standing problem to efficiently learn a halfspace using as few labels as possible in the presence of noise. In this work, we propose an efficient Perceptron-based algorithm for actively learning homogeneous halfspaces under the uniform distribution over the unit sphere. Under the bounded noise condit…
Binary PheNorm extends phenotype labeling for EHRs using binary silver labels.
Just as semantic hashing can accelerate information retrieval, binary valued embeddings can significantly reduce latency in the retrieval of graphical data. We introduce a simple but effective model for learning such binary vectors for nodes in a graph. By imagining the embeddings as independent coin flips of varying b…
In this paper, we proposed a general framework for data poisoning attacks to graph-based semi-supervised learning (G-SSL). In this framework, we first unify different tasks, goals, and constraints into a single formula for data poisoning attack in G-SSL, then we propose two specialized algorithms to efficiently solve t…
Local graph clustering improves with noisy labels, enhancing accuracy and performance.
Deep learning models have been criticized for their lack of easy interpretation, which undermines confidence in their use for important applications. Nevertheless, they are consistently utilized in many applications, consequential to humans' lives, mostly because of their better performance. Therefore, there is a great…
New insights on robust learning under strong noise models.
Proposes MLPSVM for multi-label learning, improving on binary relevance.
Privacy subsidy found in market trading with noisy direction signals.
Proposes top-label calibration and M2B framework for multiclass to binary calibration.
Recently the deep learning techniques have achieved success in multi-label classification due to its automatic representation learning ability and the end-to-end learning framework. Existing deep neural networks in multi-label classification can be divided into two kinds: binary relevance neural network (BRNN) and thre…
Label manipulation attacks are a subclass of data poisoning attacks in adversarial machine learning used against different applications, such as malware detection. These types of attacks represent a serious threat to detection systems in environments having high noise rate or uncertainty, such as complex networks and I…
Binary classification improves with a small fraction of corrupted labels.
Study flip graphs for surfaces of infinite type, finding uncountably many connected components.
This work analyzes two methods for combining multiple binary labels in bipartite ranking.
Study reveals significant performance flips in GLOD using repurposed graph classification datasets.
Instance- and Label-dependent label Noise (ILN) widely exists in real-world datasets but has been rarely studied. In this paper, we focus on Bounded Instance- and Label-dependent label Noise (BILN), a particular case of ILN where the label noise rates -- the probabilities that the true labels of examples flip into the …
Paper proposes Pcomp classification for binary classification with pairwise confidence comparisons.
After being trained, classifiers must often operate on data that has been corrupted by noise. In this paper, we consider the impact of such noise on the features of binary classifiers. Inspired by tools for classifier robustness, we introduce the same classification probability (SCP) to measure the resulting distortion…
Proposes a method to increase diversity without sacrificing meritocracy.
Two-layer ReLU networks can overfit without harm, study finds.
Finite subgraphs in flip graphs ensure unique surface embeddings.
Paper estimates FPR of Bayes classifier using soft labels.
New algorithm uses conditionally invariant components to improve domain adaptation performance.
Paper introduces DMPMs for efficient discrete data generation with sharp convergence bounds.
Early detection of breast cancer has a major contribution to curability, and using mammographic images, this can be achieved non-invasively. Supervised deep learning, the dominant CADe tool currently, has played a great role in object detection in computer vision, but it suffers from a limiting property: the need of a …
Study of flip graphs and their automorphism groups for infinite-type surfaces.
Sharp bounds on uniform generalization errors in binary linear classification.
Study of skateboard flips as continuous curves in group.
Networked sensing, where the goal is to perform complex inference using a large number of inexpensive and decentralized sensors, has become an increasingly attractive research topic due to its applications in wireless sensor networks and internet-of-things. To reduce the communication, sensing and storage complexity, t…
Flip symmetry on knot diagrams affects Khovanov homology.
In this paper, we study a classification problem in which sample labels are randomly corrupted. In this scenario, there is an unobservable sample with noise-free labels. However, before being observed, the true labels are independently flipped with a probability , and the random label noise can be class-co…