Study shows flipping a small subset of labels can severely damage machine learning models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Efficiently poisons offline RLHF models by flipping preference labels.
Deep Partition Aggregation defends against poisoning attacks with provable certificates.
Many machine learning systems rely on data collected in the wild from untrusted sources, exposing the learning algorithms to data poisoning. Attackers can inject malicious data in the training dataset to subvert the learning process, compromising the performance of the algorithm producing errors in a targeted or an ind…
A double pants decomposition of a 2-dimensional surface is a collection of two pants decomposition of this surface introduced in arXiv:1005.0073v2. There are two natural operations acting on double pants decompositions: flips and handle twists. It is shown in arXiv:1005.0073v2 that the groupoid generated by flips and h…
Machine learning algorithms are known to be susceptible to data poisoning attacks, where an adversary manipulates the training data to degrade performance of the resulting classifier. In this work, we present a unifying view of randomized smoothing over arbitrary functions, and we leverage this novel characterization t…
The paper solves pentagon equations using triangulations and edge transformations.
It has been a long-standing problem to efficiently learn a halfspace using as few labels as possible in the presence of noise. In this work, we propose an efficient Perceptron-based algorithm for actively learning homogeneous halfspaces under the uniform distribution over the unit sphere. Under the bounded noise condit…
New insights on robust learning under strong noise models.
Label manipulation attacks are a subclass of data poisoning attacks in adversarial machine learning used against different applications, such as malware detection. These types of attacks represent a serious threat to detection systems in environments having high noise rate or uncertainty, such as complex networks and I…
Study flip graphs for surfaces of infinite type, finding uncountably many connected components.
Instance- and Label-dependent label Noise (ILN) widely exists in real-world datasets but has been rarely studied. In this paper, we focus on Bounded Instance- and Label-dependent label Noise (BILN), a particular case of ILN where the label noise rates -- the probabilities that the true labels of examples flip into the …
Proposes a method to increase diversity without sacrificing meritocracy.
Two-layer ReLU networks can overfit without harm, study finds.
Finite subgraphs in flip graphs ensure unique surface embeddings.
New algorithm uses conditionally invariant components to improve domain adaptation performance.
Early detection of breast cancer has a major contribution to curability, and using mammographic images, this can be achieved non-invasively. Supervised deep learning, the dominant CADe tool currently, has played a great role in object detection in computer vision, but it suffers from a limiting property: the need of a …
Study of flip graphs and their automorphism groups for infinite-type surfaces.
Study of skateboard flips as continuous curves in group.
Neural networks have been criticized for their lack of easy interpretation, which undermines confidence in their use for important applications. Here, we introduce a novel technique, interpreting a trained neural network by investigating its flip points. A flip point is any point that lies on the boundary between two o…
Flip symmetry on knot diagrams affects Khovanov homology.
In this paper, we study a classification problem in which sample labels are randomly corrupted. In this scenario, there is an unobservable sample with noise-free labels. However, before being observed, the true labels are independently flipped with a probability , and the random label noise can be class-co…
Image classification problems are typically addressed by first collecting examples with candidate labels, second cleaning the candidate labels manually, and third training a deep neural network on the clean examples. The manual labeling step is often the most expensive one as it requires workers to label millions of im…
The flip graph and arc complex of a surface are shown to have finite rigidity.
Novel defense algorithm improves SVMs against data poisoning attacks.
New examples show flip distance and polyhedron triangulation numbers differ, with ratio close to 3/2.
We prove that every injective simplicial map between flip graphs is induced by a subsurface inclusion , except in finitely many cases. This extends a result of Korkmaz--Papadopoulos which asserts that every automorphism of the flip graph of a surface without boundary is ind…
In label-noise learning, \textit{noise transition matrix}, denoting the probabilities that clean labels flip into noisy labels, plays a central role in building \textit{statistically consistent classifiers}. Existing theories have shown that the transition matrix can be learned by exploiting \textit{anchor points} (i.e…
We prove that for a given flat surface with conical singularities, any pair of geometric triangulations can be connected by a chain of flips.
We introduce a notion of cross-flips: local moves that transform a balanced (i.e., properly -colored) triangulation of a combinatorial -manifold into another balanced triangulation. These moves form a natural analog of bistellar flips (also known as Pachner moves). Specifically, we establish the following the…
This paper is about the geometry of flip-graphs associated to triangulations of surfaces. More precisely, we consider a topological surface with a privileged boundary curve and study the spaces of its triangulations with n vertices on the boundary curve. The surfaces we consider topologically fill this boundary curve s…
We consider geometric triangulations of surfaces, i.e., triangulations whose edges can be realized by disjoint locally geodesic segments. We prove that the flip graph of geometric triangulations with fixed vertices of a flat torus or a closed hyperbolic surface is connected. We give upper bounds on the number of edge f…
Geodesics count exponentially between triangulations of surfaces with enough topology.
Identifies minimal training subset to flip a prediction.
New method improves LLM judge accuracy by accounting for dependencies in aggregated binary labels.
In order to model volatile real-world network behavior, we analyze phase-flipping dynamical scale-free network in which nodes and links fail and recover. We investigate how stochasticity in a parameter governing the recovery process affects phase-flipping dynamics, and find the probability that no more than q% of nodes…
Let be a compact surface. We prove that the set of surface cubications modulo flips, up to isotopy, is in one-to-one correspondence with .
Study finds flipped classrooms improve student self-concept, enjoyment, but not exam scores.
Study quasisymmetric maps on hyperbolic plane boundaries.
We study flip-graphs of triangulations on topological surfaces where distance is measured by counting the number of necessary flip operations between two triangulations. We focus on surfaces of positive genus with a single boundary curve and marked points on this curve; we consider triangulations up to homeomor…
Using existing technology, we prove a Masur-Minsky style distance formula for flip- graph distance between two triangulations, expressed as a sum of the distances of the projections of these triangulations into arc graphs of the suitable subsurfaces of S.
Ridge regression shows different behaviors in binary classification with noisy labels.
Local graph clustering improves with noisy labels, enhancing accuracy and performance.
Federated learning has a variety of applications in multiple domains by utilizing private training data stored on different devices. However, the aggregation process in federated learning is highly vulnerable to adversarial attacks so that the global model may behave abnormally under attacks. To tackle this challenge, …
Adversarial examples are important for understanding the behavior of neural models, and can improve their robustness through adversarial training. Recent work in natural language processing generated adversarial examples by assuming white-box access to the attacked model, and optimizing the input directly against it (E…
We investigate a type of distance between triangulations on finite type surfaces where one moves between triangulations by performing simultaneous flips. We consider triangulations up to homeomorphism and our main results are upper bounds on distance between triangulations that only depend on the topology of the surfac…
We present Noisy Student Training, a semi-supervised learning approach that works well even when labeled data is abundant. Noisy Student Training achieves 88.4% top-1 accuracy on ImageNet, which is 2.0% better than the state-of-the-art model that requires 3.5B weakly labeled Instagram images. On robustness test sets, i…
Researchers use human-in-the-loop to create counterfactually augmented data, improving model performance.