We inject undetectable backdoors into obfuscated neural networks and language models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We consider the problem of obfuscating sensitive information while preserving utility, and we propose a machine learning approach inspired by the generative adversarial networks paradigm. The idea is to set up two nets: the generator, that tries to produce an optimal obfuscation mechanism to protect the data, and the c…
This paper detects function-level obfuscation in binary code using graph-based methods.
The goal of homomorphic encryption is to encrypt data such that another party can operate on it without being explicitly exposed to the content of the original data. We introduce an idea for a privacy-preserving transformation on natural language data, inspired by homomorphic encryption. Our primary tool is {\em obfusc…
Identifying features that leak information about sensitive attributes is a key challenge in the design of information obfuscation mechanisms. In this paper, we propose a framework to identify information-leaking features via information density estimation. Here, features whose information densities exceed a pre-defined…
New flaw found in SAP defense, reducing its effectiveness to 0.1%.
Malware constitutes a major global risk affecting millions of users each year. Standard algorithms in detection systems perform insufficiently when dealing with malware passed through obfuscation tools. We illustrate this studying in detail an open source metamorphic software, making use of a hybrid framework to obtain…
It has been shown that adversaries can craft example inputs to neural networks which are similar to legitimate inputs but have been created to purposely cause the neural network to misclassify the input. These adversarial examples are crafted, for example, by calculating gradients of a carefully defined loss function w…
Improves code2vec for Java classes by obfuscating variable names.
New index principle shows indistinguishability of certain knot crossings.
The paper examines how to test if two learning algorithms produce similar outcomes.
Deep neural networks require large amounts of resources which makes them hard to use on resource constrained devices such as Internet-of-things devices. Offloading the computations to the cloud can circumvent these constraints but introduces a privacy risk since the operator of the cloud is not necessarily trustworthy.…
Executing deep neural networks for inference on the server-class or cloud backend based on data generated at the edge of Internet of Things is desirable due primarily to the limited compute power of edge devices and the need to protect the confidentiality of the inference neural networks. However, such a remote inferen…
Crowdsourced data used in machine learning services might carry sensitive information about attributes that users do not want to share. Various methods have been proposed to minimize the potential information leakage of sensitive attributes while maximizing the task accuracy. However, little is known about the theory b…
Characterizes sample complexity for outcome indistinguishability in machine learning.
In this paper we investigate the usage of adversarial perturbations for the purpose of privacy from human perception and model (machine) based detection. We employ adversarial perturbations for obfuscating certain variables in raw data while preserving the rest. Current adversarial perturbation methods are used for dat…
Deep learning models misclassify malware with added benign features.
The Internet of Things (IoT) will be a main data generation infrastructure for achieving better system intelligence. However, the extensive data collection and processing in IoT also engender various privacy concerns. This paper provides a taxonomy of the existing privacy-preserving machine learning approaches develope…
Framework uses human judgment to distinguish algorithmically indistinguishable cases.
Adversarial training is an effective methodology for training deep neural networks that are robust against adversarial, norm-bounded perturbations. However, the computational cost of adversarial training grows prohibitively as the size of the model and number of input dimensions increase. Further, training against less…
LLMs detect market patterns through causal reasoning, not just temporal association.
Survey on calibration in machine learning, viewing it as indistinguishability.
AI predicts stock winners with 2.43 Sharpe ratio, but returns are highly concentrated.
NKI integrates obfuscated datasets using nonlinear kernels for improved data collaboration.
A new metric evaluates classification algorithms at the point of indistinguishability.
We consider the problem of publicly releasing a dataset for support vector machine classification while not infringing on the privacy of data subjects (i.e., individuals whose private information is stored in the dataset). The dataset is systematically obfuscated using an additive noise for privacy protection. Motivate…
This work enhances collaborative inference privacy by minimizing conditional entropy and boosting robustness against model inversion attacks.
Efficiently generates models resistant to falsification.
New defense method inspired by encryption improves visual classification accuracy.
The paper formalizes robot environments using topological concepts.
RENNs protect input privacy by rotating d-ary features.
We investigate the possibility of statistical evaluation of the market completeness for discrete time stock market models. It is known that the market completeness is not a robust property: small random deviations of the coefficients convert a complete market model into a incomplete one. The paper shows that market inc…
Paper defends sensitive attributes in GNNs from inference attacks.
New approach shows backdoor attacks are indistinguishable from natural data features.
Transformers approximate mean-field dynamics of indistinguishable particles.
Paper proposes a method to encrypt faces while maintaining visual similarity.
The possibility of statistical evaluation of the market completeness and incompleteness is investigated for continuous time diffusion stock market models. It is known that the market completeness is not a robust property: small random deviations of the coefficients convert a complete market model into a incomplete one.…
Graph convolutional networks (GCNs) are a widely used method for graph representation learning. To elucidate the capabilities and limitations of GCNs, we investigate their power, as a function of their number of layers, to distinguish between different random graph models (corresponding to different class-conditional d…
Unsupervised learning models can be indistinguishable without identifiability, leading to unreliable representations.
Graphs indistinguishable by GNNs are fully characterized.
Time dilation and relative velocity are observationally indistinguishable in the special theory of relativity, a duality that carries over into the general theory under Fermi coordinates along a curve (in coordinate-independent language, in the tangent Minkowski space along the curve). For …
We show social events can be accurately predicted, but often undesirably.
Digital image steganalysis, or the detection of image steganography, has been studied in depth for years and is driven by Advanced Persistent Threat (APT) groups', such as APT37 Reaper, utilization of steganographic techniques to transmit additional malware to perform further post-exploitation activity on a compromised…
Efficient algorithms for deleting data from machine learning models without significantly affecting performance.
In this work we revisit gradient regularization for adversarial robustness with some new ingredients. First, we derive new per-image theoretical robustness bounds based on local gradient information. These bounds strongly motivate input gradient regularization. Second, we implement a scaleable version of input gradient…
Simplifies fair PCA with fast, efficient solution.
Study on replicability in reinforcement learning algorithms.
Two knots in three-space are S-equivalent if they are indistinguishable by Seifert matrices. We show that S-equivalence is generated by the doubled-delta move on knot diagrams. It follows as a corollary that a knot has trivial Alexander polynomial if and only if it can be undone by doubled-delta moves.