UNMIX identifies hidden buyers in darknet markets by clustering anonymized IDs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Machine learning detects phishing websites by identifying common characteristics.
A framework for anonymized risk sharing without revealing identities or preferences.
Anonymization reduces economic signal extraction from financial texts.
Recommender systems are widely used to predict personalized preferences of goods or services using users' past activities, such as item ratings or purchase histories. If collections of such personal activities were made publicly available, they could be used to personalize a diverse range of services, including targete…
Many machine intelligence techniques are developed in E-commerce and one of the most essential components is the representation of IDs, including user ID, item ID, product ID, store ID, brand ID, category ID etc. The classical encoding based methods (like one-hot encoding) are inefficient in that it suffers sparsity pr…
For a dataset of label-count pairs, an anonymized histogram is the multiset of counts. Anonymized histograms appear in various potentially sensitive contexts such as password-frequency lists, degree distribution in social networks, and estimation of symmetric properties of discrete distributions. Motivated by these app…
Calibrated ensembles improve both ID and OOD accuracy in distribution shift.
Study finds price impact follows a 'double' square-root law, suggesting mechanical origin.
IDS improves reinforcement learning with contextual information.
We study a variant of the stochastic -armed bandit problem, which we call "bandits with delayed, aggregated anonymous feedback". In this problem, when the player pulls an arm, a reward is generated, however it is not immediately observed. Instead, at the end of each round the player observes only the sum of a number…
SIGMA improves IDS robustness against new attacks using GAN and metaheuristics.
Graph matching in noisy environments with Markovian errors.
We first analyze the integrated density of states (IDS) of periodic Schrödinger operators on an amenable covering manifold. A criterion for the continuity of the IDS at a prescribed energy is given along with examples of operators with both continuous and discontinuous IDS'. Subsequently, alloy-type perturbations of th…
The social media revolution has produced a plethora of web services to which users can easily upload and share multimedia documents. Despite the popularity and convenience of such services, the sharing of such inherently personal data, including speech data, raises obvious security and privacy concerns. In particular, …
New research shows common ID estimators in neural representations are inaccurate.
New local ID estimators based on data separability.
This study analyzes VAEs using ID and II, revealing a transition in behaviour and distinct training phases.
IDS improves sparse linear bandits by balancing information and regret.
Paper investigates preserving anomalous subgroups in anonymized datasets.
IDS optimizes regret in stochastic partial monitoring with linear rewards.
ProSub uses angles in feature space to classify data as in- or out-of-distribution.
The study examines how brokers' identity affects their trading strategies on the Toronto Stock Exchange.
New bounds on IDS for RL show how to balance computation and learning efficiency.
Beta-SOD detects and corrects noisy object re-identification using cosine similarity and Beta mixtures.
Adaptive MAB algorithms handle composite, anonymous feedback without reward interval knowledge.
New IDS algorithm refines parameter norm bounds for better bandit performance.
Text-based analysis methods allow to reveal privacy relevant author attributes such as gender, age and identify of the text's author. Such methods can compromise the privacy of an anonymous author even when the author tries to remove privacy sensitive content. In this paper, we propose an automatic method, called Adver…
This work introduces a protocol to automatically select the correct range of scales for meaningful Intrinsic Dimension estimation.
Person Re-identification (re-id) faces two major challenges: the lack of cross-view paired training data and learning discriminative identity-sensitive and view-invariant features in the presence of large pose variations. In this work, we address both problems by proposing a novel deep person image generation model for…
A blindfolded LLM trading framework validates market signals without ticker memorization.
Preserving the privacy of individuals by protecting their sensitive attributes is an important consideration during microdata release. However, it is equally important to preserve the quality or utility of the data for at least some targeted workloads. We propose a novel framework for privacy preservation based on the …
Optimizes stochastic linear bandits with efficient, asymptotically optimal algorithm.
We explore a novel setting of the Multi-Armed Bandit (MAB) problem inspired from real world applications which we call bandits with "stochastic delayed composite anonymous feedback (SDCAF)". In SDCAF, the rewards on pulling arms are stochastic with respect to time but spread over a fixed number of time steps in the fut…
Paper tackles SCOD problem with optimal strategy and empirical validation.
TIPRDC anonymizes data features to protect privacy while retaining useful information.
This work synthesizes realistic data from neural excitation patterns to anonymize private data.
LoD improves model safety by integrating unlabeled wild data, reducing OOD misclassification.
The backbone of most proteins forms an open curve. To study their entanglement, a common strategy consists in searching for the presence of knots in their backbones using topological invariants. However, this approach requires to close the curve into a loop, which alters the geometry of curve. Knoto-ID allows evaluatin…
The paper proposes a privacy-preserving algorithm for decentralized learning using public-key cryptography.
High-dimensional data are ubiquitous in contemporary science and finding methods to compress them is one of the primary goals of machine learning. Given a dataset lying in a high-dimensional space (in principle hundreds to several thousands of dimensions), it is often useful to project it onto a lower-dimensional manif…
Paper tackles causal effect identification in sub-population with latent variables.
Motion sensors such as accelerometers and gyroscopes measure the instant acceleration and rotation of a device, in three dimensions. Raw data streams from motion sensors embedded in portable and wearable devices may reveal private information about users without their awareness. For example, motion data might disclose …
Network embedding has proved extremely useful in a variety of network analysis tasks such as node classification, link prediction, and network visualization. Almost all the existing network embedding methods learn to map the node IDs to their corresponding node embeddings. This design principle, however, hinders the ex…
The paper analyzes how clustering sensitive data can improve model generalization without revealing individual information.
A nonlinear differential equation of Sornette-Ide type with noise, for a complex variable, yields endogenous crashes, preceded by roughly log-periodic oscillations in the real part, and a strong increase in the imaginary part. The latter is interpreted as the trader expectation.
To reap the benefits of the Internet of Things (IoT), it is imperative to secure the system against cyber attacks in order to enable mission critical and real-time applications. To this end, intrusion detection systems (IDSs) have been widely used to detect anomalies caused by a cyber attacker in IoT systems. However, …
Analyzing large volumes of high-dimensional data is an issue of fundamental importance in data science, molecular simulations and beyond. Several approaches work on the assumption that the important content of a dataset belongs to a manifold whose Intrinsic Dimension (ID) is much lower than the crude large number of co…