A Bloom filter approach combined with Transformer models improves accuracy for machine learning tasks on opaque IDs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
TaylorPODA uses Taylor expansions to improve feature attributions for opaque models.
Many machine intelligence techniques are developed in E-commerce and one of the most essential components is the representation of IDs, including user ID, item ID, product ID, store ID, brand ID, category ID etc. The classical encoding based methods (like one-hot encoding) are inefficient in that it suffers sparsity pr…
Calibrated ensembles improve both ID and OOD accuracy in distribution shift.
LIMEADE improves AI advice for opaque models, enhancing accuracy and user satisfaction.
IDS improves reinforcement learning with contextual information.
We first analyze the integrated density of states (IDS) of periodic Schrödinger operators on an amenable covering manifold. A criterion for the continuity of the IDS at a prescribed energy is given along with examples of operators with both continuous and discontinuous IDS'. Subsequently, alloy-type perturbations of th…
An Intrusion Detection System (IDS) is a key cybersecurity tool for network administrators as it identifies malicious traffic and cyberattacks. With the recent successes of machine learning techniques such as deep learning, more and more IDS are now using machine learning algorithms to detect attacks faster. However, t…
New research shows common ID estimators in neural representations are inaccurate.
Intrinsic dimensionality (ID) is one of the most fundamental characteristics of multi-dimensional data point clouds. Knowing ID is crucial to choose the appropriate machine learning approach as well as to understand its behavior and validate it. ID can be computed globally for the whole data point distribution, or comp…
This study analyzes VAEs using ID and II, revealing a transition in behaviour and distinct training phases.
IDS improves sparse linear bandits by balancing information and regret.
IDS optimizes regret in stochastic partial monitoring with linear rewards.
ProSub uses angles in feature space to classify data as in- or out-of-distribution.
Study models opaque financial markets using multi-agent simulation.
New bounds on IDS for RL show how to balance computation and learning efficiency.
Beta-SOD detects and corrects noisy object re-identification using cosine similarity and Beta mixtures.
New IDS algorithm refines parameter norm bounds for better bandit performance.
This work introduces a protocol to automatically select the correct range of scales for meaningful Intrinsic Dimension estimation.
Person Re-identification (re-id) faces two major challenges: the lack of cross-view paired training data and learning discriminative identity-sensitive and view-invariant features in the presence of large pose variations. In this work, we address both problems by proposing a novel deep person image generation model for…
Optimizes stochastic linear bandits with efficient, asymptotically optimal algorithm.
Paper tackles SCOD problem with optimal strategy and empirical validation.
LoD improves model safety by integrating unlabeled wild data, reducing OOD misclassification.
The backbone of most proteins forms an open curve. To study their entanglement, a common strategy consists in searching for the presence of knots in their backbones using topological invariants. However, this approach requires to close the curve into a loop, which alters the geometry of curve. Knoto-ID allows evaluatin…
High-dimensional data are ubiquitous in contemporary science and finding methods to compress them is one of the primary goals of machine learning. Given a dataset lying in a high-dimensional space (in principle hundreds to several thousands of dimensions), it is often useful to project it onto a lower-dimensional manif…
New method refines model-free evaluation of complex machine learning models.
Paper tackles causal effect identification in sub-population with latent variables.
Network embedding has proved extremely useful in a variety of network analysis tasks such as node classification, link prediction, and network visualization. Almost all the existing network embedding methods learn to map the node IDs to their corresponding node embeddings. This design principle, however, hinders the ex…
A nonlinear differential equation of Sornette-Ide type with noise, for a complex variable, yields endogenous crashes, preceded by roughly log-periodic oscillations in the real part, and a strong increase in the imaginary part. The latter is interpreted as the trader expectation.
To reap the benefits of the Internet of Things (IoT), it is imperative to secure the system against cyber attacks in order to enable mission critical and real-time applications. To this end, intrusion detection systems (IDSs) have been widely used to detect anomalies caused by a cyber attacker in IoT systems. However, …
Analyzing large volumes of high-dimensional data is an issue of fundamental importance in data science, molecular simulations and beyond. Several approaches work on the assumption that the important content of a dataset belongs to a manifold whose Intrinsic Dimension (ID) is much lower than the crude large number of co…
Python package for SPD matrix distances, reproducible and extensible.
Out-of-domain (OOD) detection for low-resource text classification is a realistic but understudied task. The goal is to detect the OOD cases with limited in-domain (ID) training data, since we observe that training data is often insufficient in machine learning applications. In this work, we propose an OOD-resistant Pr…
In this paper, we study the risk bounds for samples independently drawn from an infinitely divisible (ID) distribution. In particular, based on a martingale method, we develop two deviation inequalities for a sequence of random variables of an ID distribution with zero Gaussian component. By applying the deviation ineq…
Deep neural networks progressively transform their inputs across multiple processing layers. What are the geometrical properties of the representations learned by these networks? Here we study the intrinsic dimensionality (ID) of data-representations, i.e. the minimal number of parameters needed to describe a represent…
The Sornette-Ide differential equation of herding and rational trader behaviour together with very small random noise is shown to lead to crashes or bubbles where the price change goes to infinity after an unpredictable time. About 100 time steps before this singularity, a few predictable roughly log-periodic oscillati…
Study reveals differences in medical image models' hidden representation refinement.
Rdimtools simplifies DR and IDE for high-dimensional data analysis.
New method calibrates deep models for both in-distribution and out-of-distribution samples.
Proposes a new method for feature selection using Bayesian ID with intervention.
IDS algorithm optimizes sequential decisions in various monitoring settings.
When training a deep neural network for image classification, one can broadly distinguish between two types of latent features of images that will drive the classification. We can divide latent features into (i) "core" or "conditionally invariant" features whose distribution , cond…
P-OCS detects OOD samples in a low-dimensional subspace, outperforming existing methods.
One of the founding paradigms of machine learning is that a small number of variables is often sufficient to describe high-dimensional data. The minimum number of variables required is called the intrinsic dimension (ID) of the data. Contrary to common intuition, there are cases where the ID varies within the same data…
Click-through rate (CTR) prediction has been one of the most central problems in computational advertising. Lately, embedding techniques that produce low-dimensional representations of ad IDs drastically improve CTR prediction accuracies. However, such learning techniques are data demanding and work poorly on new ads w…
Python package for estimating intrinsic dimensionality of datasets.
Rigidity theorem shows massless hyperboloidal data embeds into Minkowski space.
We consider a -dimensional Riemannian manifold equip\-ped with a circulant structure , which is an isometry with respect to the metric and $q^{4}=\id$, $q^{2}\neq \pm \id$. For such a manifold we obtain some assertions for the sectional curvatures of -planes. We construct an example of such…