SIGMA improves IDS robustness against new attacks using GAN and metaheuristics.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
IDS improves reinforcement learning with contextual information.
New IDS algorithm refines parameter norm bounds for better bandit performance.
Optimizes stochastic linear bandits with efficient, asymptotically optimal algorithm.
IDS algorithm optimizes sequential decisions in various monitoring settings.
Paper tackles causal effect identification in sub-population with latent variables.
High-dimensional data are ubiquitous in contemporary science and finding methods to compress them is one of the primary goals of machine learning. Given a dataset lying in a high-dimensional space (in principle hundreds to several thousands of dimensions), it is often useful to project it onto a lower-dimensional manif…
Many machine intelligence techniques are developed in E-commerce and one of the most essential components is the representation of IDs, including user ID, item ID, product ID, store ID, brand ID, category ID etc. The classical encoding based methods (like one-hot encoding) are inefficient in that it suffers sparsity pr…
New bounds on IDS for RL show how to balance computation and learning efficiency.
ASD algorithm maximizes model estimates by adaptively labeling points.
Proposes a new method for feature selection using Bayesian ID with intervention.
Calibrated ensembles improve both ID and OOD accuracy in distribution shift.
Rdimtools simplifies DR and IDE for high-dimensional data analysis.
We first analyze the integrated density of states (IDS) of periodic Schrödinger operators on an amenable covering manifold. A criterion for the continuity of the IDS at a prescribed energy is given along with examples of operators with both continuous and discontinuous IDS'. Subsequently, alloy-type perturbations of th…
Top-two algorithm improved for best-k-arm selection.
Random Forest outperforms other IDS algorithms in smart grids.
New research shows common ID estimators in neural representations are inaccurate.
New local ID estimators based on data separability.
New method calibrates deep models for both in-distribution and out-of-distribution samples.
One of the founding paradigms of machine learning is that a small number of variables is often sufficient to describe high-dimensional data. The minimum number of variables required is called the intrinsic dimension (ID) of the data. Contrary to common intuition, there are cases where the ID varies within the same data…
This study analyzes VAEs using ID and II, revealing a transition in behaviour and distinct training phases.
IDS improves sparse linear bandits by balancing information and regret.
IDS optimizes regret in stochastic partial monitoring with linear rewards.
ProSub uses angles in feature space to classify data as in- or out-of-distribution.
Beta-SOD detects and corrects noisy object re-identification using cosine similarity and Beta mixtures.
A critical part of multi-person multi-camera tracking is person re-identification (re-ID) algorithm, which recognizes and retains identities of all detected unknown people throughout the video stream. Many re-ID algorithms today exemplify state of the art results, but not much work has been done to explore the deployme…
New algorithm estimates intrinsic dimension of discrete datasets.
This work introduces a protocol to automatically select the correct range of scales for meaningful Intrinsic Dimension estimation.
Person Re-identification (re-id) faces two major challenges: the lack of cross-view paired training data and learning discriminative identity-sensitive and view-invariant features in the presence of large pose variations. In this work, we address both problems by proposing a novel deep person image generation model for…
Paper tackles SCOD problem with optimal strategy and empirical validation.
This paper deals with a new filter algorithm for selecting the smallest subset of features carrying all the information content of a data set (i.e. for removing redundant features). It is an advanced version of the fractal dimension reduction technique, and it relies on the recently introduced Morisita estimator of Int…
LoD improves model safety by integrating unlabeled wild data, reducing OOD misclassification.
The backbone of most proteins forms an open curve. To study their entanglement, a common strategy consists in searching for the presence of knots in their backbones using topological invariants. However, this approach requires to close the curve into a loop, which alters the geometry of curve. Knoto-ID allows evaluatin…
Program or process is an integral part of almost every IT/OT system. Can we trust the identity/ID (e.g., executable name) of the program? To avoid detection, malware may disguise itself using the ID of a legitimate program, and a system tool (e.g., PowerShell) used by the attackers may have the fake ID of another commo…
Network embedding has proved extremely useful in a variety of network analysis tasks such as node classification, link prediction, and network visualization. Almost all the existing network embedding methods learn to map the node IDs to their corresponding node embeddings. This design principle, however, hinders the ex…
A nonlinear differential equation of Sornette-Ide type with noise, for a complex variable, yields endogenous crashes, preceded by roughly log-periodic oscillations in the real part, and a strong increase in the imaginary part. The latter is interpreted as the trader expectation.
To reap the benefits of the Internet of Things (IoT), it is imperative to secure the system against cyber attacks in order to enable mission critical and real-time applications. To this end, intrusion detection systems (IDSs) have been widely used to detect anomalies caused by a cyber attacker in IoT systems. However, …
Analyzing large volumes of high-dimensional data is an issue of fundamental importance in data science, molecular simulations and beyond. Several approaches work on the assumption that the important content of a dataset belongs to a manifold whose Intrinsic Dimension (ID) is much lower than the crude large number of co…
Python package for SPD matrix distances, reproducible and extensible.
Out-of-domain (OOD) detection for low-resource text classification is a realistic but understudied task. The goal is to detect the OOD cases with limited in-domain (ID) training data, since we observe that training data is often insufficient in machine learning applications. In this work, we propose an OOD-resistant Pr…
Intelligent agents need a physical understanding of the world to predict the impact of their actions in the future. While learning-based models of the environment dynamics have contributed to significant improvements in sample efficiency compared to model-free reinforcement learning algorithms, they typically fail to g…
In this paper, we study the risk bounds for samples independently drawn from an infinitely divisible (ID) distribution. In particular, based on a martingale method, we develop two deviation inequalities for a sequence of random variables of an ID distribution with zero Gaussian component. By applying the deviation ineq…
A Bloom filter approach combined with Transformer models improves accuracy for machine learning tasks on opaque IDs.
Deep neural networks progressively transform their inputs across multiple processing layers. What are the geometrical properties of the representations learned by these networks? Here we study the intrinsic dimensionality (ID) of data-representations, i.e. the minimal number of parameters needed to describe a represent…
Study on deep learning IDS resistance against adversarial attacks.
Study reveals differences in medical image models' hidden representation refinement.
The Sornette-Ide differential equation of herding and rational trader behaviour together with very small random noise is shown to lead to crashes or bubbles where the price change goes to infinity after an unpredictable time. About 100 time steps before this singularity, a few predictable roughly log-periodic oscillati…
Efficient exploration remains a major challenge for reinforcement learning. One reason is that the variability of the returns often depends on the current state and action, and is therefore heteroscedastic. Classical exploration strategies such as upper confidence bound algorithms and Thompson sampling fail to appropri…