Visual spoofing bypasses spam filters and plagiarism detection.
problem Vulnerability in spam filters that can be exploited by visually similar but differently encoded characters.
method Replaces characters with visually similar but differently encoded characters from a different alphabet.
result Spammers can create messages that bypass existing spam filters.
Defense against spam filter attacks using mixture models.
problem Data poisoning attacks on naive Bayes spam filters.
method Mixture of naive Bayes models to isolate attacks.
result Mixture model isolates attacks in a second component, preserving original spam.
Novel spam filter improves e-mail classification accuracy.
problem Uneven class distribution, unequal error cost, frequent content change, personalized discrimination.
method TFDCR feature selection, incremental learning, dynamic feature update.
result TFDCR outperforms in feature selection, incremental model improves classification accuracy.
DeepCapture detects image spam emails using CNN and data augmentation.
problem Detecting new image spam emails with overfitting.
method CNN-XGBoost framework with data augmentation.
result DeepCapture achieves an F1-score of 88%.
DeepQuarantine detects and quarantines suspicious emails.
problem High-quality spam detection and prevention.
method Convolutional Neural Networks on MIME headers for deep feature extraction.
result DQ enhances spam detection with high precision.
Paper tackles spam filtering on forums using synthetic oversampling.
problem Imbalanced data in forums leads to poor spam detection.
method Synthetic Minority Over-sampling Technique (SMOTE) to balance data.
result Models trained with SMOTE outperform those trained on imbalanced data.
Machine learning models have been widely used in security applications such as intrusion detection, spam filtering, and virus or malware detection. However, it is well-known that adversaries are always trying to adapt their attacks to evade detection. For example, an email spammer may guess what features spam detection…
In this paper we explore the "vector semantics" problem from the perspective of "almost orthogonal" property of high-dimensional random vectors. We show that this intriguing property can be used to "memorize" random vectors by simply adding them, and we provide an efficient probabilistic solution to the set membership …
This paper tackles spam detection on Twitter by analyzing correlated features.
problem Spam detection on social media, especially Twitter, to improve user experience.
method Extracted tweet-based and user-based features, identified correlated features, and used artificial neural networks for classification.
result Achieved 97.57% accuracy in classifying tweets as spam or non-spam.
Paper proposes spamGAN to detect and generate opinion spam using limited labeled data.
problem Detecting and preventing opinion spam in online reviews with limited labeled data.
method Generative adversarial network (GAN) trained on semi-supervised data.
result spamGAN outperforms existing techniques in detecting opinion spam with limited labeled data.
New method detects evasive product reviews that evade traditional spam detection.
problem Hardened spammers evade behavior-based opinion spam detection systems.
method Proposes EMERAL to generate evasive spams and DETER to defend against them.
result DETER outperforms traditional methods in detecting evasive reviews.
Paper quarantines unreliable Yelp users by detecting review spam.
problem Unreliable and spamming users deceive Yelp's users.
method Used RSD and spam detection techniques on key features.
result More than 80% of Yelp's accounts are unreliable, and highly-rated businesses are often spammed.
To date, most studies on spam have focused only on the spamming phase of the spam cycle and have ignored the harvesting phase, which consists of the mass acquisition of email addresses. It has been observed that spammers conceal their identity to a lesser degree in the harvesting phase, so it may be possible to gain ne…
Machine learning has become an important component for many systems and applications including computer vision, spam filtering, malware and network intrusion detection, among others. Despite the capabilities of machine learning algorithms to extract valuable information from data and produce accurate predictions, it ha…
Recently, machine learning algorithms have successfully entered large-scale real-world industrial applications (e.g. search engines and email spam filters). Here, the CPU cost during test time must be budgeted and accounted for. In this paper, we address the challenge of balancing the test-time cost and the classifier …
Enhances social spam detection using multi-level dependency of relational sequences.
problem Social spam detection in multi-relation social networks.
method Developed the Multi-level Dependency Model (MDM) to exploit long-term and short-term dependencies in user relational sequences.
result MDM improves social spam detection accuracy on a real-world multi-relational social network.
New model improves classifier security against evasion attacks by selecting features.
problem Security of machine learning classifiers against evasion attacks.
method Adversary-aware feature selection model incorporating specific adversary assumptions.
result Classifier security worsens with feature selection, but new model improves it.
Discriminative classifier for compositional data using hierarchical mixture of Generalized Dirichlet models.
problem Classifying compositional data, especially in spam detection and color space identification.
method Hierarchical mixture of discriminative Generalized Dirichlet classifiers, using variational approximation for parameter learning.
result First time a variational upper-bound for Generalized Dirichlet mixture is proposed in literature.
We propose a nonparametric method for detecting nonlinear causal relationship within a set of multidimensional discrete time series, by using sparse additive models (SpAMs). We show that, when the input to the SpAM is a β-mixing time series, the model can be fitted by first approximating each unknown function with a …
AutoML tackles evolving data with drift detection.
problem Handling changing data distributions in AutoML.
method Extended Auto-Sklearn with drift detection mechanisms.
result Demonstrated effectiveness of the proposed methodology.
Developed text classification system for Azerbaijani language.
problem Text clustering problem in Azerbaijani language.
method Machine learning and embedding techniques.
result System successfully categorizes news, product reviews, and more.
In high dimensions, most machine learning methods are brittle to even a small fraction of structured outliers. To address this, we introduce a new meta-algorithm that can take in a base learner such as least squares or stochastic gradient descent, and harden the learner to be resistant to outliers. Our method, Sever, p…
A new distributed algorithm for fitting sparse additive models with feature division and decorrelation.
problem Fitting high-dimensional sparse additive models efficiently and accurately.
method Divide, decorrelate, and conquer approach.
result Effective and efficient recovery of sparsity patterns and statistical inference for each component.
New algorithm improves convergence of AUC maximization.
problem Optimizing AUC for imbalanced classes with stochastic methods.
method Variance Reduced Stochastic Proximal Algorithm for AUC Maximization (VRSPAM).
result VRSPAM converges faster than previous methods.
Consider a multi-variate time series (Xt)t=0T where Xt∈Rd which may represent spike train responses for multiple neurons in a brain, crime event data across multiple regions, and many others. An important challenge associated with these time series models is to estimate an influence network be…
New attacks bypass data sanitization defenses, increasing model errors.
problem Data poisoning attacks corrupt machine learning models trained on external data.
method Developed three attacks that coordinate poisoned points and formulate as optimization problems.
result 3% poisoned data increases test error from 3% to 24% on Enron spam detection.
A great variety of text tasks such as topic or spam identification, user profiling, and sentiment analysis can be posed as a supervised learning problem and tackle using a text classifier. A text classifier consists of several subprocesses, some of them are general enough to be applied to any supervised learning proble…
A function f:Rd→R is referred to as a Sparse Additive Model (SPAM), if it is of the form f(x)=∑l∈Sφl(xl), where S⊂[d], ∣S∣≪d. Assuming φl's and S to be unknown, the problem of estimating f from it…
Adversarial attacks against neural networks are a problem of considerable importance, for which effective defenses are not yet readily available. We make progress toward this problem by showing that non-negative weight constraints can be used to improve resistance in specific scenarios. In particular, we show that they…
Study evaluates early-stage cybersecurity firms' performance using Crunchbase data.
problem Assessing performance of early-stage cybersecurity startups.
method Empirical analysis of 19 cybersecurity sectors using Crunchbase data.
result Significant variations in capital raised and post-money valuations across cybersecurity sectors.
Machine learning improves cybersecurity by learning from data.
problem Designing effective detection algorithms for cyber threats.
method Machine learning algorithms to learn from security data.
result ML algorithms can improve threat hunting and remediation.
We introduce a new algorithm, called adaptive sparse backfitting algorithm, for solving high dimensional Sparse Additive Model (SpAM) utilizing symmetric, non-negative definite smoothers. Unlike the previous sparse backfitting algorithm, our method is essentially a block coordinate descent algorithm that guarantees to …
New methods needed to estimate individual treatment effects.
problem Estimating treatment effects on individual level.
method Machine learning for subgroup discovery under treatment effect.
result Efficient methods are needed for estimating individual treatment effects.
This paper addresses the problem of inferring a regular expression from a given set of strings that resembles, as closely as possible, the regular expression that a human expert would have written to identify the language. This is motivated by our goal of automating the task of postmasters of an email service who use r…
A function f:Rd→R is a Sparse Additive Model (SPAM), if it is of the form f(x)=∑l∈Sφl(xl) where S⊂[d], ∣S∣≪d. Assuming φ's, S to be unknown, there exists extensive work for estimating f from its sa…
Paper proves spectral filters can be transferred between graphs.
problem Proving spectral filters can be transferred between graphs.
method Introducing the Cayley smoothness space and proving filters in this space are linearly stable.
result Graph spectral filters are transferable if they are in the Cayley smoothness space.
This work prunes CNN filters based on their functionality, not just size.
problem Redundant filters in CNNs waste computation resources.
method Functionality-oriented filter pruning method.
result Pruning based on functionality optimizes computation and interprets filter importance.
A new SOHP filter improves trend estimation in economic time series.
problem Improving trend estimation in nonlinear economic time series.
method Recursive application of one-sided HP filter on updated cyclical components, combined with an incremental HP filtering algorithm.
result Better performance of SOHP filter compared to other HP-type filters on real economic data.
Deep density methods improve filtering in high-dimensional systems.
problem Nonlinear filtering in high-dimensional systems.
method Two deep density methods based on Feynman-Kac formulas and neural networks.
result Logarithmic deep backward stochastic differential equation filter outperforms classical methods in high dimensions.
Sparse coding has shown its power as an effective data representation method. However, up to now, all the sparse coding approaches are limited within the single domain learning problem. In this paper, we extend the sparse coding to cross domain learning problem, which tries to learn from a source domain to a target dom…
The paper explores modifications to filter banks for speech recognition.
problem Improving speech recognition accuracy using modified filter banks.
method The authors investigate replacing triangular filters with Gabor or Gammatone filters, and rearranging filter bank computations to integrate features over smaller time scales.
result No significant improvements in phone error rate were observed with the modifications.
Gradient filters track moving parameters under noisy data and misspecification.
problem Tracking multidimensional time-varying parameters under noisy observations and model misspecification.
method Gradient-based filters update parameters using the gradient of a postulated objective function, evaluated at either the predicted or updated parameters.
result Novel sufficient conditions for exponential stability of the filtered parameter path, and finite-sample and asymptotic mean squared error bounds.
Survey of methods for detecting fraud in networks.
problem Detecting anomalies in graph-based data for fraud detection.
method Anomaly detection techniques using graph structure and attributes.
result Survey of various methods for fraud detection in networks.
We simplify Bayesian filtering by framing it as optimization, making it practical for high-dimensional systems.
problem Bayesian filtering struggles in high-dimensional state spaces like neural networks.
method We frame Bayesian filtering as optimization, using gradient descent for nonlinear cases.
result Our method results in effective, robust, and scalable filters for high-dimensional systems.
A new model optimizes Bloom filters using machine learning.
problem Improving the efficiency of Bloom filters for data sets.
method Modeling learned Bloom filters with machine learning, optimizing with sandwiching method.
result Optimized learned Bloom filters provide improved performance.
Develops an inverse particle filter for cognitive systems.
problem Tracking cognitive adversaries in counter-adversarial applications.
method Global filtering approach using Monte Carlo methods and differentiable I-PF.
result Demonstrates convergence to optimal inverse filter and improved estimation performance.
Convolutional neural networks (CNNs) achieve state-of-the-art performance in a wide variety of tasks in computer vision. However, interpreting CNNs still remains a challenge. This is mainly due to the large number of parameters in these networks. Here, we investigate the role of compression and particularly pruning fil…
A novel method reduces dimensionality for filtering SRNs with observed variables.
problem Challenges in estimating hidden state variables in SRNs with limited observations.
method Filtered Markovian Projection (Filtered MP) for dimensionality reduction in filtering.
result Filtered MP guarantees consistency and superior computational efficiency in high dimensions.