Fast, format-agnostic web content detection for security.
problem Detecting malicious web content efficiently and accurately.
method Deep learning on static HTML tokens, avoiding complex parsing.
result 97.5% detection rate at 0.1% false positive rate.
HTMLPhish detects phishing web pages using deep learning on HTML content.
problem Growing phishing attacks on the web require new detection methods.
method HTMLPhish uses Convolutional Neural Networks (CNNs) to classify HTML documents as phishing or benign.
result HTMLPhish achieves over 93% accuracy in detecting phishing web pages.
Perceptual ad-blocking is vulnerable to attacks, creating new security risks.
problem Vulnerability of perceptual ad-blocking to attacks and new security risks.
method Analysis and creation of adversarial examples to bypass perceptual ad-blocking.
result Perceptual ad-blocking can be bypassed using adversarial examples, introducing new security risks.
Locality-sensitive hashing speeds up web app security testing.
problem Challenges in crawling Rich Internet Applications (RIAs) due to state similarity.
method Uses MinHash sketches to analyze DOM structures and estimate similarity.
result Successfully scans RIAs that traditional crawling methods cannot.
Extracts main content from web pages using neural sequence labeling.
problem Lack of generalization in existing web page content extraction models.
method Neural sequence labeling model using HTML tags and words as input.
result Model outperforms state-of-the-art and adapts to changes in web page structure.
Transforms web content for better visibility in AI-driven search engines.
problem Disruption of traditional SEO by generative AI search engines.
method Fine-tunes a BART-base transformer on synthetically generated training data.
result Significant improvements in ROUGE-L and BLEU scores, and substantial visibility gains in generative search responses.
Semi-supervised model removes noisy content from webpages.
problem Extracting relevant content from webpages with ads and noise.
method Graph representation of webpage, semi-supervised learning with Gaussian Random Fields.
result Preliminary results show successful extraction of relevant content.
Much of the data being created on the web contains interactions between users and items. Stochastic blockmodels, and other methods for community detection and clustering of bipartite graphs, can infer latent user communities and latent item clusters from this interaction data. These methods, however, typically ignore t…
A vast amount of textual web streams is influenced by events or phenomena emerging in the real world. The social web forms an excellent modern paradigm, where unstructured user generated content is published on a regular basis and in most occasions is freely distributed. The present Ph.D. Thesis deals with the problem …
With rapid development of the Internet, web contents become huge. Most of the websites are publicly available, and anyone can access the contents from anywhere such as workplace, home and even schools. Nevertheless, not all the web contents are appropriate for all users, especially children. An example of these content…
Efficient algorithm for optimizing web page layouts in real-time.
problem Optimizing web pages for conversions and click-through rates is challenging due to the large decision space and interactions between components.
method Formulated a multivariate optimization approach using bandit methodology for efficient exploration and hill-climbing for optimal selection in real-time.
result Achieved a 21% conversion increase after a single week of online optimization.
With the growth of user-generated content, we observe the constant rise of the number of companies, such as search engines, content aggregators, etc., that operate with tremendous amounts of web content not being the services hosting it. Thus, aiming to locate the most important content and promote it to the users, the…
Scalable web crawling using noisy change-indicating signals.
problem Optimizing web page freshness with limited bandwidth and noisy side information.
method Proposes a scalable crawling algorithm that uses noisy side information optimally.
result Achieves constant total rate of crawling without spikes in bandwidth usage.
Generative models predict page quality without training, useful for low-resource settings.
problem Detecting low-quality content in web articles.
method Human evaluation and analysis of 500 million web articles.
result Generative models can predict page quality without training, useful for low-resource settings.
CCAligned creates a massive web document dataset for cross-lingual research.
problem Identifying comparable documents across different languages.
method Using URL signals to label web documents and mining Common Crawl corpus.
result Release of a dataset with over 392 million URL pairs from 8144 language pairs.
Adversarial attacks can fool copyright detection systems.
problem Vulnerability of copyright detection systems to adversarial attacks.
method Used gradient methods to create adversarial music that fooled detection systems.
result Adversarial attacks can successfully deceive industrial copyright detection tools.
Most classification methods are based on the assumption that data conforms to a stationary distribution. The machine learning domain currently suffers from a lack of classification techniques that are able to detect the occurrence of a change in the underlying data distribution. Ignoring possible changes in the underly…
Much information available on the web is copied, reused or rephrased. The phenomenon that multiple web sources pick up certain information is often called trend. A central problem in the context of web data mining is to detect those web sources that are first to publish information which will give rise to a trend. We p…
New invariants detect a specific graph in spatial webs.
problem Detecting specific graphs in spatial webs.
method Introduced new invariants and used spectral sequences.
result Proved invariants detect the planar theta graph.
Linear NDCG is used for measuring the performance of the Web content quality assessment in ECML/PKDD Discovery Challenge 2010. In this paper, we will prove that the DCG error equals a new pair-wise loss.
Donut detects anomalies in web KPIs without labels.
problem Anomaly detection for seasonal KPIs with varying patterns and data quality.
method Unsupervised anomaly detection via Variational Auto-Encoder (VAE) with key techniques.
result Donut outperforms state-of-the-art approaches, achieving F-scores up to 0.9.
This paper proposes a generic classification system designed to detect security threats based on the behavior of malware samples. The system relies on statistical features computed from proxy log fields to train detectors using a database of malware samples. The behavior detectors serve as basic reusable building block…
Paper introduces fast botnet detection from web server logs.
problem Botnets cause various malicious activities; fast detection is needed.
method Inspired by PCA, uses Lanczos method to improve detection speed.
result Lanczos method significantly reduces time complexity for botnet detection.
Owners of a web-site are often interested in analysis of groups of users of their site. Information on these groups can help optimizing the structure and contents of the site. In this paper we use an approach based on formal concepts for constructing taxonomies of user groups. For decreasing the huge amount of concepts…
Machine learning detects survey validity from user behavior.
problem Detecting valid responses in web surveys.
method Uses mouse activity and machine learning models (LSTM, HMM).
result Predicts survey validity without analyzing specific answers.
Detects anomalous behavior in social media users by analyzing content and connections.
problem Identifying disruptive patterns in user behavior on social media platforms.
method Joint representation learning of content and connection to detect anomalous behavior.
result Observed densely connected users engaging in local politics and exhibiting troll-like behavior.
DEMUD-VIS detects novel image content and explains it visually.
problem Detecting and explaining novel image content in large datasets.
method Uses CNN for feature extraction, reconstruction error for novelty detection, and up-convolutional networks for image reconstruction.
result Demonstrates visual explanations of novel image content on diverse datasets.
Automated sentiment analysis and opinion mining is a complex process concerning the extraction of useful subjective information from text. The explosion of user generated content on the Web, especially the fact that millions of users, on a daily basis, express their opinions on products and services to blogs, wikis, so…
Community detection is a fundamental task in social network analysis. In this paper, first we develop an endorsement filtered user connectivity network by utilizing Heider's structural balance theory and certain Twitter triad patterns. Next, we develop three Nonnegative Matrix Factorization frameworks to investigate th…
Deep learning detects radical content on social media.
problem Detecting extremist content on social media platforms.
method Employed an LSTM based feed forward neural network to classify radical content.
result Achieved a precision of 85.9% in detecting radical content.
Detects radical content on Twitter using textual, psychological, and behavioral signals.
problem Limit the spread of extremist narratives on social media.
method Analyzed extremist material, created contextual text-based model, inferred psychological properties, evaluated on Twitter.
result Radical users exhibit distinguishable textual, psychological, and behavioral properties.
Geometric deep learning detects fake news on social media.
problem Detecting fake news on social media due to lack of context understanding.
method Propagation-based geometric deep learning model.
result Highly accurate fake news detection (92.7% ROC AUC).
A new active learning method for skewed data sets.
problem Severe class imbalance and small initial training data in sequential active learning.
method HAL: a hybrid active learning algorithm balancing labeled and unlabeled data.
result HAL makes better choices for labeling points than strong baselines.
TableQnA answers web queries about lists and superlatives from HTML tables.
problem Answer web queries about lists and superlatives from HTML tables.
method Extract intent from queries, use structure-aware matching, and train models with automatic data generation.
result Significantly higher precision and coverage for list and superlative queries.
In probabilistic approaches to classification and information extraction, one typically builds a statistical model of words under the assumption that future data will exhibit the same regularities as the training data. In many data sets, however, there are scope-limited features whose predictive power is only applicabl…
The paper generalizes abelian relations for webs of arbitrary codimension.
problem Understanding abelian relations in webs of arbitrary codimension.
method Defining p-ordinary and strongly p-ordinary webs, calculating ranks and defining tautological connections.
result For p-ordinary webs, the p-rank is finite and can be bounded.
Paper proposes detecting video manipulation using stream descriptors.
problem Misuse of manipulated video content.
method Binary classifiers on multimedia stream descriptors.
result Scalable approach can detect high-quality manipulations.
Detects fake news with fewer labels using tensor embeddings.
problem Detecting fake news with limited labeled data.
method Represent news articles as tensors, derive embeddings, create graph, propagate labels.
result 75.43% accuracy with 30% labels on a public dataset.
SLUG method detects bias and out-of-distribution content in generative models.
problem Generative models can underrepresent certain groups and fail on out-of-distribution data.
method SLUG: A new uncertainty quantification method for VAEs combining Laplace approximations and stochastic trace estimators.
result SLUG's UQ score correlates with bias and out-of-distribution content.
This study benchmarks deep learning for unsupervised near-duplicate image detection.
problem Detecting near-duplicates in large image datasets with high specificity.
method Binary classification using Receiver Operating Curve (ROC) for comparison of different descriptors.
result Fine-tuning deep convolutional networks generally outperforms off-the-shelf features, with best performance on MFND dataset.
YOLOv3 detects ships in real-time with high accuracy.
problem Real-time target detection in maritime scenarios.
method YOLOv3 model trained on a large dataset of marine vessels.
result Average Precision up to 96% for IoU of 0.5.
New method detects unusual images in large datasets.
problem Detecting new images in large image data sets.
method Combines novelty detection with CNN image features.
result Rapid discovery with interpretable explanations.
Unified machine learning framework for deep learning and web services.
problem Deep learning and web service challenges.
method MMLSpark expands Spark to deep learning, micro-service orchestration, etc.
result Unified API for deep learning and web services.
Chimera model combines link, content, and time for dynamic network analysis.
problem Community detection and prediction in evolving networks with dynamic changes.
method Shared factorization model that accounts for graph links, content, and temporal analysis.
result The approach simplifies temporal analysis and enables future community prediction.
Study uses attention-based method to detect different types of online harassment.
problem Detecting different types of online harassment in social media content.
method Multi-attention based approach using Recurrent Neural Networks to address imbalanced data.
result Demonstrates effectiveness of attention-based mechanism for detecting various types of online harassment.
Paper ranks influential Tor Darknet onion domains using content features.
problem Measuring influence of criminal onion domains in Tor Darknet.
method Content-based features from multiple sources, Learning-to-Rank approach.
result Listwise approach outperforms other methods with NDCG of 0.95 for top-10.
For the pedestrian observer, financial markets look completely random with erratic and uncontrollable behavior. To a large extend, this is correct. At first approximation the difference between real price changes and the random walk model is too small to be detected using traditional time series analysis. However, we s…
An unsupervised method clusters patient incident reports for content analysis.
problem Lack of methods to extract interpretable content from electronic healthcare records.
method Combines text-embedding with paragraph vectors and graph-theoretical multiscale community detection.
result Extracts high-intrinsic-consistency groups of patient incident reports.