Transforms web content for better visibility in AI-driven search engines.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Extracts main content from web pages using neural sequence labeling.
Scalable web crawling using noisy change-indicating signals.
Malicious web content is a serious problem on the Internet today. In this paper we propose a deep learning approach to detecting malevolent web pages. While past work on web content detection has relied on syntactic parsing or on emulation of HTML and Javascript to extract features, our approach operates directly on a …
Optimization is commonly employed to determine the content of web pages, such as to maximize conversions on landing pages or click-through rates on search engine result pages. Often the layout of these pages can be decoupled into several separate decisions. For example, the composition of a landing page may involve dec…
Perceptual ad-blocking is a novel approach that detects online advertisements based on their visual content. Compared to traditional filter lists, the use of perceptual signals is believed to be less prone to an arms race with web publishers and ad networks. We demonstrate that this may not be the case. We describe att…
With rapid development of the Internet, web contents become huge. Most of the websites are publicly available, and anyone can access the contents from anywhere such as workplace, home and even schools. Nevertheless, not all the web contents are appropriate for all users, especially children. An example of these content…
With the growth of user-generated content, we observe the constant rise of the number of companies, such as search engines, content aggregators, etc., that operate with tremendous amounts of web content not being the services hosting it. Thus, aiming to locate the most important content and promote it to the users, the…
Owners of a web-site are often interested in analysis of groups of users of their site. Information on these groups can help optimizing the structure and contents of the site. In this paper we use an approach based on formal concepts for constructing taxonomies of user groups. For decreasing the huge amount of concepts…
Web application security has become a major concern in recent years, as more and more content and services are available online. A useful method for identifying security vulnerabilities is black-box testing, which relies on an automated crawling of web applications. However, crawling Rich Internet Applications (RIAs) i…
Recently, the development and implementation of phishing attacks require little technical skills and costs. This uprising has led to an ever-growing number of phishing attacks on the World Wide Web. Consequently, proactive techniques to fight phishing attacks have become extremely necessary. In this paper, we propose H…
Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. In this paper, we exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs. We mine …
Much of the data being created on the web contains interactions between users and items. Stochastic blockmodels, and other methods for community detection and clustering of bipartite graphs, can infer latent user communities and latent item clusters from this interaction data. These methods, however, typically ignore t…
Linear NDCG is used for measuring the performance of the Web content quality assessment in ECML/PKDD Discovery Challenge 2010. In this paper, we will prove that the DCG error equals a new pair-wise loss.
A vast amount of textual web streams is influenced by events or phenomena emerging in the real world. The social web forms an excellent modern paradigm, where unstructured user generated content is published on a regular basis and in most occasions is freely distributed. The present Ph.D. Thesis deals with the problem …
Boilerplate removal refers to the problem of removing noisy content from a webpage such as ads and extracting relevant content that can be used by various services. This can be useful in several features in web browsers such as ad blocking, accessibility tools such as read out loud, translation, summarization etc. In o…
In probabilistic approaches to classification and information extraction, one typically builds a statistical model of words under the assumption that future data will exhibit the same regularities as the training data. In many data sets, however, there are scope-limited features whose predictive power is only applicabl…
Generative models predict page quality without training, useful for low-resource settings.
Proposes a model to optimize feedback for content creators on social media.
We generalize to webs of any codimension results already known in codimension one. Given a holomorphic -web of codimension in an ambiant -dimensional holomorphic manifold , we define for any integer the condition for such a web to be \emph{-ordinary} resp.…
Optimizes web publisher revenues from RTB auctions.
Document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. Such aligned data can be used for a variety of NLP tasks from training cross-lingual representations to mining parallel data for machine translation. In this paper we develop an…
Greedy algorithm nearly outperforms exploration in contextual bandits.
On many social networking web sites such as Facebook and Twitter, resharing or reposting functionality allows users to share others' content with their own friends or followers. As content is reshared from user to user, large cascades of reshares can form. While a growing body of research has focused on analyzing and c…
The web contains a vast corpus of HTML tables. They can be used to provide direct answers to many web queries. We focus on answering two classes of queries with those tables: those seeking lists of entities (e.g., `cities in california') and those seeking superlative entities (e.g., `largest city in california'). The m…
Efficiently projects points onto polytopes, especially useful in web-scale applications.
In this paper we show that two seemingly unrelated problems in economics, the hypothesis of integrability and the hypothesis of additive separability are linked by the absence of curvature of connections on webs naturally associated with each problem.
Study optimizes compute usage for LLM web agents, improving performance.
Most classification methods are based on the assumption that data conforms to a stationary distribution. The machine learning domain currently suffers from a lack of classification techniques that are able to detect the occurrence of a change in the underlying data distribution. Ignoring possible changes in the underly…
A web browser should not be only for browsing web pages but also help users to find out their target websites and recommend similar type websites based on their behavior. Throughout this paper, we propose two methods to make a web browser more intelligent about link prediction which works during typing on address-bar a…
The web is loaded with textual content, and Natural Language Processing is a standout amongst the most vital fields in Machine Learning. But when data is huge simple Machine Learning algorithms are not able to handle it and that is when Deep Learning comes into play which based on Neural Networks. However since neural …
Collaborative filtering (CF) and content-based filtering (CBF) have widely been used in information filtering applications. Both approaches have their strengths and weaknesses which is why researchers have developed hybrid systems. This paper proposes a novel approach to unify CF and CBF in a probabilistic framework, n…
In the industry of video content providers such as VOD and IPTV, predicting the popularity of video contents in advance is critical not only from a marketing perspective but also from a network optimization perspective. By predicting whether the content will be successful or not in advance, the content file, which is l…
Spotify improves content mix using contextual bandits.
RB-Modulation trains free diffusion models without external adapters.
Formula calculates MOY webs and link polynomials.
The paper studies stability of generative models trained on mixed data.
We find an invariant characterization of planar webs of maximum rank. For 4-webs, we prove that a planar 4-web is of maximum rank three if and only if it is linearizable and its curvature vanishes. This result leads to the direct web-theoretical proof of the Poincaré's theorem: a planar 4-web of maximum rank is lineari…
LOLA uses LLMs to optimize content delivery, outperforming traditional methods.
Classifies hexagonal circular 3-webs with cubic polar curves.
We construct flat 3-webs via semi-simple geometric Frobenius manifolds of dimension three and give geometric interpretation of the Chern connection of the web. These webs turned out to be biholomorphic to the characteristic webs on the solutions of the corresponding associativity equation. We show that such webs are he…
We give various results and applications using the connection associated with a -web. Precisely, we exhibit fundamental invariants of the web related to the differential equation of first order which presents the web. They cast some new lights on the connection and its construction, both conceptually an…
We investigate the linearizability problem for different classes of 4-webs in the plane. In particular, we apply a recently found in [AGL] the linearizability conditions for 4-webs in the plane to confirm that a 4-web MW (Mayrhofer's web) with equal curvature forms of its 3-subwebs and a nonconstant basic invariant is …
In the present paper we study geometric structures associated with webs of hypersurfaces. We prove that with any geodesic (n+2)-web on an n-dimensional manifold there is naturally associated a unique projective structure and, provided that one of web foliations is pointed, there is also associated a unique affine struc…
Learning from multiple-relational data which contains noise, ambiguities, or duplicate entities is essential to a wide range of applications such as statistical inference based on Web Linked Data, recommender systems, computational biology, and natural language processing. These tasks usually require working with very …
Study local invariants of divergence-free webs in geometry.
This paper has been withdrawn by the authors due to the fact that the webs considered in the paper are ``Veronese-like webs'' which are different from Veronese webs.
Improves content allocation in educational platforms with sparse data.