Scalable web crawling using noisy change-indicating signals.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Web application security has become a major concern in recent years, as more and more content and services are available online. A useful method for identifying security vulnerabilities is black-box testing, which relies on an automated crawling of web applications. However, crawling Rich Internet Applications (RIAs) i…
Optimizes web page freshness with limited crawling frequencies.
Study shows publicly available news impacts financial markets.
Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other. In this paper, we exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs. We mine …
New MAB model for online caching costs.
Pre-training text representations have led to significant improvements in many areas of natural language processing. The quality of these models benefits greatly from the size of the pretraining corpora as long as its quality is preserved. In this paper, we describe an automatic pipeline to extract massive high-quality…
AdarGCN tackles noisy web images in few-shot learning.
Embedded markup of Web pages has seen widespread adoption throughout the past years driven by standards such as RDFa and Microdata and initiatives such as schema.org, where recent studies show an adoption by 39% of all Web pages already in 2016. While this constitutes an important information source for tasks such as W…
Web crawling is the problem of keeping a cache of webpages fresh, i.e., having the most recent copy available when a page is requested. This problem is usually coupled with the natural restriction that the bandwidth available to the web crawler is limited. The corresponding optimization problem was solved optimally by …
Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for mining such data from previously obtained comparable corpora. The task is highly practical since non-para…
Real estate appraisal is a complex and important task, that can be made more precise and faster with the help of automated valuation tools. Usually the value of some property is determined by taking into account both structural and geographical characteristics. However, while geographical information is easily found, o…
In this paper, we show how using publicly available data streams and machine learning algorithms one can develop practical data driven services with no input from domain experts as a form of prior knowledge. We report the initial steps toward development of a real estate portal in Switzerland. Based on continuous web c…
Research in statistical relational learning has produced a number of methods for learning relational models from large-scale network data. While these methods have been successfully applied in various domains, they have been developed under the unrealistic assumption of full data access. In practice, however, the data …
Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our methodology for mining such data from previously obtained comparable corpora. The task is highly practical since non-parallel m…
With the growth of user-generated content, we observe the constant rise of the number of companies, such as search engines, content aggregators, etc., that operate with tremendous amounts of web content not being the services hosting it. Thus, aiming to locate the most important content and promote it to the users, the…
Online social networks (OSN) contain extensive amount of information about the underlying society that is yet to be explored. One of the most feasible technique to fetch information from OSN, crawling through Application Programming Interface (API) requests, poses serious concerns over the the guarantees of the estimat…
The tremendous increase in the amount of available research documents impels researchers to propose topic models to extract the latent semantic themes of a documents collection. However, how to extract the hidden topics of the documents collection has become a crucial task for many topic model applications. Moreover, c…
Multilayer switch networks are proposed as artificial generators of high-dimensional discrete data (e.g., binary vectors, categorical data, natural language, network log files, and discrete-valued time series). Unlike deconvolution networks which generate continuous-valued data and which consist of upsampling filters a…
Rogue is a famous dungeon-crawling video-game of the 80ies, the ancestor of its gender. Rogue-like games are known for the necessity to explore partially observable and always different randomly-generated labyrinths, preventing any form of level replay. As such, they serve as a very natural and challenging task for rei…
Designing a logo for a new brand is a lengthy and tedious back-and-forth process between a designer and a client. In this paper we explore to what extent machine learning can solve the creative task of the designer. For this, we build a dataset -- LLD -- of 600k+ logos crawled from the world wide web. Training Generati…
In this paper, we design an integrated algorithm to evaluate the sentiment of Chinese market. Firstly, with the help of the web browser automation, we crawl a lot of news and comments from several influential financial websites automatically. Secondly, we use techniques of Natural Language Processing(NLP) under Chinese…
The paper provides statistical theory and intuition for personalized PageRank (called "PPR"): a popular technique that samples a small community from a massive network. We study a setting where the entire network is expensive to obtain thoroughly or to maintain, but we can start from a seed node of interest and "crawl"…
This thesis evaluates text-based vs audio-based classification of mental health interviews.
Due to the limited resources and the scale of the graphs in modern datasets, we often get to observe a sampled subgraph of a larger original graph of interest, whether it is the worldwide web that has been crawled or social connections that have been surveyed. Inferring a global property of the original graph from such…
Techniques such as ensembling and distillation promise model quality improvements when paired with almost any base model. However, due to increased test-time cost (for ensembles) and increased complexity of the training pipeline (for distillation), these techniques are challenging to use in industrial settings. In this…
Formula calculates MOY webs and link polynomials.
We find an invariant characterization of planar webs of maximum rank. For 4-webs, we prove that a planar 4-web is of maximum rank three if and only if it is linearizable and its curvature vanishes. This result leads to the direct web-theoretical proof of the Poincaré's theorem: a planar 4-web of maximum rank is lineari…
Classifies hexagonal circular 3-webs with cubic polar curves.
Parallel texts are a relatively rare language resource, however, they constitute a very useful research material with a wide range of applications. This study presents and analyses new methodologies we developed for obtaining such data from previously built comparable corpora. The methodologies are automatic and unsupe…
We construct flat 3-webs via semi-simple geometric Frobenius manifolds of dimension three and give geometric interpretation of the Chern connection of the web. These webs turned out to be biholomorphic to the characteristic webs on the solutions of the corresponding associativity equation. We show that such webs are he…
We give various results and applications using the connection associated with a -web. Precisely, we exhibit fundamental invariants of the web related to the differential equation of first order which presents the web. They cast some new lights on the connection and its construction, both conceptually an…
We investigate the linearizability problem for different classes of 4-webs in the plane. In particular, we apply a recently found in [AGL] the linearizability conditions for 4-webs in the plane to confirm that a 4-web MW (Mayrhofer's web) with equal curvature forms of its 3-subwebs and a nonconstant basic invariant is …
In the present paper we study geometric structures associated with webs of hypersurfaces. We prove that with any geodesic (n+2)-web on an n-dimensional manifold there is naturally associated a unique projective structure and, provided that one of web foliations is pointed, there is also associated a unique affine struc…
Study local invariants of divergence-free webs in geometry.
This paper has been withdrawn by the authors due to the fact that the webs considered in the paper are ``Veronese-like webs'' which are different from Veronese webs.
Investigates webs related to cluster algebras and polylogarithms.
We present a projectively invariant description of planar linear 3-webs. For a non-hexagonal 3-web, we introduce family of projective torsion-free Cartan connections, the web leaves being geodesics for each member of the family, and give a web linearization criterion. Finally, we propose an algorithm for resolving the …
Much information available on the web is copied, reused or rephrased. The phenomenon that multiple web sources pick up certain information is often called trend. A central problem in the context of web data mining is to detect those web sources that are first to publish information which will give rise to a trend. We p…
For a four-dimensional (nonisoclinicly geodesic) three-web W (3, 2, 2), a transversal distribution is defined by the torsion tensor of the web. In general, this distribution is not integrable. The authors find necessary and sufficient conditions of its integrability and prove the existence theorem for webs W (3, 2,…
In this paper we study the linearizability problem for 3-webs on a 2-dimensional manifold. With an explicit computation based on the theory developed in the paper "On the linearizability of 3-webs" (Nonlinear analysis 47, (2001) pp. 2643-2654), we examine a 3-web whose linearizability was claimed in the same paper. We …
Veronese webs appear as the natural way of passing to the quotient of curves in the projective space. In thi paper, we give the link between classical multidimensionnal webs and veronse webs by mean of interpolation.
New hexagonal circular 3-webs with reducible curves classified.
In this paper, we reconstruct Kuperberg's web space. We introduce a new web (a trivalent diagram) and new relations between Kuperberg's web diagrams and the new diagram. Using the webs, we define crossing formulas corresponding to R-matrices associated to some irreducible representations and calculate…
Gronwall conjecture states that a planar 3-web which admits more than one distinct linearization is locally equivalent to an algebraic web. We give a partial answer to the conjecture in the affirmative for the class of planar 3-webs with the web curvature that vanishes to order three at a point. The differential relati…
We find relative differential invariants of orders eight and nine for a planar nonparallelizable 3-web such that their vanishing is necessary and sufficient for a 3-web to be linearizable. This solves the Blaschke conjecture for 3-webs. As a side result, we show that the number of linearizations in the Gronwall conject…
A deformation of the authors' instanton homology for webs is constructed by introducing a local system of coefficients. In the case that the web is planar, the rank of the deformed instanton homology is equal to the number of Tait colorings of the web.
Geodesic flows with specific integrals are linked to special 4-webs.