TailedTS dataset benchmarks heavy-tailed time series forecasting and periodicity quantification.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The production and consumption of information about Bitcoin and other digital-, or 'crypto'-, currencies have grown together with their market capitalisation. However, a systematic investigation of the relationship between online attention and market dynamics, across multiple digital currencies, is still lacking. Here,…
We present a new, efficient method for automatically detecting severe conflicts `edit wars' in Wikipedia and evaluate this method on six different language WPs. We discuss how the number of edits, reverts, the length of discussions, the burstiness of edits and reverts deviate in such pages from those following the gene…
Wikipedia is a huge opportunity for machine learning, being the largest semi-structured base of knowledge available. Because of this, many works examine its contents, and focus on structuring it in order to make it usable in learning tasks, for example by classifying it into an ontology. Beyond its textual contents, Wi…
Singular Value Decomposition (SVD) has been used successfully in recent years in the area of recommender systems. In this paper we present how this model can be extended to consider both user ratings and information from Wikipedia. By mapping items to Wikipedia pages and quantifying their similarity, we are able to use…
Study uses multiple online media to predict crude oil prices.
The importance of nodes in a network constantly fluctuates based on changes in the network structure as well as changes in external interest. We propose an evolving teleportation adaptation of the PageRank method to capture how changes in external interest influence the importance of a node. This framework seamlessly g…
Paper predicts market implied volatility using alternative data and machine learning.
We present a study on predicting the factuality of reporting and bias of news media. While previous work has focused on studying the veracity of claims or documents, here we are interested in characterizing entire news media. These are under-studied but arguably important research problems, both in their own right and …
Extracts main content from web pages using neural sequence labeling.
The paper combines Bitcoin price models with expert corrections for better predictions.
Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for mining such data from previously obtained comparable corpora. The task is highly practical since non-para…
Wiki-CS dataset benchmarks Graph Neural Networks using Wikipedia articles.
Reprint of a 1989 paper including minor corrections of misprints. Added comments (11 pages) about later related papers in the literature concerning comparison of Gabriel-Zisman calculus of (right) fractions and the use of generalized morphims in the sense of Haefliger-Skandalis-Hilsum for inverting differentiable equiv…
Study proposes learning optimal priors from data for better Bayesian inference.
Study confirms mispricing in sportsbooks but finds data issues affect results.
Alternative proof of Dynnikov's three-page index for torus links.
Recommendation algorithms are widely adopted in marketplaces to help users find the items they are looking for. The sparsity of the items by user matrix and the cold-start issue in marketplaces pose challenges for the off-the-shelf matrix factorization based recommender systems. To understand user intent and tailor rec…
In many machine learning applications, one needs to interactively select a sequence of items (e.g., recommending movies based on a user's feedback) or make sequential decisions in a certain order (e.g., guiding an agent through a series of states). Not only do sequences already pose a dauntingly large search space, but…
PAGE optimizes nonconvex problems with optimal convergence rates.
We present an LDA approach to entity disambiguation. Each topic is associated with a Wikipedia article and topics generate either content words or entity mentions. Training such models is challenging because of the topic and vocabulary size, both in the millions. We tackle these problems using a novel distributed infer…
As entity type systems become richer and more fine-grained, we expect the number of types assigned to a given entity to increase. However, most fine-grained typing work has focused on datasets that exhibit a low degree of type multiplicity. In this paper, we consider the high-multiplicity regime inherent in data source…
Grammatical Error Correction (GEC) has been recently modeled using the sequence-to-sequence framework. However, unlike sequence transduction problems such as machine translation, GEC suffers from the lack of plentiful parallel data. We describe two approaches for generating large parallel datasets for GEC using publicl…
The Wikimedia Foundation has recently observed that newly joining editors on Wikipedia are increasingly failing to integrate into the Wikipedia editors' community, i.e. the community is becoming increasingly harder to penetrate. To sustain healthy growth of the community, the Wikimedia Foundation aims to quantitatively…
We draw a formal connection between using synthetic training data to optimize neural network parameters and approximate, Bayesian, model-based reasoning. In particular, training a neural network using synthetic data can be viewed as learning a proposal distribution generator for approximate inference in the synthetic-d…
We analyze the influence and interactions of 60 largest world banks for 195 world countries using the reduced Google matrix algorithm for the English Wikipedia network with 5 416 537 articles. While the top asset rank positions are taken by the banks of China, with China Industrial and Commercial Bank of China at the f…
The paper refines the three-page index for links, proving a new bound and characterizing specific links.
Study the spectrum of Page's metric on complex projective spaces.
We describe an approach to Grammatical Error Correction (GEC) that is effective at making use of models trained on large amounts of weakly supervised bitext. We train the Transformer sequence-to-sequence model on 4B tokens of Wikipedia revisions and employ an iterative decoding strategy that is tailored to the loosely-…
Optimizes web page freshness with limited crawling frequencies.
We construct a series of finitely presented semigroups. The centers of these semigroups encode uniquely up to rigid ambient isotopy in 3-space all non-oriented spatial graphs. This encoding is obtained by using three-page embeddings of graphs into the product of the line with the cone on three points. By exploiting thr…
This research shows how to learn shared representations from unpaired data.
Extends deformation theory to higher-page analogues of manifolds.
Improved linear upper bound for ribbonlength of knots.
PAGE is a simple gradient estimator for nonconvex optimization problems.
A new scheme for FBSDEs simplifies computation without Monte Carlo.
New invariant distinguishes Legendrian surfaces in 5-manifolds.
The spectral sequence's -page is a link invariant for .
We construct a contact 5-manifold supported by infinitely many distinct open books with the identity monodromy and pairwise exotic Stein pages (i.e. pages are pairwise homeomorphic but non-diffeomorphic Stein fillings of a fixed contact 3-manifold), moreover we describe a process of generating infinitely many such exam…
Optimization is commonly employed to determine the content of web pages, such as to maximize conversions on landing pages or click-through rates on search engine result pages. Often the layout of these pages can be decoupled into several separate decisions. For example, the composition of a landing page may involve dec…
In this paper, we study some classes of submanifolds of codimension one and two in the Page space. These submanifolds are totally geodesic. We also compute their curvature and show that some of them are constant curvature spaces. Finally we give information on how the Page space is related to some other metrics on the …
Generative models predict page quality without training, useful for low-resource settings.
Model predicts web page parallelism for improved browser performance and energy.
Online knowledge repositories typically rely on their users or dedicated editors to evaluate the reliability of their content. These evaluations can be viewed as noisy measurements of both information reliability and information source trustworthiness. Can we leverage these noisy evaluations, often biased, to distill a…
The paper extends Hawking--Page solutions to various spacetimes with singularities.
Scalable web crawling using noisy change-indicating signals.
Freya PAGE optimizes nonconvex optimization with heterogeneous, asynchronous workers.
We construct, somewhat non-standard, Legendrian surgery diagrams for some Stein fillable contact structures on some plumbing trees of circle bundles over spheres. We then show how to put such a surgery diagram on the pages of an open book for with relatively low genus. Thus we produce open books with low genus p…