Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

275380106 · Jun 202019922001200920172026
48 results for Wikipedia page views

TailedTS dataset benchmarks heavy-tailed time series forecasting and periodicity quantification.

problem Benchmarking robustness of time series models under heavy-tailed distributions.
method Derived from Wikipedia page views, introduces periodicity quantification and robust loss functions.
result Standard Gaussian models degrade on high-volume page categories, while robust alternatives perform consistently.

We present a new, efficient method for automatically detecting severe conflicts `edit wars' in Wikipedia and evaluate this method on six different language WPs. We discuss how the number of edits, reverts, the length of discussions, the burstiness of edits and reverts deviate in such pages from those following the gene…

2011-07-19abs ↗pdf ↗

Singular Value Decomposition (SVD) has been used successfully in recent years in the area of recommender systems. In this paper we present how this model can be extended to consider both user ratings and information from Wikipedia. By mapping items to Wikipedia pages and quantifying their similarity, we are able to use…

2012-12-05abs ↗pdf ↗

The importance of nodes in a network constantly fluctuates based on changes in the network structure as well as changes in external interest. We propose an evolving teleportation adaptation of the PageRank method to capture how changes in external interest influence the importance of a node. This framework seamlessly g…

2012-03-27abs ↗pdf ↗

Paper predicts market implied volatility using alternative data and machine learning.

problem Predicting market implied volatility using alternative data.
method Used Google News statistics and Wikipedia site traffic as alternative data sources, and applied Logistic Regression, Support Vector Machines, and AdaBoost as machine learning models.
result Movements in market implied volatility can be predicted using machine learning techniques.

We present a study on predicting the factuality of reporting and bias of news media. While previous work has focused on studying the veracity of claims or documents, here we are interested in characterizing entire news media. These are under-studied but arguably important research problems, both in their own right and …

2018-10-02abs ↗pdf ↗

The paper combines Bitcoin price models with expert corrections for better predictions.

problem Improving Bitcoin price predictions using statistical and expert insights.
method Linear regression models combined with expert corrections, utilizing Bayesian approach for fat-tailed distributions.
result Better price prediction results compared to using either model or expert opinion alone.

Wiki-CS dataset benchmarks Graph Neural Networks using Wikipedia articles.

problem Benchmarking Graph Neural Networks on a new domain with structural differences.
method Derived from Wikipedia, nodes represent Computer Science articles, edges from hyperlinks, 10 classes for different branches, evaluated semi-supervised node classification and link prediction.
result Graph Neural Networks perform well on Wiki-CS, showing structural differences from earlier benchmarks.

Reprint of a 1989 paper including minor corrections of misprints. Added comments (11 pages) about later related papers in the literature concerning comparison of Gabriel-Zisman calculus of (right) fractions and the use of generalized morphims in the sense of Haefliger-Skandalis-Hilsum for inverting differentiable equiv…

2008-03-28abs ↗pdf ↗

Study proposes learning optimal priors from data for better Bayesian inference.

problem Challenges the use of noninformative uniform priors in Bayesian inference.
method Machine learning approach to learn optimal priors from data using a target function.
result Study models consistently outperformed baseline models in Wikipedia category classification.

Recommendation algorithms are widely adopted in marketplaces to help users find the items they are looking for. The sparsity of the items by user matrix and the cold-start issue in marketplaces pose challenges for the off-the-shelf matrix factorization based recommender systems. To understand user intent and tailor rec…

2018-09-06abs ↗pdf ↗

In many machine learning applications, one needs to interactively select a sequence of items (e.g., recommending movies based on a user's feedback) or make sequential decisions in a certain order (e.g., guiding an agent through a series of states). Not only do sequences already pose a dauntingly large search space, but…

2019-02-15abs ↗pdf ↗

We present an LDA approach to entity disambiguation. Each topic is associated with a Wikipedia article and topics generate either content words or entity mentions. Training such models is challenging because of the topic and vocabulary size, both in the millions. We tackle these problems using a novel distributed infer…

2013-09-02abs ↗pdf ↗

As entity type systems become richer and more fine-grained, we expect the number of types assigned to a given entity to increase. However, most fine-grained typing work has focused on datasets that exhibit a low degree of type multiplicity. In this paper, we consider the high-multiplicity regime inherent in data source…

2017-04-25abs ↗pdf ↗

Grammatical Error Correction (GEC) has been recently modeled using the sequence-to-sequence framework. However, unlike sequence transduction problems such as machine translation, GEC suffers from the lack of plentiful parallel data. We describe two approaches for generating large parallel datasets for GEC using publicl…

2019-04-10abs ↗pdf ↗

We draw a formal connection between using synthetic training data to optimize neural network parameters and approximate, Bayesian, model-based reasoning. In particular, training a neural network using synthetic data can be viewed as learning a proposal distribution generator for approximate inference in the synthetic-d…

2017-03-02abs ↗pdf ↗

We analyze the influence and interactions of 60 largest world banks for 195 world countries using the reduced Google matrix algorithm for the English Wikipedia network with 5 416 537 articles. While the top asset rank positions are taken by the banks of China, with China Industrial and Commercial Bank of China at the f…

2019-02-21abs ↗pdf ↗

The paper refines the three-page index for links, proving a new bound and characterizing specific links.

problem Investigating the three-page index invariant for links and proving bounds.
method Constructing three-page presentations from reduced link diagrams via binding circles and contractible subcomplexes.
result Proves a new bound for the three-page index and characterizes links achieving equality.

We construct a series of finitely presented semigroups. The centers of these semigroups encode uniquely up to rigid ambient isotopy in 3-space all non-oriented spatial graphs. This encoding is obtained by using three-page embeddings of graphs into the product of the line with the cone on three points. By exploiting thr…

2004-07-19abs ↗pdf ↗

Improved linear upper bound for ribbonlength of knots.

problem Estimating the ribbonlength of knots and links.
method Using four-page open book decompositions and spanning trees of checkerboard graphs, constructing a four-page presentation with at most 2c(K) arcs.
result Proved that ribbonlength is bounded above by the four-page index, leading to the linear bound Rib(K) ≤ 2c(K).

PAGE is a simple gradient estimator for nonconvex optimization problems.

problem Nonconvex optimization problems in machine learning.
method PAGE is a probabilistic gradient estimator that uses vanilla SGD with probability and a small adjustment with probability 1-p.
result PAGE achieves optimal convergence rates for nonconvex finite-sum and online problems.

New invariant distinguishes Legendrian surfaces in 5-manifolds.

problem Distinguishing Legendrian surfaces in closed contact 5-manifolds.
method Introduced a new Legendrian isotopy invariant, MPX(L)M\mathcal{P}_{X}(L), and extended it to an absolute invariant.
result New invariant distinguishes surfaces not distinguishable by Thurston-Bennequin invariant.

We construct a contact 5-manifold supported by infinitely many distinct open books with the identity monodromy and pairwise exotic Stein pages (i.e. pages are pairwise homeomorphic but non-diffeomorphic Stein fillings of a fixed contact 3-manifold), moreover we describe a process of generating infinitely many such exam…

2015-02-21abs ↗pdf ↗

Optimization is commonly employed to determine the content of web pages, such as to maximize conversions on landing pages or click-through rates on search engine result pages. Often the layout of these pages can be decoupled into several separate decisions. For example, the composition of a landing page may involve dec…

2018-10-22abs ↗pdf ↗

In this paper, we study some classes of submanifolds of codimension one and two in the Page space. These submanifolds are totally geodesic. We also compute their curvature and show that some of them are constant curvature spaces. Finally we give information on how the Page space is related to some other metrics on the …

2016-08-10abs ↗pdf ↗

Generative models predict page quality without training, useful for low-resource settings.

problem Detecting low-quality content in web articles.
method Human evaluation and analysis of 500 million web articles.
result Generative models can predict page quality without training, useful for low-resource settings.

The paper extends Hawking--Page solutions to various spacetimes with singularities.

problem Understanding the extensions of Hawking--Page solutions with different types of singularities.
method Kaluza--Klein reduction and Christodoulou's methods.
result Extensions of Lorentzian Hawking--Page solutions with null, spacelike singularities, and Cauchy horizons of Taub--NUT type are proven.

Freya PAGE optimizes nonconvex optimization with heterogeneous, asynchronous workers.

problem Optimizing nonconvex finite-sum problems with varying worker processing times.
method Freya PAGE, a parallel method robust to stragglers and adaptive to slow computations.
result Freya PAGE offers improved time complexity guarantees compared to previous methods.

We construct, somewhat non-standard, Legendrian surgery diagrams for some Stein fillable contact structures on some plumbing trees of circle bundles over spheres. We then show how to put such a surgery diagram on the pages of an open book for S3,S^3, with relatively low genus. Thus we produce open books with low genus p…

2006-07-14abs ↗pdf ↗