Semi-supervised model removes noisy content from webpages.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The multimodal web elements such as text and images are associated with inherent memory costs to store and transfer over the Internet. With the limited network connectivity in developing countries, webpage rendering gets delayed in the presence of high-memory demanding elements such as images (relative to text). To ove…
Many classification problems involve data instances that are interlinked with each other, such as webpages connected by hyperlinks. Techniques for "collective classification" (CC) often increase accuracy for such data graphs, but usually require a fully-labeled training graph. In contrast, we examine how to improve the…
We consider the problem of learning a high-dimensional graphical model in which certain hub nodes are highly-connected to many other nodes. Many authors have studied the use of an l1 penalty in order to learn a sparse graph in high-dimensional setting. However, the l1 penalty implicitly assumes that each edge is equall…
Recent years have witnessed a widespread increase of interest in network representation learning (NRL). By far most research efforts have focused on NRL for homogeneous networks like social networks where vertices are of the same type, or heterogeneous networks like knowledge graphs where vertices (and/or edges) are of…
Web crawling is the problem of keeping a cache of webpages fresh, i.e., having the most recent copy available when a page is requested. This problem is usually coupled with the natural restriction that the bandwidth available to the web crawler is limited. The corresponding optimization problem was solved optimally by …
Hi-RES framework extracts medical relations from articles and EHRs.
Finite mixture models have been used for unsupervised learning for some time, and their use within the semi-supervised paradigm is becoming more commonplace. Clickstream data is one of the various emerging data types that demands particular attention because there is a notable paucity of statistical learning approaches…
DAMVI algorithm improves imbalanced binary classification by adjusting weights of examples and classifiers.
This brief report (6 pages) was written in 1983 but never published. It concerns the hyperbolic 3-orbifolds obtained as quotients of hyperbolic 3-space by the group of invertible 2 by 2 matrices whose entries are integers in the imaginary quadratic extension of Q of discriminant D. For values D > -100 the topological t…
regvis.net offers a visual survey of regulatory visualization.
The multi-armed bandit problem has been extensively studied under the stationary assumption. However in reality, this assumption often does not hold because the distributions of rewards themselves may change over time. In this paper, we propose a change-detection (CD) based framework for multi-armed bandit problems und…
BootsTAP uses real-world data to improve TAP tracking performance.
In many supervised learning tasks, the entities to be labeled are related to each other in complex ways and their labels are not independent. For example, in hypertext classification, the labels of linked pages are highly correlated. A standard approach is to classify each entity independently, ignoring the correlation…
Solves online resource allocation problems with budget constraints.
We empirically verify that the market capitalisations of coins and tokens in the cryptocurrency universe follow power-law distributions with significantly different values, with the tail exponent falling between 0.5 and 0.7 for coins, and between 1.0 and 1.3 for tokens. We provide a rationale for this, based on a simpl…
New method quantifies uncertainty in denoising models.
This paper proposes a generic classification system designed to detect security threats based on the behavior of malware samples. The system relies on statistical features computed from proxy log fields to train detectors using a database of malware samples. The behavior detectors serve as basic reusable building block…
We consider the task of estimating a Gaussian graphical model in the high-dimensional setting. The graphical lasso, which involves maximizing the Gaussian log likelihood subject to an l1 penalty, is a well-studied approach for this task. We begin by introducing a surprising connection between the graphical lasso and hi…
MULTIPOLAR aggregates diverse source policies for efficient transfer RL.
We propose an alternative framework to existing setups for controlling false alarms when multiple A/B tests are run over time. This setup arises in many practical applications, e.g. when pharmaceutical companies test new treatment options against control pills for different diseases, or when internet companies test the…
GrokAlign aligns Jacobians to accelerate grokking in deep networks.
Recommendation problems with large numbers of discrete items, such as products, webpages, or videos, are ubiquitous in the technology industry. Deep neural networks are being increasingly used for these recommendation problems. These models use embeddings to represent discrete items as continuous vectors, and the vocab…
Study intrinsic motivation for synergistic tasks in reinforcement learning.
We consider the problem of estimating high-dimensional Gaussian graphical models corresponding to a single set of variables under several distinct conditions. This problem is motivated by the task of recovering transcriptional regulatory networks on the basis of gene expression data {containing heterogeneous samples, s…
We consider the estimation of heterogeneous treatment effects with arbitrary machine learning methods in the presence of unobserved confounders with the aid of a valid instrument. Such settings arise in A/B tests with an intent-to-treat structure, where the experimenter randomizes over which user will receive a recomme…
WebGUM learns web navigation from multimodal data, outperforming previous methods.
AlphaForgeBench evaluates LLMs as quantitative researchers, not trading agents, to address instability in financial decision-making.
In unsupervised learning, collecting more data is not always a costly process unlike the training. For example, it is not hard to enlarge the 40GB WebText used for training GPT-2 by modifying its sampling methodology considering how many webpages there are in the Internet. On the other hand, given that training on this…
Line graph transformation aids graph isomorphism tests by excluding challenging graph properties.
Proposes MGMN for end-to-end graph similarity learning.
The paper explores graphons of line graphs from sparse finite graphs.
MxPool learns graph features from diverse graphs using a hierarchical structure.
Study the geometry of graph product extension graphs.
Graph neural network learns graph distances effectively.
Quasi-transitive graphs quasi-isometric to planar graphs can be upgraded to Cayley graphs.
Customized-GNN generates model-specific for each graph.
GRAPH-BERT uses only attention for graph representation learning.
Graph embedding leaks sensitive graph properties and subgraphs.
Characterizes graphs with leveled embeddings and introduces new graph invariants.
The paper shows conflict graphs of Petersen family graphs are mostly unbalanced.
HGP-SL pools and learns graph structure for hierarchical representation learning.
Two new methods improve graph embedding without needing a complete graph structure.
We define a pseudo-inverse for line graphs using linear integer programming.
HaarPooling compresses graphs by Haar transforms, improving graph classification and regression.
New graph kernel scales well with graph size and number, achieving state-of-the-art performance.
Develops method to create non-Abelian Ricci-flat graphs via bundles.
MathNet uses wavelets for graph representation and learning.