Locality-sensitive hashing speeds up web app security testing.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Let $-\im\Lie_\T$ (essentially Lie derivative with respect to $\T$, a smooth nowhere zero real vector field) and be commuting differential operators, respectively of orders 1 and , the latter formally normal, both acting on sections of a vector bundle over a closed manifold. It is shown that if $P+(-i\Lie_…
Let be a countable family of rational functions of two variables with real coefficients. Each rational function can be thought as a continuous function taking values in the projective line and defined on a cofinite subset of the torus . Then t…
Building agents to interact with the web would allow for significant improvements in knowledge understanding and representation learning. However, web navigation tasks are difficult for current deep reinforcement learning (RL) models due to the large discrete action space and the varying number of actions between the s…
In a previous paper by Deroin-Tholozan, the authors construct a map from the Teichmüller space of to itself and prove that, when has sectional curvature , the image of lies (almost always) in the domain of Fuchsian representations stricly dominating . Here…
Apparently random financial fluctuations often exhibit varying levels of complexity, chaos. Given limited data, predictability of such time series becomes hard to infer. While efficient methods of Lyapunov exponent computation are devised, knowledge about the process driving the dynamics greatly facilitates the complex…
Improved object detection for scientific document images.
OLALA automates document layout annotation by selecting ambiguous regions for labeling.
Concept Relation Discovery and Innovation Enabling Technology (CORDIET), is a toolbox for gaining new knowledge from unstructured text data. At the core of CORDIET is the C-K theory which captures the essential elements of innovation. The tool uses Formal Concept Analysis (FCA), Emergent Self Organizing Maps (ESOM) and…
MARGE learns to reconstruct text by paraphrasing, achieving strong performance across multiple tasks.
CCAligned creates a massive web document dataset for cross-lingual research.
BARMPy offers a Python package for Bayesian Additive Regression Models.
Extreme multi-label text classification (XMTC) aims at tagging a document with most relevant labels from an extremely large-scale label set. It is a challenging problem especially for the tail labels because there are only few training documents to build classifier. This paper is motivated to better explore the semanti…
We propose an unsupervised object matching method for relational data, which finds matchings between objects in different relational datasets without correspondence information. For example, the proposed method matches documents in different languages in multi-lingual document-word networks without dictionaries nor ali…
In many real life problems, objects are described by large number of binary features. For instance, documents are characterized by presence or absence of certain keywords; cancer patients are characterized by presence or absence of certain mutations etc. In such cases, grouping together similar objects/profiles based o…
Paper accelerates K-means clustering for large sparse document data.
Improves document summarization by combining word embeddings and n-grams.
Many information retrieval algorithms rely on the notion of a good distance that allows to efficiently compare objects of different nature. Recently, a new promising metric called Word Mover's Distance was proposed to measure the divergence between text passages. In this paper, we demonstrate that this metric can be ex…
Unsupervised scheme ranks sentences in text documents based on semantic importance.
This paper proposes an inexpensive way to learn an effective dissimilarity function to be used for -nearest neighbor (-NN) classification. Unlike Mahalanobis metric learning methods that map both query (unlabeled) objects and labeled objects to new coordinates by a single transformation, our method learns a trans…
This paper considers extractive summarisation in a comparative setting: given two or more document groups (e.g., separated by publication time), the goal is to select a small number of documents that are representative of each group, and also maximally distinguishable from other groups. We formulate a set of new object…
The work investigates deep generative models, which allow us to use training data from one domain to build a model for another domain. We propose the Variational Bi-domain Triplet Autoencoder (VBTA) that learns a joint distribution of objects from different domains. We extend the VBTAs objective function by the relativ…
A new model CDTM improves text classification by concentrating document topics.
Study risk-sensitive reinforcement learning with optimized certainty equivalents.
We propose an algorithm for the non-negative factorization of an occurrence tensor built from heterogeneous networks. We use l0 norm to model sparse errors over discrete values (occurrences), and use decomposed factors to model the embedded groups of nodes. An efficient splitting method is developed to optimize the non…
Framework uses deep learning to analyze large documents and identify their logical structure.
A novel Python framework for Bayesian optimization known as GPflowOpt is introduced. The package is based on the popular GPflow library for Gaussian processes, leveraging the benefits of TensorFlow including automatic differentiation, parallelization and GPU computations for Bayesian optimization. Design goals focus on…
Topic modeling based on latent Dirichlet allocation (LDA) has been a framework of choice to perform scene recognition and annotation. Recently, a new type of topic model called the Document Neural Autoregressive Distribution Estimator (DocNADE) was proposed and demonstrated state-of-the-art performance for document mod…
A new method for document network embedding interprets and generalizes well.
A scalable topic model for large document collections using MapReduce.
Document clustering and topic modeling are two closely related tasks which can mutually benefit each other. Topic modeling can project documents into a topic space which facilitates effective document clustering. Cluster labels discovered by document clustering can be incorporated into topic models to extract local top…
Nonparametric Bayesian models are often based on the assumption that the objects being modeled are exchangeable. While appropriate in some applications (e.g., bag-of-words models for documents), exchangeability is sometimes assumed simply for computational reasons; non-exchangeable models might be a better choice for a…
This document contains a description of physics entirely based on a geometric presentation: all of the theory is described giving only a pseudo-riemannian manifold (M, g) of dimension n > 5 for which the g tensor is, in studied domains, almost everywhere of signature (-, -, +, ..., +). No object is added to this space-…
MOPO-LSI offers a user guide for sustainable investments.
When estimating the relevancy between a query and a document, ranking models largely neglect the mutual information among documents. A common wisdom is that if two documents are similar in terms of the same query, they are more likely to have similar relevance score. To mitigate this problem, in this paper, we propose …
A new model analyzes document structure and customer shopping patterns.
Traditional Relational Topic Models provide a way to discover the hidden topics from a document network. Many theoretical and practical tasks, such as dimensional reduction, document clustering, link prediction, benefit from this revealed knowledge. However, existing relational topic models are based on an assumption t…
Study improves document processing in banking with multimodal analytics.
This paper improves topic modeling by embedding words and topics together.
Contrastive learning helps linear models understand document topics.
Develops scalable autoencoder for document networks.
Word embedding maps words into a low-dimensional continuous embedding space by exploiting the local word collocation patterns in a small context window. On the other hand, topic modeling maps documents onto a low-dimensional topic space, by utilizing the global word collocation patterns in the same document. These two …
New seq2seq model can copy entire spans, outperforming simpler models in editing tasks.
Nonnegative matrix factorization (NMF) is a linear dimensionality reduction technique for analyzing nonnegative data. A key aspect of NMF is the choice of the objective function that depends on the noise model (or statistics of the noise) assumed on the data. In many applications, the noise model is unknown and difficu…
Proposes new listwise learning-to-rank models to address rating ties and document relevance.
A neural network method for topic modeling from few documents.
We develop a nested hierarchical Dirichlet process (nHDP) for hierarchical topic modeling. The nHDP is a generalization of the nested Chinese restaurant process (nCRP) that allows each word to follow its own path to a topic node according to a document-specific distribution on a shared tree. This alleviates the rigid, …
Latent topic models have been successfully applied as an unsupervised topic discovery technique in large document collections. With the proliferation of hypertext document collection such as the Internet, there has also been great interest in extending these approaches to hypertext [6, 9]. These approaches typically mo…