This work addresses the problem of author name homonymy in the Web of Science. Aiming for an efficient, simple and straightforward solution, we introduce a novel probabilistic similarity measure for author name disambiguation based on feature overlap. Using the researcher-ID available for a subset of the Web of Science…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Model clusters authors and topics in short texts like social media posts.
Determining the quality of the results obtained by clustering techniques is a key issue in unsupervised machine learning. Many authors have discussed the desirable features of good clustering algorithms. However, Jon Kleinberg established an impossibility theorem for clustering. As a consequence, a wealth of studies ha…
Bibliographic analysis considers the author's research areas, the citation network and the paper content among other things. In this paper, we combine these three in a topic model that produces a bibliographic model of authors, topics and documents, using a nonparametric extension of a combination of the Poisson mixed-…
This study improves child welfare risk models using clustering methods.
Survey on reproducibility and distortion issues in text clustering and topic modeling.
We introduce Grinch, a new algorithm for large-scale, non-greedy hierarchical clustering with general linkage functions that compute arbitrary similarity between two point sets. The key components of Grinch are its rotate and graft subroutines that efficiently reconfigure the hierarchy as new points arrive, supporting …
Bibliographic analysis considers author's research areas, the citation network and paper content among other things. In this paper, we combine these three in a topic model that produces a bibliographic model of authors, topics and documents using a non-parametric extension of a combination of the Poisson mixed-topic li…
Information about the contributions of individual authors to scientific publications is important for assessing authors' achievements. Some biomedical publications have a short section that describes authors' roles and contributions. It is usually written in natural language and hence author contributions cannot be tri…
Author name disambiguation in bibliographic databases is the problem of grouping together scientific publications written by the same person, accounting for potential homonyms and/or synonyms. Among solutions to this problem, digital libraries are increasingly offering tools for authors to manually curate their publica…
Resolving a conjecture of Abbe, Bandeira and Hall, the authors have recently shown that the semidefinite programming (SDP) relaxation of the maximum likelihood estimator achieves the sharp threshold for exactly recovering the community structure under the binary stochastic block model of two equal-sized clusters. The s…
Complete description of BNSR invariants for Lodha-Moore groups, proving finiteness properties.
Paper proves MS convergence for radially symmetric kernels with large bandwidths.
Distance-based hierarchical clustering (HC) methods are widely used in unsupervised data analysis but few authors take account of uncertainty in the distance data. We incorporate a statistical model of the uncertainty through corruption or noise in the pairwise distances and investigate the problem of estimating the HC…
A method for dimension reduction with clustering, classification, or discriminant analysis is introduced. This mixture model-based approach is based on fitting generalized hyperbolic mixtures on a reduced subspace within the paradigm of model-based clustering, classification, or discriminant analysis. A reduced subspac…
The classification of textual data often yields important information. Most classifiers work in a closed world setting where the classifier is trained on a known corpus, and then it is tested on unseen examples that belong to one of the classes seen during training. Despite the usefulness of this design, often there is…
In this paper, we explore the relationship between one of the most elementary and important properties of graphs, the presence and relative frequency of triangles, and a combinatorial notion of Ricci curvature. We employ a definition of generalized Ricci curvature proposed by Ollivier in a general framework of Markov p…
The following working document summarizes our work on the clustering of financial time series. It was written for a workshop on information geometry and its application for image and signal processing. This workshop brought several experts in pure and applied mathematics together with applied researchers from medical i…
Construct locally minimizing -clusters with prescribed asymptotic geometry.
We extend existing models in the financial literature by introducing a cluster-derived canonical vine (CDCV) copula model for capturing high dimensional dependence between financial time series. This model utilises a simplified market-sector vine copula framework similar to those introduced by Heinen and Valdesogo (200…
This paper concerns cluster algebras with principal coefficients A(S,M) associated to bordered surfaces (S,M), and is a companion to a concurrent work of the authors with Schiffler [MSW2]. Given any (generalized) arc or loop in the surface -- with or without self-intersections -- we associate an element of (the fractio…
We describe the Fast Greedy Sparse Subspace Clustering (FGSSC) algorithm providing an efficient method for clustering data belonging to a few low-dimensional linear or affine subspaces. The main difference of our algorithm from predecessors is its ability to work with noisy data having a high rate of erasures (missed e…
With ongoing developments and innovations in single-cell RNA sequencing methods, advancements in sequencing performance could empower significant discoveries as well as new emerging possibilities to address biological and medical investigations. In the study, we will be using the dataset collected by the authors of Sys…
It was not until the beginning of the 1990s that the effects of information and communication technology on economic growth as well as on the profitability of enterprises raised the interest of researchers. After giving a general description on the relationship between a more intense use of ICT devices and dynamic econ…
Problems of interpolation, classification, and clustering are considered. In the tenets of Radon--Nikodym approach , where the is a linear function on input attributes, all the answers are obtained from a generalized eigenproblem $|f|ψ^{[i]}\rangle =…
The Classification Literature Automated Search Service, an annual bibliography based on citation of one or more of a set of around 80 book or journal publications, ran from 1972 to 2012. We analyze here the years 1994 to 2011. The Classification Society's Service, as it was termed, has been produced by the Classificati…
Android and Facebook provide third-party applications with access to users' private data and the ability to perform potentially sensitive operations (e.g., post to a user's wall or place phone calls). As a security measure, these platforms restrict applications' privileges with permission systems: users must approve th…
Quantum trace maps for surfaces are shown to be compatible under triangulations.
Community detection, which aims to cluster nodes in a given graph into distinct groups based on the observed undirected edges, is an important problem in network data analysis. In this paper, the popular stochastic block model (SBM) is extended to the generalized stochastic block model (GSBM) that allows for ad…
This article reviews entity resolution methods and their applications.
Proposes a model for identifying 4G cells with network throughput problems.
Density Estimation is one of the central areas of statistics whose purpose is to estimate the probability density function underlying the observed data. It serves as a building block for many tasks in statistical inference, visualization, and machine learning. Density Estimation is widely adopted in the domain of unsup…
Real-world networks usually have community structure, that is, nodes are grouped into densely connected communities. Community detection is one of the most popular and best-studied research topics in network science and has attracted attention in many different fields, including computer science, statistics, social sci…
This paper outlines an agent-based model of a simple financial market in which a single asset is available for trade by three different types of traders. The model was first introduced in the PhD thesis of one of the authors, see reference [1]. The simulated log returns are examined for the presence of the stylised fac…
In the domain of technology startups, biotechnology has often been considered as specific. Their unique technology content, the type of founders and managers they have, the amount of venture capital they raise, the time it takes them to reach an exit as well as the technology clusters they belong to are seen as such un…
We introduce the author-topic model, a generative model for documents that extends Latent Dirichlet Allocation (LDA; Blei, Ng, & Jordan, 2003) to include authorship information. Each author is associated with a multinomial distribution over topics and each topic is associated with a multinomial distribution over words.…
Dirichlet Process(DP) is a Bayesian non-parametric prior for infinite mixture modeling, where the number of mixture components grows with the number of data items. The Hierarchical Dirichlet Process (HDP), is an extension of DP for grouped data, often used for non-parametric topic modeling, where each group is a mixtur…
Text-based analysis methods allow to reveal privacy relevant author attributes such as gender, age and identify of the text's author. Such methods can compromise the privacy of an anonymous author even when the author tries to remove privacy sensitive content. In this paper, we propose an automatic method, called Adver…
A genetic algorithm with community detection improves feature selection accuracy.
Optimizes portfolios using neural network approximations of asset sensitivities to common drivers.
In [5], together with J. C. Wood, the authors gave a completely explicit formula for all harmonic maps from -spheres to the unitary group in terms of freely chosen meromorphic functions on . The simplest harmonic maps are the isotropic ones. Using Morse theory Burstall and Guest [1] showed that the harmo…
Experiment shows author rankings can improve peer review scores.
NTL protects AI models by restricting their generalization ability to specific domains.
In this paper, we are interested in the location of conjugate points along a geodesic in the volumorphism group of a compact three-dimensional manifold without boundary (the configuration space of an ideal fluid). As shown in the author's previous work, these are typically pathological, i.e., they can occur in clusters…
Convex clustering can only learn convex clusters, with significant gaps between clusters.
Proposes a new clustering method based on expectiles for non-spherical clusters.
CCMM efficiently solves large-scale convex clustering problems.
Discussing issues in robust clustering, especially with Gaussian models.