Bayesian approach for multifile record linkage and duplicate detection.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New validation method prevents privacy breaches and biases in federated learning.
The paper tackles uniform sampling from databases with duplicates.
Record linkage involves merging records in large, noisy databases to remove duplicate entities. It has become an important area because of its widespread occurrence in bibliometrics, public health, official statistics production, political science, and beyond. Traditional linkage methods directly linking records to one…
Data de-duplication is the task of detecting multiple records that correspond to the same real-world entity in a database. In this work, we view de-duplication as a clustering problem where the goal is to put records corresponding to the same physical entity in the same cluster and putting records corresponding to diff…
MNIST and Fashion MNIST are extremely popular for testing in the machine learning space. Fashion MNIST improves on MNIST by introducing a harder problem, increasing the diversity of testing sets, and more accurately representing a modern computer vision task. In order to increase the data quality of FashionMNIST, this …
This article reviews entity resolution methods and their applications.
Two solutions for multi-modal record linkage using Deep Learning inspired by Visual Question Answering.
Entity resolution seeks to merge databases as to remove duplicate entries where unique identifiers are typically unknown. We review modern blocking approaches for entity resolution, focusing on those based upon locality sensitive hashing (LSH). First, we introduce -means locality sensitive hashing (KLSH), which is b…
Unsupervised near-duplicate detection has many practical applications ranging from social media analysis and web-scale retrieval, to digital image forensics. It entails running a threshold-limited query on a set of descriptors extracted from the images, with the goal of identifying all possible near-duplicates, while l…
We identify a duplicate pair in the well-known Callahan-Hildebrand-Weeks census of cusped finite-volume hyperbolic 3-manifolds. Specifically, the six-tetrahedron non-orientable manifolds x101 and x103 are homeomorphic.
Entity resolution (ER; also known as record linkage or de-duplication) is the process of merging noisy databases, often in the absence of unique identifiers. A major advancement in ER methodology has been the application of Bayesian generative models, which provide a natural framework for inferring latent entities with…
We have analyzed the risks of possible development of bubbles in the Swiss residential real estate market. The data employed in this work has been collected by comparis.ch, and carefully cleaned from duplicate records through a procedure based on supervised machine learning methods. The study uses the log periodic powe…
Robust machine learning relies on access to data that can be used with standardized frameworks in important tasks and the ability to develop models whose performance can be reasonably reproduced. In machine learning for healthcare, the community faces reproducibility challenges due to a lack of publicly accessible data…
A Python package solves source duplication in single channel LVMs using spectral regularisation.
KT models improved slightly with synthetic student data.
Dropout is a simple but efficient regularization technique for achieving better generalization of deep neural networks (DNNs); hence it is widely used in tasks based on DNNs. During training, dropout randomly discards a portion of the neurons to avoid overfitting. This paper presents an enhanced dropout technique, whic…
We are able to derive the equations of motion for forced mechanical systems in a purely variational setting, both in the context of Lagrangian or Hamiltonian mechanics, by duplicating the variables of the system as introduced by Galley [2013], Galley, Tsang, and Stein [2014]. We show that this construction is useful to…
This paper identifies duplicate questions on Quora using machine and deep learning models.
this is a duplicate submission(original is arXiv:1612.02141). Hence want to withdraw it
Beam search improves UQ in LLMs by reducing duplicates and variance.
Postprocessing reduces Bayesian optimization steps for global optima.
The papers math.QA/0403527 and math.QA/0409414 v.1 are now merged together. The final version is available at math.QA/0409414 v.2. To avoid duplication of papers, math.QA/0403527 is now removed.
Although nonnegative matrix factorization (NMF) is NP-hard in general, it has been shown very recently that it is tractable under the assumption that the input nonnegative data matrix is close to being separable (separability requires that all columns of the input matrix belongs to the cone spanned by a small subset of…
The long standing classification problem in the theory of Heegaard splittings of 3-manifolds is to exhibit for each closed 3-manifold a complete list, without duplication, of all its irreducible Heegaard surfaces, up to isotopy. We solve this problem for non Haken hyperbolic 3-manifolds.
Cost-effective feature selection improves network model choice.
Tricks adversarial attacks to target specific classes, improving classifier accuracy.
New summary measures reveal geometric structure in weighted measures on manifolds.
We develop a new method for proving regularity for small energy stationary solutions of coupled gauge field equations. Our results duplicate those of Tian--Tao [7] for the pure Yang Mills equations, but our proof is simpler, and obtains bounded curvature without the use of Coulomb gauges. It relies instead on the Weitz…
A simple Ising spin model which can describe the mechanism of price formation in financial markets is proposed. In contrast to other agent-based models, the influence does not flow inward from the surrounding neighbors to the center site, but spreads outward from the center to the neighbors. The model thus describes th…
EnsembleSVM is a free software package containing efficient routines to perform ensemble learning with support vector machine (SVM) base models. It currently offers ensemble methods based on binary SVM models. Our implementation avoids duplicate storage and evaluation of support vectors which are shared between constit…
RW-based learning is vulnerable to the Pac-Man attack, which eliminates active RWs.
We propose GraphNVP, the first invertible, normalizing flow-based molecular graph generation model. We decompose the generation of a graph into two steps: generation of (i) an adjacency tensor and (ii) node attributes. This decomposition yields the exact likelihood maximization on graph-structured data, combined with t…
Genealogy research is the study of family history using available resources such as historical records. Ancestry provides its customers with one of the world's largest online genealogical index with billions of records from a wide range of sources, including vital records such as birth and death certificates, census re…
We study the statistics of record-breaking events in daily stock prices of 366 stocks from the Standard and Poors 500 stock index. Both the record events in the daily stock prices themselves and the records in the daily returns are discussed. In both cases we try to describe the record statistics of the stock data with…
We obtain an analog of the compression of angles theorem in symmetric spaces for Bruhat--Tits buildings of the type . More precisely, consider a -adic linear space and the set of all lattices in . The complex distance in is a complete system of invariants of a pair of points of u…
We review recent advances on the record statistics of strongly correlated time series, whose entries denote the positions of a random walk or a Lévy flight on a line. After a brief survey of the theory of records for independent and identically distributed random variables, we focus on random walks. During the last few…
The study of record statistics of correlated series is gaining momentum. In this work, we study the records statistics of the time series of select stock market data and the geometric random walk, primarily through simulations. We show that the distribution of the age of records is a power law with the exponent lyi…
In this work we present a complete (no misses, no duplicates) census for closed, connected, orientable and prime 3-manifolds induced by plane graphs with a bipartition of its edge set (blinks) up to edges. Blinks form a universal encoding for such manifolds. In fact, each such a manifold is a subtle class of blin…
New approach protects privacy of deleted records in machine learning.
We present a probabilistic method for linking multiple datafiles. This task is not trivial in the absence of unique identifiers for the individuals recorded. This is a common scenario when linking census data to coverage measurement surveys for census coverage evaluation, and in general when multiple record-systems nee…
The study proves limitations on isospectral hyperbolic surfaces with discrete length spectra.
Question semantic similarity (Q2Q) is a challenging task that is very useful in many NLP applications, such as detecting duplicate questions and question answering systems. In this paper, we present the results and findings of the shared task (Semantic Question Similarity in Arabic). The task was organized as part of t…
Paper uses ensemblers to predict sepsis early from patient records.
We consider the occurrence of record-breaking events in random walks with asymmetric jump distributions. The statistics of records in symmetric random walks was previously analyzed by Majumdar and Ziff and is well understood. Unlike the case of symmetric jump distributions, in the asymmetric case the statistics of reco…
New method uses surrogate outcomes and single-record data to improve suicide risk modeling.
Feature engineering remains a major bottleneck when creating predictive systems from electronic medical records. At present, an important missing element is detecting predictive regular clinical motifs from irregular episodic records. We present Deepr (short for Deep record), a new end-to-end deep learning system that …
DNI recovers missing brain data from corrupted recordings.