Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

4285127169 · Jun 202019922001200920172026
48 results for duplicate detection

Unsupervised near-duplicate detection has many practical applications ranging from social media analysis and web-scale retrieval, to digital image forensics. It entails running a threshold-limited query on a set of descriptors extracted from the images, with the goal of identifying all possible near-duplicates, while l…

2019-07-03abs ↗pdf ↗

MNIST and Fashion MNIST are extremely popular for testing in the machine learning space. Fashion MNIST improves on MNIST by introducing a harder problem, increasing the diversity of testing sets, and more accurately representing a modern computer vision task. In order to increase the data quality of FashionMNIST, this …

2019-06-19abs ↗pdf ↗

This paper identifies duplicate questions on Quora using machine and deep learning models.

problem Detecting semantically identical questions on Quora to improve user experience.
method Applied machine learning and deep learning techniques on Quora's dataset.
result Xgboost model with character level term frequency and inverse term frequency achieved 85.82% accuracy.

We identify a duplicate pair in the well-known Callahan-Hildebrand-Weeks census of cusped finite-volume hyperbolic 3-manifolds. Specifically, the six-tetrahedron non-orientable manifolds x101 and x103 are homeomorphic.

2013-11-29abs ↗pdf ↗

Detects AI-synthesized speech using cepstral and bispectral analysis.

problem Validating the authenticity of speech from AI-generated content.
method Integrates cepstral and bispectral analysis for distinguishing human from AI-synthesized speech.
result Higher-order statistics show less correlation for human speech compared to AI-synthesis, and cepstral analysis reveals unique power components.

A Python package solves source duplication in single channel LVMs using spectral regularisation.

problem Source duplication in LVMs hampers their practical use in single channel applications.
method Spectral regularisation term added to address source duplication issue.
result Spectral regularisation framework enables easier investigation and utilisation of LVMs.

Data de-duplication is the task of detecting multiple records that correspond to the same real-world entity in a database. In this work, we view de-duplication as a clustering problem where the goal is to put records corresponding to the same physical entity in the same cluster and putting records corresponding to diff…

2018-10-10abs ↗pdf ↗

We are able to derive the equations of motion for forced mechanical systems in a purely variational setting, both in the context of Lagrangian or Hamiltonian mechanics, by duplicating the variables of the system as introduced by Galley [2013], Galley, Tsang, and Stein [2014]. We show that this construction is useful to…

2017-12-26abs ↗pdf ↗

Question semantic similarity (Q2Q) is a challenging task that is very useful in many NLP applications, such as detecting duplicate questions and question answering systems. In this paper, we present the results and findings of the shared task (Semantic Question Similarity in Arabic). The task was organized as part of t…

2019-09-12abs ↗pdf ↗

Beam search improves UQ in LLMs by reducing duplicates and variance.

problem Peaked distributions in multinomial sampling lead to duplicates and high variance in uncertainty estimates.
method Employ beam search to generate candidates for consistency-based UQ, providing a theoretical lower bound and empirical evaluation.
result Beam search achieves smaller error than multinomial sampling, leading to state-of-the-art UQ performance.

The long standing classification problem in the theory of Heegaard splittings of 3-manifolds is to exhibit for each closed 3-manifold a complete list, without duplication, of all its irreducible Heegaard surfaces, up to isotopy. We solve this problem for non Haken hyperbolic 3-manifolds.

2015-09-19abs ↗pdf ↗

New metric for disentangling multivariate representations, accounting for more complex entanglements.

problem Current disentanglement metrics fail to detect entanglements involving more than two variables.
method Partial Information Decomposition framework to analyze information sharing and propose a new disentanglement metric.
result The proposed metric correctly identifies entanglements in high-dimensional spaces.

Cost-effective feature selection improves network model choice.

problem Selecting informative features from noisy candidates in network models.
method Adapted feature selection methods to account for feature costs and used pilot simulations.
result Reduced computational cost by two orders of magnitude without sacrificing model accuracy.

Tricks adversarial attacks to target specific classes, improving classifier accuracy.

problem Recent adversarial defense approaches have failed to protect classifiers from untargeted attacks.
method Target Training defense tricks untargeted attacks into targeted attacks on designated classes, then derives the real class.
result 86.2% accuracy for CW-L2 (confidence=0) in CIFAR10, outperforming unsecured classifiers.

New summary measures reveal geometric structure in weighted measures on manifolds.

problem Lack of geometric information in standard weight-only summaries.
method Heat-kernel entropy profiles, tracking nonuniformity across scales.
result Geometric effective sample size discounts nearby or duplicate particles.

A simple Ising spin model which can describe the mechanism of price formation in financial markets is proposed. In contrast to other agent-based models, the influence does not flow inward from the surrounding neighbors to the center site, but spreads outward from the center to the neighbors. The model thus describes th…

2000-12-30abs ↗pdf ↗

We propose GraphNVP, the first invertible, normalizing flow-based molecular graph generation model. We decompose the generation of a graph into two steps: generation of (i) an adjacency tensor and (ii) node attributes. This decomposition yields the exact likelihood maximization on graph-structured data, combined with t…

2019-05-28abs ↗pdf ↗

We obtain an analog of the compression of angles theorem in symmetric spaces for Bruhat--Tits buildings of the type AA. More precisely, consider a pp-adic linear space VV and the set Lat(V)Lat(V) of all lattices in VV. The complex distance in Lat(V)Lat(V) is a complete system of invariants of a pair of points of Lat(V)Lat(V) u…

2004-10-09abs ↗pdf ↗

In this work we present a complete (no misses, no duplicates) census for closed, connected, orientable and prime 3-manifolds induced by plane graphs with a bipartition of its edge set (blinks) up to k=9k=9 edges. Blinks form a universal encoding for such manifolds. In fact, each such a manifold is a subtle class of blin…

2013-05-24abs ↗pdf ↗

The study proves limitations on isospectral hyperbolic surfaces with discrete length spectra.

problem Characterizing isospectral hyperbolic surfaces with discrete length spectra.
method Utilizing Sunada's method and topological self-duplicating ends, the study explores isospectral families and their cardinality.
result Finite groups can be realized as full isometry groups of hyperbolic structures with discrete spectrum on surfaces with self-duplicating ends.

Entity resolution seeks to merge databases as to remove duplicate entries where unique identifiers are typically unknown. We review modern blocking approaches for entity resolution, focusing on those based upon locality sensitive hashing (LSH). First, we introduce kk-means locality sensitive hashing (KLSH), which is b…

2018-10-11abs ↗pdf ↗

Sampling is a fundamental technique, and sampling without replacement is often desirable when duplicate samples are not beneficial. Within machine learning, sampling is useful for generating diverse outputs from a trained model. We present an elegant procedure for sampling without replacement from a broad class of rand…

2020-02-21abs ↗pdf ↗

We have analyzed the risks of possible development of bubbles in the Swiss residential real estate market. The data employed in this work has been collected by comparis.ch, and carefully cleaned from duplicate records through a procedure based on supervised machine learning methods. The study uses the log periodic powe…

2013-03-19abs ↗pdf ↗

Framework for automatically assessing and correcting data quality issues without domain knowledge.

problem Ensuring data quality in datasets across various domains.
method Hybrid approach combining statistical and machine learning methods.
result Effective detection and correction of missing values, duplicates, and typographical errors.

Two solutions for multi-modal record linkage using Deep Learning inspired by Visual Question Answering.

problem Matching records from multiple sources representing the same entity.
method Two fusion modules: Recurrent Neural Network + Convolutional Neural Network and Stacked Attention Network. A Siamese Neural Network computes similarity.
result Recurrent Neural Network + Convolutional Neural Network fusion module outperforms a simple model.

RGRR allocates between QQQ and DIA based on relative states, improving Sharpe and CAGR.

problem Optimizing ETF allocation between QQQ and DIA for better risk-adjusted returns.
method Screened relative and macro states, globally screened interactions, fixed position mapping, walk-forward validation.
result RGRR improves Sharpe and CAGR compared to 100% QQQ and 50/50 QQQ-DIA allocations.

Record linkage involves merging records in large, noisy databases to remove duplicate entities. It has become an important area because of its widespread occurrence in bibliometrics, public health, official statistics production, political science, and beyond. Traditional linkage methods directly linking records to one…

2017-03-08abs ↗pdf ↗