Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

8152330 · Feb 202019922001200920182026
48 results for duplicate items

This report clarifies the distribution of unique items in bootstrap samples.

problem Understanding the role of duplicate items in bootstrap samples.
method Analyzes the distribution of unique items in bootstrap samples and derives a heuristic for normal approximation.
result Derives a heuristic for when a normal approximation is permissible for the distribution of unique items in bootstrap samples.

This paper removes near-duplicates from Fashion-MNIST to improve testing accuracy.

problem Near-duplicate images in Fashion-MNIST increase testing accuracy, reducing dataset quality.
method Identified and removed near-duplicate images between training and testing sets.
result Improved dataset quality for better testing accuracy in machine learning models.

This study benchmarks deep learning for unsupervised near-duplicate image detection.

problem Detecting near-duplicates in large image datasets with high specificity.
method Binary classification using Receiver Operating Curve (ROC) for comparison of different descriptors.
result Fine-tuning deep convolutional networks generally outperforms off-the-shelf features, with best performance on MFND dataset.

We identify a duplicate pair in the well-known Callahan-Hildebrand-Weeks census of cusped finite-volume hyperbolic 3-manifolds. Specifically, the six-tetrahedron non-orientable manifolds x101 and x103 are homeomorphic.

2013-11-29abs ↗pdf ↗

A Python package solves source duplication in single channel LVMs using spectral regularisation.

problem Source duplication in LVMs hampers their practical use in single channel applications.
method Spectral regularisation term added to address source duplication issue.
result Spectral regularisation framework enables easier investigation and utilisation of LVMs.

This paper identifies duplicate questions on Quora using machine and deep learning models.

problem Detecting semantically identical questions on Quora to improve user experience.
method Applied machine learning and deep learning techniques on Quora's dataset.
result Xgboost model with character level term frequency and inverse term frequency achieved 85.82% accuracy.

Beam search improves UQ in LLMs by reducing duplicates and variance.

problem Peaked distributions in multinomial sampling lead to duplicates and high variance in uncertainty estimates.
method Employ beam search to generate candidates for consistency-based UQ, providing a theoretical lower bound and empirical evaluation.
result Beam search achieves smaller error than multinomial sampling, leading to state-of-the-art UQ performance.

The long standing classification problem in the theory of Heegaard splittings of 3-manifolds is to exhibit for each closed 3-manifold a complete list, without duplication, of all its irreducible Heegaard surfaces, up to isotopy. We solve this problem for non Haken hyperbolic 3-manifolds.

2015-09-19abs ↗pdf ↗

New validation method prevents privacy breaches and biases in federated learning.

problem Privacy breaches and data leakage in federated learning.
method Stratified cross-validation for unbiased and privacy-preserving federated learning.
result Stratified cross-validation prevents data leakage without demanding deduplication algorithms.

Framework learns item representations from text data for complementary and similar items.

problem Generating accurate complementary item recommendations from textual data.
method Quadruplet network learning framework for latent space representation of items.
result Items are placed closer together in latent space for similar and complementary items compared to non-complementary items.

WCF uses Wasserstein distance to recommend cold-start items based on content similarity.

problem Recommendation performance drops for new items with little interaction history.
method Applies Wasserstein distance to map interaction history to contents, inferring user preferences.
result WCF outperforms state-of-the-art methods in cold-start recommendation.

Bayesian method improves adaptive testing item selection, ensuring full item exposure.

problem Adaptive testing selects items to estimate ability, but must also ensure diverse item exposure.
method Formulated as Bayesian model averaging, deriving optimal item sampling probabilities.
result Stochastic method achieves full item bank exposure without sacrificing accuracy.

FBSM improves item recommendation for cold-start users by modeling feature interactions.

problem Cold-start item recommendation for new users.
method Factorized bilinear similarity model learning interactions among item features.
result Improves TOP-n recommendation performance compared to traditional methods.

Cost-effective feature selection improves network model choice.

problem Selecting informative features from noisy candidates in network models.
method Adapted feature selection methods to account for feature costs and used pilot simulations.
result Reduced computational cost by two orders of magnitude without sacrificing model accuracy.

The paper analyzes and optimizes recommendation systems using user-user and item-item collaborative filtering.

problem Optimizing recommendation systems to minimize disliked recommendations.
method Proposes algorithms inspired by user-user and item-item collaborative filtering, proving performance guarantees in terms of expected regret.
result Information-theoretic lower bounds on regret match upper bounds up to logarithmic factors in two model parameter regimes.

Tricks adversarial attacks to target specific classes, improving classifier accuracy.

problem Recent adversarial defense approaches have failed to protect classifiers from untargeted attacks.
method Target Training defense tricks untargeted attacks into targeted attacks on designated classes, then derives the real class.
result 86.2% accuracy for CW-L2 (confidence=0) in CIFAR10, outperforming unsecured classifiers.

NNMs improve item recommendation by providing interpretable user and item representations.

problem Creating recommender systems that are both accurate and understandable.
method Normalized nonnegative models (NNMs) for item recommendation.
result NNM-based recommender systems provide high predictive power, computational tractability, and expressive user and item representations.

Active learning improves ordering of items with contextual attributes.

problem Learning accurate item orderings from pairwise comparisons, especially when exhaustive comparisons are impractical.
method Proposes an active learning strategy that samples items to minimize expected ordering error, accounting for uncertainty in comparisons.
result Superior sample efficiency and generalization compared to non-contextual ranking approaches and active preference learning baselines.

The study measures similarity in introductory programming items, offering a method and evaluation.

problem Measuring similarity in a diverse pool of programming items for personalized learning.
method General approach to measuring similarity, specific measures for introductory programming, three levels of abstraction evaluation.
result Evaluation of similarity measures using diverse programming environments.

New summary measures reveal geometric structure in weighted measures on manifolds.

problem Lack of geometric information in standard weight-only summaries.
method Heat-kernel entropy profiles, tracking nonuniformity across scales.
result Geometric effective sample size discounts nearby or duplicate particles.

This paper optimizes the number of comparisons needed to find the best k items from pairwise comparisons.

problem Finding the best k items from pairwise comparisons with limited comparisons.
method Developed algorithms for finding probably approximately correct and exact best k items under stochastic conditions.
result Upper and lower bounds on the number of comparisons for finding the best k items, with matching upper bounds for PAC best k items.

Two methods improve 10-K item segmentation using large language models.

problem Challenges in extracting specific items from 10-K reports due to variations in document formats and item presentation.
method Two advanced item segmentation methods: GPT4ItemSeg and BERT4ItemSeg.
result BERT4ItemSeg achieves a macro-F1 of 0.9825, surpassing other methods.

Much of the data being created on the web contains interactions between users and items. Stochastic blockmodels, and other methods for community detection and clustering of bipartite graphs, can infer latent user communities and latent item clusters from this interaction data. These methods, however, typically ignore t…

2015-05-25abs ↗pdf ↗

New model improves recommendation systems by analyzing user-item interactions.

problem Improving recommendation systems for better user-item interactions.
method Sliced Anti-symmetric Decomposition (SAD) model using tensor decomposition.
result SAD produces the most consistent personalized preferences compared to SOTA models.

Eigenvalue analogy explains item-based recommender system accuracy.

problem Lack of theoretical explanation for item-based recommender system success.
method Formalized as an eigenvalue problem, estimating ratings as true ratings multiplied by user-specific eigenvalues.
result Eigenvalue magnitude correlates with user's recommendation accuracy and can measure confidence.

ProbMinHash improves Jaccard similarity hashing for big data applications.

problem Efficiently estimating set similarities in big data with weighted elements.
method Locality-sensitive hash algorithms that calculate signatures collectively.
result Significantly faster than the original approach, with improved estimation error.

Algorithm identifies best item from subsets with random utility model feedback.

problem PAC learning the best item from subsets with random utility model feedback.
method Pairwise relative counts and hierarchical elimination for learning algorithm.
result Near-optimal PAC sample complexity guarantee for identifying ε-optimal item.

Adaptive cascade submodular maximization tackles sequential selection under uncertainty.

problem Maximizing expected utility from a set of items with unknown states and continuation probabilities.
method Proposed adaptive cascade submodular functions and a 0.12 approximation algorithm.
result Identified a class of functions (adaptive cascade submodular) that many practical applications satisfy.

In this paper, we consider decentralized sequential decision making in distributed online recommender systems, where items are recommended to users based on their search query as well as their specific background including history of bought items, gender and age, all of which comprise the context information of the use…

2013-09-26abs ↗pdf ↗

Etsy uses novel embeddings to improve user recommendations based on item interactions.

problem Improving personalized recommendations for users based on diverse item interactions.
method Learning interaction-based item embeddings to encode co-occurrence patterns of item and interaction types.
result Taking interaction type into account improves user shopping behavior modeling accuracy.