Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Jan 199519922001200920182026
48 results for data products

Model learns product vectors from baskets and browsing sessions for better complementary product recommendations.

problem Inferring complementary products from basket and browsing data.
method Proposes BB2vec model that learns product vectors from both baskets and browsing sessions.
result The BB2vec model improves complementary product recommendations and alleviates the cold start problem.

Unified product embeddings improve cross-task performance in e-commerce.

problem Training product embeddings in isolation limits cross-task performance.
method Combining text, clickstream, and image data using denoising auto-encoders, BPR, and Siamese neural networks.
result Unified product embeddings uniformly outperform isolated embeddings across three e-commerce tasks.

Dynamic pricing improves sales of low-sale products using online clustering.

problem Low-sale products lack sufficient data for traditional dynamic pricing algorithms.
method Online clustering of product demand and dynamic pricing decisions based on cluster analysis.
result The proposed algorithms significantly outperform traditional single-product pricing policies in increasing revenue.

Linear classifiers in product space forms improve scRNA-seq data classification.

problem Linear classification in products of Euclidean, spherical, and hyperbolic spaces.
method Novel formulations of linear classifiers on Riemannian manifolds, proving expressive power, and formalizing perceptron and SVM classifiers.
result Linear classifiers in product space forms have the same expressive power as in Euclidean space of the same dimension.

Study uses network analysis to rank countries based on food production diversity and specialization.

problem Ranking countries based on their food production diversity and specialization.
method Network analysis of country-food production data, transforming into overlap matrices, identifying subsets, and ranking based on fitness and specialization.
result Countries with high fitness produce highly specialized food commodities, while those with low fitness produce diverse but low specialized food products.

ASOS improves fashion product recommendations by learning from unstructured data.

problem Lack of consistent product information in e-commerce.
method Developed a hybrid recommender system to learn product attributes from unstructured data.
result Quantitative understanding of products improves recommendation accuracy.

New model predicts global oil production and consumption through 2050.

problem Inaccurate past oil production forecasts leading to public interest.
method Analyzes past regional oil production data to predict future production and consumption.
result Predicts global oil production and consumption through 2050, highlighting limited potential for unconventional oil.

Improved 3D LiDAR data classification using product coefficients.

problem Enhancing accuracy in 3D LiDAR data classification.
method Introducing product coefficients derived from measure theory as additional features in the classification process, alongside PCA.
result Significant improvement in classification accuracy with product coefficients.

Improved sales forecasting for new products using transfer learning.

problem Insufficient training data for new products leads to inaccurate sales forecasts.
method Network-based Transfer Learning approach for deep neural networks.
result Deep neural networks' prediction accuracy for food sales forecasting can be effectively increased.

The paper tackles imbalance in production data by proposing sampling methods to improve model performance on underrepresented observations.

problem Imbalance in production data negatively impacts model predictive performance on underrepresented observations.
method Three sampling approaches are investigated to adjust for imbalance in training data and improve model performance.
result Fitting a model using sampled data yields a small reduction in overall predictive performance but a better performance on underrepresented observations.

Framework for pricing data products in data-poor markets.

problem Challenges in pricing advanced data products due to lack of transaction data.
method Prior-predictive Monte Carlo framework for generating probabilistic price bands.
result Stable probabilistic price bands for data products in data-poor markets.

Proposes using entity embedding vectors to improve Gaussian Process models for knowledge transfer across cell lines.

problem Lack of reuse of experimental data for predicting novel processes.
method Hybrid Gaussian Process models with entity embedding vectors to represent product identity.
result Improved performance in predicting novel processes compared to traditional methods.

The paper improves consumer preference modeling by considering multiple product categories.

problem Estimating consumer preferences across multiple product categories with varying attributes and price sensitivity.
method Extends matrix factorization techniques to account for time-varying product attributes and out-of-stock products, pooling information across categories to estimate heterogeneity in preferences.
result The model improves over traditional approaches, accurately estimating consumer preferences and price sensitivity.

Paper tackles product categorization with structured and unstructured attributes for large-scale eCommerce.

problem Challenges in categorizing products with thousands of classes and millions of products.
method Compares hierarchical and flat models, uses Deep Learning for feature extraction, combines structured and unstructured attributes.
result Flat models perform better in specific cases, and the proposed approach handles faulty attribute names and values.

The paper proposes a machine learning approach for production forecasting without model calibration.

problem Generating accurate production forecasts for reservoir development.
method Sequential model aggregation using machine learning algorithms without model calibration.
result The proposed method provides robust multi-step-ahead production forecasts.

Adaptive time decay functions improve financial product recommendation accuracy.

problem Inaccurate recommendations due to static historical data in finance.
method Time-dependent collaborative filtering with personalized decay functions.
result Significant improvements over state-of-the-art benchmarks in financial product recommendation.

Combines RNNs and tensor products for sequential data, outperforming state-of-the-art.

problem Improving symbolic interpretation and systematic generalization in natural language reasoning.
method End-to-end training of a recurrent neural network architecture with tensor product representations.
result Significantly outperforms state-of-the-art models in natural language reasoning tasks.

New method uses product embeddings to predict bundle success.

problem Designing effective product bundles in large retail settings.
method Leverage historical purchases and clickstream data to generate product embeddings, then use heuristics for complementarity and substitutability.
result Embeddings-based heuristics predict bundle success, robust across categories and retailers.

ProductNet curates high-quality product datasets for better product understanding.

problem Lack of high-quality product datasets for product representation learning.
method Curated high-quality product datasets with a multi-modal deep neural network and active learning.
result Master model yields high categorization accuracy (94.7% top-1 accuracy for 1240 classes).

Framework for monitoring ML model performance in production without labels.

problem Monitoring real-time prediction quality of ML models in production without labels.
method ML Health framework using diagnostic methods to generate alerts for further investigation.
result The method outperforms standard distance metrics at detecting issues with mismatched data sets.

Paper proves translating solutions for a specific flow in a product manifold.

problem Existence of translating solutions for nonparametric mean curvature flow with Neumann boundary data.
method Proves existence using product manifold MnimesRM^{n} imes\mathbb{R} with specific conditions.
result Existence of translating solutions for the flow in the product manifold.

This study analyses, through cross-section estimation methods, the influence of spatial effects in the conditional product convergence in the parishes' economies of mainland Portugal between 1991 and 2001 (the last year with data available for this spatial disaggregation level). To analyse the data, Moran's I statistic…

2011-10-25abs ↗pdf ↗

Link prediction improves future product export forecasts in the international trade network.

problem Predicting future evolution of complex systems like international trade networks.
method Applied link prediction algorithms based on heat and mass diffusion processes, improved with country fitness and product similarity metrics.
result Best results achieved with a new product similarity metric that considers causality.

Paper proposes a new method to optimize feature coordinates for better image classification.

problem Improving feature extraction for better machine learning classification.
method Mutual-energy inner product optimization method.
result The method enhances low-frequency features and suppresses high-frequency noise, leading to better classification results.

New product structures encode superintegrable Hamiltonian systems in Euclidean spaces.

problem Encoding superintegrable Hamiltonian systems using product structures.
method Introducing commutative and associative product structures on Euclidean spaces of dimension at least three, satisfying specific conditions.
result All abundant superintegrable Hamiltonian systems on Euclidean space of dimension at least three arise from these product structures.

The paper proposes a new method for product recommendation that considers revenue contributions and user similarity.

problem High dimensionality and sparsity in user-item data, especially in terms of revenue contributions.
method The approach encodes revenue contributions in the user-item matrix and computes customer similarity using suitable distance measures.
result The method segments users based on revenue-based similarity and supports recommendations aligned with profitability objectives.

Unified theory for neural scaling laws in hierarchically compositional data.

problem Understanding neural scaling laws in hierarchically compositional data.
method Probabilistic context-free grammars and power-law distributed production rules.
result Unified learning curve behavior for classification and next-token prediction tasks.

Paper proposes a new method for approximating high-dimensional matrices using Kronecker products.

problem Discovering low-dimensional structure in high-dimensional data.
method Hybrid Kronecker Product Approximation (hKoPA) and estimation procedures.
result The proposed methods provide flexible and effective dimension reduction.

Bayesian networks improve product risk assessment by handling uncertainty and causality.

problem Limited handling of uncertainty and inability to incorporate causal explanations in existing methods.
method Bayesian Networks (BNs) for improved systematic product risk assessment.
result BN approach provides more powerful and flexible risk assessments.

This work characterizes topological descriptors of graph products and their expressive power.

problem Capturing multiscale structural information in graph products using topological descriptors.
method Analysis of various filtrations on graph products, including Euler characteristic and persistent homology.
result Persistent homology of graph products contains more information than individual graphs.

The paper optimizes exceptions in a statistical production system using machine learning.

problem Lack of curated and labeled training data for machine learning in data quality assurance.
method Explainable supervised machine learning to identify and prioritize exceptions.
result Improvement in the quality and efficiency of exceptions generated and authenticated by users.

SHOPPER models consumer choices with substitutes and complements.

problem Understanding how consumers choose products with interactions.
method Sequential probabilistic model with interpretable components for price interventions.
result SHOPPER accurately predicts consumer choices and identifies product pairs.

Modified deep LSTM model with EnKF improves gas flow rate predictions in mature gas wells.

problem Predicting gas production from mature gas wells with high accuracy and robustness.
method Modified deep LSTM model for flow rate prediction and EnKF for updating predictions.
result EnKF updated model leads to better flow predictions with lower Jeffreys' J-divergences.