System optimizes product images for e-commerce, enhancing customer engagement.
problem Optimizing product images for e-commerce to improve customer engagement.
method Machine learning, deep learning, and computer vision techniques applied to large e-commerce catalogs.
result System produces superior image sets tailored to customer preferences.
System solves author name ambiguity in e-commerce catalogs.
problem Finding correct author names in e-commerce catalogs with abbreviations and spelling variants.
method Composite system using open data sources and machine learning techniques for natural language processing.
result Top proposal of the system is the normalized author name with 72% accuracy.
Method retrieves similar fashion items from images and text, enabling style refinement.
problem Lack of intuitive, interactive refinement in search engines for fashion items.
method Joint visual-textual embedding training, Mini-Batch Match Retrieval, attribute extraction.
result Improved performance in multimodal style search, demonstrated through benchmark.
AI classifies tourist events for better service.
problem Classifying tourist events efficiently across diverse sources.
method CRISP-DM, supervised machine learning, NLP.
result Automatic classification tool for consistent event catalogs.
Method finds reference products for a given item.
problem Finding relevant products for a given item.
method Product representation learning and fingerprint-type vector searching.
result The method outperforms peer services in search return rate and precision.
Paper tackles circularity issues in machine learning predictions.
problem Circularity problems in machine learning predictions.
method Not specified in the abstract.
result Not specified in the abstract.
Making mock simulated catalogs is an important component of astrophysical data analysis. Selection criteria for observed astronomical objects are often too complicated to be derived from first principles. However the existence of an observed group of objects is a well-suited problem for machine learning classification.…
We present an automatic classification method for astronomical catalogs with missing data. We use Bayesian networks, a probabilistic graphical model, that allows us to perform inference to pre- dict missing values given observed data and dependency relationships between variables. To learn a Bayesian network from incom…
The study uses historical revenue data to forecast music catalog cashflows and multipliers.
problem Valuation of music catalogs based on historical revenue data.
method Risk-neutral approach using discounted cashflows formula.
result Ask prices are close to multipliers justified by median song cashflows, while best bids are near multipliers justified by bottom decile cashflows.
OpenTag extracts missing attribute values from product descriptions.
problem Extract missing attribute values from product descriptions.
method Developed a deep tagging model OpenTag using LSTM and CRF, with an attention mechanism and active learning.
result OpenTag discovers new attribute values with minimal human annotation, achieving high F-score.
We present a new, fully generative model for constructing astronomical catalogs from optical telescope image sets. Each pixel intensity is treated as a random variable with parameters that depend on the latent properties of stars and galaxies. These latent properties are themselves modeled as random. We compare two pro…
ProductNet curates high-quality product datasets for better product understanding.
problem Lack of high-quality product datasets for product representation learning.
method Curated high-quality product datasets with a multi-modal deep neural network and active learning.
result Master model yields high categorization accuracy (94.7% top-1 accuracy for 1240 classes).
We consider the "intrinsic" symmetry group of a two-component link L, defined to be the image Σ(L) of the natural homomorphism from the standard symmetry group $\MCG(S^3,L)$ to the product $\MCG(S^3) \cross \MCG(L)$. This group, first defined by Whitten in 1969, records directly whether L is isotopic to a link $L…
CHARM creates mock halo catalogs from dark matter density fields using neural networks.
problem Creating accurate mock halo catalogs for cosmological studies is computationally expensive.
method CHARM uses multi-stage neural spline flow networks to learn the mapping from dark matter density fields to halo catalogs.
result Mock halo catalogs have the same statistical properties as those from high-resolution N-body simulations.
Machine learning detects type Ia supernovae from photometric data.
problem Detecting type Ia supernovae accurately from photometric data.
method Machine learning approach using only real observation data.
result Good results on real data from the Open Supernovae Catalog.
Unified product embeddings improve cross-task performance in e-commerce.
problem Training product embeddings in isolation limits cross-task performance.
method Combining text, clickstream, and image data using denoising auto-encoders, BPR, and Siamese neural networks.
result Unified product embeddings uniformly outperform isolated embeddings across three e-commerce tasks.
BLISS detects and separates astronomical sources quickly and accurately.
problem Detecting and separating overlapping astronomical sources in large images.
method Bayesian Light Source Separator (BLISS) using deep generative models and variational inference.
result BLISS can process megapixel images in seconds and produce highly accurate catalogs.
Atlas dataset categorizes clothing products with high accuracy.
problem Lack of real-world datasets for e-commerce clothing product categorization.
method Collected and labeled a dataset of 186,150 images, established a benchmark for image classification and sequence models.
result Benchmark model achieved a micro f-score of 0.92.
Celeste is a procedure for inferring astronomical catalogs that attains state-of-the-art scientific results. To date, Celeste has been scaled to at most hundreds of megabytes of astronomical images: Bayesian posterior inference is notoriously demanding computationally. In this paper, we report on a scalable, parallel v…
RecoBERT uses a language model to recommend items from catalogs.
problem Harnessing language models for text-based item recommendations.
method RecoBERT is a BERT-based approach that learns specialized language models for item recommendations without requiring labeled data.
result RecoBERT outperforms other techniques in inferring item similarities from textual catalogs.
EQShapelets detect earthquakes with high accuracy and interpretability.
problem Automated detection and cataloging of earthquakes.
method Time-series shape-based approach embedded in machine learning.
result EQShapelets detected all cataloged and 281 uncataloged events with lower false detection rate.
Machine learning separates Gaia DR2 stars into accreted and in-situ categories.
problem Identifying accreted stars from Gaia DR2 data using limited 5D kinematics.
method Transfer learning on mock Gaia catalogs, then cross-matched with real data.
result Identifies 767,000 accreted stars within Gaia DR2.
We present a large catalog of optically selected galaxy clusters from the application of a new Gaussian Mixture Brightest Cluster Galaxy (GMBCG) algorithm to SDSS Data Release 7 data. The algorithm detects clusters by identifying the red sequence plus Brightest Cluster Galaxy (BCG) feature, which is unique for galaxy c…
Determinantal point processes (DPPs) have garnered attention as an elegant probabilistic model of set diversity. They are useful for a number of subset selection tasks, including product recommendation. DPPs are parametrized by a positive semi-definite kernel matrix. In this work we present a new method for learning th…
Analog methods improve forecast accuracy in complex models.
problem Improving forecast accuracy in complex models like Lorenz-96.
method Constructing analogs using variational autoencoders for ensemble data assimilation.
result Constructed analogs perform as well as a full ensemble square root filter.
New algorithm reduces cold-start costs in multi-armed bandits for many products.
problem High burn-in costs in multi-armed bandits for new products.
method Two-phase bandit algorithm using subsampling and low-rank matrix estimation.
result Reduces burn-in costs and expedites experiment in large product sets.
Learning attribute applicability of products in the Amazon catalog (e.g., predicting that a shoe should have a value for size, but not for battery-type at scale is a challenge. The need for an interpretable model is contingent on (1) the lack of ground truth training data, (2) the need to utilise prior information abou…
The paper evaluates the probability distributions of analog-to-target distances for multiple analogs.
problem Understanding the performance of analog applications through the distribution of distances to target states.
method Theoretical analysis and numerical experiments using dynamical systems theory.
result The size of the catalog and dimensionality affect the probability distributions of the K-best analogs.
Bayesian method finds voids in galaxy surveys with deep neural networks.
problem Finding genuine matter underdensities in sparse galaxy surveys is underconstrained.
method Deep graph neural network evolves 'test particles' to sample from stochastic void definitions.
result Trained model performs well and finds Bayes-optimal void mappings.
Kolmogorov-Arnold network improves GW catalog posterior construction.
problem Efficiently constructing posterior distributions for GW catalogs.
method Using the Kolmogorov-Arnold network to create lightweight neural density estimators.
result Kolmogorov-Arnold network achieves superior interpretability and accuracy in posterior construction.
Developing a visual platform for faster astronomical source cataloging.
problem Speeding up cataloging of large area surveys in radio astronomy.
method Integration of advanced source finding and classification tools into a visual analytic platform.
result Improvement and acceleration of cataloging process in astronomical surveys.
Paper speeds up policy optimization for large recommendation systems.
problem Offline optimization of large-scale recommendation systems is computationally expensive.
method Derives an approximation of policy learning algorithms that scales logarithmically with the catalogue size.
result Our algorithm is an order of magnitude faster than naive approaches while producing equally good policies.
Results of the application of pattern recognition techniques to the problem of identifying Giant Radio Sources (GRS) from the data in the NVSS catalog are presented and issues affecting the process are explored. Decision-tree pattern recognition software was applied to training set source pairs developed from known NVS…
Deep models predict missing product attributes from text and images.
problem Incomplete or missing product attributes in e-commerce catalogs.
method Combining textual and visual data with a novel modality-merging method.
result Our approach improves attribute prediction on Rakuten-Ichiba and other datasets.
TUV Austria proposes certification for ML applications to ensure reliability.
problem Ensuring trust in AI applications to meet societal reliance requirements.
method Holistic approach analyzing security, functionality, data quality, ethics, and criticality levels.
result Certification process for low-risk ML applications in supervised learning.
PICZL improves photometric redshifts for AGN in all-sky surveys.
problem Challenges in accurately computing photo-z for AGN due to interplay of SMBH and host galaxy emissions.
method PICZL uses an ensemble of CNNs with cross-channel integration of image and catalog data, leveraging Gaussian mixture models.
result PICZL achieves a photo-z variance of 4.5% and outlier fraction of 5.6% on a validation sample of 8098 AGN, outperforming previous methods.
Paper tackles invoice line item matching in P2P processes using agent feedback.
problem Matching product/service descriptions in invoices with purchase orders.
method Two approaches using agent feedback data: similarity ranking and classification.
result Proposed approaches outperform benchmarks and real-world data sets.
Determinantal point processes (DPPs) are an elegant model for encoding probabilities over subsets, such as shopping baskets, of a ground set, such as an item catalog. They are useful for a number of machine learning tasks, including product recommendation. DPPs are parametrized by a positive semi-definite kernel matrix…
Researchers analyze tagging patterns on Stack Exchange communities.
problem Understanding the structure and evolution of tags in Q&A platforms.
method Empirical analysis and development of a generative model for tag co-occurrence.
result The model can reproduce statistical properties of co-tagging graphs.
Framework learns item representations from text data for complementary and similar items.
problem Generating accurate complementary item recommendations from textual data.
method Quadruplet network learning framework for latent space representation of items.
result Items are placed closer together in latent space for similar and complementary items compared to non-complementary items.
This paper introduces two-dimensional diagrams that are slight generalizations of moment map images for toric four-manifolds and catalogs techniques for reading topological and symplectic properties of a symplectic four-manifold from these diagrams. The paper offers a purely topological approach to toric manifolds as w…
Automatically identifies RRLyrae stars from VVV survey data.
problem Classifying RRLyrae stars from a large dataset of light curves.
method Developed an automatic ML-based procedure to identify RRLs, using features like period and intensity, and pseudo-colors.
result Constructed an ensemble classifier with Recall of 0.48 and Precision of 0.86 over 15 tiles.
Paper determines 2-adjacent knots up to 12 crossings.
problem Identifying 2-adjacent knots with up to 12 crossings.
method Used Heegaard Floer d-invariants and Alexander polynomial to obstruct 2-adjacency.
result Proved conjectures about 2-adjacent knots by Ito and Kato.
Improved cover detection in music datasets with novel triplet loss.
problem Challenging task of automatically detecting covers in audio datasets.
method Convolutional neural network mapping melodic features to embeddings, training to minimize cover distance and maximize non-cover distance.
result New prototypical triplet loss improves accuracy for large datasets and live songs.
This paper maps the insurability of AI risks across various insurance products.
problem Emerging AI risks and their implications for insurance coverage.
method Coding 55 AI threat classes against 26 insurance products using public carrier materials and threat catalogs.
result Identification of a four-tier insurability frontier: affirmatively insured, silent-AI exposures, actively excluded, and unstructured perils.
Taxonomy for ML in simulations, covering patterns and algorithms.
problem Enhancing simulations using machine learning.
method Presentation of eight patterns and three algorithmic areas.
result Catalog of activities and patterns for ML integration in simulations.
In this paper, we work to construct mosaic representations of knots on the torus, rather than in the plane. This consists of a particular choice of the ambient group, as well as different definitions of contiguous and suitably connected. We present conditions under which mosaic numbers might decrease by this projection…
Identifies LA-groups via VB-group structure and complementary actions.
problem Understanding the structure and integrability of LA-groups.
method Identifies LA-groups via VB-group structure and complementary actions up to homotopy.
result Establishes an equivalence between LA-groups and LA-matched pairs.