Grapevine clusters wine reviews for personalized recommendations.
problem Providing personalized wine recommendations based on user preferences.
method Multi-dimensional clustering and unsupervised learning on wine reviews.
result Optimal wine recommendations based on user preference clusters and price-quality ratio.
RecoBERT uses a language model to recommend items from catalogs.
problem Harnessing language models for text-based item recommendations.
method RecoBERT is a BERT-based approach that learns specialized language models for item recommendations without requiring labeled data.
result RecoBERT outperforms other techniques in inferring item similarities from textual catalogs.
Compact E-Nose detects wine spoilage by acetic acid quickly.
problem Early detection of wine spoilage by acetic acid.
method Portable E-Nose with SnO2 sensors and deep MLP neural network trained approach.
result Classifies wine spoilage levels in 2.7 seconds.
Investment behavior in wine industry influenced by profitability and capitalization.
problem Exploring investment dynamics in wine industry from EU largest producers.
method Firm-level data from France, Italy, and Spain (2007-2014). Difference-and system-GMM estimators used.
result Profitability positively impacts investment dynamics, while capitalization negatively impacts only in France and Spain.
There are few papers about the consumption pattern of the Portuguese wine, using econometrics techniques. This work, pretend to analyze the consumers behavior of the wine produced in Portugal, determining the demand equation with panel data methods. There were used statistical data available in the Alentejo Regional Wi…
Proposes a method to cluster data with missing entries using non-convex penalties.
problem Challenges of clustering data with missing entries.
method Uses a relaxation of a ℓ0 fusion penalty optimization problem with saturating non-convex fusion penalties. result The proposed method produces solutions that degrade gradually with increasing missing feature values.
A deep learning method reduces chemometric data size and improves analysis accuracy.
problem The curse of dimensionality in chemometric data analysis.
method L2 regularized sparse autoencoder with automatic node selection and Gaussian process regression.
result Significant improvement in regression accuracy compared to state-of-the-art methods.
In the presence of weak overall correlation, it may be useful to investigate if the correlation is significantly and substantially more pronounced over a subpopulation. Two different testing procedures are compared. Both are based on the rankings of the values of two variables from a data set with a large number n of o…
Proposes a method to ensure low losses across all subpopulations in large datasets.
problem Standard practice of minimizing average loss fails to guarantee low losses across all subpopulations in heterogeneous datasets.
method Convex procedure that controls worst-case performance over all subpopulations of a given size with finite-sample convergence guarantees.
result Empirically, the worst-case procedure learns models that do well against unseen subpopulations.
Clustering is a separation of data into groups of similar objects. Every group called cluster consists of objects that are similar to one another and dissimilar to objects of other groups. In this paper, the K-Means algorithm is implemented by three distance functions and to identify the optimal distance function for c…
Develops a tool to measure gradual internationalization performance.
problem Lack of objective performance indicators for gradual internationalization.
method Quantitative tool based on export data, tested in Spanish wine sector.
result Creation of an international priority index for analyzing geographically differentiated strategies.
Regularization helps protect machine learning models from poisoning attacks.
problem Mitigating the impact of poisoned data on machine learning models.
method Distributionally-robust optimization using Wasserstein distance to find an upper bound for worst-case fitness.
result The regularizer is equal to the dual norm of the model parameters for regression models.
Proposes a method to generate counterfactual and contrastive explanations using SHAP.
problem Need for explainable AI and legal requirement for model interpretability.
method Model agnostic method using SHAP to generate contrastive and counterfactual explanations.
result Demonstrates effectiveness of the method on various datasets.
Develops a semi-analytic resampling method for Lasso regression.
problem Statistical fluctuations and high computational cost in Lasso resampling.
method Semi-analytic message passing algorithm based on state evolution analysis.
result Significant reduction in computational time and improved inference accuracy.
Quantum machine learning: Adiabatic quantum SVM outperforms classical methods.
problem Training support vector machines efficiently on large datasets.
method Adiabatic quantum computing for SVM training.
result Quantum approach outperforms classical methods in accuracy and scalability.
CA-NN uses neural networks to scale correspondence analysis.
problem Scaling correspondence analysis to large datasets.
method Reinterpreting CA as a functional optimization problem over finite variance functions, approximated by neural networks.
result CA-NN enables scalable correspondence analysis.
3D RadViz improves 3D data visualization of multidimensional datasets.
problem Tackles the challenge of visualizing multidimensional datasets in 3D.
method Develops RadViz3D, a 3D radial visualization tool with uniform anchor points.
result Improves the display of multidimensional datasets, especially for uncorrelated variables.
A new method detects and displays pairwise dependence between variates.
problem Detecting and visualizing dependence between variates of different types.
method Recursive random binning with approximations to Pearson's statistic.
result The method is well-calibrated and powerful against common test alternatives.
Python package for ordinal regression using gradient boosting.
problem Handling ordinal variables in machine learning.
method Gradient boosting with latent variable framework.
result Performs joint optimization of latent function and threshold vector.
Paper presents a novel approach to train deep neural networks using geometric and topological methods.
problem Training deep neural networks efficiently and effectively.
method Uses topological coverings and linear matrix inequalities to define neural network architecture.
result Constructive algorithm trains deep neural networks in one shot with equal or superior accuracy.
New financial ratios using compositional data improve analysis of firm health.
problem Statistical issues with standard financial ratios, especially skewness and outliers.
method Compositional data (CoDa) methodology to analyze financial statements.
result Outliers and skewness reduced, results invariant to numerator and denominator permutation.
Algorithm finds best Dirac mass approximation of target measure.
problem Finding optimal Dirac mass approximation of target measure.
method Minimizes statistical distance between original measure and quantized version using Huber-energy kernel.
result HEMQ algorithm robust and versatile, matches intuitive behavior.
The paper develops a method to estimate consumer preferences from observed rankings.
problem Estimating consumer preferences from partial ranking information.
method Interpreting observed rankings as pairwise comparisons, modeling latent utility, and correcting for selection bias.
result The method improves recommendation performance, especially for previously unconsumed products.
Dependent MMD coresets help compare multiple related datasets.
problem Comparing multiple related datasets for insights into model generalization.
method Dependent MMD coresets for collections of datasets.
result Dependent MMD coresets facilitate comparison and understanding of multiple related datasets.
Synthetic dataset for deep learning with known Gaussian distribution.
problem Lack of datasets with known distribution for deep learning verification.
method Proposes a method to generate a synthetic dataset with Gaussian distribution.
result Synthetic dataset mimics MNIST and can be used with DNNs.
Proposes a method to analyze distributed datasets without sharing original data.
problem Difficulty in centralizing large, distributed datasets due to size and privacy concerns.
method Centralizes intermediate representations instead of original datasets.
result Achieves higher prediction performance compared to individual analyses.
SCARY dataset generates complex causal scenarios for causality research.
problem Lack of complexity in existing causal datasets.
method Synthetic dataset with 40 scenarios, three seeds, and two data generation mechanisms.
result Provides a valuable resource for realistic causal discovery.
MusPy is a toolkit for symbolic music generation, providing tools for dataset management and analysis.
problem Facilitating the creation and analysis of symbolic music datasets.
method Development of an open-source Python library (MusPy) with features for dataset management, data I/O, preprocessing, and model evaluation. Demonstrated through statistical analysis and cross-dataset generalizability experiments.
result MusPy's dataset analysis reveals varying degrees of cross-genre representation across different music datasets.
New handwritten digits dataset for Kannada script.
problem Lack of datasets for Kannada numeral digits.
method Developed Kannada-MNIST and Dig-MNIST datasets.
result Initial CNN accuracy is lower than MNIST, indicating a challenge in generalization.
StyleDiff compares unlabeled datasets using disentangled image spaces.
problem Mismatches between development and real-world datasets lead to inaccurate predictions.
method Uses disentangled image spaces and focuses on attributes to compare datasets.
result Accurately detects and presents differences between datasets.
Paper introduces a new pedestrian dataset for adverse weather conditions.
problem Lack of controlled and annotated pedestrian datasets for adverse weather conditions.
method Presentation of a new dataset and baseline results for various machine learning tasks.
result Baseline results for various tasks on the Cerema AWP dataset.
New framework assesses graph-learning datasets for better evaluation.
problem Insufficient evaluation of graph-learning datasets and methods.
method Introduces Rings framework for dataset ablations and proposes performance separability and mode complementarity measures.
result Demonstrates utility of Rings framework for graph-learning dataset evaluation.
Improved dataset distillation for images and texts boosts model accuracy.
problem Reducing dataset size for faster and more energy-efficient model training.
method Simultaneous distillation of images and soft labels, extending to text datasets.
result 2-4% increase in accuracy for image classification tasks, 20% reduction in distilled samples.
Paper introduces ToyADMOS dataset for detecting anomalous machine sounds.
problem Lack of large-scale datasets for ADMOS anomaly detection.
method Collected anomalous sounds of miniature machines by deliberate damage.
result Released dataset includes over 180 hours of normal and 4,000 anomalous sounds.
Meta-Dataset benchmarks few-shot learning models with diverse datasets.
problem Lack of diverse and realistic datasets for evaluating few-shot learning models.
method Meta-Dataset: a new benchmark with diverse datasets and realistic tasks.
result Meta-Dataset uncovers important research challenges in few-shot learning.
Fairness GAN generates fair datasets for decision making.
problem Creating fair datasets for allocative decision making.
method Novel auxiliary classifier GAN aiming for demographic parity or equality of opportunity.
result Improves demographic parity and equality of opportunity in generated images.
MTL method uses unlabeled data with pseudo labels to improve classification with disjoint datasets.
problem Improving classification performance with disjoint labeled datasets using unlabeled data.
method Proposes MTL-SA method to select and augment unlabeled data with confident pseudo labels and close distribution to labeled data.
result Extensive experiments show the effectiveness of MTL-SA method in improving classification performance.
Study shows pruning datasets can improve machine learning model performance.
problem Improving machine learning model performance through dataset pruning.
method Comparison of different algorithms on unpruned and iteratively pruned datasets.
result Algorithms that perform better on unpruned datasets also perform better on pruned datasets.
The study examines dataset usage patterns in machine learning research.
problem Lack of attention to dataset dynamics in machine learning research.
method Analysis of dataset usage patterns across machine learning subcommunities and time periods (2015-2020).
result Increasing concentration on fewer and fewer datasets, significant adoption from other tasks, and concentration across the field on datasets introduced by elite institutions.
We create synthetic Morse code datasets for machine learning.
problem Creating challenging datasets for neural networks.
method Algorithm to generate synthetic Morse code datasets of varying difficulty.
result Network performance is affected by noise and feature set expansion.
Method embeds numeric tabular datasets into a shared vector space for similarity and retrieval.
problem Lack of meaningful representation for numeric tabular datasets in large language models.
method Structured exploratory data analysis descriptors, sentence transformer embedding, CCA for cross-dataset alignment.
result Total P@1 score of 0.9 across 15 datasets, robust nearest-neighbor retrieval and cluster structure.
Two large medical dialogue datasets for improving healthcare.
problem Improving healthcare through better dialogue systems.
method Building two large-scale medical dialogue datasets: MedDialog-EN and MedDialog-CN.
result The datasets are the largest medical dialogue datasets to date.
New framework transforms labeled datasets for various machine learning tasks.
problem Lack of principled methods to transform labeled datasets.
method Wasserstein gradient flows in probability space for optimization of data-generating distributions.
result Framework can impose constraints, adapt for transfer learning, or re-purpose models.
New dataset for industrial machine sounds to aid maintenance.
problem Lack of public datasets for industrial machine sounds.
method Recorded normal and anomalous sounds of industrial machines.
result Assists in automated facility maintenance development.
A dataset of 10 molecule types for machine learning studies.
problem Lack of suitable datasets for machine learning in molecular imaging.
method Generated 2D cross-sectional projections of 10 molecule types from Molecular Dynamics trajectories.
result Benchmark dataset for machine learning, deep learning, and image processing in scattering, imaging, and microscopy.
KIP meta-learning compresses datasets significantly.
problem Training data size and quality issues in machine learning.
method Kernel Inducing Points (KIP) for dataset compression.
result Significant reduction in dataset size with similar model performance.
Partial-input models fail to detect dataset artifacts, even when they perform poorly.
problem The effectiveness of partial-input models in detecting dataset artifacts is questionable.
method Design artificial datasets and identify trivial patterns in the SNLI dataset.
result Partial-input models can solve examples previously considered hard, indicating potential dataset artifacts.
Elastic co-clustering improves clustering of single-cell genomic data.
problem Improving clustering performance of single-cell genomic datasets.
method Elastic coupled co-clustering in an unsupervised transfer learning framework.
result Our algorithm significantly improves clustering performance over traditional methods.