Transforms complex tactile data into compact representations.
problem Complexity of tactile data prevents widespread use of tactile sensors.
method Uses unsupervised learning to transform data into a compact latent representation.
result Compact representations can be used directly in controllers or for calibration.
Proposes a deep Auto-Encoder-like framework for visual-tactile fusion object clustering.
problem Combining visual and tactile information for better object clustering.
method Deep Auto-Encoder-like Non-negative Matrix Factorization framework, graph regularizer, modality-level consensus regularizer, alternating minimization strategy.
result Improves object clustering performance by leveraging both visual and tactile modalities.
Graph neural network predicts grasp stability from tactile sensor data.
problem Predicting grasp stability from tactile sensor data.
method Graph Convolutional Network (GCN) trained on tactile sensor data.
result Graph neural network effectively predicts grasp stability.
Robot learns to grasp and adjust using vision and touch.
problem Robotic grasping relies solely on visual input, missing tactile feedback.
method End-to-end action-conditional model that learns from visuo-tactile data.
result Model predicts grasp adjustment outcomes and selects efficient actions.
Emotion classification improved using brain signals from tactile enhanced multimedia.
problem Classifying viewer emotions in tactile enhanced multimedia.
method Frequency domain features from EEG data analyzed using SVM.
result Increased accuracy (76.19%) compared to time domain features (63.41%).
Robotic grasp stability improved with fingertip slippage detection.
problem Improving grasp stability in robotic manipulation.
method Task-relevant feature extraction and efficient classifier design for fingertip slippage detection.
result The proposed method effectively detects object slippage with fingertips in an online fashion.
DIGIT is a low-cost tactile sensor for in-hand manipulation.
problem Difficulty in sensing contact forces limits robotic manipulation.
method DIGIT miniaturizes and improves a vision-based tactile sensor.
result DIGIT enables better control of interactions with the environment.
Touch sensing improves grasp prediction accuracy.
problem Predicting grasp outcomes from indirect measurements like vision is challenging.
method Investigated touch sensing's value in multimodal grasping using visuo-tactile deep neural networks.
result Tactile readings significantly improve grasp prediction accuracy.
TACTO simulates high-resolution touch sensing for robotics.
problem Accurate simulation of touch sensing in robotics.
method Fast, flexible, open-source simulator for vision-based tactile sensors.
result Demonstrated TACTO's effectiveness in grasping stability prediction and marble manipulation control.
Maximizes mutual info across views for better image representations.
problem Improving image representation learning through multiple views.
method Maximizing mutual information between features from multiple views.
result ImageNet accuracy of 68.1% using linear evaluation, significantly outperforming prior methods.
Deep learning framework predicts surface texture parameters and their uncertainties.
problem Predicting surface texture parameters and their uncertainties from multi-instrument datasets.
method Reproducible deep learning framework using multi-instrument dataset, quantile and heteroscedastic heads for uncertainty modeling, and post-hoc conformal calibration.
result High fidelity predictions (R2: Ra 0.9824, Rz 0.9847, RONt 0.9918) and well-modelled uncertainty targets (Ra_uncert 0.9899, Rz_uncert 0.9955).
Proposes RSP model for efficient big data analysis.
problem Efficiently partitioning big data sets for analysis.
method Random sample partition (RSP) data model and block-level sampling.
result RSP data blocks can estimate statistics and build models equivalent to whole data set.
Data preprocessing improves data quality for robust data mining.
problem Noisy and incomplete data hinders data mining models.
method Overview of data cleaning, transformation, and preprocessing methods.
result Preprocessing significantly affects data mining model performance.
A new method for handling imbalanced big data using ensembles and smart data.
problem Imbalanced data distribution in big data scenarios.
method Smart Data driven Decision Trees Ensemble (SD_DeTE) methodology.
result SD_DeTE outperforms Random Forest in handling imbalanced binary classification problems in big data.
Prevents sensitive data generation in diffusion models using labeled and unlabeled data.
problem Generating sensitive data in diffusion models using unlabeled data.
method Positive-Unlabeled Diffusion Models, approximating ELBO with labeled and unlabeled data.
result Prevents the generation of sensitive data without compromising image quality.
Study reveals Data Shapley's inconsistent performance in data selection tasks.
problem Inconsistency of Data Shapley's performance in data selection across different settings.
method Hypothesis testing framework and identification of utility functions.
result Data Shapley's performance is no better than random selection without specific constraints.
Survey on data collection challenges in machine learning.
problem Data scarcity and need for labeled data in machine learning.
method Comprehensive study of data acquisition, labeling, and improvement techniques.
result Identification of research challenges in data collection.
PRRO generates synthetic tabular data that improves SL performance and class distribution.
problem Low SL utility of synthetic data due to class imbalance and overlooked data relationships.
method Data pruning and column reordering to optimize SL utility.
result Synthetic data generated with PRRO enhances predictive performance and class distribution.
Defines data science as a natural ecosystem with challenges and missions.
problem Challenges and missions in data science due to 5D complexities and data life cycle phases.
method Systemic and data-centric view of data science as a fusion of data universe and its challenges, formalizing a general-purpose architecture.
result Essential data science as a natural ecosystem integrating specific disciplines and high-impact applications.
Synthetic data enhances analytics but requires careful volume management.
problem Accuracy of statistical methods on synthetic data vs. raw data.
method Synthetic Data Generation for Analytics framework using tabular diffusion models.
result Error rate decreases with more synthetic data but may stabilize or increase.
Data science redefines causal inference from observational data, classifying tasks into description, prediction, and counterfactual prediction.
problem Widespread misunderstandings about data science's role in causal inference from observational data.
method Organizing data science tasks into three classes: Description, prediction, and counterfactual prediction (including causal inference).
result The necessity of subject-matter expert knowledge for causal analyses in data science.
This paper evaluates how dirty data affects data mining and machine learning results.
problem Negative impacts of dirty data on data mining and machine learning results.
method Experimental comparison of missing, inconsistent, and conflicting data on classification and clustering algorithms.
result Guidelines for algorithm selection and data cleaning based on experimental findings.
DPASF stream preprocesses Big Data streams efficiently.
problem Efficient preprocessing of streaming Big Data.
method Implemented six preprocessing algorithms in Apache Flink.
result Preprocessing improves data accuracy in streaming Big Data.
This paper introduces C-DSL to improve data mining outcomes by considering context.
problem Data collection ambiguities, data imbalance, hidden biases, lack of domain info, and data incompleteness.
method Developed Context-Driven Data Science Lifecycle (C-DSL) to address data quality issues.
result Tangible improvements to data mining outcomes were achieved through C-DSL.
Proposes using probabilistic models for privacy-preserving synthetic data.
problem Designing high-quality synthetic data for privacy preservation.
method Formulate the problem through probabilistic modelling, choosing a model for the data.
result Statistical discoveries can be reliably reproduced from synthetic data.
Unlabeled data helps stop active learning better than labeled data.
problem Reducing the need for manual annotation in text classification.
method Compared stopping methods based on labeled, unlabeled, and training data.
result Stopping methods using unlabeled data are more effective.
New test ensures quality of shared data in machine learning.
problem Ensuring quality of external data in machine learning tasks.
method Distribution-free two-sample testing procedures grounded in conformal outlier detection.
result Identifies valuable external data agents for model personalization.
Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.
problem Ensuring fair predictions across sensitive attributes in synthetic data.
method Equalizing target probability distributions across sensitive attributes in synthetic data generation.
result Synthetic data provides strong fair predictions, equal across all thresholds.
A new method classifies multiple correlated data streams simultaneously.
problem Classifying multiple correlated data streams in practical scenarios.
method Double-Coupling Support Vector Machines (DC-SVM) considers both internal and external correlations.
result The proposed method outperforms traditional methods on artificial and real-world data streams.
This paper improves neural machine translation training by selecting and denoising data.
problem Reduces negative impact of noisy data on neural machine translation training.
method Measures and selects domain data, applies denoising curriculum using online data selection.
result Significant effectiveness for training on noisy data.
DPA preserves data distribution in reduced dimensions.
problem Loss of data distribution in dimension reduction.
method DPA combines encoder and decoder to match data distribution.
result DPA successfully reconstructs data distribution.
Framework captures missing data in sparse data sets.
problem Capturing missing data in extremely sparse data sets.
method Coupled compound Poisson factorization with stochastic variational inference.
result Explicitly modeling missing data improves results in clustering, prediction, and matrix factorization.
Efficient synthetic data generation improves model performance on tabular data.
problem Improving model robustness and performance with scarce or low-quality data.
method Hardness characterization to identify high-value training points, generating synthetic data only from these points.
result Synthetic data generated from hardest points outperforms non-targeted methods on tabular datasets.
For most problems in science and engineering we can obtain data sets that describe the observed system from various perspectives and record the behavior of its individual components. Heterogeneous data sets can be collectively mined by data fusion. Fusion can focus on a specific target relation and exploit directly ass…
DAERNN models censored data using neural networks with data augmentation.
problem Handling censored data in expectile regression.
method Data augmentation based Expectile Regression Neural Networks (ERNNs).
result DAERNN outperforms existing censored ERNNs methods and achieves comparable predictive performance to fully observed data.
This paper quantifies uncertainty in Data Shapley using statistical inference.
problem Uncertainty in data valuation due to dynamic data distribution.
method Established relationship with U-statistics and quantified uncertainty using statistical inference.
result Confidence intervals for Data Shapley estimations are provided.
Generative Adversarial Networks create time series data from images.
problem Generating realistic time series data from images.
method Wasserstein GANs with gradient penalty for stability, synthesizing sinusoidal, PPG, and ECG data.
result Successfully generated time series data using image-based GANs.
DCoM uses deep neural networks to detect semantic data types from raw column values.
problem Detecting semantic data types from dirty and unseen data.
method DCoM employs multi-input NLP-based deep neural networks trained on 686,765 data columns.
result DCoM outperforms existing methods significantly on 78 different semantic data types.
Model refines coarse spatial data using diverse auxiliary data sets.
problem Tackles the challenge of refining coarse spatial data with varying auxiliary data granularities.
method Proposes a probabilistic model using Gaussian processes to hierarchically incorporate auxiliary data sets of various granularities.
result Can effectively refine coarse-grained spatial data using auxiliary data sets of different granularities.
GANs generate training data for machine learning tasks.
problem Imbalanced data sets and sensitive information.
method Generative Adversarial Networks (GANs) to create artificial training data.
result A Decision Tree classifier trained on GAN-generated data achieved similar or better accuracy and recall than on original data.
Task-agnostic data valuation without validation requirements.
problem Valuing data without specific task assumptions.
method Estimating data diversity and relevance through queries without raw data.
result Estimates capture the diversity and relevance of seller's data for the buyer.
Framework generates private synthetic data for unlabeled mixed-type data.
problem Generating private synthetic data for unlabeled mixed-type data.
method Combines autoencoders and GANs for differential privacy.
result Learned model generates synthetic data with similar statistical properties.
New algorithm improves data imputation for complex multimodal data sets.
problem Artifacts in imputation methods for multimodal distributions.
method Combines kNN and KDE for probabilistic estimates. result Lower imputation errors and higher likelihood estimates.
New algorithms for clustering and synthetic data generation of heterogeneous tabular datasets.
problem Clustering and generating synthetic data from heterogeneous tabular datasets with hidden cluster structure.
method Developed MMM and MMMsynth algorithms for clustering and synthetic data generation.
result MMMsynth algorithm outperforms other literature tabular-data generators and approaches real data performance.
WeMix improves data augmentation by correcting bias in deep learning.
problem Data augmentation's effectiveness is limited by data bias.
method Developed AugDrop and MixLoss algorithms to correct data bias.
result WeMix improves data augmentation performance through bias correction.
A new method reduces data valuation variance for more trustworthy data trading.
problem Data valuation and trustworthy data trading in algorithmic prediction.
method Variance reduced Shapley value estimation using stratified sampling.
result VRDS method reduces estimation variance and improves data marketplace development.
Develops a new metric to equitably value data for machine learning models.
problem Equitable valuation of individual data in machine learning predictions.
method Data Shapley framework, Monte Carlo and gradient-based methods.
result Data Shapley uniquely satisfies properties of equitable data valuation.
Adapts data analysis for growing data, improving generalization guarantees.
problem Challenges of overfitting and statistical validity in adaptive workflows with growing data.
method Generalizes adaptive analysis on dynamic data, incorporating time-varying empirical accuracy bounds and mechanisms.
result First generalization bounds for adaptive analysis on dynamic data, matching prior works' improvement over data splitting.