Study shows dataset properties impact adversarial machine learning robustness.
problem Vulnerability of DNNs to adversarial attacks.
method Examined five datasets, analyzed input size and contrast effects.
result Input size and contrast significantly influence adversarial success.
Deep models generate geometric objects with global properties.
problem Comparing neural models' global properties from generated samples.
method Training on datasets of reflexive polytopes, comparing different representations.
result Models learn non-trivial global properties of geometric objects.
Diffusion Variational Autoencoders capture topological properties of datasets.
problem Standard VAEs struggle with topological properties of certain datasets.
method Introduces Diffusion VAEs with transition kernels of Brownian motion on arbitrary manifolds.
result Diffusion VAEs can capture topological properties of synthetic datasets.
Study links different drug property prediction methods and datasets.
problem Unclear which method is best for drug property prediction.
method Empirical study linking multiple datasets and methods.
result Best method depends on dataset, classical methods often outperform deep learning.
Multi-party machine learning leaks global dataset properties even with black-box access.
problem Leakage of global dataset properties in multi-party machine learning.
method Demonstrated leakage of sensitive attribute distributions in pooled data.
result A curious party can infer sensitive attribute distributions in other parties' data with high accuracy.
Study on protecting sensitive properties of datasets during analysis.
problem Ensuring privacy of sensitive properties in datasets.
method Proposes definitions and mechanisms for attribute privacy using the Pufferfish framework.
result Developed efficient and inefficient mechanisms for attribute privacy.
Meta-learning symbolic default hyperparameters from dataset properties.
problem Empirical hyperparameter optimization is slow and requires manual configuration.
method Evolutionary algorithm to learn symbolic hyperparameter formulas from dataset properties.
result Meta-learning finds viable symbolic defaults for ML algorithms.
Paper introduces attacks to infer GAN training dataset properties.
problem Security and privacy risks of generative models like GANs.
method Proposes a general attack pipeline for two attack scenarios.
result Demonstrates strong performance in inferring GAN training dataset properties.
Paper improves molecular property prediction using denoising autoencoders.
problem Limited data for molecular property prediction from 3D structures.
method Pre-training via denoising for learning molecular force fields.
result Achieves new state-of-the-art performance on QM9 dataset.
This paper analyzes how data and model properties affect membership inference attacks.
problem Understanding and mitigating the vulnerability of machine learning models to membership inference attacks.
method Empirical analysis of data and model properties on MIA success.
result Data and model properties, not just model overfitting, influence MIA success.
This paper investigates how dataset properties affect GAN training outcomes.
problem GANs often fail to reach equilibrium due to instability or mode collapse.
method Experiments to identify patterns in dataset properties and their effects on GAN training.
result Patterns in dataset properties influence GAN training dynamics and outcomes.
Artificial datasets can serve as a form of regularization for deep learning.
problem Real data shortage in deep learning.
method Injecting noise to high-level features in artificial data generation.
result Artificial data generation can be treated as a form of 'deep' regularization.
New dataset abla2DFT for drug-like molecules benchmarks neural network potentials.
problem Lack of large, diverse datasets for training neural network potentials in quantum chemistry.
method Developed a new dataset abla2DFT containing energies, forces, and molecular properties for drug-like molecules. result First dataset with relaxation trajectories for drug-like molecules.
MCML uses ML to study learnability of Alloy properties, showing simple models can perform well but fail on full input space.
problem Empirical study of learnability of relational properties in Alloy.
method MCML combines ML with model counting to evaluate performance on bounded input spaces.
result Simple ML models can achieve high accuracy and F1-score on training/test datasets but fail on full input space, highlighting complexity of learning relational properties.
Hölder-DPO aligns models robustly with noisy human feedback.
problem No existing alignment methods can handle severe label noise.
method Proposes Hölder-DPO, a principled alignment loss with provable redescending property.
result Hölder-DPO enables scalable human feedback valuation and improves model alignment.
FlowMO uses Gaussian Processes for molecular property prediction with uncertainty.
problem Predicting molecular properties with uncertainty for small datasets.
method Gaussian Processes implemented in FlowMO, built on GPflow and RDKit.
result Comparable predictive performance to deep learning but superior uncertainty calibration.
This paper creates a comprehensive BTC transaction network dataset spanning 15 years.
problem Lack of a full-history BTC graph and network property dataset.
method Thorough analysis of BTC transaction network, creating a dataset and investigating decentralization.
result First systematic investigation of BTC's asset decentralization and design of decentralization degrees.
Study explores calibration properties in neural architectures.
problem Calibration issues in deep neural networks despite improved accuracy.
method Leverages Neural Architecture Search (NAS) to evaluate 117,702 neural networks.
result Identifies key architectural designs beneficial for calibration.
PropEn uses matching to create a larger dataset for efficient design optimization.
problem Limited data and complex landscapes in scientific applications.
method PropEn uses a matching approach to implicitly guide design without a discriminator.
result PropEn efficiently approximates the gradient of property improvement within the data distribution.
Partial Wasserstein Covering aims to identify missing patterns in datasets.
problem Identifying missing patterns in datasets compared to actual applications.
method Formulated as a discrete optimization problem with partial Wasserstein divergence. Proved submodular, allowing greedy approximation. Proposed quasi-greedy algorithms with acceleration techniques.
result Efficiently fills gaps and finds missing scenes in real driving scenes datasets.
Study shows explanation disparities in machine learning models are influenced by data and model properties.
problem Disparities in post-hoc machine learning explanation methods across race and gender.
method Simulations and experiments on a real-world dataset to assess challenges to explanation disparities.
result Increased covariate shift, concept shift, and omission of covariates increase explanation disparities, especially for neural network models.
New benchmarks for offline RL from diverse datasets.
problem Measuring progress in offline RL due to lack of suitable benchmarks.
method Developed benchmarks tailored for offline RL, focusing on diverse dataset properties.
result Revealed deficiencies in existing offline RL algorithms.
We introduce the Hierarchically Interacting Particle Neural Network (HIP-NN) to model molecular properties from datasets of quantum calculations. Inspired by a many-body expansion, HIP-NN decomposes properties, such as energy, as a sum over hierarchical terms. These terms are generated from a neural network--a composit…
Research proposes a test case generation system for deep learning models using dataset properties.
problem Automated generation of extensive test cases for deep learning models is challenging.
method Measures dataset quality and proposes a test case generation system guided by dataset properties.
result Systematic test case generation for deep learning models is effective.
We present a novel condition, which we term the net- work nullspace property, which ensures accurate recovery of graph signals representing massive network-structured datasets from few signal values. The network nullspace property couples the cluster structure of the underlying network-structure with the geometry of th…
One significant challenge to scaling entity resolution algorithms to massive datasets is understanding how performance changes after moving beyond the realm of small, manually labeled reference datasets. Unlike traditional machine learning tasks, when an entity resolution algorithm performs well on small hold-out datas…
We introduce a new parameter to measure the inhomogeneity of training datasets.
problem The need for non-stationary models in supervised learning.
method We introduce a new parameter, the inhomogeneity parameter, to measure the inhomogeneity of training datasets.
result A training set with a non-zero inhomogeneity parameter requires a non-stationary model for accurate predictions.
The PAC-Bayesian approach is a powerful set of techniques to derive non- asymptotic risk bounds for random estimators. The corresponding optimal distribution of estimators, usually called the Gibbs posterior, is unfortunately intractable. One may sample from it using Markov chain Monte Carlo, but this is often too slow…
Five simple soft sensor methodologies with two update conditions were compared on two experimentally-obtained datasets and one simulated dataset. The soft sensors investigated were moving window partial least squares regression (and a recursive variant), moving window random forest regression, the mean moving window of…
CHILI datasets tackle inorganic nanomaterials, advancing graph machine learning.
problem Challenges in modelling inorganic crystalline materials and nanomaterials with graph ML.
method Presented two large-scale datasets of inorganic nanomaterials, defined property and structure prediction tasks.
result Benchmarked performance of graph ML methods on inorganic nanomaterials, highlighting areas for future work.
RIn-Close_CVC2 improves biclustering efficiency by reducing memory usage.
problem Mining maximal biclusters in numerical datasets efficiently and without redundancy.
method Proposes RIn-Close_CVC2, a new version of RIn-Close_CVC that eliminates redundant biclusters without a symbol table.
result RIn-Close_CVC2 reduces memory usage and improves runtime compared to RIn-Close_CVC.
TimeGraph creates synthetic datasets for robust time-series causal discovery.
problem Lack of reliable synthetic benchmark datasets for robust time-series causal discovery.
method Developed comprehensive synthetic datasets with temporal properties, including trends, seasonality, and noise.
result Demonstrated significant variations in algorithm performance under realistic temporal conditions.
Improved molecular property prediction using multitask learning.
problem Predicting molecular properties from chemical data is challenging.
method Multitask learning applied to graph neural networks.
result Multitask learning significantly improves model performance and reduces variance.
New online feature selection method handles streaming data with concept drift.
problem Handling streaming data with concept drift and sparsity.
method Online feature screening method with model adaptation.
result Online screening methods with model adaptation outperform without model adaptation on data streams with concept drift.
The paper explores how neural networks generalize differently from natural and medical images.
problem Discrepancies in generalization error between natural and medical images.
method Established and empirically validated a generalization scaling law with respect to intrinsic dataset properties.
result Higher intrinsic 'label sharpness' of medical images leads to higher adversarial vulnerability.
One of the biggest challenges in the research of generative adversarial networks (GANs) is assessing the quality of generated samples and detecting various levels of mode collapse. In this work, we construct a novel measure of performance of a GAN by comparing geometrical properties of the underlying data manifold and …
METASET selects diverse unit cells for efficient data-driven metamaterial design.
problem Imbalanced datasets in unit cells can bias data-driven metamaterial design.
method METASET uses similarity metrics and DPPs to select diverse subsets of unit cells.
result Smaller, diverse subsets improve search process and structural performance.
Novel chaotic neurons improve AI with minimal training data.
problem Limited training data for AI algorithms.
method Intrinsically chaotic neurons inspired by chaos theory.
result Classification accuracy up to 95.8% with just 2 training samples per class.
Neural message passing on molecular graphs is one of the most promising methods for predicting formation energy and other properties of molecules and materials. In this work we extend the neural message passing model with an edge update network which allows the information exchanged between atoms to depend on the hidde…
Machine learning predicts band gaps for large organic crystals.
problem Predicting band gaps for complex organic crystal structures.
method Released a dataset of 12,500 crystal structures and their band gaps. Trained two state-of-the-art models to achieve a mean absolute error of 0.388 eV.
result Trained models predict band gaps with 13% error for an average gap of 3.05 eV.
Evaluation of hydrocarbon reservoir requires classification of petrophysical properties from available dataset. However, characterization of reservoir attributes is difficult due to the nonlinear and heterogeneous nature of the subsurface physical properties. In this context, present study proposes a generalized one cl…
With access to large datasets, deep neural networks (DNN) have achieved human-level accuracy in image and speech recognition tasks. However, in chemistry, data is inherently small and fragmented. In this work, we develop an approach of using rule-based knowledge for training ChemNet, a transferable and generalizable de…
Cormorant learns molecular properties via rotationally covariant neural networks.
problem Learning molecular potential energy surfaces and properties.
method Rotationally covariant neural network architecture with tensor products and Clebsch-Gordan decomposition.
result Significantly outperforms competing algorithms in learning molecular Potential Energy Surfaces.
SGLD protects membership privacy in Bayesian deep learning.
problem Preventing membership attack in Bayesian deep learning.
method Theoretical framework based on information leakage analysis.
result SGLD prevents information leakage to a certain extent.
Large learning rates enhance model robustness and compressibility.
problem Achieving robustness and resource-efficiency in machine learning models.
method Identifying and utilizing large learning rates as a facilitator for robustness and compressibility.
result Large learning rates produce desirable representation properties and compare favorably to other methods.
Publish a core-set of data to protect against adversarial use.
problem Protecting datasets from adversarial use.
method Construct a fair core-set for linear and neural models.
result Core-sets improve primary task performance while hindering unwanted tasks.
Wiki-CS dataset benchmarks Graph Neural Networks using Wikipedia articles.
problem Benchmarking Graph Neural Networks on a new domain with structural differences.
method Derived from Wikipedia, nodes represent Computer Science articles, edges from hyperlinks, 10 classes for different branches, evaluated semi-supervised node classification and link prediction.
result Graph Neural Networks perform well on Wiki-CS, showing structural differences from earlier benchmarks.
We propose a molecular generative model based on the conditional variational autoencoder for de novo molecular design. It is specialized to control multiple molecular properties simultaneously by imposing them on a latent space. As a proof of concept, we demonstrate that it can be used to generate drug-like molecules w…