System detects overfitting in ML apps, improving quality and efficiency.
problem Overfitting in ML applications during continuous development.
method ease. ml/meter system for automated overfitting detection and measurement.
result Probabilistic overfitting signals for developers to take actions.
ML .NET simplifies integrating machine learning into software development.
problem Difficulty and cost of incorporating machine learning in software development.
method Developed a framework (ML .NET) and introduced DataView for efficient data abstraction.
result ML .NET outperforms recent machine learning frameworks in performance studies.
Machine learning enhances fuzzing for better software testing.
problem Challenges in traditional fuzzing.
method Applications of machine learning to improve fuzzing.
result Machine learning tools have successfully addressed fuzzing bottlenecks.
Machine learning models emulate and approximate complex mappings in model physics.
problem Developing and ensuring accurate physical parameterizations.
method Machine learning tools to emulate and approximate mappings.
result ML can improve parameterizations and enforce physical constraints.
The use of machine learning (ML) is on the rise in many sectors of software development, and automotive software development is no different. In particular, Advanced Driver Assistance Systems (ADAS) and Automated Driving Systems (ADS) are two areas where ML plays a significant role. In automotive development, safety is…
The rise of Big Data has led to new demands for Machine Learning (ML) systems to learn complex models with millions to billions of parameters, that promise adequate capacity to digest massive datasets and offer powerful predictive analytics thereupon. In order to run ML algorithms at such scales, on a distributed clust…
This paper explores ML in power line communications, from modeling to diagnostics.
problem Improving efficiency and diagnostics in power line communications.
method Discusses classical ML models and their application to PLC at various layers.
result Demonstrates how ML can enhance various aspects of PLC.
Paper bridges AI/ML and causal modeling to reduce bias.
problem Difficulty in combining methods from different assumptions.
method Integrates system dynamics and structural equation modeling.
result Unified mathematical framework for AI/ML and causal modeling.
Developed R package for creating nomograms for any ML algorithms.
problem Creating nomograms for any machine learning algorithms.
method Formulated a function to transform ML prediction models into nomograms, requiring specific datasets.
result Created 5 types of nomograms for various ML algorithms and predictor types.
Develops a framework to quantify uncertainties in multiple ML models.
problem Uncertainty in ML model predictions and model inputs.
method Develops a theoretical framework to decouple and transform uncertainties.
result Generates joint distribution of ML predictions considering uncertainties.
Paper discusses challenges in deploying ML models for structural engineering.
problem Challenges in deploying machine learning models for structural engineering applications.
method Illustrates challenges through two examples, focusing on model overfitting, underspecification, training data representativeness, variable omission bias, and cross-validation.
result Highlights the importance of rigorous model validation techniques.
Study analyzes carbon footprint of 1,417 ML models on Hugging Face.
problem Scarce knowledge on measuring and reporting carbon footprint of ML models.
method Repository mining study on Hugging Face Hub API.
result Stalled carbon emissions-reporting models, slight decrease in carbon footprint over 2 years.
With few exceptions, the field of Machine Learning (ML) research has largely ignored the browser as a computational engine. Beyond an educational resource for ML, the browser has vast potential to not only improve the state-of-the-art in ML research, but also, inexpensively and on a massive scale, to bring sophisticate…
New method for valid inference from ML-predicted data.
problem Invalid scientific conclusions from ML-predicted outcomes.
method Assumption-Lean and Data-Adaptive Post-Prediction Inference (PSPA)
result Valid and powerful inference based on ML-predicted data.
Develops Active Fourier Auditor to estimate ML model properties without reconstructing them.
problem Verifying and auditing properties of Machine Learning models in real-world applications.
method A new framework that quantifies ML model properties using Fourier coefficients, without reconstructing the model.
result Active Fourier Auditor (AFA) is more accurate and sample-efficient than baselines for estimating robustness, individual fairness, and group fairness.
Interactive tool helps understand ML model performance.
problem Understanding ML model performance across various inputs.
method Open-source What-If Tool for interactive probing and analysis.
result Allows testing performance in hypothetical situations and measuring fairness.
Data application developers and data scientists spend an inordinate amount of time iterating on machine learning (ML) workflows -- by modifying the data pre-processing, model training, and post-processing steps -- via trial-and-error to achieve the desired model performance. Existing work on accelerating machine learni…
Industry lacks tools to secure ML systems, study finds.
problem Insufficient security tools for ML systems in industry.
method Interviews with 28 organizations to identify gaps.
result Need revised Security Development Lifecycle for ML.
Review of mathematical representations for biomolecular data.
problem Complexity and high dimensionality of biomolecular datasets hinder ML applications.
method Developed low-dimensional and scalable mathematical representations using algebraic topology, differential geometry, and graph theory.
result Mathematical representations improve protein-ligand binding predictions and other biomolecular applications.
Molecular Dynamics (MD) simulation is widely used to analyze the properties of molecules and materials. Most practical applications, such as comparison with experimental measurements, designing drug molecules, or optimizing materials, rely on statistical quantities, which may be prohibitively expensive to compute from …
GRETEL unifies GCE evaluation across various settings.
problem Lack of standardized evaluation for Graph Counterfactual Explanations.
method Unified framework for testing GCE methods in diverse settings.
result GRETEL promotes reproducible evaluations of GCE techniques.
Develops a measure for subjective explainability of ML predictions.
problem Ensuring transparency and trust in automated decision-making.
method Information-theoretic concepts applied to conditional entropy of predictions given user feedback.
result EERM principle balances subjective explainability and risk.
Survey examines challenges of ML in avionic systems certification.
problem Challenges in current certification standards for ML in avionic systems.
method Literature review focusing on robustness and explainability of ML results.
result Current certification standards do not support ML in avionic systems.
The development of accurate and transferable machine learning (ML) potentials for predicting molecular energetics is a challenging task. The process of data generation to train such ML potentials is a task neither well understood nor researched in detail. In this work, we present a fully automated approach for the gene…
Gradio simplifies sharing and testing ML models for non-experts.
problem Accessibility and collaboration challenges in machine learning.
method Developed an open-source Python package, Gradio, to create visual interfaces for ML models.
result Gradio makes it easy for non-technical users to test and provide feedback on ML models.
Machine learning reveals hidden features in knot classification.
problem Classifying the topology of closed curves.
method Investigating shortcut methods used by ML for knot classification.
result Developed a dataset and code to remove non-topological features.
Unified framework for analyzing model stealing attacks and defenses.
problem Vulnerability of ML applications to model stealing attacks.
method Developed a rigorous threat model and evaluation criteria, proposed methods to quantify attack and defense strategies.
result Demonstrated the importance of attack-specific perturbations for effective defenses.
A human-in-the-loop ML framework for precision dosing reduces expert workload and removes bias.
problem High cost of data annotation and lack of appropriate data for ML models.
method Incorporates human experts into the model learning loop to improve interpretability and reduce bias.
result The approach learns interpretable rules from data and potentially lowers expert workload.
New attacks show ML models can be compromised even when targeting one concept.
problem Vulnerability of ML models to multi-concept attacks.
method Developed novel multi-concept attack techniques for deep learning.
result Successfully attacked one set of classifiers without impacting others.
Develops fair machine learning models resistant to sensitive perturbations.
problem Ensuring model performance is invariant to sensitive attributes like gender and ethnicity.
method Distributionally robust optimization to enforce individual fairness.
result Demonstrates effectiveness on tasks prone to bias.
ML weather forecasts lack physical consistency, but add value.
problem Lack of physical consistency in ML weather forecasts.
method Examined Pangu-Weather forecasts for fidelity and physical consistency.
result ML forecasts lack physical consistency and accuracy can be partly due to this.
A review of ML and DL for ecological data analysis.
problem Understanding the strengths and limitations of ML and DL in ecological research.
method Historical overview, algorithm families, differences, universal principles, and emerging trends.
result ML and DL excel in prediction tasks but are still debated for causal inference.
fintech-kMC simulates financial platforms for AI/ML model validation.
problem Validation of AI/ML models in real-world financial applications.
method Agent-based model with kinetic Monte Carlo engine.
result Generates realistic synthetic data for testing AI/ML models.
Paper robustifies reinforcement learning agents against action space perturbations.
problem Vulnerability of reinforcement learning agents to action space perturbations (e.g. actuator attacks).
method Adversarial training to robustify DRL agents against perturbations.
result DRL agents can be robustified against action space perturbations through adversarial training.
Develops ML-DQA for healthcare data quality assurance.
problem Inconsistent use of real-world data in machine learning projects.
method Develops ML-DQA framework based on RWD best practices.
result Five generalizable practices emerge from ML-DQA implementation.
Develops a machine learning pipeline for learning causal structure in time-series data.
problem Current ML algorithms fail to learn causal structure in time-series data due to lack of temporal order consideration.
method Integrates machine learning with chaos theory using ChaosFEX feature extractor to learn generalized causal structure.
result Successfully learns generalized causal structure in time-series data.
Sage platform protects ML models trained on sensitive data from leakage.
problem Protecting sensitive data in machine learning models exposed to untrusted domains.
method Develops block composition for privacy accounting and privacy-adaptive training to manage privacy budget and utility tradeoff.
result Enables continuous training of models on sensitive data streams while maintaining global DP guarantees.
Systematic review of ML explainability in process mining.
problem Understanding the black-box nature of ML models in process mining.
method Systematic literature review using PRISMA framework.
result Identification of key trends and challenges in interpretability.
Community-based system dynamics improves ML fairness by involving excluded stakeholders.
problem Bias in ML system development during problem formulation.
method Community-based system dynamics (CBSD) for stakeholder participation.
result CBSD facilitates deeper problem understanding and bias mitigation.
DC-Check helps guide ML development by considering data-centric aspects.
problem Lack of standardized framework for data-centric considerations in ML.
method DC-Check is a checklist-style framework for data-centric AI at ML pipeline stages.
result Promotes thoughtfulness and transparency in ML development.
This paper uses ML to identify prey handling in seals.
problem Automatically classify prey handling activity in seals for monitoring.
method Developed and compared three ML algorithms: Input Delay Neural Networks, Support Vector Machines, and Echo State Networks.
result Echo State Networks outperformed other algorithms in terms of accuracy and F1score.
This tutorial covers methods for handling missing data in SP and ML.
problem Dealing with missing data in signal processing and machine learning.
method Grouping strategies into three tasks: imputation, estimation, and prediction.
result Promising and future research directions are discussed.
This paper tackles hidden technical debts in fair ML systems for Fintech.
problem Building fair machine learning systems in financial services.
method Examining key stages of ML system development and deployment.
result Technical debts exist in deploying fair ML systems in Fintech.
Gaussian processes help in modeling complex, nonlinear relationships in signal processing.
problem Modeling complex, nonlinear relationships in signal processing.
method Sequential inference for Gaussian processes.
result Gaussian processes enable efficient and accurate modeling of complex relationships.
Deep ReLU networks can approximate matrix-vector products with error bounds.
problem Can deep ReLU networks accurately approximate matrix-vector products?
method Derived error bounds in Lebesgue and Sobolev norms for deep ReLU FNNs.
result Developed deep approximation theory with successful applications.
A framework for developing ML systems from ML primitives.
problem Complexity and fragmentation in ML systems development.
method Unified API, ML primitives, AutoML strategies, and pipelines.
result General-purpose, multi-task AutoML system for various data modalities and problem types.
We present SmartChoices, an approach to making machine learning (ML) a first class citizen in programming languages which we see as one way to lower the entrance cost to applying ML to problems in new domains. There is a growing divide in approaches to building systems: on the one hand, programming leverages human expe…
New ML approach predicts glare in open-plan offices with high accuracy.
problem Challenging to predict discomfort glare in open-plan offices.
method Used Machine Learning algorithms to predict glare.
result ML model achieved 83.8% accuracy in predicting glare.