torchbearer simplifies deep learning model fitting in PyTorch.
problem Simplifying model fitting for researchers in deep learning.
method High-level metric and callback API for model persistence, learning rate decay, logging, and data visualization.
result Extensive documentation and built-in callbacks for various applications.
fastai simplifies deep learning with a layered API.
problem Creating state-of-the-art deep learning models efficiently.
method Carefully layered architecture, Python type dispatch, GPU optimization, and novel callbacks.
result fastai achieves high performance with minimal compromise in usability and flexibility.
Study detects illegal discrimination by employers using correspondence experiments.
problem Detecting illegal discrimination by individual employers based on protected characteristics.
method Correspondence experiments, bounding higher moments of causal effects, decision rules for investigation.
result 85% of jobs contacting both white and black applicants are likely to discriminate.
This paper investigates how machine learning APIs change over time and proposes an efficient method to monitor these changes.
problem Understanding and assessing changes in machine learning APIs over time.
method Systematic investigation of ML API shifts, proposing a principled adaptive sampling algorithm (MASA) for efficient estimation of confusion matrix shifts.
result MASA can accurately estimate confusion matrix shifts using up to 90% fewer samples compared to random sampling.
FrugalML optimizes API selection for cost and accuracy.
problem Challenges in choosing the best ML prediction APIs.
method Proposes FrugalML, a framework that learns API strengths and weaknesses.
result Achieves up to 90% cost reduction with matching accuracy.
Automates API mapping across languages with minimal manual effort.
problem Manual effort required to identify parallel program corpora and API mappings.
method Combines domain adaptation and code embedding with GANs to align vector spaces of APIs.
result Automatically identifies cross-language API mappings with less prior knowledge.
OpenML-Python API simplifies access to OpenML for Python users.
problem Limited access to OpenML for Python users.
method Developed a Python API (OpenML-Python) to integrate OpenML with Python-based tools.
result Facilitates easy access to OpenML's datasets, tasks, and experiments.
Generative GAN improves classifier performance with limited training data.
problem Limited training data for black-box API attacks.
method Generative adversarial network (GAN) to generate synthetic training data.
result Improves classifier performance with synthetic data.
Paper studies deep learning attacks on online APIs with limited data.
problem Adversarial machine learning threats on online APIs with strict rate limitations.
method Develops an active learning approach to build adversarial classifiers with limited training data.
result Active learning can build adversarial classifiers with small statistical difference from target classifiers using limited data.
Paper proposes OpenAPI to interpret PLM models without access to parameters.
problem Interpreting hidden predictive models without access to parameters.
method Closed-form solution using overdetermined linear equation systems.
result Exact and consistent interpretations for PLM models.
VarDetect monitors API queries to detect model extraction attacks.
problem Attackers can extract ML models through API queries.
method VarDetect uses a modified variational autoencoder to track and detect attacker samples.
result VarDetect successfully detects and raises alarms for model extraction attacks.
I present a web service for querying an embedding of entities in the Wikidata knowledge graph. The embedding is trained on the Wikidata dump using Gensim's Word2Vec implementation and a simple graph walk. A REST API is implemented. Together with the Wikidata API the web service exposes a multilingual resource for over …
CK simplifies ML model deployment and reproducibility with open APIs and DevOps.
problem Making ML models reproducible and deployable across different environments.
method Decompose complex systems into reusable sub-components with unified APIs and DevOps principles.
result Automatically co-design and optimize ML models for speed, accuracy, energy, and size.
Open-source Vizier optimizes complex systems for Google and beyond.
problem Optimizing large-scale systems with multiple objectives and constraints.
method Distributed, fault-tolerant, flexible API for blackbox optimization.
result OSS Vizier supports a wide range of optimization problems and is available as open-source.
Research evaluates model extraction attacks on complex ML models and introduces a defense.
problem Model extraction attacks steal functionality of ML models through prediction APIs.
method Evaluation of Knockoff nets and introduction of a defense.
result Realistic adversaries can effectively steal complex ML models and evade known defenses.
We consider the infinite-horizon discounted optimal control problem formalized by Markov Decision Processes. We focus on several approximate variations of the Policy Iteration algorithm: Approximate Policy Iteration, Conservative Policy Iteration (CPI), a natural adaptation of the Policy Search by Dynamic Programming a…
Method generates adversarial examples to improve classifier robustness.
problem Addressing adversarial examples in machine learning.
method MDL principle for generating adversarial examples.
result Evasion rate of 78.24% for adversarial examples compared to 8.16% for original malware.
PettingZoo library accelerates multi-agent reinforcement learning research.
problem Challenges in multi-agent reinforcement learning, especially conceptual models of games.
method Developed PettingZoo library with AEC games model to address multi-agent reinforcement learning challenges.
result AEC games model addresses conceptual issues in multi-agent reinforcement learning environments.
Proposes an online model for LLM cascading with adaptive API selection.
problem Adaptive querying and selection of LLM APIs in a context-dependent environment.
method Develops a learning approach combining GMM estimation and UCB-style bounds.
result Achieves cumulative regret of O ~ ( T ) \widetilde O(\sqrt T) O ( T ) over T T T periods. tf_geometric simplifies graph deep learning in TensorFlow.
problem Efficient graph deep learning in TensorFlow.
method Kernel libraries and infrastructures for GNNs.
result tf_geometric supports various graph tasks and provides efficient GNN models.
Paper uses dynamic analysis to detect malware with PHMMs.
problem Malware detection using static and dynamic analysis techniques.
method Hidden Markov Models (HMMs) and Profile Hidden Markov Models (PHMMs) trained on API call sequences.
result PHMMs outperform HMMs in malware detection.
Regression or classification? This is perhaps the most basic question faced when tackling a new supervised learning problem. We present an Evolutionary Deep Learning (EDL) algorithm that automatically solves this by identifying the question type with high accuracy, along with a proposed deep architecture. Typically, a …
Paper explores how non-neural simulators can enhance DP synthetic data generation.
problem Generating differentially private synthetic data without access to foundation models.
method Private Evolution (PE) framework using inference APIs and simulators.
result Sim-PE framework improves downstream classification accuracy and FID scores.
InferPy simplifies probabilistic modeling with deep neural networks in Python.
problem Complex probabilistic models with deep neural networks.
method User-friendly API for defining, learning, and evaluating models.
result Compact and simple way to define general hierarchical probabilistic models.
Deep learning models misclassify malware with added benign features.
problem Detecting malware with deep learning when it's mixed with benign code.
method Trained a deep neural network classifier using benign and malware features. Demonstrated the impact of adding benign features to malware. Used data augmentation to improve classifier robustness.
result Adding benign features to malware significantly increases false negatives.
Python library for integrating TDA with machine learning.
problem Data exploration and interpretability in machine learning.
method Integrates TDA with scikit-learn API, uses C++ for performance.
result Enhanced data exploration and interpretability in machine learning.
LIBTwinSVM offers a free library for efficient Twin Support Vector Machines.
problem Large-scale classification problems.
method Efficient implementation of Twin Support Vector Machines.
result Effectiveness demonstrated through benchmarks.
KL-constrained API shows optimization issues and improved with regularization.
problem Optimization issues in KL-constrained API algorithms.
method Comparison of KL divergence as a constraint vs. regularizer, empirical evaluation.
result KL-constrained API is not guaranteed to converge and incurs linear regret.
Paper develops a framework to test deep RL traffic controllers under various uncertainties.
problem Developing robust deep RL traffic controllers for dynamic urban areas.
method Open-source callback-based framework for evaluating deep RL configurations in a traffic simulation.
result Deep RL controllers perform well under demand surges, incidents, and sensor failures.
This paper addresses the model-free nonlinear optimal problem with generalized cost functional, and a data-based reinforcement learning technique is developed. It is known that the nonlinear optimal control problem relies on the solution of the Hamilton-Jacobi-Bellman (HJB) equation, which is a nonlinear partial differ…
GANs visualize malware behavior for proactive protection.
problem Malware authors' advantage in testing and augmenting malicious code.
method GAN trained on distributed image representation of malware behaviors.
result Generated synthetic malware for adversarial training of anti-malware models.
PHOTONAI simplifies machine learning model development in Python.
problem Rapid and efficient machine learning model development.
method Unified framework combining algorithms from different toolboxes, automating repetitive tasks.
result Achieves state-of-the-art solution in medical machine learning.
TorchKM: A GPU-Oriented Library for Kernel Learning and Model Selection
problem Kernel learning and model selection
method GPU acceleration
result Competitive predictive performance with speedups
API technique improves multiclass GAM interpretability.
problem Interpretability breakdown in multiclass GAMs.
method Developed two axioms and API technique for multiclass GAMs.
result API technique transforms multiclass GAMs to be visually interpretable.
Transformer models improve malware classification, especially for imbalanced datasets.
problem Imbalanced multiclass malware classification.
method Bagging-based random transformer forest (RTF) ensemble of pre-trained transformer models (BERT or CANINE).
result Ensemble of pre-trained transformer models achieves state-of-the-art F1-score of 0.6149 on a benchmark dataset.
Math model helps bees decide between winter survival and raising young.
problem Deciding between winter survival and raising young bees in honeybee colonies.
method Mathematical model considering resource geometry around the hive.
result Optimal resource allocation strategy for honeybee colonies.
torchsom simplifies SOMs in PyTorch with GPU acceleration and scikit-learn API.
problem Efficient implementation and usability of SOMs in PyTorch.
method PyTorch backend, GPU acceleration, scikit-learn API, 90% test coverage.
result Ease of use and scalability for SOMs in PyTorch.
Online social networks (OSN) contain extensive amount of information about the underlying society that is yet to be explored. One of the most feasible technique to fetch information from OSN, crawling through Application Programming Interface (API) requests, poses serious concerns over the the guarantees of the estimat…
Uses simulations to detect AI bias in face detection.
problem Bias in machine learning models can lead to poor performance on minority groups.
method Bayesian parameter search to identify weaknesses in ML classifiers.
result Identifies demographic biases in commercial face APIs.
DAWN dynamically embeds watermarks in neural network predictions to prevent model extraction.
problem Preventing theft of machine learning models through model extraction.
method Dynamic adversarial watermarking at prediction API level.
result DAWN effectively resists two state-of-the-art model extraction attacks.
LSTM networks improve stock price prediction accuracy.
problem Enhancing stock price forecasting accuracy.
method LSTM networks with hyperparameter tuning and feature selection.
result 53% improvement in predictive accuracy.
MAX simplifies access to DL models for non-experts.
problem Difficulty for non-experts in adopting latest DL models.
method Proposes MAX, a Python library that wraps DL models and provides RESTful APIs.
result Maximizes ease of using state-of-the-art DL models for inference.
BoFire optimizes chemistry experiments using Bayesian Optimization.
problem Effective deployment of Bayesian Optimization in the chemical industry.
method Combines Bayesian Optimization with DoE strategies, providing a rich feature-set.
result BoFire enables seamless integration into RESTful APIs for real-world use.
In this paper, we propose a novel policy iteration method, called dynamic policy programming (DPP), to estimate the optimal policy in the infinite-horizon Markov decision processes. We prove the finite-iteration and asymptotic l\infty-norm performance-loss bounds for DPP in the presence of approximation/estimation erro…
Recently, a novel class of Approximate Policy Iteration (API) algorithms have demonstrated impressive practical performance (e.g., ExIt from [2], AlphaGo-Zero from [27]). This new family of algorithms maintains, and alternately optimizes, two policies: a fast, reactive policy (e.g., a deep neural network) deployed at t…
Denoised smoothing defends pretrained classifiers against adversarial attacks.
problem Adversarial attacks on pretrained classifiers.
method Prepending a denoiser to any off-the-shelf classifier using randomized smoothing.
result Guaranteed ℓ p \ell_p ℓ p -robustness to adversarial examples without modifying the pretrained classifier. Efficiently distills white-box adversarial attacks into black-box models.
problem Generating efficient adversarial examples for robustness.
method Train a model to emulate white-box attack behavior and distill it into a more efficient black-box model.
result Reduces adversarial example generation time by 19x-39x and transfers to black-box settings.
Optimizes leveraged staking strategies in decentralized finance.
problem Maximizing returns on staked assets in decentralized lending platforms.
method Developed a mathematical framework to optimize leveraged staking strategies, reducing the multi-market problem to convex allocation over market exposures.
result Rebalanced leveraged positions can achieve up to 6.2% APY, significantly higher than unleveraged staking.