New causal models perform poorly when evaluated on biased training sets.
problem Sample selection bias affects the evaluation of causal models' prediction performance.
method Re-evaluated prediction performance of causal models on a genetic perturbation data set, proposing a less-biased evaluation set.
result Causal models have similar or worse performance when evaluated on a less-biased set compared to standard association-based estimators.
Plug-in method improves performative prediction accuracy.
problem Learning under performative feedback with slow convergence rates.
method Plug-in performative optimization using models.
result Plug-in method can be superior to model-agnostic strategies.
Proposes a new cross-validation method to estimate model performance.
problem The standard cross-validation method does not accurately estimate the performance of the recommended model.
method Develops a new random-effects model framework to improve naive cross-validation estimators.
result Proposed estimators outperform conventional and naive methods in estimating model performance.
The study evaluates AI model performance measures for medical use.
problem Selecting appropriate performance measures for AI models in medical practice.
method Assessed 32 performance measures across five domains for binary outcomes.
result 17 measures are both proper and reflect decision-analytic performance.
XPER methodology decomposes credit scoring model performance.
problem Monitoring and understanding the key drivers of credit scoring model performance.
method XPER methodology based on Shapley values, decomposing performance metrics into feature contributions.
result A small number of features explain a large part of model performance.
Partially performative prediction studies how predictive models influence future data.
problem Distribution shift in predictive models due to endogenous and exogenous factors.
method Generalizing performative prediction to capture both endogenous and exogenous sources of distribution shift.
result Developed online analogues of performative stability and optimality for partially performative environments.
New method reduces variance in subpopulation model performance estimates.
problem High variance in subpopulation performance metrics for small groups.
method Using an evaluation model to form model-based metric (MBM) estimates.
result MBMs produce more accurate and lower variance estimates for small subpopulations.
Training deep learning models on mobile devices recently becomes possible, because of increasing computation power on mobile hardware and the advantages of enabling high user experiences. Most of the existing work on machine learning at mobile devices is focused on the inference of deep learning models (particularly co…
Dynamic model pruning improves performance on deep neural networks without retraining.
problem High memory and latency requirements for deep neural networks on low-end devices.
method Dynamic allocation of sparsity pattern and feedback signal to reactivate pruned weights.
result Sparse models achieve state-of-the-art performance with no additional retraining.
This work introduces a method to attribute model performance drops to distribution shifts.
problem Attributing performance drops of machine learning models to distribution shifts.
method Formulated as a cooperative game, value of a set of distributions is defined as the change in model performance when only that set of distributions changes. Importance weighting method for computing the value of an arbitrary set of distributions is derived. Quantifying the contribution of each distribution as its Shapley value.
result Demonstrated the effectiveness of the method on various case studies.
Optimal allocation between explainable and black box models for high performance and explainability.
problem Balancing explainability and performance in model ensembles.
method Optimal allocation of observations between explainable and black box models to maximize ensemble performance and explainability.
result Learned allocations maintain high ensemble performance and explainability, sometimes outperforming individual models.
New framework models algorithmic decisions affecting data distribution, enabling efficient learning.
problem Algorithmic decisions can alter data distribution, affecting model performance.
method Model performative effects as push-forward measures, estimating gradients under shift operators.
result Prove convexity of performative risk, allowing more accurate models to be harder to classify.
Gaussian processes are a flexible Bayesian nonparametric modelling approach that has been widely applied but poses computational challenges. To address the poor scaling of exact inference methods, approximation methods based on sparse Gaussian processes (SGP) are attractive. An issue faced by SGP, especially in latent …
MO-PaDGAN generates diverse, high-performance designs with multiple metrics.
problem Challenges in generating diverse, high-performance designs with multiple metrics.
method MO-PaDGAN uses a new Determinantal Point Processes based loss function for probabilistic modeling of diversity and performances.
result MO-PaDGAN expands the design space towards high-performance regions and generates new designs with high diversity and performances.
Bayesian approach models match and non-match score distributions over continuous covariates.
problem Complex evaluation of model performance over continuous covariates in biometric verification.
method Generative model of score distributions, mixture models, local basis functions, Bayesian inference.
result Accurate and effective method for studying model performance over continuous covariates.
This paper optimizes performative risk by focusing on convex properties and developing efficient algorithms.
problem Performative risk, the loss experienced by decision makers, is not optimized by stable models.
method Identifying convex properties of loss function and model-induced distribution shift, developing algorithms for optimization.
result Optimization of performative risk with better sample efficiency than generic methods.
Proposes a method to compute valid lower confidence bounds for multiple models selected based on their performance.
problem Model selection and evaluation in machine learning.
method Interprets model selection as a simultaneous inference problem, uses bootstrap tilting and maxT-type multiplicity correction.
result Yields valid lower confidence bounds that are at least as good as standard approaches and reliably reach nominal coverage probability.
This research examines how model explanations change under distribution shifts in tabular data.
problem Detecting distribution shifts in tabular data affecting model performance and explanations.
method Investigates the relationship between model performance and explanation characteristics under distribution shifts.
result Explanation shifts are a better indicator for detecting predictive performance changes than traditional distribution shift techniques.
Study on multitask learning performance factors.
problem Mixed results in multitask learning performance.
method Task simulator and symbolic regression to learn performance factors.
result Empirical formulas relating model performance to sqrt(n), sqrt(T), and sqrt(AMI).
SHIFT framework identifies subgroups with large ML model performance decay.
problem Large model performance decay in subgroups when deployed.
method Subgroup-scanning Hierarchical Inference Framework (SHIFT) for performance drift.
result SHIFT identifies interpretable subgroups with large performance decay and suggests targeted actions to mitigate it.
Paper analyzes impact of PRM on binary random variables and distribution shifts.
problem Impact of performative risk minimization on binary random variables and distribution shifts.
method Formulated two measures of impact, derived explicit formulas for full information, and provided estimators for partial information.
result PRM can have amplified side effects compared to methods that do not model data shift.
PUMA augments models to remove unique data points without performance loss.
problem Preserving model performance while removing unique training data points.
method Explicitly models data influence, reweights remaining data optimally.
result PUMA effectively removes unique data points without performance degradation.
Study identifies negative data externalities affecting model performance on specific groups.
problem Negative data externalities on group performance in machine learning models.
method Characterized and detected data-model inefficiencies, focusing on specific types of externalities.
result Negative data externalities can lower model performance on specific sub-groups, even with larger datasets.
Method diagnoses model performance under distribution shifts.
problem Understanding and improving model performance under distribution shifts.
method DIstribution Shift DEcomposition (DISDE) method to attribute performance drop to distribution shifts.
result Shows how model performance can be improved across different distribution shifts.
Paper explores how black box models can deviate from average performance.
problem Understanding and interpreting predictions from sophisticated black box models.
method Two general approaches to provide interpretable descriptions of black box classification model performance.
result Identifies regions where black box models deviate significantly from their average performance.
Paper tackles performative prediction without convexity assumptions.
problem Performative prediction where data distribution changes with model deployment.
method Reparameterization framework to transform non-convex objective into convex one.
result Provably sublinear regret guarantees for learnable model.
Study on optimizing model updates in performative prediction.
problem Optimizing model updates influenced by model predictions.
method Stochastic optimization with greedy and lazy deploy approaches.
result Rates of convergence for both greedy and lazy deploy methods.
Investigates the use of Information Coefficient as a stock selection model performance measure.
problem The adequacy and effectiveness of Information Coefficient (IC) for evaluating stock selection models is unclear.
method Simulation and simple statistical modeling to examine IC behavior statically and dynamically.
result Proposes two practical procedures for IC-based ongoing performance monitoring of stock selection models.
This paper analyzes forecasting models for COVID-19 cases and deaths.
problem Reliable forecasting of COVID-19 cases and deaths is crucial for managing the disease.
method Quantitative analysis of forecasting models across different regions in the US, evaluating model selection, hyperparameter tuning, and training time.
result Model selection is the most influential factor in forecasting performance.
Paper proposes GP-NAS-ensemble for fast neural architecture performance prediction.
problem Estimating neural network performance without training time-consuming evaluations.
method GP-NAS-ensemble framework using ensemble learning improvements.
result Ranked second in a NAS performance prediction challenge.
This study evaluates zero-shot LLMs in finance, finding ChatGPT performs well but fine-tuned models are better.
problem Evaluating zero-shot LLMs in financial tasks.
method Comparison of ChatGPT and fine-tuned models on annotated data.
result Fine-tuned models generally outperform zero-shot LLMs.
A new framework for performative prediction robust to distributional misspecification.
problem Performative prediction models can be influenced by their own predictions, leading to suboptimal outcomes.
method Introduces distributionally robust performative prediction (DRPO) to approximate the true performative optimum (PO) robustly.
result DRPO provides provable guarantees as a robust approximation to the true PO when the nominal distribution map is misspecified.
AI benchmarks evaluate football team performance using generative models.
problem Evaluating human performance in complex interactive tasks is error-prone and unreliable.
method Trained Conditional VRNN Model on player and ball tracking data to imitate and predict team interactions.
result Trained model as a useful benchmark for evaluating team performance in football.
Study compares Islamic banks' accounting and market performance.
problem Assessing the relationship between Islamic banks' accounting and market performance.
method Selected six Islamic banks, collected data from 2009-2013, used random-effect models.
result Superior accounting performance does not correlate with superior market performance.
New risk theory for 'Pay-for-Performance' models.
problem How to price and hedge operational and financial risks in new business models.
method Developed a new risk theory and calculation method for 'Pay-for-Performance' models.
result Presented a model for determining risk premiums including both financial and operational risks.
We study large-scale kernel methods for acoustic modeling in speech recognition and compare their performance to deep neural networks (DNNs). We perform experiments on four speech recognition datasets, including the TIMIT and Broadcast News benchmark tasks, and compare these two types of models on frame-level performan…
The paper proposes a method to predict the performance of data-driven algorithms using surrogate models.
problem Improving the performance prediction of data-driven knowledge discovery algorithms.
method Surrogate-assisted performance prediction using evolutionary modeling of clinical pathways.
result The proposed approach provides interpretable prediction of algorithm performance and quality.
The paper explores how machine learning models can be learnable despite label shifts.
problem Learnability of binary classification models in the presence of label shifts.
method Developed a performative empirical risk function that is an unbiased estimate of the true risk on the shifted distribution.
result PAC-learnable hypothesis spaces remain PAC-learnable for performative scenarios.
One of the important measures of quality of education is the performance of students in the academic settings. Nowadays, abundant data is stored in educational institutions about students which can help to discover insight on how students are learning and how to improve their performance ahead of time using data mining…
This paper improves combine harvester performance using ANN-PSO hybrid model.
problem Improving performance of combine harvesters to minimize waste and reduce maintenance.
method Proposes a hybrid machine learning model combining artificial neural networks and particle swarm optimization.
result Demonstrates higher accuracy and stability in predicting optimal performance of combine harvesters.
Sharp bounds on binary model inference performance.
problem High-dimensional inference in binary models.
method Convex empirical risk minimization, sharp asymptotics, optimal performance bounds.
result Sharp predictions and optimal performance bounds for binary models.
Improves tree model performance by considering future node splits.
problem Improving tree model performance.
method Next-Depth Lookahead Tree (NDLT) model that evaluates future node splits.
result Enhanced tree model performance.
This paper compares preprocessing techniques for XGBoost models on various data sets.
problem Improving predictive performance of XGBoost models through optimal data preprocessing.
method Comparison of feature selection, categorical handling, and null imputation methods.
result XGBoost importance by gain is the most consistent and highest-performing method for feature selection.
We present a supervised neural network model for polyphonic piano music transcription. The architecture of the proposed model is analogous to speech recognition systems and comprises an acoustic model and a music language model. The acoustic model is a neural network used for estimating the probabilities of pitches in …
Study shows cross-domain X-ray prediction performance discrepancies and label shifts.
problem Quantifying generalization limits across different X-ray datasets.
method Large-scale study on multiple X-ray datasets, focusing on performance and label shifts.
result Interesting discrepancies found between model performance and agreement, and concept similarity across tasks.
Bayesian stacking improves model performance with varying model weights.
problem Improving model predictions with heterogeneous input performance.
method Bayesian hierarchical stacking with varying model weights inferred via Bayesian inference.
result Hierarchical stacking yields better predictions than linear averaging.
Using a low-dimensional parametrization of signals is a generic and powerful way to enhance performance in signal processing and statistical inference. A very popular and widely explored type of dimensionality reduction is sparsity; another type is generative modelling of signal distributions. Generative models based o…
This paper shows feature importance remains valid even in low-performing models.
problem Feature importance validity in low-performing machine learning models for biomedical data.
method Experiments with synthetic and real biomedical datasets to compare feature rank stability under different data reductions.
result Feature importance can be maintained even at low performance levels if data size is adequate.