Study compares different scoring rules for machine-learned weather forecasts, finding scale-awareness improves forecast realism.
problem Improving the accuracy of machine-learned probabilistic weather forecasts.
method Comparison of scoring rules (CRPS, fair global energy score, graph energy score) and analysis of their impact on forecast field spectra.
result Scale-awareness improves forecast realism, particularly in the tropics.
Large learning rates prevent memorization in denoising score matching.
problem Memorization of training data in diffusion-based generative models.
method Investigating the role of large learning rates in the small-noise regime, proving that they prevent convergence to the empirical optimal score.
result Large learning rates prevent memorization by making it impossible for the learned score to be arbitrarily close to the empirical optimal score.
ScoreMatchingRiesz improves debiased machine learning and policy effects estimation.
problem Improving debiased machine learning and policy effects estimation.
method Score matching and Riesz representer estimation.
result Estimates policy path for continuous treatments, improving interpretability.
CNN scoring function predicts protein-ligand interactions.
problem Scoring protein-ligand interactions for drug discovery.
method Convolutional Neural Networks (CNN) for 3D protein-ligand interactions.
result CNN scoring function outperforms AutoDock Vina in ranking poses.
New algorithm learns bridged diffusion processes without time-reversals.
problem Learning bridged diffusion processes efficiently and accurately.
method Score matching with Doob's h-transform, avoiding time-reversals.
result Outperforms existing methods in learning bridged diffusion processes.
Paper tackles transparency and auditability of machine learning in credit scoring.
problem Missed potential in using modern machine learning for credit scoring due to lack of transparency.
method Develops a framework for making black box machine learning models transparent, auditable, and explainable.
result Comparable interpretability can be achieved with machine learning while maintaining predictive power.
Paper bridges score estimation to parameter and density estimation in DDPMs.
problem Efficiently estimating scores for generative models.
method Introduces a framework linking score estimation to parameter and density estimation.
result Denoising score-matching in DDPMs is asymptotically efficient for parameter estimation.
Two-stage scoring approach for P2P lending improves loan profitability prediction.
problem Class imbalance and lack of profitability prediction in existing scoring methods.
method Integrates credit scoring and profit scoring using wide and deep learning.
result Two-stage scoring approach outperforms existing methods in loan profitability prediction.
Study examines fairness in machine learning for credit scoring.
problem Bias in machine learning models for credit scoring.
method Comprehensive experimental study of fairness-aware machine learning models.
result Fairness-aware models improve fairness while maintaining accuracy.
A new method for estimating complex models and high-dimensional data.
problem Difficulty in computing Hessian of log-density functions for complex models and high-dimensional data.
method Sliced score matching, which projects scores onto random vectors before comparison.
result Sliced score matching can learn deep energy-based models and produce accurate score estimates.
DEEN learns energy and score functions from complex data.
problem Challenges in density estimation for high-dimensional data.
method Inference-free hierarchical framework using score matching and multilayer perceptrons.
result DEEN successfully learns energy and score functions from synthetic and high-dimensional data.
Credit scores misclassify borrowers, especially minorities, leading to inequitable access.
problem Misclassification of borrowers by credit scores, particularly minorities.
method Benchmarked a widely used credit score against a machine learning model.
result Machine learning model improves predictive accuracy for low-quality data, leading to more equitable access.
New approach learns risk scores efficiently and optimally.
problem Learning risk scores from data is challenging due to calibration, sparsity, and operational constraints.
method Formulated as a mixed integer nonlinear program and solved using a cutting plane algorithm with specialized techniques.
result Improves risk score learning efficiency and optimality, providing a feasible solution without post-processing.
Paper develops machine learning algorithms to learn optimal integer weights for clinical risk scores.
problem Deriving optimal integer weights for clinical risk scores without computational burden.
method Flexible greedy optimization strategy to directly optimize a value function.
result Constructed an integer-weighted comorbidity score for measuring post-discharge mortality risk.
The article reviews scoring rules for estimating and evaluating forecasts.
problem Evaluating probabilistic forecasts and estimating probability distributions.
method Mathematical foundations and characterization of scoring rules.
result Important families of scoring rules and their applications in statistics and machine learning.
The paper tackles fairness in scoring functions for binary classification.
problem Fairness in scoring functions for binary classification tasks.
method Introduces ROC-based fairness constraints and learning algorithms.
result Generalization bounds and practical learning algorithms for fair scoring functions.
Paper develops efficient methods for leverage score sampling and kernel ridge regression.
problem Efficiently sampling leverage scores for large matrices.
method Novel algorithm for leverage score sampling and kernel ridge regression solver.
result Proposed algorithms are the most efficient and accurate for leverage score sampling and kernel ridge regression.
SCORE improves tree-based predictions with boosted residual extraTrees.
problem Improving tree-based prediction models with reduced errors.
method Inspired by representation learning, SCORE uses boosting, regularized regression, and variable selection.
result SCORE provides comparable or superior performance compared to other models.
The paper addresses boundary term learning in reflected diffusion models.
problem Boundary term learning in reflected diffusion models to ensure correct boundary behavior.
method Integration by parts and reflection masking techniques to enforce boundary conditions.
result The conormal trace of the diffusion-weighted normal component is crucial for boundary term learning.
New active learning methods use statistical leverage scores to select examples efficiently.
problem Efficiently selecting labeled examples for high model accuracy with limited labeled data.
method Proposes ALEVS and DBALEVS methods based on statistical leverage scores.
result DBALEVS selects diverse, representative examples efficiently.
This paper introduces and develops a novel variable importance score function in the context of ensemble learning and demonstrates its appeal both theoretically and empirically. Our proposed score function is simple and more straightforward than its counterpart proposed in the context of random forest, and by avoiding …
New bounds for score matching in polynomial exponential families.
problem Understanding the sample complexity of score matching for polynomial exponential families.
method Non-asymptotic sample complexity analysis for score matching.
result First finite sample bounds for score matching in polynomial exponential families.
The H-score can bias bicluster comparisons, but a correction is provided.
problem Bias in H-score for comparing biclusters.
method Analytical proof and simulation to demonstrate bias; correction method provided.
result The H-score can be biased towards small clusters, but a correction is possible.
New machine learning models improve credit scoring in banks.
problem Improving credit scoring models in heavily regulated financial institutions.
method Gradient Boosting Machines (XGBoost) and Shapley Values.
result Improved performance and default capture rate compared to current models.
The paper investigates how calibrating propensity scores improves DML estimates of average treatment effects.
problem Improving the accuracy of DML estimates in finite samples.
method Propensity score calibration within the Double/debiased machine learning framework.
result Calibrating propensity scores reduces the root mean squared error of DML estimates of average treatment effects in finite samples.
Adaptive learning of SPDE solutions using score-based diffusion models.
problem Model errors and reduced accuracy in SPDE solutions due to incomplete physical knowledge and environmental variability.
method Score-based diffusion models with recursive Bayesian inference, incorporating simulation data and observational information.
result Accuracy and robustness of the proposed method demonstrated on benchmark SPDEs.
New analysis shows scores learn data manifolds better than distributions.
problem Learning the full distribution vs. just the data manifold.
method Novel analysis of scores in the small-σ regime.
result Scores learn data manifold information Θ(σ−2) stronger than distribution information. Score matching is a recently developed parameter learning method that is particularly effective to complicated high dimensional density models with intractable partition functions. In this paper, we study two issues that have not been completely resolved for score matching. First, we provide a formal link between maxim…
BSAC improves credit scoring models by leveraging autoencoders and addressing imbalanced datasets.
problem Imbalanced and heterogeneous credit scoring datasets.
method Bagging Supervised Autoencoder Classifier (BSAC) that uses autoencoders and undersampling.
result BSAC improves classification of loan applicants, demonstrating robustness and effectiveness.
Paper proposes a self-learning framework for reject inference in credit scoring.
problem Sample bias in credit scoring models due to training on accepted cases only.
method Develops a self-learning framework considering distinct training regimes for iterative labeling and model training, introduces a new evaluation measure.
result Demonstrates the superiority of the adjusted self-learning framework over regular self-learning and previous reject inference strategies.
Study shows how diffusion models learn on low-dimensional manifolds.
problem Learning efficiency of diffusion models on manifolds.
method Analyzes denoising score matching with random feature neural networks.
result Sample complexity scales linearly with intrinsic dimension, not ambient dimension.
A novel framework quantifies uncertainty using proper scores for various tasks.
problem Uncertainty quantification in machine learning for reliable applications.
method Proposes a general framework based on proper scores for epistemic, aleatoric uncertainty, and model calibration.
result Achieves state-of-the-art uncertainty estimation for large language models and generative models.
Automated denoising score matching handles nonlinear diffusion processes.
problem Nonlinear diffusion processes limit generative modeling and property estimation.
method Local-DSM using local increments and Taylor expansions.
result Tractable training and score estimation for nonlinear diffusion processes.
This paper improves random feature sampling using empirical leverage scores.
problem Optimizing the number of features for kernel approximation and supervised learning.
method Uses empirical leverage scores to optimize feature sampling.
result Empirical sampling of random features using leverage scores outperforms vanilla Monte Carlo sampling.
Generalizes leverage score sampling for neural networks, accelerating kernel methods and deep learning.
problem Accelerating kernel methods and deep learning training.
method Generalizes leverage score sampling to neural networks and proves equivalence to neural tangent kernel ridge regression.
result Equivalence between regularized neural network and neural tangent kernel ridge regression under leverage score sampling initialization.
The paper argues for using Neyman orthogonal score for balancing in debiased machine learning.
problem Debiased machine learning requires a proper approach to balance covariates.
method The paper advocates for using Riesz regression with basis functions of X for balancing.
result Covariate balancing is only valid when the score-relevant regression error is a function of covariates alone.
Bayesian scores improve structure learning in probabilistic circuits.
problem Improper structure learning in probabilistic circuits based on heuristics.
method Developed Bayesian structure scores for deterministic PCs, using them in a greedy cutset algorithm.
result Effective protection against overfitting and fast, almost hyper-parameter-free structure learner.
New machine learning methods for inference from simulated data.
problem Modeling score and likelihood ratio functions from sampled data.
method InferoStatic Networks (ISN), Kernel Score Estimation (KSE), Kernel Likelihood Ratio Estimation (KLRE).
result Improved inference methods for complex models.
Score function estimators improve k-subset sampling efficiency.
problem Efficiently sampling k-subsets in machine learning tasks. method Revisit score function estimators, using discrete Fourier transform and control variates.
result Efficient and unbiased gradient estimates for k-subset sampling. Deep networks can approximate score functions in high-dimensional graphical models efficiently.
problem Approximation efficiency of score functions by deep neural networks in high-dimensional graphical models like Markov random fields.
method Variational inference denoising algorithms and efficient neural network representation.
result Efficient sample complexity bound for diffusion-based generative modeling when score functions are learned by deep neural networks.
Simplifies denoising score matching for manifold learning.
problem Learning distributions on manifolds is computationally intensive.
method Modifies denoising score matching to implicitly account for the manifold.
result Reduces computational burden while maintaining efficiency.
New method improves IL from imperfect demos using confidence scores.
problem Learning optimal policies from imperfect demonstrations is challenging.
method Proposes two confidence-based IL methods: 2IWIL and IC-GAIL.
result Confidence scores from sub-optimal demos significantly improve IL performance.
Study improves credit scoring model calibration using machine learning techniques.
problem Improving accuracy of probability of default predictions.
method Exploring calibration techniques (Platt Scaling, Isotonic Regression) and machine learning models (Logistic Regression, Random Forest, Gradient Boosting) on real-world datasets.
result Isotonic Regression re-calibration improves long-term model accuracy.
Interactive tool helps analyze classifier prediction scores.
problem Understanding and assessing the reliability of multi-class classifiers.
method Interactive visualization tool (Classilist) to analyze prediction scores.
result Varying classifier behavior revealed through analysis of prediction scores.
SBMs learn manifold-like structures by mixing samples with a non-conservative field.
problem How SBMs learn data distributions on low-dimensional manifolds.
method Investigating linear approximations and subspaces of local feature vectors during diffusion.
result SBMs mix samples by a non-conservative field within the manifold, maintaining manifold-like structure.
Framework scores DeFi users based on liquidity and trading behavior.
problem Distinguishing between liquidity provision and active trading in DeFi.
method Rule-based decomposition, deep residual neural network, pool-level context.
result Deep residual neural network improves user scoring and risk assessment.
Develops Hamiltonian Score Matching and Generative Flows for machine learning.
problem Estimating score functions and designing generative models.
method Introduces Hamiltonian velocity predictors (HVPs) for score matching and generative flows.
result Hamiltonian Generative Flows (HGFs) rival leading generative modeling techniques.
New geometric analysis shows L2 score error is flawed for diffusion models.
problem Score matching errors in diffusion models do not fully capture distributional quality.
method Decomposed score errors into gradient and solenoidal components, focusing on gradient's role in Fokker-Planck dynamics.
result Only gradient component affects marginal distributional quality; solenoidal component is structurally invisible.