Efficiently learns Ising models with missing data.
problem Learning Ising models with missing data due to independent failures.
method Developed a novel unbiased estimator for the ISO gradient and applied stochastic multiplicative gradient descent.
result Matches optimal runtime and sample complexity bounds for learning Ising models.
A new method normalizes activations to match batch normalization without batch dependence.
problem Performance degradation with batch-independent normalization techniques.
method Proxy-Normalizing Activations
result Proxy-Normalization technique emulates batch normalization's behavior and performance.
Unified model predicts multi-mode failure with multi-sensor data.
problem Independent failure mode and RUL prediction ignores inherent relationship.
method Hierarchical Bayesian framework with Cox model, Gaussian process, and multinomial distributions.
result Robust uncertainty quantification and accurate prediction of multi-mode failure.
Machine learning predicts circulatory failure in ICU patients.
problem Limited ability of clinicians to recognize early signs of patient deterioration.
method Developed an early warning system using machine learning on ICU data.
result Predicts 90.0% of circulatory failure events with 81.8% identified more than two hours in advance.
New method predicts RUL and failure modes from partial data.
problem Predicting RUL and failure modes from incomplete data.
method Formulated as vector General Value Function (GVF) prediction on an absorbing degradation process, using TD(n,λ) for estimation. result TD improves RUL and failure-mode prediction compared to Monte Carlo methods, especially under scarce complete labels.
New framework tests AVs as a black box, prioritizing rare failure modes.
problem Lack of rigorous and scalable testing methods for AVs.
method Developed a simulation testing framework that learns to identify and rank failure scenarios via adaptive importance-sampling methods.
result First independent evaluation of a full-stack commercial AV system (Comma AI's OpenPilot).
Feedback alignment methods need to be evaluated for accuracy and gradient cosine similarity.
problem Evaluating feedback alignment methods
method Proposed diagnostic evaluation protocol
result Identified silent failures in standard reporting pair
Unified model predicts equipment failure and remaining useful life.
problem Predicting equipment failure and remaining useful life separately is sub-optimal.
method Two methods: Deep Weibull model (DW-RNN) and multi-task learning (MTL-RNN).
result Our methods consistently outperform baseline RUL methods and produce consistent results for RUL and FP.
Network analysis improves risk assessment for surety bonds.
problem Network effects in surety bonds increase risk assessment complexity.
method Modelled contractor network as directed graph, extended Friedkin-Johnsen model with stochastic process.
result Network effects increase average risk for surety organizations.
A new boosting model handles dependent censoring in time-to-event data.
problem Independent censoring assumption leads to biased predictions in time-to-event analysis.
method Clayton-boost, a boosting approach using Clayton copula.
result Clayton-boost outperforms other methods in handling dependent censoring.
Paper develops streaming algorithms to estimate classifier accuracy on unlabeled data.
problem Estimating classifier accuracy on unlabeled data with noisy decisions.
method Two algebraic evaluators: majority voting and a novel method to handle correlated classifiers.
result The novel method can be as accurate as 1% when handling small amounts of correlation.
Develops a method for causal inference in recurrent event data with terminal failure.
problem Causal inference in recurrent event data with a terminal event.
method Multiply robust estimation framework for causal inference.
result Proposes an estimator for the expected number of recurrent events and failure survival function.
The study diagnoses fairness issues in healthcare models under distribution shifts.
problem Understanding and diagnosing fairness changes in machine learning models under distribution shifts in healthcare.
method Causal framing and conditional independence tests to characterize distribution shifts.
result Knowledge of distribution shifts helps diagnose fairness transfer failures, including complex cases.
Proposes a federated learning approach for industrial asset failure prediction.
problem Lack of data and privacy concerns in industrial prognostics.
method Two-stage federated learning: dimension reduction and parameter estimation.
result Validated the approach using simulated and real data.
A new algorithm selects independent coordinates for complex manifolds.
problem Embedding algorithms fail with large aspect ratio manifolds.
method IES algorithm selects smooth embeddings using carefully chosen eigenfunctions of the Laplace-Beltrami operator.
result The IES algorithm successfully embeds synthetic and real data.
Survival models predict component failures using neural networks and resampled data.
problem Accurately predicting component failure times for maintenance planning.
method Neural network-based survival models trained on non-independent, homogeneously sampled data.
result Random resampling during training reduces dataset size and improves efficiency.
Differentiable sorting and rank normalization are incompatible, with specific conditions for admissibility.
problem Incompatibility between differentiable sorting and rank normalization.
method Formalized admissibility through monotone invariance, batch independence, and rank-space stability conditions.
result Different gap-sensitive and batchwise relaxations of rank normalization violate the conditions for admissibility.
Study shows E2E training fails for over-parameterized models.
problem When does E2E training fail for complex Deep Network architectures?
method Blend gradient between independent training and E2E training.
result Optimum can lie between ensemble and E2E, challenging traditional approaches.
Economics tool predicts failure times in reliability systems.
problem Predicting optimal failure times in weighted k-out-of-n reliability systems with heterogeneous component failure.
method Using rational expectations to analyze and predict failure times in reliability systems with heterogeneous component failure.
result Different measures are optimal for predicting system failure depending on component failure distributions.
New method uses probabilistic independence to discover disease signatures from medical records.
problem Insufficiently precise diagnosis of clinical disease leading to treatment failures.
method Unsupervised machine learning using probabilistic independence to disentangle disease patterns.
result Inferred 2000 clinical disease signatures from medical records, improving cancer prediction.
ECCIT improves conditional independence tests by calibrating for miscalibration.
problem Inaccurate frequentist guarantees in CITs, especially in small samples and misspecified models.
method Empirically Calibrated Conditional Independence Tests (ECCIT) that optimize and correct for miscalibration.
result ECCIT achieves valid FDR with higher power than existing calibration strategies.
Framework classifies machine learning failures into intentional and unintentional.
problem Understanding and preventing failures in machine learning systems.
method Developed a taxonomy of machine learning failure modes.
result Stakeholders found the framework useful for discussing machine learning failures.
GAN-FP uses GANs to predict equipment failures from imbalanced data.
problem Accurately predicting equipment failures with limited data and high imbalance.
method GAN-FP employs two GAN networks to generate and classify imbalanced data, optimizing a weighted loss objective and a consistency GAN.
result GAN-FP outperforms traditional methods in imbalanced failure prediction.
PAGER detects failures in deep regression models using a new framework.
problem Detecting failures in deep regression models.
method PAGER uses a combination of epistemic uncertainty and manifold non-conformity scores.
result PAGER accurately characterizes and detects failures in deep regressors.
New insights into CI tests reveal key factors for practical performance.
problem Understanding and improving CI tests in practical applications.
method Investigation of the Kernel-based Conditional Independence (KCI) test and analysis of its practical behavior.
result Errors in conditional mean embedding estimates and appropriate conditioning kernel selection are crucial for CI tests.
Predict and explain service failures in supply-chain networks using data models.
problem Predict and explain service failures in supply-chain networks, particularly last-mile pickup and delivery.
method Used supervised classification with Random Forests and Association Rules on a dataset of 500,000 services.
result Classifier reaches an average sensitivity of 0.7 and specificity of 0.7 for 5 types of failure.
Risk Advisor predicts and mitigates ML deployment failures.
problem Predicting and mitigating test-time failure risks of ML systems.
method Post-hoc meta-learner for estimating failure risks and uncertainties.
result Reliably predicts deployment-time failure risks across various ML models.
CalNF models rare failures with limited data, improving safety in autonomous systems.
problem Challenges in modeling and debugging rare safety-critical failures due to limited data.
method CalNF, a self-regularized framework for posterior learning from limited data.
result Achieves state-of-the-art performance on data-limited failure modeling and inverse problems.
New method predicts rare failures in aerospace systems.
problem Rare failure prediction in aerospace applications.
method Event matching based on technical system peculiarities.
result Illustrated the method's applicability on aircraft operations.
This work improves safety validation of autonomous vehicles by finding interpretable failures.
problem Finding interpretable failures of autonomous systems in simulation.
method Signal temporal logic expressions optimized for high likelihood and human interpretability.
result Our methodology finds more interpretable failures with higher likelihood compared to baseline approaches.
Framework predicts remaining useful life of DSH subsystems under unknown failure modes.
problem Predicting remaining useful life of DSH subsystems with unknown failure modes.
method Unsupervised framework using mixture of Gaussian regressions and Expectation-Maximization algorithm.
result Improved prediction accuracy and interpretability of RUL.
Credit networks represent a way of modeling trust between entities in a network. Nodes in the network print their own currency and trust each other for a certain amount of each other's currency. This allows the network to serve as a decentralized payment infrastructure---arbitrary payments can be routed through the net…
Adversarial method finds rare catastrophic failures in safety-critical agents.
problem Evaluating safety-critical learning systems for catastrophic failures.
method Adversarial evaluation approach focusing on rare adversarial situations.
result Adversarial evaluation finds catastrophic failures and estimates failure rates faster.
Alternative to likelihood-based LSNM model selection, residual independence testing is more robust to noise misspecification.
problem Cause-effect inference in location-scale noise models with misspecified noise distributions.
method Residual independence testing as an alternative to likelihood-based model selection.
result Residual independence testing is more robust to noise misspecification.
LLMs struggle to generate random numbers from statistical distributions, leading to biased results in applications.
problem LLMs' inability to generate random numbers accurately from specified distributions.
method Dual-protocol design: Batch Generation and Independent Requests, benchmarking 11 models across 15 distributions.
result Sampling fidelity degrades with distributional complexity and horizon, leading to systematic biases in downstream applications.
The paper calculates the likelihood of a financial market failure involving multiple major banks.
problem Estimating the probability of a market failure involving multiple globally important banks.
method Multivariate Cox process across G-SIBs, deriving various theorems on market failure probabilities.
result The probability of a market failure increases with the number of G-SIBs and is inevitable if there are too many.
SSMs can be poisoned with clean labels, leading to generalization failure.
problem The implicit bias of SSMs can be manipulated by including special training examples with clean labels.
method Formal proof and empirical demonstration of the phenomenon.
result SSMs can fail to generalize even with clean labels, due to the inclusion of special training examples.
Paper proposes a method to predict disk failures using multi-layer domain adaptive learning.
problem Traditional machine learning models struggle to predict disk failures due to limited data.
method Multi-layer domain adaptive learning with source and target domains.
result The proposed method improves failure prediction accuracy on disk data with few failure samples.
GE finds failures in autonomous systems without domain heuristics.
problem Finding failures in autonomous systems without domain-specific heuristics.
method Adaptive stress testing using go-explore (GE) algorithm.
result GE finds failures in scenarios other RL techniques cannot solve.
Improved AST method finds more useful failure scenarios for autonomous vehicles.
problem Finding useful failure scenarios for autonomous vehicle validation is challenging.
method Adaptive Stress Testing with reward augmentation, modified to encode domain information.
result The modified AST method discovers a larger and more expressive subset of failure scenarios.
Method distinguishes between failures and domain shifts in industrial data streams.
problem Confusing domain shifts with failures in industrial data.
method Modified Page-Hinkley changepoint detector and supervised domain-adaptation-based anomaly detection.
result Allows differentiation between failures and domain shifts.
Predicts failure of autonomous vehicle steering control models.
problem Evaluating and predicting failure of machine learning models in safety-critical applications.
method Trains a student model to predict the main model's error based on saliency maps.
result Preliminary results show the failure predictor model works on autonomous vehicle steering control systems.
System predicts respiratory failure up to 8 hours early.
problem Early detection of respiratory failure in ICU patients.
method Machine learning on ICU patient monitoring data.
result System outperforms traditional clinical decision-making.
New method avoids failures in physics-constrained systems using active learning.
problem Handling fatal failures in systems governed by physics constraints.
method Develops a novel active learning method that considers implicit physics constraints.
result Achieves zero-failure in composite fuselage assembly process without explicit failure regions.
DeepSIP predicts network failures' impact using CNN from syslog and traffic data.
problem Predicting service impact from network failures.
method Temporal multimodal CNN for predicting time to recovery and traffic loss.
result DeepSIP reduced prediction error by approximately 50%.
RODMAN improves ML-based disk failure prediction accuracy in cloud environments.
problem Imperfect data quality in real-world cloud environments degrades ML-based disk failure prediction accuracy.
method RODMAN uses three data preprocessing techniques: failure-type filtering, spline-based data filling, and automated pre-failure backtracking.
result RODMAN significantly improves prediction accuracy compared to no preprocessing.
Two BO methods improve reliability optimization for rare failures.
problem Maximizing reliability of designs subject to random perturbations.
method Bayesian optimization with Thompson sampling and knowledge gradient.
result Proposed methods outperform existing techniques in extreme failure probability scenarios.
New method finds failures in high-fidelity simulators with fewer steps.
problem Finding failures in high-fidelity simulators is expensive and impractical.
method Adaptive stress testing with backward algorithm adaptation from low-fidelity to high-fidelity.
result Significantly fewer high-fidelity simulation steps needed to find failures.