CorrCA identifies reliable dimensions in multivariate data across repetitions.
problem Finding consistent dimensions in multivariate data across trials, subjects, or raters.
method Maximizes the ratio of between-repetition to within-repetition covariance.
result CorrCA leads to repeat-reliability maximization and is equivalent to Linear Discriminant Analysis for zero-mean signals.
SQAIR generates videos of moving objects, tracking and predicting them reliably.
problem Generating videos of moving objects with reliable object tracking and prediction.
method Explicitly encoding object presence, locations, and appearances in latent variables.
result SQAIR reliably discovers and tracks objects throughout sequences of frames and generates future frames.
This paper tackles reliability analysis for stochastic systems using surrogate models.
problem Traditional reliability analysis relies on deterministic models, which are not suitable for stochastic systems with non-repeatable outcomes.
method The paper introduces reliability analysis for stochastic models by using generalized lambda models and stochastic polynomial chaos expansions as surrogate models to lower computational cost.
result The surrogate models enable efficient uncertainty quantification at a lower cost than traditional Monte Carlo simulation.
This work studies how an AI-controlled dog-fighting agent with tunable decision-making parameters can learn to optimize performance against an intelligent adversary, as measured by a stochastic objective function evaluated on simulated combat engagements. Gaussian process Bayesian optimization (GPBO) techniques are dev…
New model accounts for sequential dependence in LLM reliability.
problem Uncertainty in LLM reliability assessment due to sequential interactions.
method Extended Bayesian framework with Hidden Markov Model for sequential dependence.
result Ignoring sequential dependence leads to overconfident reliability estimates.
Ribbon: Scalable Approximation and Robust Uncertainty Quantification
problem Reliably quantifying predictive uncertainty for complex models
method Ribbon, a scalable approximation to Dirichlet-reweighted bootstrap uncertainty
result Asymptotically equivalent to a flat-prior Laplace approximation under correct likelihood specification, recovers robust sandwich covariance under misspecification
AL-SPCE improves reliability analysis for complex systems with active learning and SPCE.
problem Efficiently analyzing reliability of complex, computationally expensive models with intrinsic randomness.
method Active learning framework using stochastic polynomial chaos expansions (SPCE) to reduce computational burden.
result AL-SPCE maintains high accuracy in reliability estimates while significantly improving efficiency.
CoNBONet improves reliability analysis of complex systems with fast, energy-efficient predictions.
problem Time-dependent reliability analysis of nonlinear systems under stochastic excitations is computationally demanding.
method CoNBONet combines deep operator networks with neuroscience-inspired neuron models for fast, energy-efficient inference.
result CoNBONet provides reliable coverage of failure probabilities with theoretical guarantees.
Adaptive Quantum Conformal Prediction improves reliability of quantum machine learning predictions.
problem Quantum machine learning lacks robust uncertainty quantification methods.
method Adaptive Conformal Inference applied to quantum conformal prediction to maintain validity over time.
result AQCP achieves target coverage levels and is more stable than standard quantum conformal prediction.
Paper tackles machine performance testing under uncertain inputs.
problem Guarantee machine performance under input uncertainty.
method Formulates as IU-rLSE problem, proposes active learning method.
result Efficient algorithm for reliable level set estimation.
LIME method provides stability indices to ensure reliable explanations for machine learning models.
problem Difficulty in understanding machine learning model decisions.
method LIME method with stability indices.
result Stability indices ensure reliable explanations for machine learning models.
Small sample sizes lead to unreliable error bars in predictive models.
problem Unreliable error bars in cross-validation due to small sample sizes.
method Simple experiments show that sample sizes of neuroimaging studies lead to large error bars.
result Error bars from cross-validation underestimate true variability.
Advocates a local feedback approach for RL in unknown systems.
problem Finding optimal feedback laws in unknown nonlinear dynamical systems.
method Searches over a local feedback representation consisting of an open-loop sequence and an optimal linear feedback law.
result Results in highly efficient training and superior performance compared to global methods.
MONSTOR estimates influence in unseen networks with high accuracy.
problem Estimating and maximizing influence in social networks.
method Inductive machine learning approach replacing Monte Carlo simulations.
result Highly accurate estimates with strong correlation to actual influence.
New MRI method maps tissue parameters more accurately by ignoring voxel independence.
problem Voxel independence assumption limits model fitting reliability and repeatability.
method Self-supervised deep variational approach with Gaussian mixture prior.
result Our method outperforms current techniques in dMRI simulations and real data.
Study improves reliability of reinforcement learning with real-world robots.
problem Difficulty in setting up reinforcement learning tasks with real-world robots.
method Developed a learning task with a UR5 robotic arm to study task setup.
result Learning performance is highly sensitive to task setup details.
Sequence probability predicts correctness in LLMs, but not for repeated prompts
problem Predicting correctness in large language models
method Quantifying sequence probability and correctness across different levels
result Higher sequence probability often predicts correctness across prompt-answer pairs
A key drawback of the current generation of artificial decision-makers is that they do not adapt well to changes in unexpected situations. This paper addresses the situation in which an AI for aerial dog fighting, with tunable parameters that govern its behavior, will optimize behavior with respect to an objective func…
TSDS framework reduces edge LLM agent compute by 43%-73% while maintaining safety and reliability.
problem Managing reasoning budget and uncertainty in edge LLM agents.
method Integrates a lightweight convergence probe and a perplexity-based deferral rule calibrated via multi-objective LTT.
result Reduces per-episode thinking compute by 43%-73% over deferral-only baselines.
Flexible framework assesses multilevel data group heterogeneity.
problem Multilevel data structure complicates model selection.
method Flexible framework for assessing differences between levels of grouping variables.
result Framework reliably identifies relevant multilevel components.
The paper studies how to allocate human validation in AI-assisted tasks to minimize errors.
problem Heterogeneous reliability of AI-generated signals across tasks, products, and customer segments.
method Tuned prediction-powered inference, upper confidence bounds policy, Neyman square-root rule.
result The proposed policy outperforms uniform and epsilon-greedy allocation, closing most of the gap to the oracle when reliability is heterogeneous.
DIGIT is a low-cost tactile sensor for in-hand manipulation.
problem Difficulty in sensing contact forces limits robotic manipulation.
method DIGIT miniaturizes and improves a vision-based tactile sensor.
result DIGIT enables better control of interactions with the environment.
A new method improves feature importance and model stress-testing reliability.
problem Estimating feature contributions in machine learning models for trust and transparency.
method Replacing multiple random permutations with a single, deterministic, and optimal permutation.
result Improved bias-variance tradeoffs and accuracy in challenging scenarios.
New method predicts language model scaling risks with limited data.
problem Accurately predicting model capabilities and risks with limited data.
method Introduces a robust estimation framework using beta-binomial distribution and dynamic sampling strategy.
result More accurate predictions of rare risks and capabilities at lower computational cost.
Study compares crowdsourcing with lab experiments using comparison-based psychophysics.
problem Improving data quality in crowdsourcing psychophysics experiments.
method Comparison-based psychophysics, machine learning for triplet prediction.
result Accuracy of crowdsourcing psychophysics close to lab experiments.
Proposes using probabilistic models for privacy-preserving synthetic data.
problem Designing high-quality synthetic data for privacy preservation.
method Formulate the problem through probabilistic modelling, choosing a model for the data.
result Statistical discoveries can be reliably reproduced from synthetic data.
TDA improves accuracy of machine learning models for repeated measurements.
problem Limited accuracy of machine learning models for repeated measurements.
method Samples from data space, builds network graph based on data topology.
result TDA classifier achieves high accuracy (up to 96.8%) in repeated measurement datasets.
Algorithm learns to bid in auctions with shilling, masking real bids.
problem Learning to bid in auctions manipulated by shilling.
method Combines interval-elimination and optimistic branches, debiases losing-side reports.
result Achieves dynamic-pricing rate i l d e O ( T 2 / 3 ) ilde{\mathcal{O}}(T^{2/3}) i l d e O ( T 2/3 ) and first-price auctions rate i l d e O ( T ) ilde{\mathcal{O}}(\sqrt{T}) i l d e O ( T ) . Optimal strategies are found for a repeated betting game using diffusion approximation.
problem Finding optimal strategies for a repeated betting game with i.i.d. outcomes.
method Constructing a diffusion approximation of the repeated game and analyzing the wealth share process.
result Necessary and sufficient conditions for the wealth share process to be transient or recurrent are derived.
Over the past two decades, several consistent procedures have been designed to infer causal conclusions from observational data. We prove that if the true causal network might be an arbitrary, linear Gaussian network or a discrete Bayes network, then every unambiguous causal conclusion produced by a consistent method f…
MixFlows uses a mixture of flows for efficient variational inference.
problem Efficient and reliable variational inference for complex models.
method A new variational family of mixed flows with efficient algorithms and convergence guarantees.
result MixFlows provides more reliable posterior approximations and comparable sample quality to MCMC methods.
Convolutional networks struggle with repeating patterns in ECGs.
problem Modeling repeating patterns in electrocardiogram signals.
method Demonstrated through ECG examples, highlighting systemic issues in deep learning.
result Counterintuitive effects on generalization in deep networks.
New method reduces uncertainty in AI-driven Monte Carlo simulations.
problem Epistemic uncertainty in AI surrogate models affects Monte Carlo sampling outcomes.
method Penalty Ensemble Method (PEM) modifies Metropolis acceptance rule to increase rejection probability in uncertain regions.
result PEM enhances reliability of Monte Carlo simulations by reducing uncertainty propagation.
Study optimizes auction pricing for strategic bidders in repeated auctions.
problem Optimizing revenue in auctions with multiple strategic bidders.
method Proposes a novel algorithm with strategic regret bound of O(log log T).
result Algorithm learns strategic buyer's valuation with theoretical guarantees.
Study geometric properties of symmetric matrices with repeated eigenvalues.
problem Investigate geometric properties of symmetric matrices with repeated eigenvalues.
method Explicitly compute the volume of the intersection with the sphere and prove an Eckart-Young-Mirsky-type theorem.
result Prove connections to Real Algebraic Geometry and Random Matrix Theory.
Study explores algorithmic collusion in repeated games using various learning dynamics.
problem Understanding algorithmic collusion in repeated games with different learning dynamics.
method Examines Q Q Q -learning, gradient learning, and other dynamics in a general repeated game setting. result Characterizes the set of payoff vectors achievable by these dynamics, revealing possibilities for collusion.
We study two systems of tangle equations that arise when modeling the action of the Integrase family of proteins on DNA. These two systems--direct and inverted repeats--correspond to two different possibilities for the initial DNA sequence. We present one new class of solutions to the tangle equations. In the case of i…
Scorio.jl ranks systems from repeated tasks using various methods.
problem Evaluating and ranking systems from repeated responses to shared tasks.
method Common tensor-based interface for multiple ranking methods.
result Pilot experiments show stability and runtime scaling.
We investigate the question of when distinct branched surfaces in the complement of a 2-bridge knot support essential surfaces with identical boundary slopes. We determine all instances in which this occurs and identify an infinite family of knots for which no boundary slopes are repeated.
CM algorithm improves MMI classifications for unseen instances.
problem Improving classification accuracy for unseen instances using MMI criterion.
method Introduces CM algorithm for MMI classifications, combining semantic and Shannon channels for matching.
result Achieves high mutual information (99%) with minimal iterations in low-dimensional feature spaces.
LAFF algorithm balances adaptability and non-exploitability in repeated games.
problem Low regret in repeated games against unknown opponent classes.
method LAFF algorithm searches within sub-algorithms optimal for each opponent class and uses a punishment policy for exploitation.
result LAFF guarantees sublinear regret uniformly over possible opponents, except exploitative ones, for which it guarantees linear regret.
Over the past years Robust PCA has been established as a standard tool for reliable low-rank approximation of matrices in the presence of outliers. Recently, the Robust PCA approach via nuclear norm minimization has been extended to matrices with linear structures which appear in applications such as system identificat…
Link concordance and Whitney towers linked to Milnor invariants.
problem Link concordance and Whitney towers classification.
method Clasper surgeries, Whitney towers, and Milnor invariants.
result Link concordance and Whitney towers classified in terms of Milnor invariants.
Investigates the number of experiments needed for statistical significance in medication testing.
problem Determining the number of experiments needed for a statistically significant result.
method Examines binomial and general probability distributions, considering placebo efficacy and varying distributions.
result The number of experiments needed can be significantly higher when placebo efficacy is considered.
UA-SABI uses surrogates to speed up Bayesian inference for expensive models.
problem Inference for computationally expensive models is slow and uncertain.
method Combines surrogate modeling with Amortized Bayesian Inference (ABI) to propagate uncertainties.
result Reliable, fast, and repeated Bayesian inference for expensive models is achieved.
This paper shows hedging algorithms improve performance in repeated matrix games.
problem Improving multi-agent learning algorithms in repeated matrix games.
method Develops and experiments with hedging algorithms combining a top-level and a set of basic algorithms.
result Well-selected hedging algorithms outperform previous MAL algorithms on repeated matrix games.
Algorithm learns to prune search space for repeated computations.
problem Exploit common structure in repeated similar problems.
method Exploit explore-exploit technique for pruning search space.
result Reduces runtime while provably outputting correct solutions.
Risk is part of the fabric of every business; surprisingly, there is little work on establishing best practices for systematic, repeatable risk identification, arguably the first step of any risk management process. In this paper, we present a proposal that constitutes a more holistic risk management approach, a method…