The study improves off-policy learning by smoothing IPS and provides a generalization bound.
problem Improving off-policy learning from logged bandit data.
method Smooth regularization for IPS, deriving a two-sided PAC-Bayes generalization bound.
result The bound is valid for standard IPS and provides insights into when regularization is useful.
Pessimistic estimator improves multi-objective policy optimization.
problem Optimizing multi-objective policies from existing data.
method Pessimistic estimator based on inverse propensity scores (IPS).
result Pessimistic estimator outperforms naive IPS estimator in theory and experiments.
New estimator SWITCH improves off-policy evaluation in contextual bandits.
problem Estimating value of a target policy using data from another policy in contextual bandits.
method Proposes SWITCH estimator that uses an existing reward model to outperform IPS and DR.
result Switch estimator achieves better MSE than IPS and DR, often outperforming prior work.
AIPS improves ranking policy evaluation by adapting to diverse user behavior.
problem Inaccurate Off-Policy Evaluation of ranking policies due to high variance under diverse user behavior.
method Developed Adaptive IPS (AIPS) that adapts to different user behaviors and minimizes MSE.
result AIPS achieves minimum variance among unbiased estimators and provides significant empirical accuracy improvement.
Contextual bandit methods fail with deficient support data.
problem Learning from support-deficient data in contextual bandits.
method Three approaches to IPS-based learning: action space restriction, reward extrapolation, and policy space restriction.
result Systematic analysis and empirical evaluation of approaches to IPS-based learning.
A new estimator reduces variance in slate bandit OPE.
problem Large action spaces in slate bandits cause high variance in OPE.
method Develops Latent IPS (LIPS) to optimize slate abstractions for low variance and bias.
result LIPS substantially outperforms existing estimators in scenarios with non-linear rewards and large slate spaces.
Modeling shows stronger IP protection speeds tech advancement.
problem Effect of intellectual property policy on tech advancement speed.
method Agent-based modeling with varying IP protection levels.
result Stronger IP protection leads to faster technological progress.
New DR-IC estimator reduces bias and variance in OPE.
problem Estimating value of a target policy using logged data from a different policy.
method DR-IC estimator that combines parametric reward model and context-based switching rule.
result DR-IC estimator outperforms state-of-the-art OPE algorithms.
CAEL-MIPS learns embeddings to improve MIPS for better OPE in contextual bandits.
problem High variance in IPS weighting for OPE in large action spaces.
method Context-Action Embedding Learning (CAEL) for MIPS to minimize MSE.
result CAEL-MIPS outperforms baselines in MSE for OPE in contextual bandits.
CAB estimator improves evaluation and learning performance in policy contexts.
problem Offline A/B-testing and off-policy learning using logged contextual bandit feedback.
method Continuous Adaptive Blending (CAB) estimator, subsuming most counterfactual estimators.
result CAB estimator is less biased and has less variance than other estimators.
New MLIPS method reduces error in batch bandit learning.
problem Learning from logged bandit feedback with policy discrepancy.
method Estimates a surrogate policy to reduce error in inverse propensity weighting.
result MLIPS has smaller mean squared error than IPS.
New estimator uses clustering to improve off-policy evaluation accuracy.
problem Improving off-policy evaluation accuracy when logging and evaluation policies differ.
method Proposes an estimator that shares information across similar contexts using clustering.
result Clustering contexts improves estimation accuracy, especially in deficient information settings.
New method tunes prior IP to data for flexible predictive distributions.
problem Challenges in approximate inference for large models with high parameter dependencies.
method Inducing-point representation of prior IP to approximate posterior process.
result Scalable method that tunes prior IP to data and provides accurate non-Gaussian predictive distributions.
RL improves IP solver performance by learning to select cutting planes.
problem Improving the performance of IP solvers through heuristic optimization.
method Employing reinforcement learning to intelligently select cutting planes in the Cutting Plane Method.
result Trained RL agent significantly outperforms human-designed heuristics across various IP tasks.
The paper improves IPS for modern optimization, scaling and regularizing it.
problem Improving iterative proportional scaling for modern optimization.
method Coordinate descent, majorization-minimization, optimization techniques, regularized variants.
result IPS can deliver coefficient estimates and handle log-affine models.
DVIP improves on IP-based methods by using IPs as priors over latent functions.
problem Limited expressiveness of IP-based models, especially in function space.
method Proposes DVIP, a multi-layer generalization of IPs, and scalable variational inference.
result DVIP outperforms previous IP-based methods and deep GPs in regression and classification tasks.
A new model for fair IP monetization.
problem Inequitable sharing between IP creators and users.
method A mutual exchange model where users pay for use of IP, and creators receive payment when IP is no longer needed.
result Proposes a new model that aims to improve the coexistence of IP creators and users.
Study generalizes property elicitation to imprecise probabilities.
problem Minimizing risk over imprecise probability distributions.
method Maximin risk minimization over a set of imprecise probabilities.
result Conditions for elicitability of IP-properties.
A new model SIPS improves graph embedding by approximating non-PD similarities.
problem Improving neural network-based graph embedding by approximating non-positive definite similarities.
method Shifted Inner Product Similarity (SIPS) model that approximates Conditionally Positive Definite (CPD) similarities.
result SIPS significantly improves graph embedding without configuring the similarity function.
Extends V-IP framework to use LLMs for generating task-relevant concepts, improving interpretability and performance.
problem Limited applicability of V-IP to small-scale tasks due to manual data annotation.
method Integrates Foundational Models with Large Language and Multimodal Models to generate and annotate concepts.
result FM+V-IP achieves better test performance with fewer concepts/queries compared to other frameworks.
MEC-IP uses IP to efficiently find MECs in BNs from observational data.
problem Discovering Markov Equivalent Classes (MECs) in Bayesian Networks (BNs) efficiently.
method Clique-focusing strategy and EMSG for MEC discovery via Integer Programming.
result Significant reduction in computational time and improved accuracy.
A new method for making interpretable predictions by sequentially asking questions, faster and more efficient.
problem Developing interpretable machine learning models for complex tasks.
method Variational Information Pursuit (V-IP) that bypasses the need for learning generative models.
result V-IP is 10-100x faster and finds shorter query chains compared to IP and reinforcement learning.
New estimator GMIPS reduces variance in ranking policy evaluation.
problem High variance in off-policy evaluation for ranking policies.
method GMIPS estimator with user behavior model on ranking embedding spaces.
result GMIPS achieves lowest MSE and balances bias-variance trade-off.
Two machine learning applications for IP/Optical networks: traffic prediction and optical path performance.
problem Agile resource management and optical path performance prediction in IP/Optical networks.
method Machine learning for traffic prediction and optical performance prediction using SDN controllers.
result Efficient implementation of SDN controllers for agile resource management and optical path performance prediction.
A new method combines online and offline learning to tackle contextual bandits with missing action support.
problem Learning optimal policies with logged data when the logging policy has deficient support.
method Hybrid approach using online exploration to exploit supported actions and offline learning to avoid unnecessary explorations.
result Determines an optimal policy with theoretical guarantees using minimal online explorations.
A Riemannian manifold is called IP, if the eigenvalues of its skew-symmetric curvature operator are pointwise constant. It was previously shown that for all n\ge 4, except n=7, any IP manifold either has constant curvature, or is a warped product, with some specific function, of a line and a space of constant curvature…
VIPs use IPs for efficient inference in flexible models.
problem Efficient inference in flexible models like Bayesian neural networks and Gaussian processes.
method Variational Implicit Processes (VIPs) using generalised wake-sleep updates.
result VIPs provide better uncertainty estimates and lower errors compared to existing methods.
Paper shows non-integrality of dike building model and provides conditions for integrality.
problem Determining the integrality of a dike building model for flood protection.
method Analyzes experimental data and mathematical proofs to establish conditions for integrality.
result Established non-integrality of the polytope and conditions for linear programming relaxation to be integral.
C-IP improves LLMs' query selection for interactive tasks by estimating uncertainty robustly.
problem Minimizing the number of queries for interactive LLMs.
method Conformal Information Pursuit (C-IP) using conformal prediction sets.
result C-IP achieves better predictive performance and shorter query-answer chains.
Let M be a pseudo-Riemannian manifold with a pseudo-Hermitian complex structure J. We give necessary and sufficient conditions that the curvature operator R(π) is complex linear when π is a J invariant real 2 plane. Under this assumption, we study when M is complex IP - i.e. the spectrum, or more generally the …
Paper designs optimal ECOCs using IP for robust multiclass classification.
problem Designing robust ECOCs for multiclass classification.
method Integer Programming formulation to minimize codebooks with desirable error-correcting properties, leveraging graph-theoretic structure and edge clique covers.
result IP-generated codebooks achieve high nominal and robust adversarial accuracy.
Paper proposes CounterSample to improve convergence in LTR models.
problem Large variance in IPS weights slows convergence in LTR models.
method Introduces CounterSample algorithm with provably better convergence.
result CounterSample converges faster than standard IPS-weighted methods.
New sampling method uses gradient-free IPS with RKHS velocity field.
problem Efficient sampling from unnormalized target densities.
method Gradient-free interacting particle systems (IPS) with RKHS velocity field.
result IPS produce high-quality samples from various target distributions.
New IP analysis for deep neural networks using Rényi's entropy and tensor kernels.
problem Estimating mutual information in high-dimensional hidden layers of deep neural networks.
method Matrix-based Rényi's entropy coupled with tensor kernels for convolutional layers.
result First comprehensive IP analysis of large-scale DNNs and CNNs.
New methods improve off-policy evaluation for survival outcomes with censoring.
problem Systematic underestimation of policy performance due to censoring bias in survival outcomes.
method Proposes IPCW-IPS and IPCW-DR to handle censoring bias in survival outcomes.
result The proposed methods are unbiased and achieve double robustness.
The Information Plane theory predicts autoencoders do not compress input information.
problem Understanding the training dynamics of hidden layers in autoencoders.
method Derive a theoretical convergence for the Information Plane of autoencoders using a Gram-matrix based mutual information estimator.
result Ideal autoencoders with a large bottleneck layer size do not compress input information, while a small size causes compression only in the encoder layers.
We construct a family of pseudo-Riemannian manifolds so that the skew-symmetric curvature operator, the Jacobi operator, and the Szabo operator have constant eigenvalues on their domains of definition. This provides new and non-trivial examples of Osserman, Szabo, and IP manifolds. We also study when the associated Jor…
Method identifies IPS governing equations from particle data efficiently.
problem Identify governing equations of interacting particle systems efficiently.
method Combines mean-field theory and WSINDy for large N and M. result Converges with rate O(N−1/2) for N≥100. DeepPeep attacks DNN architectures to reveal design details, posing IP theft risks.
problem Protecting DNN architecture from IP theft in cloud-based services.
method Two-stage attack methodology exploiting design characteristics.
result DeepPeep successfully reverses-engineers compact DNN architectures.
Study examines two topologies on future causal completion of spacetimes.
problem Characterizing differences between two topologies on future causal completion.
method Systematic examination of the stronger topology τ+ on Geroch-Kronheimer-Penrose future completion IP(X) of spacetimes X. result Complete characterization of the difference in convergence between τ+ and the weaker topology. Conformal Prediction Regions match Imprecise Highest Density Regions under consonance.
problem Matching conformal prediction regions with highest density regions.
method Using consonance and the Imprecise Probability theory of clouds.
result Imprecise Highest Density Regions are equivalent to Conformal Prediction Regions under consonance.
Improved CAEs reduce training time and enhance generalization.
problem Stability issues in Concrete Autoencoders (CAEs) for feature selection.
method Indirectly Parameterized Concrete Autoencoders (IP-CAEs) learn parameters of Gumbel-Softmax distributions.
result IP-CAEs achieve significant improvements in generalization and training time.
Adaptive selection of IPs improves online GP performance.
problem Efficiently training GPs in streaming data.
method Adaptive selection of inducing points (IPs) based on GP properties and data structure.
result Adaptive IPs enhance online GP performance.
Paper tackles BNSL with IP, improving quality of solutions.
problem Bayesian Network Structure Learning (BNSL) with IP formulations.
method Inexact column generation using difference-of-submodular optimization.
result Improved solutions quality compared to state-of-the-art approaches.
This paper investigates the statistical properties of within-country GDP and industrial production (IP) growth rate distributions. Many empirical contributions have recently pointed out that cross-section growth rates of firms, industries and countries all follow Laplace distributions. In this work, we test whether als…
This work introduces a new metric for comparing imprecise probability models.
problem Quantifying differences between imprecise probability models.
method Integral imprecise probability metric framework based on Choquet integral.
result IIPM enables comparison across different imprecise probability models and quantifies epistemic uncertainty.
We consider an equity-linked contract whose payoff depends on the lifetime of policy holder and the stock price. We assume the limited capital for hedging and we provide with the best strategy for an insurance company in the meaning of so called succes factor $\IE^\IP\left[{\mathbf 1}_{\{V_T \geq D)}+{\mathbf 1}_{\{V_T…
Adaptive IP approach optimizes intervention design for causal graph recovery.
problem Designing efficient interventions to recover causal relationships from data.
method Iterative integer programming approach for optimizing information gain.
result Adaptive IP approach achieves full causal graph recovery with fewer interventions.