Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

2.0%4.0%5.9%7.9% · Jun 201919922001200920182026
48 results for IP policy

New estimator SWITCH improves off-policy evaluation in contextual bandits.

problem Estimating value of a target policy using data from another policy in contextual bandits.
method Proposes SWITCH estimator that uses an existing reward model to outperform IPS and DR.
result Switch estimator achieves better MSE than IPS and DR, often outperforming prior work.

AIPS improves ranking policy evaluation by adapting to diverse user behavior.

problem Inaccurate Off-Policy Evaluation of ranking policies due to high variance under diverse user behavior.
method Developed Adaptive IPS (AIPS) that adapts to different user behaviors and minimizes MSE.
result AIPS achieves minimum variance among unbiased estimators and provides significant empirical accuracy improvement.

A new estimator reduces variance in slate bandit OPE.

problem Large action spaces in slate bandits cause high variance in OPE.
method Develops Latent IPS (LIPS) to optimize slate abstractions for low variance and bias.
result LIPS substantially outperforms existing estimators in scenarios with non-linear rewards and large slate spaces.

New DR-IC estimator reduces bias and variance in OPE.

problem Estimating value of a target policy using logged data from a different policy.
method DR-IC estimator that combines parametric reward model and context-based switching rule.
result DR-IC estimator outperforms state-of-the-art OPE algorithms.

CAEL-MIPS learns embeddings to improve MIPS for better OPE in contextual bandits.

problem High variance in IPS weighting for OPE in large action spaces.
method Context-Action Embedding Learning (CAEL) for MIPS to minimize MSE.
result CAEL-MIPS outperforms baselines in MSE for OPE in contextual bandits.

CAB estimator improves evaluation and learning performance in policy contexts.

problem Offline A/B-testing and off-policy learning using logged contextual bandit feedback.
method Continuous Adaptive Blending (CAB) estimator, subsuming most counterfactual estimators.
result CAB estimator is less biased and has less variance than other estimators.

New estimator uses clustering to improve off-policy evaluation accuracy.

problem Improving off-policy evaluation accuracy when logging and evaluation policies differ.
method Proposes an estimator that shares information across similar contexts using clustering.
result Clustering contexts improves estimation accuracy, especially in deficient information settings.

New method tunes prior IP to data for flexible predictive distributions.

problem Challenges in approximate inference for large models with high parameter dependencies.
method Inducing-point representation of prior IP to approximate posterior process.
result Scalable method that tunes prior IP to data and provides accurate non-Gaussian predictive distributions.

RL improves IP solver performance by learning to select cutting planes.

problem Improving the performance of IP solvers through heuristic optimization.
method Employing reinforcement learning to intelligently select cutting planes in the Cutting Plane Method.
result Trained RL agent significantly outperforms human-designed heuristics across various IP tasks.

The paper improves IPS for modern optimization, scaling and regularizing it.

problem Improving iterative proportional scaling for modern optimization.
method Coordinate descent, majorization-minimization, optimization techniques, regularized variants.
result IPS can deliver coefficient estimates and handle log-affine models.

DVIP improves on IP-based methods by using IPs as priors over latent functions.

problem Limited expressiveness of IP-based models, especially in function space.
method Proposes DVIP, a multi-layer generalization of IPs, and scalable variational inference.
result DVIP outperforms previous IP-based methods and deep GPs in regression and classification tasks.

A new model SIPS improves graph embedding by approximating non-PD similarities.

problem Improving neural network-based graph embedding by approximating non-positive definite similarities.
method Shifted Inner Product Similarity (SIPS) model that approximates Conditionally Positive Definite (CPD) similarities.
result SIPS significantly improves graph embedding without configuring the similarity function.

Extends V-IP framework to use LLMs for generating task-relevant concepts, improving interpretability and performance.

problem Limited applicability of V-IP to small-scale tasks due to manual data annotation.
method Integrates Foundational Models with Large Language and Multimodal Models to generate and annotate concepts.
result FM+V-IP achieves better test performance with fewer concepts/queries compared to other frameworks.

MEC-IP uses IP to efficiently find MECs in BNs from observational data.

problem Discovering Markov Equivalent Classes (MECs) in Bayesian Networks (BNs) efficiently.
method Clique-focusing strategy and EMSG for MEC discovery via Integer Programming.
result Significant reduction in computational time and improved accuracy.

A new method for making interpretable predictions by sequentially asking questions, faster and more efficient.

problem Developing interpretable machine learning models for complex tasks.
method Variational Information Pursuit (V-IP) that bypasses the need for learning generative models.
result V-IP is 10-100x faster and finds shorter query chains compared to IP and reinforcement learning.

Two machine learning applications for IP/Optical networks: traffic prediction and optical path performance.

problem Agile resource management and optical path performance prediction in IP/Optical networks.
method Machine learning for traffic prediction and optical performance prediction using SDN controllers.
result Efficient implementation of SDN controllers for agile resource management and optical path performance prediction.

A new method combines online and offline learning to tackle contextual bandits with missing action support.

problem Learning optimal policies with logged data when the logging policy has deficient support.
method Hybrid approach using online exploration to exploit supported actions and offline learning to avoid unnecessary explorations.
result Determines an optimal policy with theoretical guarantees using minimal online explorations.

Paper shows non-integrality of dike building model and provides conditions for integrality.

problem Determining the integrality of a dike building model for flood protection.
method Analyzes experimental data and mathematical proofs to establish conditions for integrality.
result Established non-integrality of the polytope and conditions for linear programming relaxation to be integral.

C-IP improves LLMs' query selection for interactive tasks by estimating uncertainty robustly.

problem Minimizing the number of queries for interactive LLMs.
method Conformal Information Pursuit (C-IP) using conformal prediction sets.
result C-IP achieves better predictive performance and shorter query-answer chains.

Let M be a pseudo-Riemannian manifold with a pseudo-Hermitian complex structure JJ. We give necessary and sufficient conditions that the curvature operator R(π)R(π) is complex linear when ππ is a JJ invariant real 2 plane. Under this assumption, we study when M is complex IP - i.e. the spectrum, or more generally the …

2002-05-08abs ↗pdf ↗

Paper designs optimal ECOCs using IP for robust multiclass classification.

problem Designing robust ECOCs for multiclass classification.
method Integer Programming formulation to minimize codebooks with desirable error-correcting properties, leveraging graph-theoretic structure and edge clique covers.
result IP-generated codebooks achieve high nominal and robust adversarial accuracy.

New IP analysis for deep neural networks using Rényi's entropy and tensor kernels.

problem Estimating mutual information in high-dimensional hidden layers of deep neural networks.
method Matrix-based Rényi's entropy coupled with tensor kernels for convolutional layers.
result First comprehensive IP analysis of large-scale DNNs and CNNs.

New methods improve off-policy evaluation for survival outcomes with censoring.

problem Systematic underestimation of policy performance due to censoring bias in survival outcomes.
method Proposes IPCW-IPS and IPCW-DR to handle censoring bias in survival outcomes.
result The proposed methods are unbiased and achieve double robustness.

The Information Plane theory predicts autoencoders do not compress input information.

problem Understanding the training dynamics of hidden layers in autoencoders.
method Derive a theoretical convergence for the Information Plane of autoencoders using a Gram-matrix based mutual information estimator.
result Ideal autoencoders with a large bottleneck layer size do not compress input information, while a small size causes compression only in the encoder layers.

We construct a family of pseudo-Riemannian manifolds so that the skew-symmetric curvature operator, the Jacobi operator, and the Szabo operator have constant eigenvalues on their domains of definition. This provides new and non-trivial examples of Osserman, Szabo, and IP manifolds. We also study when the associated Jor…

2002-05-08abs ↗pdf ↗

Method identifies IPS governing equations from particle data efficiently.

problem Identify governing equations of interacting particle systems efficiently.
method Combines mean-field theory and WSINDy for large NN and MM.
result Converges with rate O(N1/2)\mathcal{O}(N^{-1/2}) for N100N \geq 100.

Study examines two topologies on future causal completion of spacetimes.

problem Characterizing differences between two topologies on future causal completion.
method Systematic examination of the stronger topology τ+τ_+ on Geroch-Kronheimer-Penrose future completion IP(X)IP(X) of spacetimes XX.
result Complete characterization of the difference in convergence between τ+τ_+ and the weaker topology.

Conformal Prediction Regions match Imprecise Highest Density Regions under consonance.

problem Matching conformal prediction regions with highest density regions.
method Using consonance and the Imprecise Probability theory of clouds.
result Imprecise Highest Density Regions are equivalent to Conformal Prediction Regions under consonance.

Improved CAEs reduce training time and enhance generalization.

problem Stability issues in Concrete Autoencoders (CAEs) for feature selection.
method Indirectly Parameterized Concrete Autoencoders (IP-CAEs) learn parameters of Gumbel-Softmax distributions.
result IP-CAEs achieve significant improvements in generalization and training time.

Paper tackles BNSL with IP, improving quality of solutions.

problem Bayesian Network Structure Learning (BNSL) with IP formulations.
method Inexact column generation using difference-of-submodular optimization.
result Improved solutions quality compared to state-of-the-art approaches.

We consider an equity-linked contract whose payoff depends on the lifetime of policy holder and the stock price. We assume the limited capital for hedging and we provide with the best strategy for an insurance company in the meaning of so called succes factor $\IE^\IP\left[{\mathbf 1}_{\{V_T \geq D)}+{\mathbf 1}_{\{V_T…

2014-05-04abs ↗pdf ↗

Adaptive IP approach optimizes intervention design for causal graph recovery.

problem Designing efficient interventions to recover causal relationships from data.
method Iterative integer programming approach for optimizing information gain.
result Adaptive IP approach achieves full causal graph recovery with fewer interventions.