Develops a reinforcement learning algorithm for learning deterministic equilibrium policies in time-inconsistent control problems.
problem Learning equilibrium policies in time-inconsistent control problems.
method Continuous-time model-free reinforcement learning algorithm using deterministic policy gradient approach.
result Learned equilibrium policies in general time-inconsistent control problems.
We study consistency properties of machine learning methods based on minimizing convex surrogates. We extend the recent framework of Osokin et al. (2017) for the quantitative analysis of consistency properties to the case of inconsistent surrogates. Our key technical contribution consists in a new lower bound on the ca…
Paper tackles inconsistent CATE estimation across group assignments.
problem Inconsistent learning behavior for the same instance across different group assignments.
method CLAGA method to eliminate inconsistency.
result Significant performance improvements with CLAGA method.
DAGnosis uses DAGs to identify and localize data inconsistencies.
problem Handling data inconsistencies in machine learning models at deployment time.
method Directed acyclic graphs (DAGs) to encode feature probability distribution and independencies.
result Localization of inconsistencies and insights into their causes.
Aux-Net model handles dynamic systems with inconsistent inputs.
problem Inconsistent or unreliable input data in real-world scenarios.
method Aux-Net uses a weighted ensemble of classifiers and online gradient descent.
result Aux-Net provides scalable and agile online learning for dynamic systems.
Proposes a new multi-view graph learning framework to model consistency and inconsistency.
problem Graph learning methods often neglect inconsistency across multiple views, making them vulnerable to noisy datasets.
method Proposes a unified objective function to simultaneously model consistency and inconsistency, iteratively learning consistent and unified graphs.
result Demonstrates robustness and efficiency of the proposed approach on twelve multi-view datasets.
Paper analyzes venture capital exit decisions under inconsistent preferences.
problem Time-inconsistent preferences in venture capital exit timing.
method Modeling four types of venture capitalists with varying levels of inconsistency.
result Time-inconsistent venture capitalists exit earlier than consistent ones.
Paper addresses inconsistency between offline and online LTR performance.
problem Inconsistency between offline and online LTR performance in E-commerce.
method Proposes an evaluator-generator framework to maximize evaluator score using reinforcement learning.
result Significant improvement in Conversion Rate (CR) over existing models.
Paper solves time-inconsistent control problems with BSDEs.
problem Time-inconsistent stochastic control in continuous time.
method Probabilistic representation via BSDEs.
result Equilibrium value function resolved for inconsistent cases.
Empirical study shows consistent meta-RL algorithms adapt to OOD tasks.
problem Theoretical consistency of meta-RL algorithms and its practical implications.
method Empirical investigation of representative meta-RL algorithms, focusing on consistency and adaptation to out-of-distribution tasks.
result Theoretical consistent algorithms can adapt to OOD tasks, while inconsistent ones cannot, but can still fail for poor exploration.
New method detects inconsistencies in AHP matrices using triadic preference reversals.
problem Challenges in assessing consistency in AHP pairwise comparison matrices.
method Triadic preference reversals to detect inconsistencies between pairs of elements.
result 97% accuracy in detecting inconsistencies, significantly surpassing traditional methods.
Researchers improve Gaussian processes to model inconsistent preferences.
problem Model inconsistent preferences and clusters of comparable items.
method Generalized Gaussian processes with spectral decomposition and universal RKHS.
result Competitive with state-of-the-art methods on simulated and real-world data.
New theory extends LQ control to non-exponential discount scenarios.
problem Time-inconsistent deterministic LQ control problems.
method Extended equivalent relationship to non-exponential discount functions, studied Riccati equation solvability.
result Existence and uniqueness of linear equilibrium for time-inconsistent LQ problem.
Proposes CI-GMVC to improve graph-based multi-view clustering performance.
problem Inconsistency in multi-view data affects clustering performance.
method Integrates consistent and inconsistent parts of multiple views using a unified matrix.
result Demonstrates improved clustering performance on real-world datasets.
Study solves HJB equations for time-inconsistent control problems.
problem Time-inconsistent deterministic linear quadratic control problems.
method Characterized solutions using Riccati equations with integral terms, proving uniqueness.
result Uniqueness of solutions to equilibrium HJB equations proved.
Enhances predictive performance in Bayesian deep learning via generalized Laplace approximation.
problem Inconsistency in Bayesian deep learning.
method Interprets posterior tempering as a correction for model misspecification and recalibration of priors. Introduces generalized Laplace approximation.
result Generalized Laplace approximation enhances predictive performance.
This paper tackles objective inconsistency in federated optimization with heterogeneous clients.
problem Objective inconsistency due to heterogeneity in clients' datasets and computation speeds.
method General framework for analyzing federated heterogeneous optimization algorithms, including FedAvg and FedProx, and proposing FedNova.
result FedNova eliminates objective inconsistency while preserving fast error convergence.
This paper considers a time-inconsistent stopping problem in which the inconsistency arises from non-constant time preference rates. We show that the smooth pasting principle, the main approach that has been used to construct explicit solutions for conventional time-consistent optimal stopping problems, may fail under …
New method resolves inconsistency in learning directed acyclic graphs using penalized likelihood.
problem Inconsistency of ℓ1-penalized likelihood in learning directed acyclic graphs. method Developed a hybrid differentiable structure learning method based on ℓ0-penalized likelihood with hard acyclicity constraint. result Demonstrated and explained why ℓ1-penalized likelihood is fundamentally inconsistent in identifying true structure up to Markov equivalence classes. Study time-inconsistent portfolio optimization for competitive agents with relative performance criteria.
problem Time-inconsistent mean field and n-agent games under relative performance criteria.
method Construct open-loop equilibrium strategies for n-agent games and mean field games.
result Explicit solutions for n-agent games and mean field games, unique in a special class of equilibria.
Proposes a method to improve multi-output Gaussian process for transfer learning.
problem Negative transfer and domain inconsistency in multi-output Gaussian process.
method Regularized MGP with convolution process and domain adaptation.
result Outperforms state-of-the-art benchmarks in simulation and real-world studies.
A new approach to MV portfolio optimization with jumps and RL.
problem Continuous-time Mean-Variance portfolio optimization with jumps.
method Jump-diffusion process, Reinforcement Learning, time-inconsistent control (TIC), Actor-Critic RL algorithm.
result The proposed RL model is profitable in real-world market data.
Paper solves a complex portfolio selection problem with time-inconsistent preferences.
problem Time-inconsistent preferences in portfolio selection.
method Unified framework with minimal assumptions, proving existence and uniqueness of solution.
result Existence and uniqueness of square-integrable solution for the integral equation.
Kernel interpolation is inconsistent for norms with smoothness above a constant.
problem Inconsistency of kernel interpolation in reproducing kernel Hilbert spaces.
method Lower bounds for generalization error in Sobolev norms.
result Kernel interpolation is always inconsistent for norms with smoothness above a constant.
Existing ordinal embedding methods usually follow a two-stage routine: outlier detection is first employed to pick out the inconsistent comparisons; then an embedding is learned from the clean data. However, learning in a multi-stage manner is well-known to suffer from sub-optimal solutions. In this paper, we propose a…
Despite strong performance on a variety of tasks, neural sequence models trained with maximum likelihood have been shown to exhibit issues such as length bias and degenerate repetition. We study the related issue of receiving infinite-length sequences from a recurrent language model when using common decoding algorithm…
New method improves consistency of reinforcement learning performance evaluations.
problem Inconsistent performance results in reinforcement learning due to flawed evaluation metrics.
method Proposes a new comprehensive evaluation methodology for reinforcement learning algorithms.
result Demonstrates improved reliability of performance measurements for reinforcement learning algorithms.
In this paper we study a class of time-inconsistent terminal Markovian control problems in discrete time subject to model uncertainty. We combine the concept of the sub-game perfect strategies with the adaptive robust stochastic to tackle the theoretical aspects of the considered stochastic control problem. Consequentl…
In this paper, we study a time-inconsistent consumption-investment problem with random endowments in a possibly incomplete market under general discount functions. We provide a necessary condition and a verification theorem for an open-loop equilibrium consumption-investment pair in terms of a coupled forward-backward …
Time inconsistency leads to intra-personal conflict and reconciliation strategies.
problem Time inconsistency in dynamic choice problems.
method Rigorous treatment of intra-personal equilibrium in continuous-time settings.
result A new approach to understanding and reconciling intra-personal conflicts.
Separable losses are inconsistent for structured prediction models.
problem Inconsistency of separable losses in structured prediction models.
method Analysis of separable negative log-likelihood losses for structured prediction.
result Separable losses are not Bayes consistent and may not predict the most probable structure.
The paper shows cross-validation fails in learning Gaussian graphical model structures.
problem Cross-validation's failure in learning Gaussian graphical model structures.
method Finite-sample bounds on misidentification probability of Lasso estimator.
result Cross-validation is inconsistent for learning Gaussian graphical model structures.
Paper proves existence and uniqueness of solutions to nonlocal systems, generalizing stochastic game theory.
problem Time inconsistency in stochastic differential games.
method Proves existence and uniqueness of solutions to nonlocal fully-nonlinear parabolic systems.
result Generalizes stochastic game theory to include time-inconsistent preferences.
The paper tackles inconsistency in removal-based explanations and proposes methods to reduce it.
problem Inconsistency in removal-based explanations.
method Established the Impossible Trinity Theorem and proposed two novel algorithms to minimize interpretation error.
result The proposed methods achieve a substantial reduction in interpretation error, up to 31.8 times lower.
We consider a general time-inconsistent stochastic linear-quadratic differential game. The time-inconsistency arises from the presence of quadratic terms of the expected state as well as state-dependent term in the objective functionals. We define an equilibrium strategy, which is different from the classical one, and …
The paper solves TIC LQ control problems using stochastic differential games.
problem Time-inconsistent linear-quadratic stochastic control problems.
method Stochastic differential games, spike variation approach.
result Achieves Nash equilibrium for TIC problems, demonstrating impact of ambiguity aversion.
The paper analyzes optimal dividend and capital injection strategies under time-inconsistent preferences.
problem Optimal dividend and capital injection strategies under time-inconsistent preferences.
method Diffusion risk model with general discount functions, weak equilibrium definition, HJB equation system.
result Explicit solutions and threshold types of optimal strategies derived under different discount functions.
Investigates time-inconsistent portfolio selection under MMV preferences.
problem Time-inconsistent optimal strategies for MMV preferences.
method Nash equilibrium controls for MMV and MV preferences, solving FBSDE and HJB equations.
result MMV optimal strategies lead to higher investment amounts than MV strategies, narrowing over time.
PINNs struggle with data-to-PDE inconsistencies, limiting their accuracy.
problem Data inconsistency in PINNs affects their accuracy and convergence.
method Systematic analysis of PINNs with varying data fidelity and residual errors.
result PINNs saturate at an error level dictated by data inconsistency.
Proposes a deep learning approach for optimizing portfolios with stocks and options.
problem Optimizing portfolios with time-inconsistent objectives and trading constraints.
method Neural networks with adaptive activation functions for asset allocation and option strike prices.
result Adding options leads to more stable and consistent stock allocations.
Markov networks (MNs) are a powerful way to compactly represent a joint probability distribution, but most MN structure learning methods are very slow, due to the high cost of evaluating candidates structures. Dependency networks (DNs) represent a probability distribution as a set of conditional probability distributio…
Proposes a neural network method to improve consistencies in high dimensional data analysis.
problem Inconsistencies among dimensionality reduction, clustering, and visualization tasks in high dimensional data analysis.
method Consistent Representation Learning (CRL) neural network that performs NLDR transformations to satisfy LGP constraints.
result Improves consistencies in data interpretation through end-to-end task execution.
New research finds many coreset methods for logistic regression are not better than simple sampling.
problem Evaluation of coreset methods for reducing data size in logistic regression.
method Comparison of multiple coreset and optimal subsampling methods for logistic regression.
result Many coreset methods do not outperform simple uniform subsampling.
Large-scale classification of data where classes are structurally organized in a hierarchy is an important area of research. Top-down approaches that exploit the hierarchy during the learning and prediction phase are efficient for large scale hierarchical classification. However, accuracy of top-down approaches is poor…
Energy-based diffusion models improve molecular sampling and simulation.
problem Inconsistency between diffusion model scores and equilibrium distributions.
method Fokker-Planck regularization to enforce consistency.
result Improved consistency and efficient sampling of biomolecular systems.
The paper solves a portfolio selection problem in incomplete markets by balancing utility and risk.
problem Time-inconsistent portfolio selection in incomplete markets.
method Characterizes equilibrium via a coupled quadratic BSDE system, introduces approximate equilibrium for general cases.
result Established existence theory for equilibrium strategies in special and general cases.
Bandit algorithms struggle with consistent performance and robustness.
problem Achieving consistent and robust performance in stochastic multi-armed bandit settings.
method Analyzing regret minimization trade-offs and proposing distribution-oblivious algorithms.
result Logarithmic regret is inconsistent and super-logarithmic regret is necessary for consistent learning.
The task of unsupervised domain adaptation is proposed to transfer the knowledge of a label-rich domain (source domain) to a label-scarce domain (target domain). Matching feature distributions between different domains is a widely applied method for the aforementioned task. However, the method does not perform well whe…