Dual behavior policy improves reinforcement learning across various environments.
problem Challenges in reinforcement learning with inconsistent or unknown environments.
method Dual, advantage-based behavior policy using counterfactual regret minimization.
result Demonstrated improved performance over baseline models in diverse environments.
PH-VAE models heavy-tailed data with flexible Phase-Type distributions.
problem Standard VAEs fail to capture heavy-tailed behavior in real-world data.
method PH-VAE uses Phase-Type distributions defined by continuous-time Markov chains to adaptively model tail behavior.
result PH-VAE significantly outperforms existing heavy-tail-aware VAEs in approximating diverse heavy-tailed distributions.
Imitation Learning (IL) is an appealing approach to learn desirable autonomous behavior. However, directing IL to achieve arbitrary goals is difficult. In contrast, planning-based algorithms use dynamics models and reward functions to achieve goals. Yet, reward functions that evoke desirable behavior are often difficul…
GARIM theory explains how conscious manipulation of internal representations enhances goal-directed behavior.
problem Limited understanding of how consciousness supports flexible goal-directed cognition.
method Extending a three-component theory of flexible cognition, proposing GARIM theory.
result Conscious states actively manipulate internal representations to align with goals, enhancing flexibility.
Unified framework for human-like decision making in various sequential tasks.
problem Real-life decision-making involves diverse strategies leading to similar outcomes.
method Two-stream reward processing mechanism for flexible and unified models.
result Framework unified MAB, CB, and RL with comparable performance.
A new framework for adaptive behavior using reusable value profiles.
problem Adaptive behavior in changing environments requires switching among value-control regimes, but maintaining separate parameters for each situation is impractical.
method Introduces value profiles: reusable bundles of parameters assigned to hidden states, allowing for state-conditional strategy recruitment without independent parameters for each context.
result Profile-based models outperform simpler alternatives in probabilistic reversal learning, suggesting belief-dependent control of adaptive behavior.
Agent-based simulation assesses tradable credit schemes for congestion reduction.
problem Simplistic modeling of TCS impacts in transportation research.
method Agent- and activity-based simulation framework within SimMobility.
result TCS stabilizes network and market performance over time, reducing congestion.
Adma proposes a flexible loss function for neural networks.
problem Static loss functions limit neural network performance.
method Introduces a flexible loss function that adapts to ANN complexity and data distribution.
result Flexible loss function achieves state-of-the-art performance.
The majority of real-world networks are dynamic and extremely large (e.g., Internet Traffic, Twitter, Facebook, ...). To understand the structural behavior of nodes in these large dynamic networks, it may be necessary to model the dynamics of behavioral roles representing the main connectivity patterns over time. In th…
This paper proposes an agent-based model that combines both spot and balancing electricity markets. From this model, we develop a multi-agent simulation to study the integration of the consumers' flexibility into the system. Our study identifies the conditions that real-time prices may lead to higher electricity costs,…
The Hull-Strominger system for supersymmetric vacua of the heterotic string allows general unitary Hermitian connections with torsion and not just the Chern unitary connection. Solutions on unimodular Lie groups exploiting this flexibility were found by T. Fei and S.T. Yau. The Anomaly flow is a flow whose stationary p…
GPDFlow models extreme threshold exceedance with flexible dependence using normalizing flows.
problem Challenges in modeling multivariate threshold exceedance probabilities due to infinite parametrizations.
method GPDFlow uses normalizing flows to flexibly represent dependence without explicit parametric assumptions.
result GPDFlow significantly improves modeling accuracy and flexibility compared to traditional parametric methods.
RegFlow models future states with flexible probability distributions.
problem Predicting future states under complex, non-deterministic scenarios.
method Hypernetwork architecture and continuous normalizing flow model.
result RegFlow achieves state-of-the-art results on benchmark datasets.
Study evaluates how much knowledge LLMs have by comparing their prediction accuracy to flexible models.
problem Evaluating the predictive power of LLMs without access to their training data.
method Equivalent sample size measure, comparing LLM's prediction error to flexible models trained on varying amounts of domain-specific data.
result LLMs encode varying amounts of predictive information across different economic variables.
Graph convolutional networks adapt the architecture of convolutional neural networks to learn rich representations of data supported on arbitrary graphs by replacing the convolution operations of convolutional neural networks with graph-dependent linear operations. However, these graph-dependent linear operations are d…
Bayes-TrEx finds in-distribution examples for model inspection.
problem Challenges in interpreting neural networks, especially high-confidence failures and ambiguous classifications.
method Bayesian sampling approach to find in-distribution examples with specified prediction confidence.
result Bayes-TrEx enables more flexible holistic model analysis than just inspecting the test set.
Localized debiased machine learning simplifies estimating quantile treatment effects.
problem Estimating quantile treatment effects in causal inference with many covariates and flexible relationships.
method Localized debiased machine learning (LDML) avoids learning the full nuisance function by estimating only at a single initial guess.
result LDML enables practically-feasible and theoretically-grounded efficient estimation of quantile treatment effects.
New Hawkes processes model spatiotemporal events with triggering and clustering.
problem Modeling self-excitatory behavior in spatiotemporal data.
method Developed a new class of spatiotemporal Hawkes processes with efficient inference method.
result Efficiently modeled and inferred spatiotemporal events with triggering and clustering.
A new framework for offline RL improves policy flexibility and regularity.
problem Lack of environmental interactions in offline RL leads to poor policy performance.
method Proposes a behavior-regularized implicit policy framework with modified policy-matching methods.
result The framework improves policy effectiveness and robustness beyond static datasets.
Study optimizes health incentives to balance efficiency and fairness.
problem Designing health incentives to balance efficiency and fairness.
method Inverse behavioral optimization framework integrating QALY-based incentives and adaptive learning.
result Modern health systems operate near an efficiency-saturated frontier, with small fairness adjustments yielding diminishing returns.
At the core of interpretable machine learning is the question of whether humans are able to make accurate predictions about a model's behavior. Assumed in this question are three properties of the interpretable output: coverage, precision, and effort. Coverage refers to how often humans think they can predict the model…
This work improves Gaussian process inference using mixtures of experts and nested SMC samplers.
problem High computational and memory costs of Gaussian processes.
method Mixtures of Gaussian process experts with nested SMC samplers.
result Significantly improved inference compared to importance sampling.
Empirical analysis serves as an important complement to theoretical analysis for studying practical Bayesian optimization. Often empirical insights expose strengths and weaknesses inaccessible to theoretical analysis. We define two metrics for comparing the performance of Bayesian optimization methods and propose a ran…
Paper improves normalizing flows to better capture distribution tails.
problem Difficult to learn tail behavior of distributions.
method Develops a new type of flows using flexible base distributions and data-driven linear layers.
result Improves accuracy, especially on distribution tails, and generates heavy-tailed data.
Study models human investors' sub-rational behavior in financial markets.
problem Lack of a comprehensive model for human sub-rationality in financial markets.
method Flexible reinforcement learning model incorporating five human sub-rational aspects.
result Model accurately reproduces human behavior and reveals insights into market dynamics.
Paper formalizes Simon's satisficing through FFSD, proving its equivalence to expected utility theory.
problem Formalizing Herbert Simon's bounded rationality concept in economic decision-making.
method Developed FFSD framework using Lean 4 theorem prover, proving equivalence to expected utility theory.
result Equivalence theorem linking FFSD to expected utility maximization for approximate indicator functions.
This work trains a model to predict human driving directions from road scenes.
problem Defining implicit rules of human behavior for autonomous vehicles.
method Self-supervised learning of probabilistic network model.
result Model successfully generalizes to new road scenes.
Agents need world models to generalize multi-step tasks.
problem The necessity of world models for flexible, goal-directed behavior.
method Formal analysis and demonstration of the necessity of world models for agents to generalize multi-step tasks.
result World models are necessary for agents to generalize to multi-step goal-directed tasks.
Paper develops fast, flexible Hawkes process inference for space-time data.
problem Capturing self-exciting, clustering spatio-temporal data.
method Finite support kernels, discretization, precomputations, ℓ2 gradient-based solver. result Statistically accurate and fast inference for space-time Hawkes processes.
Flexible nonstationary Gaussian process with neural network parameters.
problem Limited expressiveness of stationary Gaussian processes.
method Nonstationary kernels with neural network parameters trained jointly.
result Better accuracy and log-score compared to stationary and hierarchical models.
Study improves choice model accuracy and heterogeneity representation using mixture models.
problem Improving prediction accuracy and heterogeneity representation in choice models.
method Semi-nonparametric Latent Class Choice Model with mixture models and EM algorithm.
result Mixture models enhance prediction accuracy and heterogeneity representation without sacrificing interpretability.
Predicting fine-grained interests of users with temporal behavior is important to personalization and information filtering applications. However, existing interest prediction methods are incapable of capturing the subtle degreed user interests towards particular items, and the internal time-varying drifting attention …
Consider a smooth closed surface M of fixed genus ⩾2 with a hyperbolic metric σ of total area A. In this article, we study the behavior of geometric and dynamical characteristics (e.g., diameter, Laplace spectrum, Gaussian curvature and entropies) of nonpositively curved smooth metrics with total area …
Learning goal-directed behavior in environments with sparse feedback is a major challenge for reinforcement learning algorithms. The primary difficulty arises due to insufficient exploration, resulting in an agent being unable to learn robust value functions. Intrinsically motivated agents can explore new behavior for …
Unified DICE estimators as regularized Lagrangians for improved off-policy evaluation.
problem Improving off-policy evaluation from behavior-agnostic data.
method Unified derivation of DICE estimators as regularized Lagrangians of a linear program.
result Dual solutions offer greater flexibility and provide superior estimates in practice.
How can deep learning systems flexibly reuse their knowledge? Toward this goal, we propose a new class of challenges, and a class of architectures that can solve them. The challenges are meta-mappings, which involve systematically transforming task behaviors to adapt to new tasks zero-shot. The key to achieving these c…
For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems. In this work, we explore goals defined in terms of (non-expert) human preferences between pairs of trajectory segments. We show that this approach can effective…
This paper surveys imitation learning methods and challenges.
problem Difficulty in programming AI systems in complex environments.
method Learning from expert demonstrations to adapt behaviors.
result Overview of recent advances and emerging areas in IL.
Meta-causal states group equivalent qualitative causal dynamics, useful for analyzing system changes.
problem Qualitative changes in causal relationships due to agent actions or environmental tipping points.
method Propose meta-causal states to group causal models based on equivalent qualitative behavior and parameterize specific mechanisms.
result Meta-causal states can be inferred from observed agent behavior and disentangled from unlabeled data.
AR-Flow VAE improves blind source separation with flexible autoregressive priors.
problem Unsupervised blind source separation of latent signals from mixtures.
method AR-Flow VAE uses autoregressive flows to model latent sources, enhancing flexibility and capturing complex dependencies.
result AR-Flow VAE effectively separates latent sources, demonstrating improved performance over conventional methods.
Recent developments in machine-learning algorithms have led to impressive performance increases in many traditional application scenarios of artificial intelligence research. In the area of deep reinforcement learning, deep learning functional architectures are combined with incremental learning schemes for sequential …
Proposes a new model for complex multivariate event data.
problem Modeling complex multivariate event data with spatio-temporal dynamics.
method Integrates spatial information into latent state evolution through learned temporal and spatial decay dynamics.
result Successfully recovers sensible temporal and spatial intensity structure in multivariate spatio-temporal point patterns.
New method for efficient maximum likelihood estimation of p-generalized probit regression.
problem Efficient estimation of p-generalized probit regression models. method Combining sketching techniques with importance subsampling to obtain a coreset.
result Maximum likelihood estimator can be approximated efficiently up to a factor of (1+ε) on large data. A new framework for mobile authentication using deep metric learning.
problem Challenges in mobile authentication using behavioral biometrics.
method Deep metric learning, private data protection, flexible training scheduling.
result 95% authentication accuracy on public datasets, robust against attacks.
Modified model prevents volatility from approaching zero.
problem Volatility in the Gatheral model can approach zero, making it statistically indistinguishable.
method Proposed a modified model with Skorokhod reflection to prevent volatility from approaching zero.
result The modified model prevents volatility from approaching zero, preserving the model's flexibility.
The study models mortgage prepayment risk, accounting for behavioral uncertainty, and provides replication strategies.
problem Modeling and replicating the prepayment option of mortgages with behavioral uncertainty.
method Modeling behavioral uncertainty as a non-hedgeable risk factor, proving its impact on exposure value, and using IRSs and swaptions for replication.
result Including behavioral uncertainty reduces the exposure's value, and swaptions are necessary for optimal replication.
Study finds financial constraints explain zero-leverage firms.
problem Why some firms have zero leverage despite various explanations.
method Examined three measures of financial constraints; analyzed firms' behavior before and after levering.
result Firms are financially constrained, not due to managerial entrenchment or market valuation.
Deep learning improves analysis of complex natural processes.
problem Simplistic dynamics in regression analyses of complex natural processes.
method Flexible function approximation using deep learning, relaxing standard assumptions.
result Substantial improvements in behavioral and neuroimaging data.