DiCE uses diverse agents to explore and learn, avoiding local minima.
problem Local minima in RL due to limited exploration and correlated behavior.
method DiCE employs a group of heterogeneous agents to explore simultaneously and share experiences, with a diversity regularization mechanism.
result DiCE achieves substantial improvement over baselines in MuJoCo locomotion tasks.
New technique improves imitation learning by preventing local minima and exploring states.
problem Behavioral cloning gets stuck in local minima and lacks effective exploration.
method Two-phase model with sampling mechanisms and self-attention modules.
result Significantly outperforms previous state-of-the-art in various environments.
The study confirms conditions for Q-learning with persistent exploration.
problem Formulating conditions for Q-learning with persistent exploration. method Formulated assumptions for Q-learning with local and global clocks, ensuring persistent exploration. result The Robbins-Monro conditions are confirmed for Q-learning with persistent exploration. New meta-RL method avoids exploration-exploitation trade-off.
problem Learning to explore and exploit simultaneously in meta-RL.
method Developed new objectives for exploration and exploitation.
result DREAM outperforms existing methods on complex tasks.
Scientific discovery is limited by hypothesis redundancy, and hybrid methods can exploit non-local exploration.
problem Limitation of scientific discovery due to hypothesis redundancy.
method Hybrid discovery systems combining structured local search with LLM-generated non-local proposals.
result Hybrid methods can exploit non-local exploration when three geometric conditions co-occur.
POLO framework enables efficient learning and exploration in model-based control.
problem Efficient learning and exploration in model-based control settings.
method Combines local model-based control, global value function learning, and exploration.
result POLO framework accelerates value function learning and enables better policies.
Hill-ADAM optimizes loss landscapes by exploring state space deterministically.
problem Escaping local minima in loss landscapes.
method Hill-ADAM alternates between minimizing and maximizing error to explore the loss space.
result Hill-ADAM finds the global minimum state in loss landscapes.
Enhances exploration in reinforcement learning with diversity-driven approach.
problem Challenges in efficient exploration in reinforcement learning, especially in large state spaces.
method Diversity-driven exploration strategy combining off- and on-policy reinforcement learning algorithms.
result Significantly enhances exploratory behaviors, preventing local optima.
The paper explores local-correlation models for pricing complex financial contracts.
problem Calibrating synthetic quanto forward contracts and composite options.
method Design on-line calibration procedures for local and stochastic volatility models.
result Calibration performance of local-correlation models compared to simpler approximations.
Algorithm achieves optimal pricing with minimal exploration for dynamic markets.
problem Optimal pricing in dynamic markets with contextual information.
method Localized exploration-then-commit (LetC) algorithm with pure exploration, refinement, and exploitation stages.
result Achieves minimax optimal, dimension-free regret bound.
Central to robot exploration and mapping is the task of persistent localization in environmental fields characterized by spatially correlated measurements. This paper presents a Gaussian process localization (GP-Localize) algorithm that, in contrast to existing works, can exploit the spatially correlated field measurem…
The paper connects machine learning interpretability with learning theory.
problem Performance and explanation generalization in local machine learning models.
method Theoretical analysis and empirical validation of local approximation explanations.
result Theoretical bounds on test-time accuracy and explanation generalization.
Study local exploration on dynamic graphs with time-varying edges.
problem Learning optimal actions in a network with changing connections.
method Local explore-then-commit algorithms under a structural condition ensuring intrinsic walk stability.
result Sublinear expected regret for reward-aware strategies.
The paper analyzes the intrinsic exploration terms in policy-gradient algorithms.
problem Exploration in policy-gradient algorithms and its impact on policy optimization.
method Numerical optimization criteria and stochastic gradient analysis.
result Exploration techniques improve policy optimization by smoothing the learning objective and modifying gradient estimates.
A new method captures higher-order interactions in data clusters.
problem Accurately characterizing complex higher-order variable interactions.
method Local Correlation Explanation (CorEx) method: clustering and total correlation.
result Captures higher-order interactions at a local scale.
Local SGD outperforms conventional methods in LLM training.
problem Training large language models on distributed devices.
method Local SGD algorithm applied to distributed training.
result Local SGD achieves competitive results in LLM training.
Novelty search learns attentional layers to quickly explore and guess numbers.
problem Exploring and guessing numbers quickly in structured spaces.
method Attentional neural network layers trained on supervised learning of local sensory-motor contingencies.
result Greedy local policies can quickly explore structured spaces and guess numbers.
Explores local structure of morphisms and formal submanifolds in formal manifolds theory.
problem Understanding the local structure of morphisms and formal submanifolds in formal manifolds.
method Study of formal manifolds, including local structure of constant rank morphisms and formal submanifolds.
result Developed the local structure of constant rank morphisms and formal submanifolds.
Paper uses deep learning to improve thermal-hydraulic simulations.
problem Limited credibility of thermal-hydraulic codes in real plant conditions.
method Feature Similarity Measurement (FSM) and deep learning.
result Deep learning constructs relationships between local physical features and simulation errors.
Simple mode exploration methods do not improve performance in neural networks.
problem Improving predictive probabilities in neural networks.
method Exploring local regions around diverse solutions using simple methods.
result Simple mode exploration methods do not improve performance.
New Thompson sampling uses local uncertainty for better decision making.
problem Sequential decision making with exploration-exploitation dilemma.
method Proposes a new probabilistic modeling framework using local latent variable uncertainty for Thompson sampling, with variational inference and semi-implicit structure.
result Thompson sampling guided by local uncertainty achieves state-of-the-art performance with low computational complexity.
cKAM improves adaptive sampling by incorporating a cyclical stepsize scheme.
problem Adaptive Metropolis algorithms can get stuck in local modes.
method cKAM uses a cyclical stepsize scheme to encourage exploration and escape from local modes.
result cKAM successfully escapes local modes and converges to the true posterior distribution.
The study explores how different Grothendieck topologies and functors between categories preserve locality.
problem Exploring relationships between different Grothendieck topologies and functors.
method Using Grothendieck topologies and functors to relate categories and geometric objects.
result Objects like sheaves, groupoids, and functors are invariant under equivalences of Grothendieck topologies and certain functors.
New method uses reward prediction error for efficient exploration.
problem Efficient exploration in reinforcement learning, especially in complex environments.
method Reward prediction error (RPE) as an intrinsic motivation for exploration, combined with a deep reinforcement learning method (QXplore).
result QXplore outperforms state-novelty methods in diverse tasks, especially when state novelty is not correlated with improved reward.
LGnet jointly models local and global dynamics for MTS forecasting with missing values.
problem Missing values in multivariate time series data.
method LGnet framework using memory network and adversarial training.
result LGnet effectively forecasts MTS with missing values and robust under various missing ratios.
LNUCB-TA improves MAB performance by dynamically adjusting exploration rates and recognizing spatiotemporal patterns.
problem Suboptimal performance in environments with rapidly changing reward structures and static exploration rates.
method Hybrid model combining linear and nonlinear estimation, with adaptive k-NN for temporal attention.
result Significantly outperforms state-of-the-art algorithms in cumulative and mean reward, convergence, and robustness.
Adaptor 'E' extends gradient-based optimizers to explore loss landscapes, improving generalization.
problem Finding lower and better-generalizing minima in deep learning.
method Proposes an adaptor 'E' to extend gradient-based optimizers, encouraging exploration along landscape valleys.
result Adapted optimizers increase test accuracy by an average of 2.5% in large-batch training tasks.
The paper explores stable surfaces in Einstein-Maxwell theory, proving mass bounds and nonexistence results.
problem Exploring stable surfaces in static Einstein-Maxwell space-time.
method Using mean-stable surfaces theory to prove properties of lapse functions and mass bounds.
result Proves ADM mass is bounded by Hawking quasi-local mass.
Improves reinforcement learning by combining off-policy data and exploration.
problem Data inefficiency and local optima in policy gradient methods.
method Combines off-policy data reuse, exploration, and deterministic policies with stochastic optimization.
result Successfully learns solutions using fewer interactions than standard methods.
VCSMC improves efficiency in Bayesian phylogenetic inference.
problem Inefficient exploration of phylogenetic state space.
method Variational Combinatorial Sequential Monte Carlo (VCSMC) and nested CSMC.
result VCSMC and VNCSMC explore higher probability spaces efficiently.
Local search improves GFlowNets' ability to generate high-reward samples.
problem GFlowNets struggle with over-exploration in high-reward space.
method Local search focusing on high-reward samples via backtracking and reconstruction.
result Significant performance improvement in biochemical tasks.
Normalizing flows policy improves trust region policy optimization.
problem Improving exploration and avoiding local optima in policy optimization.
method Constructing trust region with KL divergence constraints and using normalizing flows policy.
result Normalizing flows policy significantly improves policy optimization, especially on high-dimensional tasks.
Bundle Networks explore many-to-one maps using fiber bundles and local trivializations.
problem Exploring and sampling from fibers of many-to-one maps in machine learning.
method Introducing Bundle Networks based on fiber bundles and local trivializations, using invertible components.
result Exploring fibers of many-to-one maps becomes natural in Bundle Networks.
The paper explores connections between topological dynamics and groupoids.
problem Understanding connections between topological dynamics and groupoids.
method Investigates connections between Topological Dynamics, G-Principal Bundles, and Locally Trivial Groupoids.
result Connections between topological dynamics and groupoids are explored.
A new active learning method for Gaussian process models.
problem Efficiently exploring unbounded state spaces for accurate models.
method Maximizes mutual information with respect to a bounded region using model predictive control.
result Our method yields a better model within the region of interest than entropy-based methods.
In this paper, we explore the virtual technique that is very useful in studying moduli problem from differential geometric point of view. We introduce a class of new objects "virtual manifolds/orbifolds", on which we develop the integration theory. In particular, the virtual localization formula is obtained.
LCBO tackles constrained optimization in high dimensions, offering a polynomial convergence rate.
problem Bayesian optimization for high-dimensional constrained problems.
method LCBO uses local descent and uncertainty-driven exploration, proving polynomial convergence rate.
result LCBO achieves a polynomial convergence rate for KKT residuals in high dimensions.
Meta-SAGE improves deep RL scalability for CO tasks by adapting pre-trained models to larger-scale problems.
problem Improving scalability of deep reinforcement learning models for combinatorial optimization tasks.
method Meta-SAGE combines a scale meta-learner and scheduled adaptation with guided exploration to adjust model parameters for larger-scale problems.
result Meta-SAGE outperforms previous methods and significantly improves scalability in CO tasks.
New Q-Learning algorithm reduces switching cost in MDPs.
problem Reducing adaptivity in real-world applications.
method Q-Learning with UCB2 exploration, quantified by local switching cost.
result Achieves sublinear regret with low switching cost.
This thesis explores symplectic foliations and local Lie groupoids, with applications in current theory.
problem Understanding calibratable symplectic foliations and local Lie groupoids.
method Applying de Rham's and Sullivan's theories to symplectic foliations and generalizing Mal'cev and Olver's theorems.
result Generalizations of theorems by Mal'cev and Olver for local Lie groupoids and algebroids.
Paper improves stochastic collocation for local volatility models.
problem Improving local volatility models for assets with boundaries.
method Applied stochastic collocation to lognormal distributions, derived analytical local volatility.
result Simple analytical Dupire local volatility derived from option prices.
Develops an SSBO algorithm for global optimization of expensive models.
problem Global optimization of expensive black-box models.
method Asynchronous hybrid-criterion with interval reduction.
result Improves global search ability and local search efficiency.
The paper explores moduli space of heterotic system using two deformation paths.
problem Exploring the moduli space of the heterotic system.
method Considering two dual deformation paths starting from a Kähler solution, one along Bott-Chern cohomology class and the other along Aeppli cohomology class. Using the implicit function theorem to prove local existence of heterotic solutions.
result Established an initial step to construct local moduli coordinates around a Kähler solution.
This thesis explores GNNs, categorizing them into local and global approaches.
problem Understanding the convergence of global GNNs and connecting local and global approaches.
method Categorization of GNNs into local and global, study of Invariant Graph Networks, connecting local and global approaches, and using local MPNN for graph coarsening.
result Established a connection between local and global GNN approaches.
Paper explores supervised learning methods to approximate ideal observer for joint signal detection and localization.
problem Optimizing medical imaging systems by assessing their performance using the Ideal Observer model.
method Uses supervised learning methods, specifically convolutional neural networks, to approximate the Ideal Observer for joint signal detection and localization tasks.
result Supervised learning-based methods can approximate the Ideal Observer for joint signal detection and localization tasks, as shown by comparisons to MCMC and analytical methods.
The paper explores risk-minimization for exponential additive models, providing mathematical expressions and numerical examples.
problem Risk-minimization in incomplete markets for exponential additive models.
method Derive explicit mathematical expressions for local risk-minimization strategies in exponential additive models.
result Provide necessary conditions for deriving expressions and confirm integrability conditions for specific models.
NTP struggles to learn relationships without increased exploration.
problem NTP's performance in extracting true relationships among data is poor.
method Created synthetic logical datasets with injected relationships to test NTP's performance and identify algorithmic issues.
result Increasing exploration in NTP's algorithm improves its performance in recovering relationships.
Simpler ε-greedy with longer action durations improves exploration.
problem Limited exploration capability of ε-greedy in complex domains.
method Temporally extended ε-greedy with repeated actions for random durations.
result Temporally extended ε-greedy outperforms sophisticated methods on various domains.