Paper proposes an active multi-step TD algorithm for reinforcement learning.
problem Challenging decision making and control tasks in reinforcement learning.
method Active stepsize learning and adaptive multi-step TD algorithm with context-aware mechanism.
result Competitive results compared to other reinforcement learning baselines on discrete and continuous space tasks.
Transformers learn multi-step reasoning through gradient descent.
problem Understanding how transformers solve symbolic multi-step reasoning tasks.
method Theoretical analysis of gradient descent dynamics and multi-phase training.
result Trained one-layer transformers can solve both backward and forward reasoning tasks with generalization guarantees.
Robots learn tasks from a single demonstration using auxiliary video context.
problem Learning from demonstrations is challenging due to ambiguity and lack of labeled data.
method Metalearning to localize actions in auxiliary videos, learning reward functions, and reinforcement learning.
result Robots can learn multi-step tasks more effectively with auxiliary video context.
Stratify unifies and improves multi-step forecasting strategies.
problem Lack of unified frameworks for multi-step forecasting strategies.
method Proposes Stratify, a parameterized framework for multi-step forecasting.
result Novel strategies in Stratify outperform existing ones in over 84% of experiments.
STG2Seq predicts multi-step passenger demand with graph and hierarchical structure.
problem Predicting passenger demand over multiple time horizons is challenging due to nonlinear and dynamic spatial-temporal dependencies.
method Proposes a graph-based model with a hierarchical graph convolutional structure to capture spatial and temporal correlations.
result Consistently outperforms baseline and state-of-the-art models on real-world datasets.
The paper explores dynamic ensembles for multi-step forecasting.
problem Lack of research on dynamic ensembles for multi-step forecasting.
method Extensive experiments with 3568 time series and an ensemble of 30 multi-output models.
result Dynamic ensembles based on arbitrating and windowing perform best.
Improved multi-step TD learning with control variates reduces variance and improves performance.
problem Variance in multi-step TD learning causes divergence in off-policy settings.
method Per-decision control variates for multi-step TD algorithms.
result Control variates significantly improve performance in both on and off-policy tasks.
JANET improves time series prediction with adaptive uncertainty regions.
problem Time series data's lack of exchangeability and multi-step prediction challenges.
method Proposes JANET, a framework for joint adaptive prediction regions with controlled error rates.
result Demonstrates superior performance in multi-step prediction tasks across diverse datasets.
Diffusion-VAE tackles multi-step stock price prediction with stochastic noise.
problem Challenges in multi-step stock price prediction due to stochasticity and target price sequence.
method Combines hierarchical VAE and diffusion probabilistic techniques for seq2seq stock prediction.
result D-Va model outperforms state-of-the-art solutions in prediction accuracy and variance.
Proposes a multi-stream RNN model for predicting merchant transactions.
problem Predicting future transaction statistics of merchants.
method Multi-stream RNN model tailored for multivariate time series and multi-step predictions.
result Outperforms existing state-of-the-art methods in merchant transaction predictions.
New RL algorithms improve performance using multi-step greedy policies.
problem Improving model-free reinforcement learning performance.
method Developed multi-step greedy κ-Policy Iteration and κ-Value Iteration algorithms for model-free RL. result Multi-step greedy algorithms outperform DQN and TRPO on various benchmark tasks.
Proposes QDF to improve multi-step time-series forecasting.
problem Ignoring label autocorrelation and unequal task weights in training objectives.
method Quadratic-form weighted training objective and QDF learning algorithm.
result Improves performance of various forecast models, achieving state-of-the-art results.
Agents need world models to generalize multi-step tasks.
problem The necessity of world models for flexible, goal-directed behavior.
method Formal analysis and demonstration of the necessity of world models for agents to generalize multi-step tasks.
result World models are necessary for agents to generalize to multi-step goal-directed tasks.
Multi-step ahead forecasting is still an open challenge in time series forecasting. Several approaches that deal with this complex problem have been proposed in the literature but an extensive comparison on a large number of tasks is still missing. This paper aims to fill this gap by reviewing existing strategies for m…
The paper introduces a multi-step loss function to improve model-based reinforcement learning.
problem Compounding of one-step prediction errors in long trajectories.
method A multi-step objective function combining MSE losses at various future horizons.
result Models trained with the multi-step loss achieve significant improvement in future prediction.
This work analyzes CoT prompting methods from a statistical estimation perspective.
problem Improving the effectiveness of LLMs in solving multi-step reasoning problems.
method Introducing a multi-step latent variable model to characterize CoT prompting from a statistical estimation viewpoint.
result The CoT estimator is equivalent to a Bayesian estimator when the pretraining dataset is large.
This study evaluates multi-step reinforcement learning methods in the Mountain Car environment.
problem Lack of statistical significance in evaluating multi-step reinforcement learning methods.
method Combines n-step action-value algorithms with DQN architecture, tests in Mountain Car environment.
result Performance varies with off-policy correction, backup length, and target network update frequency.
The paper calculates prices for multi-step barrier options under the Black-Scholes model.
problem Calculating prices for multi-step barrier options with varying barriers and time steps.
method Derives a general, explicit expression for option prices using the Black-Scholes model and a multi-step reflection principle.
result Derives a multi-step reflection principle that generalizes the reflection principle of Brownian motion.
Efficiently optimizes expensive functions with multi-step lookahead using one-shot optimization.
problem Optimizing expensive functions with long-term impacts using myopic approaches.
method Formulated as nested optimization problems within a multi-step scenario tree, optimized in one-shot fashion.
result Multi-step expected improvement is computationally tractable and outperforms existing methods.
A new multi-step model improves model-based reinforcement learning efficiency.
problem Expensive environmental interaction in reinforcement learning.
method Proposes a multi-step model for predicting action sequences with variable length.
result Multi-step model outperforms one-step model in preliminary tests.
Paper presents a copula-based method to efficiently generate correlated sample paths from multi-step time series models.
problem Generating realistic correlation structures in multi-step forecast sample paths is expensive and time-consuming.
method Copula-based approach to generate correlated sample paths in one forward pass.
result Improved sample path quality and significant speedup over autoregressive sampling.
A multi-step model reduces compounding errors in reinforcement learning.
problem Compounding errors in one-step models lead to inaccurate predictions in reinforcement learning.
method Introduced a multi-step model that directly outputs the outcome of a sequence of actions.
result The multi-step model yields better action selection and more accurate value-function estimation.
A novel meta-learning method using ES for efficient reinforcement learning.
problem Sample inefficiency in reinforcement learning.
method Evolution strategies (ES) for exploration in parameter space, deterministic policy gradients for adaptation.
result Demonstrates improved performance in high-dimensional control tasks compared to gradient-based methods.
New multi-step approach improves fiber nonlinearity compensation efficiency.
problem Fewer steps are traditionally considered better for fiber nonlinearity compensation.
method Carefully designed multi-step machine learning approaches.
result Multi-step approaches lead to better performance-complexity trade-offs.
Paper adapts ACI for online multi-step time-series forecasting with coverage guarantees.
problem Achieving reliable error bounds in online multi-step time-series forecasting.
method Adaptive conformal inference (ACI) adapted for multi-step forecasting with dynamic significance levels.
result Proposes a multi-step ACI algorithm with finite-sample coverage guarantees for non-exchangeable data.
Improved multi-step prediction of drivable space for autonomous vehicles.
problem Accurate prediction of drivable space for safer, more comfortable navigation.
method Recurrent Neural Network (RNN) architectures trained on KITTI dataset, incorporating motion features.
result Significant improvement in prediction accuracy over state-of-the-art methods.
Study shows how a strong model can learn a task's feature while retaining other capabilities.
problem How to align superhuman AI systems using weak-to-strong generalization.
method Two-layer neural networks, reward-model learning, multi-step SGD, feature learning.
result The strong model efficiently learns task features while retaining general capabilities.
Model predicts stock price changes and forecasts using tokenized data.
problem Challenges in stock price forecasting and prediction due to dynamic data and statistical differences.
method Introduces PCIE model with tokenization to handle both forecasting and prediction.
result PCIE model outperforms state-of-the-art models in forecast and prediction tasks.
Improved model-based RL for 2-agent tasks reduces error accumulation.
problem Accumulating errors in model-based reinforcement learning for multi-agent systems.
method Disentangled variational auto-encoder for latent variable models of multi-step trajectory segments.
result Our approach achieves better sample efficiency and learns both cooperative and adversarial behavior.
Looped Transformers learn to implement multi-step gradient descent for in-context learning.
problem Understanding the learnability of multi-step algorithms in multi-layer Transformers.
method Training weight-sharing looped Transformers for in-context linear regression, proving gradient dominance condition for convergence.
result Looped Transformers implement multi-step preconditioned gradient descent, converging to global minimizer.
New approaches improve uncertainty quantification in autoregressive models for sequence data.
problem Uncertainty quantification in autoregressive models for exchangeable sequences.
method Study of inferential and architectural biases for autoregressive models, focusing on multi-step inference.
result Custom architectures are necessary for multi-step inference to ensure exchangeability.
In many forecasting applications, it is valuable to predict not only the value of a signal at a certain time point in the future, but also the values leading up to that point. This is especially true in clinical applications, where the future state of the patient can be less important than the patient's overall traject…
Paper proposes a dual-level approach for multi-step forecasting of dynamical systems.
problem Accurate multi-step forecasting of time series systems for automatic control and optimization.
method Hybrid input forecasting using LSTM-STMs and physics-informed neural networks (PINNs).
result Hybrid models achieve higher log-likelihood and lower MSE compared to conventional methods.
New framework guarantees convergence of multi-step MAML.
problem Convergence of multi-step MAML in nonconvex settings.
method Developed a theoretical framework for two types of MAML objective functions.
result Guaranteed convergence rate and computational complexity for multi-step MAML.
Paper explains DRL strategies for portfolio management using linear models.
problem Difficulty in understanding DRL-based trading strategies.
method Empirical approach using linear models and integrated gradients.
result DRL agents show stronger multi-step prediction power than machine learning methods.
Deep neural networks approximate unknown governing equations from data.
problem Approximating unknown governing equations from observational data.
method Residual network (ResNet) and multi-step methods (RT-ResNet, RS-ResNet) for equation recovery.
result Deep neural networks can recover governing equations without time derivative data.
In its simplest form, the traffic flow prediction problem is restricted to predicting a single time-step into the future. Multi-step traffic flow prediction extends this set-up to the case where predicting multiple time-steps into the future based on some finite history is of interest. This problem is significantly mor…
ForecastNet uses a time-variant deep feed-forward neural network for better multi-step-ahead time series forecasting.
problem Time-invariant architectures limit multi-step-ahead forecasting.
method ForecastNet employs a deep feed-forward architecture with time-variant parameters and interleaved outputs.
result ForecastNet outperforms other models on multi-step-ahead time series forecasting tasks.
Achieving artificial visual reasoning - the ability to answer image-related questions which require a multi-step, high-level process - is an important step towards artificial general intelligence. This multi-modal task requires learning a question-dependent, structured reasoning process over images from language. Stand…
A multi-step framework tackles online unsupervised domain adaptation with novel mean-target subspace computation.
problem Online unsupervised domain adaptation with unlabelled target data arriving sequentially.
method Multi-step framework with a novel mean-target subspace computation and temporal coherency consideration.
result Improved performance over previous approaches on four datasets.
Quantile deep learning improves time series prediction accuracy and uncertainty quantification.
problem Uncertainty in multi-step time series prediction.
method Developed a novel quantile regression deep learning framework for multi-step time series prediction.
result Integrating quantile loss function with deep learning provides additional predictions for selected quantiles without loss in accuracy.
A two-step defense method generates strong adversarial examples at low cost.
problem Vulnerability of deep neural networks to adversarial attacks.
method Develops a two-step defense approach that generates strong adversarial examples using FGSM at a lower computational cost compared to traditional multi-step adversarial training.
result Demonstrates effectiveness of the two-step defense approach against various attack methods with comparable robustness to traditional multi-step adversarial training.
Proposes a framework to quantify uncertainty in multi-step decision-making by LLMs.
problem Uncertainty quantification in multi-step decision-making scenarios of LLMs.
method A principled, information-theoretic framework decomposing uncertainty into internal and extrinsic components, and proposing UProp for efficient extrinsic uncertainty estimation.
result UProp significantly outperforms existing single-turn UQ baselines in multi-step decision-making benchmarks.
AEnbMIMOCQR generates robust multi-step ahead prediction intervals for time series data.
problem Generating reliable multi-step ahead prediction intervals for time series data.
method Adaptive ensemble batch multi-input multi-output conformalized quantile regression (AEnbMIMOCQR) based on conformal prediction principles.
result AEnbMIMOCQR provides close to exact coverage and robustness to distribution shifts.
TSDS framework reduces edge LLM agent compute by 43%-73% while maintaining safety and reliability.
problem Managing reasoning budget and uncertainty in edge LLM agents.
method Integrates a lightweight convergence probe and a perplexity-based deferral rule calibrated via multi-objective LTT.
result Reduces per-episode thinking compute by 43%-73% over deferral-only baselines.
A method to reduce Hessian matrix calculation cost in gradient-based meta-learning.
problem High memory footprint in calculating Hessian matrix for large-scale applications.
method Multi-step estimation of gradients to reuse the same gradient in a window of inner steps.
result Significant reduction in training time and memory usage with competitive or improved accuracies.
PCHID improves sample efficiency in reinforcement learning tasks.
problem Sparse rewards make learning policies difficult in reinforcement learning.
method Hindsight Inverse Dynamics with Hindsight Experience Replay and Policy Continuation.
result PCHID significantly improves sample efficiency and final performance on multi-goal tasks.
Survey of machine learning methods for spatiotemporal sequence forecasting.
problem Forecasting multi-step future of spatiotemporal systems based on past observations.
method Defined STSF problem, classified into subcategories, identified challenges, reviewed existing methods.
result No unified perspective on machine learning for STSF previously existed.