Optimizes experimental design using synthetic controls for better outcomes.
problem Estimating average treatment effects in studies with pre-treatment data.
method Mixed-integer programming for selecting treated and control units and weights.
result Improves mean squared error and statistical power compared to simple alternatives.
Contrastive ICA identifies features in experimental groups relative to controls.
problem Jointly analyzing experimental and control datasets to identify salient features.
method Developed contrastive ICA (cICA) using tensor decomposition.
result cICA identifies patterns and visualizes data effectively, outperforming existing methods.
Combines active learning and logged data for better classifier learning.
problem Learning classifier on entire population from logged data.
method Combines active learning and controlled random experimentation, modifies disagreement-based algorithms.
result Achieves best of active learning and logged data approaches.
New method estimates model and uncertainty for efficient control design.
problem Optimal control of systems with unknown dynamics.
method Coarse-ID control procedure using random matrix theory and convex optimization.
result End-to-end bounds on control cost relative error, nearly optimal.
New method minimizes experiment cost while maintaining accuracy.
problem Minimizing cost in experiments with interference or other concerns.
method Synthetically Controlled Thompson Sampling (SCTS).
result Minimizes regret and maintains inferential ability.
Survey of reinforcement learning in continuous control, focusing on LQR.
problem Optimizing control in uncertain environments with models and cost of generality.
method Survey and case study of LQR, merging learning theory and control.
result Theoretical and experimental characterizations match, showing the role of models.
TAD efficiently finds optimal settings for advanced manufacturing.
problem Optimizing high-dimensional process control parameters for optimal design features.
method TAD uses Gaussian process surrogate models and optimizes log-predictive likelihood to find optimal settings.
result TAD efficiently locates optimal settings with quantified uncertainty.
New method infers human sensorimotor costs from behavior.
problem Inferring human sensorimotor costs from observed behavior.
method Inverse optimal control with signal-dependent noise.
result Recovering costs and benefits in sensorimotor behavior.
ADDIS improves power in online FDR control for conservative nulls.
problem Lack of power in adaptive FDR control algorithms for conservative nulls.
method ADDIS: adaptive discarding algorithm for online FDR control.
result ADDIS achieves best of both worlds: high power for conservative nulls and no loss for uniformly distributed nulls.
New method discovers discrepancies between simplified models and experimental data.
problem Model discrepancies in nonlinear systems lead to significant deviations from true behavior.
method Sparse Identification of Nonlinear Dynamics (SINDy) algorithm to discover sparse model terms.
result Improvement in performance with a discrepancy model in simulations.
Greedy policy maximizes information in unknown linear systems.
problem Exploration in unknown linear dynamical systems.
method Online greedy policy maximizing information.
result Competitive performance compared to gradient-based methods.
New RL algorithms improve control tasks with data reuse.
problem Real-world control requires performance guarantees and data efficiency.
method Generalized Policy Improvement combining on-policy guarantees and sample reuse.
result Extensive experimental analysis shows benefits of new algorithms.
New algorithm for adaptive experimental design in scientific settings.
problem Identifying true positives while controlling false discoveries in adaptive experimental design.
method Provably sample efficient adaptive algorithm for FDR control.
result First provably sample efficient adaptive algorithm for adaptive experimental design.
Study improves voice conversion model with Mel-spectrogram augmentation.
problem Insufficient speech pairs data for training sequence-to-sequence voice conversion models.
method Experimented with Mel-spectrogram augmentation using SpecAugment policies and proposed new augmentation policies.
result Time axis warping policies showed better performance in training the voice conversion model.
Data-driven control of robotic systems using Koopman operators with error bounds.
problem Real-time control of nonlinear robotic systems with unknown dynamics.
method Constructing a Koopman operator-based linear representation using higher-order derivatives of nonlinear dynamics, with error bounds derived from Taylor series accuracy analysis.
result The Koopman model provides marginally better performance than competing nonlinear modeling methods and can be efficiently controlled using linear control design tools.
A method for control of complex systems using sparse data and reinforcement learning.
problem Control of complex systems with limited and streaming data.
method Discrete embedding space, Markov process model, reinforcement learning.
result The method performs well on experimental systems.
A new Q-learning controller improves line follower robot control.
problem Challenges in controlling line follower robots due to unknown mechanical characteristics and uncertainties.
method Simulated annealing based Q learning method to address controller performance issues.
result The proposed controller outperforms conventional P controllers in line follower robots.
Agents acting in the natural world aim at selecting appropriate actions based on noisy and partial sensory observations. Many behaviors leading to decision mak- ing and action selection in a closed loop setting are naturally phrased within a control theoretic framework. Within the framework of optimal Control Theory, o…
Paper proposes a method to test and verify control systems with machine learning components.
problem Testing and verifying control systems with machine learning components is challenging.
method Gradient-based method combined with randomized search to find adversarial samples.
result Method outperforms Simulated Annealing optimization in finding adversarial samples.
Study uses multi-agent reinforcement learning to control self-assembly with high-resolution external control.
problem Designing effective external control protocols for self-assembly with high-resolution control.
method Investigated a multi-agent reinforcement learning approach, comparing fully decentralized and partially decentralized strategies.
result Partially decentralized approach outperforms fully decentralized in controlling self-assembly towards target structures.
HiDe learns hierarchical control for complex tasks by separating planning and control.
problem Solving long horizon control tasks with generalization to unseen scenarios.
method Functional decomposition of state-action spaces, RL-based planner, modular transfer of policy layers.
result Generalizes across unseen test environments and scales to longer horizons.
A method for user-controlled semantic image filling.
problem Generating coherent images with user-specified semantics.
method Deep generative model combining encoder, latent variables, and PixelCNN.
result User can control the inpainting of unobserved pixels while maintaining semantic coherence.
Maximizes robustness in Bayesian experimental design under model uncertainty.
problem Brittleness of Bayesian experimental design under model misspecification.
method Formulates as a max--min game, uses Sibson's α-MI, and adopts PAC-Bayes framework.
result Establishes robust belief update and conditional information gain measure.
Safe offline RL for chemical reactors using input convex neural networks.
problem Safe control of exothermic polymerization reactors using historical data.
method Gymnasium-compatible simulation, behaviour cloning, implicit Q-learning, input convex neural networks (PICNNs).
result Offline RL with convex action correction outperforms traditional control approaches.
New LQR kernels improve controller learning from data.
problem Optimal controller design for nonlinear systems from data is challenging.
method Developed LQR kernels for Bayesian optimization.
result LQR kernels lead to superior learning performance on uncertain systems.
Estimates treatment effects in time series data with always-missing controls.
problem Lack of control group in time series data, especially during specific events.
method Recover control group in event period, account for confounders and temporal dependencies.
result Robust estimation of control group's potential outcome and accurate predicted holiday effect.
Estimates brain connectivity networks from natural stimuli, controlling type I error.
problem Estimating brain connectivity networks from natural stimuli with nuisance signals.
method Estimating stimulus-locked brain network by treating non-stimulus-induced signals as nuisance parameters. Testing maximum degree of network using inferential method.
result Proves type I error can be controlled and power increases asymptotically.
New algorithms control FDX while achieving more power in online multiple testing.
problem Problems with previous online multiple testing methods, including high FDX and low power.
method Developed new dynamic algorithms that adjust testing levels based on accumulated wealth.
result SupLORD algorithm achieves higher power and FDR control in synthetic experiments.
New method combines experimental and observational data for causal inference.
problem Combining internal validity of experiments and larger sample sizes of observations.
method Empirical risk minimization (ERM) framework with cross-validation.
result Efficacy and reliability demonstrated on real and synthetic data.
CNN identifies nonlinear human posture control models efficiently.
problem Identifying nonlinear human posture control models.
method Convolutional Neural Networks (CNN) for model identification.
result Efficiently identifies nonlinear human posture control models.
A controller optimizes sampling from unknown distributions to maximize a score function.
problem Optimizing sampling from unknown distributions to maximize a score function.
method Uniformly Fast (UF) sampling policies and UCB policy.
result UCB policy is asymptotically optimal and achieves the lower bound for sub-optimal activations.
Paper proposes E/PD-Control for better neural network training.
problem Training efficiency and robustness of CNNs in online data flows.
method E/PD-Control combines feedback PD controller with exponential signal.
result Better learning efficiency and robustness demonstrated experimentally.
Combines experimental and historical data for robust policy evaluation.
problem Policy evaluation with mixed data sources, especially experimental vs historical.
method Linear integration of estimators from experimental and historical data, optimized for MSE minimization.
result Proposed estimators outperform traditional methods in ridesharing company data.
Revisits VIC method to correct intrinsic reward bias in stochastic environments.
problem Intrinsic reward bias in VIC leading to suboptimal solutions.
method Proposes two methods based on transitional probability model and Gaussian mixture model to correct bias.
result Achieves maximal empowerment through corrected intrinsic reward.
A new reinforcement learning method uses mutual information to encourage agents to control their environment.
problem Learning from internal drives instead of external rewards.
method Formulate an intrinsic objective as mutual information between goal states and controllable states, derive a surrogate objective for efficient optimization.
result Demonstrated the efficacy of the approach in robotic tasks.
A new method for stochastic optimal control improves accuracy over existing techniques.
problem Improving the accuracy of stochastic optimal control for noisy systems.
method Stochastic Optimal Control Matching (SOCM) using Iterative Diffusion Optimization (IDO) with path-wise reparameterization trick.
result SOCM achieves lower error than existing techniques for three out of four control problems, sometimes by an order of magnitude.
A new policy gradient estimator reduces variance for clipped actions in continuous control tasks.
problem Policy gradient methods struggle with bounded action spaces.
method Proposes a new policy gradient estimator that accounts for clipped actions.
result The new estimator achieves lower variance and outperforms conventional methods.
Batch-normalized RHN improves gradient control in recurrent networks.
problem Gradient vanishing or exploding in recurrent networks.
method Batch normalization applied at each recurrence loop in RHN.
result Batch-normalized RHN converges faster and performs better.
D4PG combines distributional reinforcement learning with distributed learning for control tasks.
problem Continuous control tasks in reinforcement learning.
method Adapting distributional reinforcement learning to continuous control, using a distributed framework, N-step returns, and prioritized experience replay.
result D4PG achieves state-of-the-art performance across various control tasks.
A new approach to reinforcement learning improves policy performance by adjusting control frequency.
problem Improving reinforcement learning performance by optimizing control frequency.
method Introducing action persistence and a novel algorithm, PFQI, to learn optimal value function at a given persistence.
result PFQI effectively learns optimal value function with action persistence, improving reinforcement learning performance.
Paper proposes DigMA to generate controllable financial market orders.
problem Generating realistic financial market orders with controllability.
method DigMA model using conditional diffusion and meta agent.
result DigMA achieves superior controllability and generation fidelity.
Proposes neural networks for variance reduction in Monte Carlo estimation.
problem High variance in Monte Carlo estimations for complex functions.
method Uses neural networks to learn control variates from auxiliary random variables.
result Significant variance reduction in thermodynamic integration and reinforcement learning.
Safe learning of stochastic dynamics with safety constraints.
problem Learning controlled stochastic dynamics with safety constraints.
method Iterative expansion of a safe control set using kernel-based confidence bounds.
result The method ensures safe exploration and efficient estimation of system dynamics.
Paper proposes a deep RL approach for traffic signal control balancing efficiency and equity.
problem Inefficient and inflexible traffic signal controllers.
method Deep reinforcement learning with a novel reward function combining efficiency and equity.
result The proposed algorithm achieves state-of-the-art performance on various traffic scenarios.
Backdoor attacks on DRL-based traffic controllers cause stop-and-go waves or crashes.
problem Vulnerability of DRL-based traffic controllers to machine learning attacks.
method Developed a trigger design methodology based on traffic physics principles.
result Backdoored models can cause stop-and-go traffic waves or AV crashes when triggered.
Trust-region methods and natural gradients are equivalent in certain policy search scenarios.
problem Improving policy search methods in continuous control tasks.
method Introducing compatible policy search (COPOS) that uses natural parameterization and compatible value function approximation to control entropy loss.
result COPOS yields state-of-the-art results in challenging tasks and reduces entropy loss.
A new method for learning controlled dynamical systems efficiently and avoiding local minima.
problem Learning controlled dynamical systems with efficient and robust methods.
method Predictive State Representation with Random Fourier Features (RFFPSR) combining moment-matching, kernel embedding, and local optimization.
result The method avoids local minima and efficiently models controlled dynamical systems.
Action chunking and data exploration improve behavior cloning in robotics.
problem Exponential errors in learning from demonstrations for continuous control tasks.
method Action chunking and exploratory data collection.
result Control-theoretic stability is key to improving imitation learning.