Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

113225338450 · Jun 202019922001200920182026
48 results for controlled experimentation

Optimizes experimental design using synthetic controls for better outcomes.

problem Estimating average treatment effects in studies with pre-treatment data.
method Mixed-integer programming for selecting treated and control units and weights.
result Improves mean squared error and statistical power compared to simple alternatives.

Contrastive ICA identifies features in experimental groups relative to controls.

problem Jointly analyzing experimental and control datasets to identify salient features.
method Developed contrastive ICA (cICA) using tensor decomposition.
result cICA identifies patterns and visualizes data effectively, outperforming existing methods.

Survey of reinforcement learning in continuous control, focusing on LQR.

problem Optimizing control in uncertain environments with models and cost of generality.
method Survey and case study of LQR, merging learning theory and control.
result Theoretical and experimental characterizations match, showing the role of models.

ADDIS improves power in online FDR control for conservative nulls.

problem Lack of power in adaptive FDR control algorithms for conservative nulls.
method ADDIS: adaptive discarding algorithm for online FDR control.
result ADDIS achieves best of both worlds: high power for conservative nulls and no loss for uniformly distributed nulls.

New method discovers discrepancies between simplified models and experimental data.

problem Model discrepancies in nonlinear systems lead to significant deviations from true behavior.
method Sparse Identification of Nonlinear Dynamics (SINDy) algorithm to discover sparse model terms.
result Improvement in performance with a discrepancy model in simulations.

New algorithm for adaptive experimental design in scientific settings.

problem Identifying true positives while controlling false discoveries in adaptive experimental design.
method Provably sample efficient adaptive algorithm for FDR control.
result First provably sample efficient adaptive algorithm for adaptive experimental design.

Study improves voice conversion model with Mel-spectrogram augmentation.

problem Insufficient speech pairs data for training sequence-to-sequence voice conversion models.
method Experimented with Mel-spectrogram augmentation using SpecAugment policies and proposed new augmentation policies.
result Time axis warping policies showed better performance in training the voice conversion model.

Data-driven control of robotic systems using Koopman operators with error bounds.

problem Real-time control of nonlinear robotic systems with unknown dynamics.
method Constructing a Koopman operator-based linear representation using higher-order derivatives of nonlinear dynamics, with error bounds derived from Taylor series accuracy analysis.
result The Koopman model provides marginally better performance than competing nonlinear modeling methods and can be efficiently controlled using linear control design tools.

A new Q-learning controller improves line follower robot control.

problem Challenges in controlling line follower robots due to unknown mechanical characteristics and uncertainties.
method Simulated annealing based Q learning method to address controller performance issues.
result The proposed controller outperforms conventional P controllers in line follower robots.

Agents acting in the natural world aim at selecting appropriate actions based on noisy and partial sensory observations. Many behaviors leading to decision mak- ing and action selection in a closed loop setting are naturally phrased within a control theoretic framework. Within the framework of optimal Control Theory, o…

2014-06-27abs ↗pdf ↗

Paper proposes a method to test and verify control systems with machine learning components.

problem Testing and verifying control systems with machine learning components is challenging.
method Gradient-based method combined with randomized search to find adversarial samples.
result Method outperforms Simulated Annealing optimization in finding adversarial samples.

Study uses multi-agent reinforcement learning to control self-assembly with high-resolution external control.

problem Designing effective external control protocols for self-assembly with high-resolution control.
method Investigated a multi-agent reinforcement learning approach, comparing fully decentralized and partially decentralized strategies.
result Partially decentralized approach outperforms fully decentralized in controlling self-assembly towards target structures.

HiDe learns hierarchical control for complex tasks by separating planning and control.

problem Solving long horizon control tasks with generalization to unseen scenarios.
method Functional decomposition of state-action spaces, RL-based planner, modular transfer of policy layers.
result Generalizes across unseen test environments and scales to longer horizons.

Safe offline RL for chemical reactors using input convex neural networks.

problem Safe control of exothermic polymerization reactors using historical data.
method Gymnasium-compatible simulation, behaviour cloning, implicit Q-learning, input convex neural networks (PICNNs).
result Offline RL with convex action correction outperforms traditional control approaches.

Estimates treatment effects in time series data with always-missing controls.

problem Lack of control group in time series data, especially during specific events.
method Recover control group in event period, account for confounders and temporal dependencies.
result Robust estimation of control group's potential outcome and accurate predicted holiday effect.

Estimates brain connectivity networks from natural stimuli, controlling type I error.

problem Estimating brain connectivity networks from natural stimuli with nuisance signals.
method Estimating stimulus-locked brain network by treating non-stimulus-induced signals as nuisance parameters. Testing maximum degree of network using inferential method.
result Proves type I error can be controlled and power increases asymptotically.

New algorithms control FDX while achieving more power in online multiple testing.

problem Problems with previous online multiple testing methods, including high FDX and low power.
method Developed new dynamic algorithms that adjust testing levels based on accumulated wealth.
result SupLORD algorithm achieves higher power and FDR control in synthetic experiments.

New method combines experimental and observational data for causal inference.

problem Combining internal validity of experiments and larger sample sizes of observations.
method Empirical risk minimization (ERM) framework with cross-validation.
result Efficacy and reliability demonstrated on real and synthetic data.

A controller optimizes sampling from unknown distributions to maximize a score function.

problem Optimizing sampling from unknown distributions to maximize a score function.
method Uniformly Fast (UF) sampling policies and UCB policy.
result UCB policy is asymptotically optimal and achieves the lower bound for sub-optimal activations.

Combines experimental and historical data for robust policy evaluation.

problem Policy evaluation with mixed data sources, especially experimental vs historical.
method Linear integration of estimators from experimental and historical data, optimized for MSE minimization.
result Proposed estimators outperform traditional methods in ridesharing company data.

Revisits VIC method to correct intrinsic reward bias in stochastic environments.

problem Intrinsic reward bias in VIC leading to suboptimal solutions.
method Proposes two methods based on transitional probability model and Gaussian mixture model to correct bias.
result Achieves maximal empowerment through corrected intrinsic reward.

A new reinforcement learning method uses mutual information to encourage agents to control their environment.

problem Learning from internal drives instead of external rewards.
method Formulate an intrinsic objective as mutual information between goal states and controllable states, derive a surrogate objective for efficient optimization.
result Demonstrated the efficacy of the approach in robotic tasks.

A new method for stochastic optimal control improves accuracy over existing techniques.

problem Improving the accuracy of stochastic optimal control for noisy systems.
method Stochastic Optimal Control Matching (SOCM) using Iterative Diffusion Optimization (IDO) with path-wise reparameterization trick.
result SOCM achieves lower error than existing techniques for three out of four control problems, sometimes by an order of magnitude.

D4PG combines distributional reinforcement learning with distributed learning for control tasks.

problem Continuous control tasks in reinforcement learning.
method Adapting distributional reinforcement learning to continuous control, using a distributed framework, N-step returns, and prioritized experience replay.
result D4PG achieves state-of-the-art performance across various control tasks.

A new approach to reinforcement learning improves policy performance by adjusting control frequency.

problem Improving reinforcement learning performance by optimizing control frequency.
method Introducing action persistence and a novel algorithm, PFQI, to learn optimal value function at a given persistence.
result PFQI effectively learns optimal value function with action persistence, improving reinforcement learning performance.

Paper proposes a deep RL approach for traffic signal control balancing efficiency and equity.

problem Inefficient and inflexible traffic signal controllers.
method Deep reinforcement learning with a novel reward function combining efficiency and equity.
result The proposed algorithm achieves state-of-the-art performance on various traffic scenarios.

Backdoor attacks on DRL-based traffic controllers cause stop-and-go waves or crashes.

problem Vulnerability of DRL-based traffic controllers to machine learning attacks.
method Developed a trigger design methodology based on traffic physics principles.
result Backdoored models can cause stop-and-go traffic waves or AV crashes when triggered.

Trust-region methods and natural gradients are equivalent in certain policy search scenarios.

problem Improving policy search methods in continuous control tasks.
method Introducing compatible policy search (COPOS) that uses natural parameterization and compatible value function approximation to control entropy loss.
result COPOS yields state-of-the-art results in challenging tasks and reduces entropy loss.

A new method for learning controlled dynamical systems efficiently and avoiding local minima.

problem Learning controlled dynamical systems with efficient and robust methods.
method Predictive State Representation with Random Fourier Features (RFFPSR) combining moment-matching, kernel embedding, and local optimization.
result The method avoids local minima and efficiently models controlled dynamical systems.

Action chunking and data exploration improve behavior cloning in robotics.

problem Exponential errors in learning from demonstrations for continuous control tasks.
method Action chunking and exploratory data collection.
result Control-theoretic stability is key to improving imitation learning.