The paper analyzes how SGD visits different regions of a non-convex problem's state space.
problem Understanding the long-run distribution of stochastic gradient descent in non-convex problems.
method Large deviations theory and randomly perturbed dynamical systems.
result The long-run distribution of SGD resembles the Boltzmann-Gibbs distribution with temperature equal to the step-size.
New RL algorithm learns state aggregation architecture adaptively.
problem Adapting reinforcement learning value function architectures.
method Adapts state aggregation architecture using state visit frequency feedback.
result Improves RL performance on various test problems.
Study automates detection of visitation disruptions in ICU patients.
problem Difficulty in detecting frequent visitation disruptions in ICU patients.
method Used DensePose R-CNN model to count people in video frames, analyzed disruptions and patient outcomes.
result Automated method detects visitation disruptions, impacts on pain and length of stay examined.
Adapts reinforcement learning architectures using state visit frequency.
problem Determining an optimal approximation architecture for reinforcement learning.
method Adapts state aggregation approximation architecture based on state visit frequency.
result Guarantees VF estimate arbitrarily close to zero with large S. RECODE uses clustering and embedding to track state visitation counts in RL.
problem Efficient novelty-based exploration in nonstationary RL environments.
method Non-parametric clustering, online density estimation, inverse dynamics loss.
result RECODE achieves state-of-the-art performance in challenging RL tasks.
The paper introduces a new intrinsic reward method for exploration in reinforcement learning.
problem Improving exploration in reinforcement learning agents.
method Intrinsic rewards proportional to the entropy of future state-action features.
result The new objective leads to improved visitation of features within individual trajectories.
New RL approach uses future state and action visitation measures for better exploration.
problem Improving exploration in reinforcement learning.
method Intrinsic reward based on future state and action visitation measures, using contraction operators.
result Policies achieve good state-action space coverage and high performance.
MedGraph learns patient visit embeddings from EMRs, capturing both attributes and temporal sequences.
problem Limited EMR embedding methods fail to capture patient demographics, utilisation, and code descriptions.
method MedGraph constructs an attributed bipartite graph and uses a point process to model temporal sequences.
result MedGraph outperforms state-of-the-art methods in medical risk prediction tasks.
HDT improves MCMC on graphs with history-dependent sampling.
problem Efficient sampling from target distributions on general graphs with low computational overhead.
method History-driven target (HDT) framework that replaces the original target distribution with a history-dependent one.
result Near-zero variance performance and scalability to large graphs with memory-efficient implementation.
BeBold improves exploration in sparse-reward tasks by regulating visitation counts.
problem Efficient exploration in deep reinforcement learning under sparse rewards.
method Regulated difference of inverse visitation counts.
result BeBold solves 12 challenging tasks in MiniGrid with fewer steps than previous state-of-the-art.
Novel approach to high frequency trading microstructure.
problem Understanding high frequency trading and its impact on market dynamics.
method Introducing a new concept of informed traders and rigorous limit order book analysis.
result Market frictions misrepresent wealth by 50% and alter trading strategies.
FLARe model forecasts Alzheimer's progression using learned latent representations.
problem Forecasting Alzheimer's disease progression at the patient level.
method Generates a sequence of latent representations from longitudinal data across multiple modalities, incorporating time horizon.
result Outperforms baseline in forecasting accuracy and F1 score, robustly handling missing visits.
Efficiently learns uniform state distributions in unknown MDPs.
problem Learning uniform state distributions in unknown MDPs without rewards.
method Uses conditional gradient method with approximate MDP solver.
result Provable efficiency in sample and computational complexities.
This paper analyzes the multi-armed bandit problem using frequency-domain methods.
problem The exploration-exploitation trade-off in sequential decision-making.
method Proposes a frequency-domain analysis framework, reformulating the bandit process as a signal processing problem.
result Confidence bound term in UCB algorithm is equivalent to a time-varying gain in frequency domain.
The paper introduces return parity for fairness in MDPs, addressing delayed and adverse effects.
problem Fairness in MDPs for dynamic domains with delayed and adverse effects.
method Proposes return parity, decomposes return disparity, and develops algorithms for state visitation distributional alignment.
result The proposed algorithms can successfully close the disparity gap while maintaining policy performance.
Unsupervised model detects healthcare fraud from patient visit data.
problem Detecting fraudulent healthcare bills from patient visit data.
method Uses LSTM and seq2seq models for anomaly detection, normalizes scores with EDF.
result Improves anomaly detection for high class imbalance problems.
BRPO optimizes batch RL policies to better exploit state-action differences.
problem Batch RL's conservatism limits exploitation of state-action differences.
method Proposes residual policies and derives BRPO to maximize policy performance.
result BRPO achieves state-of-the-art performance in various tasks.
Deep learning predicts asthma ED visits better than traditional methods.
problem Predicting asthma-related ED visits to improve patient management.
method Deep learning (Artificial Neural Networks) compared to Lasso logistic regression.
result Deep learning model (ANN) outperforms traditional Lasso logistic regression (AUC = 0.845 vs. AUC = 0.842).
Simplifies complex RL policies by ranking important decisions.
problem Complexity in RL policies makes them hard to analyze and interpret.
method Statistical fault localisation to rank states and prune unimportant decisions.
result Pruned policies can perform similarly to original policies, improving interpretability.
The paper develops personalized DAG models for web user behavior.
problem Understanding user behavior transitions between websites with user heterogeneity and network dependency.
method Personalized Binomial DAG models with network-structured covariates, embedding network structure into a dimension-reduced covariate, learning node neighborhoods, and exploring variance-mean relation.
result The proposed algorithm outperforms state-of-the-art competitors in heterogeneous data.
Framework identifies comorbidities for frequent ED and inpatient visits.
problem Reducing resource usage and costs in frequent patients.
method Developed MSAR algorithm to identify comorbidities.
result MSAR identifies conditions most associated with reoccurring ED and inpatient visits.
Study online RL with mismatched dynamics, achieving sublinear regret.
problem Exploration challenges in online RL with mismatched training and deployment dynamics.
method Introduce supremal visitation ratio, propose efficient algorithm with f-divergence. result Achieves sublinear regret in online RMDPs with optimal dependence on supremal visitation ratio and interaction episodes.
Study how untrained policies explore in RL environments.
problem Challenges in reinforcement learning, especially sparse or adversarial reward structures.
method Theoretical and empirical analysis of untrained deep neural policies in a toy model.
result Untrained policies generate correlated actions and non-trivial state-visitation distributions.
Transforming sparse outcomes into dense process rewards for efficient reinforcement learning.
problem Training RL policies to maximize sparse outcomes.
method Incentivizing policy matching state-action visitations of successful episodes.
result Significantly faster RL finetuning performance.
Paper uses NLP to cluster patient visits for diagnosis validation.
problem Validating if similar patients receive similar diagnoses.
method Representation of medical visits using word embeddings, clustering patients' visits.
result Stable and separated segments of visits positively validated against diagnoses.
Algorithm optimizes constrained reinforcement learning with dual variables.
problem Minimizing convex functional subject to convex constraint in large state spaces.
method VPDPO algorithm using Lagrangian and Fenchel duality.
result Achieves sublinear regret and constraint violation, globally optimal policy.
Self-guided ALPs improve MDP policies without domain knowledge.
problem Improving MDP policies with minimal domain knowledge.
method Self-guided sequence of ALPs with random basis functions and state-relevance distribution.
result High probability error bounds and improved policy performance.
The paper introduces a new volatility model for state heterogeneous financial markets using high-frequency data.
problem State heterogeneity in financial volatility processes.
method Developed a state heterogeneous GARCH-Ito (SG-Ito) model based on continuous Ito diffusion process.
result Empirical studies reveal various state heterogeneities in S&P 500 index volatility.
UAIL uses uncertainty estimation to improve control systems in safety-critical tasks.
problem Improving control systems in safety-critical domains like autonomous driving.
method UAIL applies Monte Carlo Dropout to estimate uncertainty in control output and selectively acquire new training data.
result UAIL can reliably predict infractions and outperforms existing algorithms.
New algorithm stabilizes RL policy learning through divergence regularization.
problem Stabilize policy learning and improve performance in RL.
method Proximity term constraining discounted state-action visitation distributions to be close to each other.
result Proposed algorithm improves stability and final performance in RL tasks.
A new method improves reinforcement learning by directing exploration towards new knowledge.
problem Efficient exploration in reinforcement learning, especially in complex environments.
method Proposed E-values, a generalization of visit-counters, for directed exploration in model-free reinforcement learning. result Improves learning and performance in continuous Markov Decision Processes (MDPs) compared to traditional methods.
Study analyzes data breach reporting patterns and frequency across U.S. states, finding increasing trends after 2020.
problem Contradictory conclusions in data breach frequency trends due to inconsistent data collection and reporting standards.
method Joint analysis of state Attorneys General's publications on data breaches across eight states with established notification laws.
result Frequency of data breaches is increasing after 2020, with commonalities and heterogeneities across states.
SSMs have a built-in bias towards low-frequency components, which can be adjusted.
problem Frequency bias in SSMs affects their performance on long-range sequences.
method Proposed two mechanisms to tune frequency bias: scaling initialization or applying a Sobolev-norm-based filter.
result Tuning frequency bias improves SSMs' performance on long-range sequence learning tasks.
Novel non-parametric tree model learns tree distributions.
problem Learning distributions for tree-structured data.
method Bottom-up hidden tree Markov model with infinite states.
result Novel non-parametric generalization of hidden tree Markov model.
VisitHGNN predicts visit probabilities between neighborhoods and POIs using graph neural networks.
problem Estimating visit probabilities between neighborhoods and POIs for urban planning.
method Heterogeneous, relation-specific graph neural network (VisitHGNN) trained on mobility data.
result Strong predictive performance with high fidelity to observed travel behavior.
New method improves music transcription by treating frequency distributions holistically.
problem Small frequency shifts and variations in sound timbre harm traditional fit measures.
method Optimal transportation and new holistic frequency distribution measure.
result Simplified note templates lead to faster, state-of-the-art performance.
Complex frequency generalizes eigenvalues in LTI systems.
problem Characterizing dynamics of signals with complex values.
method Geometric frequency interpretation and transformation analysis.
result Complex frequencies in LTI systems match eigenvalues.
Low frequency perturbations improve model robustness, contrary to high frequency attacks.
problem Improving model robustness against adversarial attacks.
method Systematic control of frequency components in perturbations.
result Low frequency perturbations improve model robustness, especially in white-box and black-box settings.
Optimizes RTB campaigns by selecting user profiles and website configurations.
problem Maximizing impressions and profitability in RTB campaigns.
method Optimizes user profiles and website configurations, combines with other strategies.
result As the required number of visits increases, average profitability decreases.
New results on inferring hidden states in trackable weak models.
problem Inferring hidden states in trackable weak models.
method Analyzing strongly-connected trackable weak models and reconstructing branch choices.
result The number of hypotheses in strongly-connected trackable models is bounded by a constant.
Proposes intelligent web browser features for link prediction and category-wise recommendation.
problem Improving web browser intelligence for better link prediction and recommendation.
method Proposes frecency prediction for link prediction and URL classification for recommendation, with hyperparameter optimization.
result Improves browser recommendation accuracy by 10% through hyperparameter optimization.
A new search-control strategy improves Dyna's efficiency.
problem Improving sample efficiency in model-based reinforcement learning.
method Proposes a novel search-control strategy by sampling high frequency regions of the value function.
result Empirically shows that high frequency regions require more samples to approximate, suggesting a better search-control strategy.
New algorithms speed up inverse reinforcement learning by solving MDPs once.
problem Slow convergence in Maximum Entropy Inverse Reinforcement Learning.
method Deep Inverse Q-learning with constraints exploiting Q-learning.
result Up to several orders of magnitude speedup compared to existing methods.
New method learns frequency-dependent partial correlations.
problem Learning dependencies across distinct frequency bands.
method Formulate and solve two nonconvex learning problems.
result Proposed methods outperform existing state of the art.
Paper proposes machine learning model for early Alzheimer's diagnosis.
problem Early and accurate diagnosis of Alzheimer's Disease.
method Machine learning models, demographic, biomarker, and cognitive test data.
result 90% accuracy and 87% accuracy in predicting Alzheimer's development.
HyFAD improves time series imputation by combining time and frequency diffusion.
problem Improve time series imputation by handling frequency-sensitive denoising and balancing global and local dynamics.
method HyFAD is a hybrid time-frequency diffusion model with frequency-aware embedding, built on DDPM paradigm.
result HyFAD achieves state-of-the-art performance in time series imputation.
A new neural network improves frequency estimation from noisy signals.
problem Estimating frequencies of sinusoidal components in noisy signals.
method A novel neural network architecture combined with a module to detect the number of frequencies.
result Significantly more accurate frequency estimation at medium-to-high noise levels.
Predicts traffic flow using reinforcement learning and sensor data.
problem Accurately predict expanding and evolving long-term streaming traffic networks.
method Formulates the problem as a continuous reinforcement learning task, where the agent predicts future traffic based on sensor data.
result The approach improves accuracy in predicting traffic flow by updating the agent's state representation over time.