Paper tackles reinforcement learning with complex observations and simple latent dynamics.
problem Understanding reinforcement learning with complex observations and simple latent dynamics.
method Statistical and algorithmic analysis of reinforcement learning under general latent dynamics.
result Identifies latent pushforward coverability as a condition for statistical tractability.
Unified framework for singular statistical models using observable charts.
problem Non-identifiability and breakdown of classical asymptotic theory in singular models.
method Invariant framework based on observable charts to define local coordinate systems in model space.
result Observable order provides a lower bound on KL divergence vanishing rate in singular models.
Sig-PCA integrates model outputs and observations to correct model biases.
problem Improving model accuracy and reliability by correcting biases and numerical approximations.
method Sig-PCA framework that combines summary statistics from model outputs with localized observations via a neural network.
result Corrects model outputs to align closely with observational data, preserving essential statistical information.
New algorithm improves reinforcement learning from partial observations.
problem Inferior performance of algorithms in real-world reinforcement learning due to partial observability.
method Representation-based approach to POMDPs, leading to a tractable algorithm.
result Empirically demonstrates superior performance with partial observations.
This research uses machine learning to approximate ideal and hotelling observer performance for binary signal detection.
problem Approximating the Ideal and Hotelling Observers for binary signal detection tasks.
method Supervised learning methods, including CNNs and SLNNs, are employed to approximate the IO and HO test statistics.
result The proposed supervised learning methods provide accurate approximations of the IO and HO test statistics.
Paper uses learned summary statistics for Bayesian inference with difficult likelihood functions.
problem Difficult to obtain exact likelihood function for observation data and simulation model.
method Simulation-based inference with learned summary statistics, using Cressie-Read discrepancy criterion.
result Effective inference performed over selected sample sets of observation data.
EFI automates statistical inference for big data.
problem Statistical inference for model parameters based on observations.
method EFI uses stochastic gradient Markov chain Monte Carlo and sparse deep neural networks.
result EFI provides higher fidelity in parameter estimation and automates the inference process.
New method uses sufficient statistics to infer causal relationships from observational data.
problem Inferring causal relationships from observational data with hidden variables.
method Information Bottleneck method applied to find functional sufficient statistics.
result New causal rules not obtainable from standard methods, validated on simulated and real data.
Efficiently learns Ising model parameters with limited statistics.
problem Learning Ising model parameters with limited sample configurations.
method Examines trade-offs between computation and observation, using Ising model as example.
result Reconstructs model parameters with statistics up to order O(γ) for ℓ1 width γ. Consider observing an undirected network that is `noisy' in the sense that there are Type I and Type II errors in the observation of edges. Such errors can arise, for example, in the context of inferring gene regulatory networks in genomics or functional connectivity networks in neuroscience. Given a single observed ne…
Reduced model helps estimate parameters from partially observed multiscale data.
problem Estimating parameters from high-dimensional partially observed multiscale data.
method Established convergence of filter from high-dimensional to reduced dimension model.
result Statistical estimation of parameters is possible in lower dimensions.
Stochastic methods improve data assimilation with high-frequency sensor data.
problem Computational challenges in data assimilation with high-frequency sensor data.
method Adapted stochastic approximation methods to handle high-frequency observations.
result Produces high-quality estimates using all observations without compromising statistical accuracy.
Horse racing odds match random division statistics.
problem Understanding the distribution of horses' winning abilities.
method Comparing horse racing data with the 'randomly broken stick' problem.
result Horses' winning abilities are exponentially distributed.
Efficient RL in large POMDPs with latent determinism and embeddings.
problem Efficient reinforcement learning in large-scale POMDPs with latent states and observations.
method Conditional Hilbert space embeddings, linear optimal Q-function, deterministic latent transitions, gap assumption. result Computationally and statistically efficient algorithm for exact optimal policy.
The paper introduces a method to incorporate expert opinion on observable quantities into statistical models.
problem Tackling the challenge of integrating expert knowledge on observable quantities into statistical models.
method The approach involves updating a prior belief using a loss function that reflects expert opinion on observable quantities.
result The method allows for a flexible specification of expert opinion and is straightforward to implement.
Study interpolating estimators for causal learning from observational data.
problem Learning causal models from observational data in complex model classes.
method Investigate min-norm interpolators and ridge-regularized regressors in a linearly confounded model.
result Interpolators cannot be optimal for causal learning under the principle of independent causal mechanisms, requiring stronger regularization.
Guarantees for third-person imitation learning from offline data.
problem Improving generalizability in imitation learning.
method Problem-dependent statistical learning guarantees for third-person imitation from offline observation.
result Strong performance guarantees for transferred policies in the offline setting.
StatEcoNet models species distribution using neural networks to correct observation errors.
problem Correcting observation errors in wildlife surveys for accurate species distribution modeling.
method StatEcoNet integrates a graphical generative model with neural networks to address SDM challenges.
result StatEcoNet outperforms traditional methods on simulated and real datasets.
A new framework for scalable uncertainty quantification in statistical models.
problem Computational bottleneck in uncertainty quantification for statistical machine learning.
method Predictive-matching Generative Parameter Sampler (GPS) framework.
result The GPS framework provides successful uncertainty quantification and additional flexibility.
Complicated generative models often result in a situation where computing the likelihood of observed data is intractable, while simulating from the conditional density given a parameter value is relatively easy. Approximate Bayesian Computation (ABC) is a paradigm that enables simulation-based posterior inference in su…
Proposes Population Difference Criterion for visually observed subpopulation differences.
problem Statistical significance of visually observed subpopulation differences in high-dimensional and high-signal contexts.
method Balanced permutation approach and bootstrap confidence interval for quantifying uncertainty.
result Balanced permutation approach is more powerful in high-signal contexts.
A new method for statistical inference using SGD under φ-mixing data.
problem Valid statistical inference for time series data with general correlation.
method Proposes a mini-batch SGD estimator and associated mini-batch bootstrap procedure for φ-mixing data. result The proposed method constructs valid confidence intervals for φ-mixing data. Approximate Bayesian Computation (ABC) are likelihood-free Monte Carlo methods. ABC methods use a comparison between simulated data, using different parameters drew from a prior distribution, and observed data. This comparison process is based on computing a distance between the summary statistics from the simulated da…
The paper introduces a statistical distance matrix for better feature representation and clustering.
problem Lack of detailed distance representation between feature elements.
method Extended traditional statistical distance to a matrix form (statistical distance matrix) and applied hierarchical clustering.
result The statistical distance matrix with clustering (Information Mandala) provides clearer and geometrically arranged feature representations.
New algorithm for partially observable contexts in finance.
problem Decision making based on partially observable, correlated market information.
method EMKF-Bandit algorithm integrating system identification, filtering, and bandit algorithms.
result Sub-linear regret under conditions on filtering.
Algorithm detects unmeasured confounding in observational data.
problem Estimating treatment effects in observational studies with untestable conditions.
method Two-stage procedure that detects dependencies between causal mechanisms.
result Algorithm efficiently detects confounding on simulated and semi-synthetic data.
In this paper we present detailed simulation results on the wealth distribution model with quenched saving propensities. Unlike other wealth distribution models where the saving propensities are either zero or constant, this model is not found to be ergodic and self-averaging. The wealth distribution statistics with a …
The level crossing and inverse statistics analysis of DAX and oil price time series are given. We determine the average frequency of positive-slope crossings, να+, where Tα=1/να+ is the average waiting time for observing the level α again. We estimate the probability P(K,α), which provides us the probab…
The study provides statistical theory for WGANs in time series forecasting.
problem Statistical analysis of WGANs for time series forecasting.
method Statistical theory and upper bounds for excess Bayes risk, weak convergence, and confidence intervals.
result Developed confidence intervals for time series forecasting using WGANs.
Statistical framework for inexact graph matching with errorfully observed graphs.
problem Finding a correspondence between graphs with errors.
method Introducing a corrupting channel model and using maximum likelihood estimation.
result Maximum likelihood estimation is a solution to the inexact graph matching problem.
New criteria distinguish cause from effect in data, overcoming statistical limitations.
problem Determining causal direction from statistical dependence alone.
method Intuitive criteria based on simplicity of prediction, tested on synthetic data.
result Criteria accurately distinguish cause from effect in various scenarios.
Estimates proportions of LLM-generated text in mixed documents.
problem Estimating the proportion of text generated by a pre-specified LLM in mixed documents.
method Developed estimators for two observation regimes: full observation and pivotal reduction, and established sample complexity bounds.
result Full observation estimators require fewer samples than pivotal reduction estimators.
New statistics improve feature importance detection with false discovery guarantees.
problem Identifying truly correlated features from observational data.
method Developed efficient knockoff generation from Bayesian Networks and new statistics.
result Improved power and efficiency of feature importance detection.
A new method uses maximum entropy for time series analysis.
problem Challenges in testing statistical properties of multivariate time series.
method Statistical mechanical approach for ensembles of time series.
result Shows possible applications in financial portfolio selection.
Novel strategy benchmarks observational studies against randomized trials.
problem Benchmarking observational studies for treatment effect bias.
method Statistical test for null hypothesis of treatment effect difference.
result Valid lower bound on maximum bias strength for any subgroup.
We introduce RSE to measure robustness in estimation problems.
problem Estimating statistical models from observed data.
method Developed theory for spectral functions of measures to compute RSE.
result RSE reveals a reciprocal relationship with problem complexity.
A new imputation method estimates missing values by matching observed marginals from masked data.
problem Missing values in data undermine statistical and machine learning analysis.
method Estimates a distribution from masked observations using positive semi-definite kernel density estimation.
result The method yields both single and multiple imputations from the same fitted density, with statistical consistency and fast adaptive excess risk.
Study on estimating Gumbel--Max watermark proportions in edited documents.
problem Estimating the proportion of a document generated from a watermarked LLM.
method Comparison of full observation and pivotal reduction observation regimes; development of estimators and information-theoretic lower bounds.
result Full observation yields a substantially smaller sample complexity compared to pivotal reduction.
Paper uses statistical depth to create DP estimators for regression.
problem Creating differentially private estimators in high dimensions.
method Uses halfspace and regression depth to analyze maximum influence and construct DP estimators.
result New DP estimators for location and regression show favorable performance.
The paper develops a method to learn robust decision policies from observational data, reducing high-cost outcomes.
problem Learning safe decision policies from observational data with high-risk outcomes.
method Develops a method to learn policies that reduce high-cost outcomes, valid under finite samples and uneven feature overlap.
result Validates the method with real and synthetic data, providing statistical bounds on decision costs.
The paper improves generative models to avoid replicating observed examples.
problem Improving generative models to avoid replicating observed examples.
method Theoretical insights into the Wasserstein GAN, constrained to left-invertible push-forward maps, generating distributions that avoid replication and significantly deviate from the empirical distribution.
result Left-invertibility achieves this without compromising statistical optimality.
A review of statistical SSL methods showing improved classifier performance.
problem Forming classifiers from limited labeled data and many unlabeled data.
method Statistical approaches to semi-supervised learning.
result A classifier from partially labeled data can have lower expected error rate.
Paper proposes a scoring function for detecting anomalies in large datasets.
problem Detecting outliers in large, feature-rich datasets.
method Binary classification problem with a two-sample linear rank statistic.
result Empirical results show the effectiveness of the proposed scoring function.
A new measure of model complexity based on Fisher Information.
problem Model complexity measurement in statistical models.
method Effective dimension defined by the number of cubes needed to cover the model space.
result The effective dimension is scale-dependent and measures model complexity.
Survey on causal inference methods for observational data.
problem Estimating causal effects from observational data.
method Comprehensive review of causal inference methods under the potential outcome framework.
result Various causal effect estimation methods have been developed and compared.
Paper proposes a method to validate statistical models based on data consistency.
problem Validation of statistical modeling assumptions in scientific inference problems.
method Automatic evaluation of model consistency with observed data.
result The proposed criterion assesses models' ability to generate similar data.
We select n stocks traded in the New York Stock Exchange and we form a statistical ensemble of daily stock returns for each of the k trading days of our database from the stock price time series. We analyze each ensemble of stock returns by extracting its first four central moments. We observe that these moments are fl…
New theory allows ICA without assuming non-Gaussian sources.
problem Traditional ICA struggles with Gaussian sources.
method Developed identifiability theory based on second-order statistics and sparsity.
result Identifiability theory and estimation methods validated experimentally.