PSDs improve RNN performance by predicting future observations.
problem Modeling dynamic processes with unknown latent states.
method Augmenting RNNs with Predictive-State Decoders (PSDs) that target predicting future observations.
result PSDs improve statistical performance of state-of-the-art RNNs with fewer iterations and less data.
RPSP networks combine PSRs and RNNs for reinforcement learning in POE.
problem Learning in partially observable environments.
method Recurrent filter with PSR, reactive policy, gradient descent.
result RPSP networks outperform memory-preserving models.
New method uses Cantor embeddings and Wasserstein distances to analyze predictive states in time series data.
problem Analyzing predictive states in stochastic processes using time series data.
method Wasserstein distances for detecting predictive equivalences in symbolic data, using Cantor embeddings for finite-dimensional representation.
result Exploratory analysis of temporal structure in various processes reveals insights.
This paper advances sample-efficient learning for partially observable RL by introducing B-stability and new algorithms.
problem Hard sample complexity for learning near-optimal policies in partially observable RL.
method Proposes B-stability as a unified structural condition and develops new algorithms for sample-efficient learning.
result Any B-stable PSR can be learned with polynomial samples, improving over current best complexities.
PSRNNs combine RNN and PSR insights for system filtering and prediction.
problem Modeling dynamical systems efficiently and accurately.
method Combines insights from RNNs and PSRs using bilinear transfer functions and tensor decomposition.
result PSRNNs outperform other models in filtering and prediction tasks across multiple datasets.
Improves PSR learning by refining spectral initialization with PSIM-style updates.
problem Inference performance of PSRs is poor despite good theoretical guarantees.
method Combines spectral algorithms for PSRs with PSIM-style updates for inference-based loss optimization.
result Inference Gradients outperforms PSRs and PSIMs on real and synthetic data.
New UCB algorithm for learning PSRs with tractable computation and accuracy.
problem Learning predictive state representations in sequential decision-making problems.
method Proposes a novel UCB-type algorithm with a bonus term to estimate PSRs accurately and efficiently.
result First known UCB-type approach for PSRs with guaranteed model accuracy and computational tractability.
Predictive State Representations (PSRs) are an expressive class of models for controlled stochastic processes. PSRs represent state as a set of predictions of future observable events. Because PSRs are defined entirely in terms of observable data, statistically consistent estimates of PSR parameters can be learned effi…
This paper proposes a method to select bases for spectral learning of PSRs using model entropy.
problem Learning PSR models with limited data and computational resources.
method Adopting model entropy to select columns for spectral learning of PSRs.
result The proposed method can effectively select bases for spectral learning of PSRs.
A new method for learning controlled dynamical systems efficiently and avoiding local minima.
problem Learning controlled dynamical systems with efficient and robust methods.
method Predictive State Representation with Random Fourier Features (RFFPSR) combining moment-matching, kernel embedding, and local optimization.
result The method avoids local minima and efficiently models controlled dynamical systems.
Predictive state representations (PSRs) offer an expressive framework for modelling partially observable systems. By compactly representing systems as functions of observable quantities, the PSR learning approach avoids using local-minima prone expectation-maximization and instead employs a globally optimal moment-base…
Reservoir computers and RNNs fall short of optimal prediction for stochastic PDFA.
problem Predicting stochastic processes generated by probabilistic deterministic finite-state automata.
method Generalized linear models, Reservoir computers, and Long Short-Term Memory (LSTM) RNNs were tested.
result Each method can fall short of maximal predictive accuracy by up to 50% after training.
We introduce 'mixed LICORS', an algorithm for learning nonlinear, high-dimensional dynamics from spatio-temporal data, suitable for both prediction and simulation. Mixed LICORS extends the recent LICORS algorithm (Goerg and Shalizi, 2012) from hard clustering of predictive distributions to a non-parametric, EM-like sof…
OMLE combines optimism and MLE for efficient sequential decision making.
problem Efficiently solving sequential decision making problems, especially in partially observable settings.
method Combines optimism for exploration and maximum likelihood estimation for model learning.
result OMLE learns near-optimal policies for a wide range of sequential decision making problems.
Temporal-difference (TD) networks are a class of predictive state representations that use well-established TD methods to learn models of partially observable dynamical systems. Previous research with TD networks has dealt only with dynamical systems with finite sets of observations and actions. We present an algorithm…
Spatio-temporal data is intrinsically high dimensional, so unsupervised modeling is only feasible if we can exploit structure in the process. When the dynamics are local in both space and time, this structure can be exploited by splitting the global field into many lower-dimensional "light cones". We review light cone …
As a technology to read brain states from measurable brain activities, brain decoding are widely applied in industries and medical sciences. In spite of high demands in these applications for a universal decoder that can be applied to all individuals simultaneously, large variation in brain activities across individual…
Novel low-rank neural decoder improves μ μ μ -ECoG neural decoding.
problem Challenging neural decoding from high-dimensional μ μ μ -ECoG data. method Low-rank structure in neural network decoder.
result Low-rank decoder outperforms standard PCA.
The paper introduces FMCI and hybrid decoding for hidden Markov models.
problem Computing distributions and decoding hidden state sequences in HMMs.
method Finite Markov chain imbedding (FMCI) and hybrid decoding.
result Hybrid decoding improves performance over traditional methods.
Study examines how decoding algorithms affect fairness in language generation models.
problem Impact of decoding algorithms on fairness in open-ended language generation.
method Systematic analysis of top- p p p , top- k k k , and temperature decoding algorithms. result Decoding algorithms significantly impact fairness across demographic groups.
Paper uses RL to optimize bit-flipping decoding for binary codes.
problem Improving bit-flipping decoding for binary linear codes.
method Mapped iterative decoding algorithms to MDPs for reinforcement learning.
result Learned BF decoders offer performance-complexity trade-offs and near-optimal performance.
Deep invertible networks decode EEG signals better than chance.
problem Decoding brain signals from EEG data.
method Deep invertible networks for generating and classifying brain signals.
result Deep invertible networks generate realistic EEG signals and classify novel signals above chance.
RL-VAE uses RL to decode molecular graphs from latent embeddings.
problem Efficiently decoding molecular graphs from latent embeddings.
method Repurposed simple graph generator for efficient decoding.
result Decoding molecular graphs from latent embeddings is possible with a simple graph generator.
Iterative BP-CNN improves channel decoding under correlated noise.
problem Channel decoding under correlated noise.
method Concatenates CNN with BP decoder, iteratively improving SNR.
result Iterative BP-CNN achieves better BER with lower complexity.
Deep learning aids ADMM-based decoding for binary linear codes.
problem Improving decoding efficiency for binary linear codes.
method Designing a decoding network based on ADMM and deep learning.
result Numerical results show improved performance compared to original ADMM.
Neural networks improve error correction in topological codes.
problem Finding optimal correction of errors in generic stabilizer codes is computationally hard.
method Systematic study of versatile neural-network decoders for topological codes.
result Neural decoders significantly improve error-correction threshold over leading efficient decoders.
Paper introduces deep neural decoders for near-term fault-tolerant quantum experiments.
problem Efficient decoders for quantum error correction under realistic noise.
method Deep neural decoders complemented by traditional algorithms.
result Deep neural decoders perform well in low noise regimes.
This paper analyzes speculative decoding, a method to speed up large language model inferences.
problem Theoretical understanding of speculative decoding is lacking.
method Conceptualizes speculative decoding as a markov chain problem and studies its key properties.
result Reveals fundamental connections between LLM components and their impact on decoding efficiency.
Neural decoder improves topological code performance.
problem Improving error correction for topological codes.
method Two-step neural network using pseudo-inverse of parity check matrix.
result Outperforms state-of-the-art non-neural decoders for 2D hexagonal color codes.
Machine learning improves neural decoding performance.
problem Traditional neural decoding methods are inefficient.
method Apply modern machine learning algorithms (neural networks, gradient boosting) for neural decoding.
result Modern methods significantly outperform traditional approaches.
Deep learning improves decoding of constrained sequence codes, reducing errors and increasing throughput.
problem Errors during transmission of constrained sequence codes.
method Deep learning, specifically MLP and CNN networks.
result Achieved low bit error rates close to MAP decoding and improved system throughput.
Improving the interpretability of brain decoding approaches is of primary interest in many neuroimaging studies. Despite extensive studies of this type, at present, there is no formal definition for interpretability of brain decoding models. As a consequence, there is no quantitative measure for evaluating the interpre…
CARDS improves decoding efficiency and alignment quality for LLMs.
problem Efficiency bottlenecks in decoding-time alignment for LLMs.
method Cascade Reward Sampling (CARDS) with segment-level rejection sampling and uncertainty-based segmentation.
result Significant improvement in decoding efficiency and alignment quality.
The paper develops a theory for speculative decoding acceptance criteria.
problem Speculative decoding's acceptance criteria and their rejection regions.
method Characterization of rejection regions as lower level sets of the target distribution, derivation of exact and margin-based certificates.
result Relaxed and tree-based acceptance criteria substantially enlarge the region of certified acceptance.
This work proposes an efficient autoregressive model for text generation.
problem The challenge of generating high-quality text with autoregressive models.
method Introduces a cascaded decoding approach using Markov transformers to achieve sub-linear parallel time generation.
result Shows competitive accuracy/speed tradeoff compared to existing methods on five machine translation datasets.
Proposes a secure communication method independent of eavesdropper's decoder.
problem Lack of practical security constraints in existing methods.
method Dual MINE-based neural secure communications model.
result Security performance is not affected by eavesdropper's decoding means.
This work prevents variational autoencoders from collapsing by adding an auxiliary decoder.
problem Variational autoencoders can collapse into autodecoders, losing semantic information.
method Adding an auxiliary decoder to regularize the latent space.
result Auxiliary decoders increase semantic information in the latent space and reconstructions.
DD-VAE uses deterministic decoding for better latent code utilization in discrete data.
problem Inflexible decoders in VAEs lead to poor utilization of latent codes in discrete data.
method Proposed DD-VAE with deterministic decoding and new proposal distributions.
result DD-VAE improves latent code utilization and structure of learned manifold.
New method aligns brain data across individuals for better brain decoding.
problem Inter-individual variability in brain response patterns limits decoder generalization.
method SpectralOT method that embeds cortical geometry into Laplace-Beltrami eigenmodes.
result SpectralOT strikes balance between aligning functional features and preserving anatomical structure.
A new method for effective VAE training using calibrated decoders.
problem Training VAEs requires hyperparameter tuning, leading to inefficiency.
method Calibrated decoders that learn uncertainty and automatically determine information retention.
result Calibrated decoders can simplify VAE training without heuristic modifications.
Dual-decoder model generates responses with targeted sentiment.
problem Generating human-like responses with specific sentiment.
method Simple dual-decoder model with two sentiment decoders connected to one encoder.
result Significant performance gain in sentiment accuracy and word diversity.
Decoding, ie prediction from brain images or signals, calls for empirical evaluation of its predictive power. Such evaluation is achieved via cross-validation, a method also used to tune decoders' hyper-parameters. This paper is a review on cross-validation procedures for decoding in neuroimaging. It includes a didacti…
New insights into how encoder-decoder networks generate attention matrices.
problem Understanding how encoder-decoder networks use attention matrices.
method Decomposing hidden states into temporal and input-driven components.
result Attention matrices are formed based on task requirements, not architecture type.
Deep neural networks decode natural visual scenes from neural spikes.
problem Decoding visual scenes from neural spikes for brain-machine interfaces.
method Developed a novel spike-image decoder (SID) using deep neural networks.
result SID reconstructs natural visual scenes from neural spikes with high accuracy.
Sparse superposition codes were recently introduced by Barron and Joseph for reliable communication over the AWGN channel at rates approaching the channel capacity. The codebook is defined in terms of a Gaussian design matrix, and codewords are sparse linear combinations of columns of the matrix. In this paper, we prop…
New insights into encoder-decoder structures using information measures.
problem Understanding the role of encoder-decoder design in machine learning.
method Using information sufficiency and mutual information loss concepts.
result Characterizes the expressiveness loss in encoder-decoder designs.
Deep learning improves error detection in intracranial EEG.
problem Improving error detection in intracranial EEG.
method Employed convolutional neural networks (CNNs) for classification and characterization of error-related brain responses.
result CNNs outperformed traditional methods in classifying and decoding errors in intracranial EEG.
Entropy-based decoding improves DLM sampling efficiency.
problem Decoding strategy challenges in flexible DLMs.
method Entropy sum-based confidence-based decoding.
result Entropy sum-based decoding achieves ε \varepsilon ε -accuracy with O ~ ( H ( X 0 ) / ε ) \widetilde O(H(X_0)/\varepsilon) O ( H ( X 0 ) / ε ) iterations.