Paper proposes an efficient online learning method using an offline dataset for infinite horizon MDPs.
problem Efficient online reinforcement learning in infinite horizon MDPs with an unknown expert policy.
method Bayesian approach to model the expert's policy and minimize cumulative regret.
result Upper bound on regret of i l d e O ( T ) ilde{O}(\sqrt{T}) i l d e O ( T ) for the Informed PSRL algorithm. Theory explains neural network scaling with dataset and model size.
problem Neural network scaling laws with dataset and model size.
method Identified variance-limited and resolution-limited scaling behaviors.
result Four scaling regimes explained: infinite data, infinite width, resolution-limited, and large width.
Paper introduces variance reduction for infinite datasets with finite-sum structure.
problem Optimizing composite and strongly convex objectives with stochastic perturbations.
method Variance reduction approach for stochastic optimization with composite and strongly convex objectives.
result Convergence rate outperforms SGD with a smaller constant factor.
Infinite-dimensional SBDMs improve image generation across multiple resolutions.
problem Efficient image generation at high resolutions and across different levels.
method Developed SBDMs in infinite-dimensional setting, using trace class operators and operator networks.
result Improved efficiency and generalization across resolution levels.
This work improves understanding of neural network reconstruction attacks and distillation.
problem Understanding and mitigating reconstruction attacks on neural networks.
method Developed a stronger dataset reconstruction attack and studied its characteristics.
result Reconstruction attacks can recover entire training sets in the infinite width regime.
UDN adapts depth to data complexity, outperforming standard neural networks.
problem Adapting neural network depth to data complexity.
method Variational inference for infinitely deep neural networks with a novel algorithm.
result UDN outperforms standard neural networks and other infinite-depth approaches.
The paper analyzes Laplace learning for Gaussian measure data in infinite dimensions, proving convergence.
problem Analyzing Laplace learning for infinite-dimensional Gaussian measure data.
method Minimizes Dirichlet energy on a graph constructed from the full dataset.
result Proves pointwise convergence of the graph Dirichlet energy for Gaussian measure data.
HIRM models noisy, sparse, heterogeneous relational data using hierarchical clustering and Dirichlet processes.
problem Modeling noisy, sparse, and heterogeneous relational data.
method Hierarchical Chinese restaurant process and Dirichlet process mixture for clustering and modeling relation values.
result HIRM generalizes standard models and discovers relational structure in real-world datasets.
Study evaluates initialization strategies for infinite hidden Markov models.
problem Limited attention to initialization in infinite hidden Markov models.
method Systematically evaluated distance-based clustering, model-based, and uniform initializations.
result Distance-based clustering initializations consistently outperform other methods.
Infinite-horizon Gaussian processes reduce computational complexity for long datasets.
problem Cubic computational cost in state dimensionality for Gaussian processes.
method Single-sweep EP inference scheme for GPs with general likelihoods, reducing cost to O(m^2) per data point.
result Reduced computational complexity from cubic to quadratic in state dimensionality.
InfiniteBoost builds an infinite ensemble using gradient descent.
problem Building high accuracy ensembles for various machine learning tasks.
method Combines gradient boosting and random forests properties using gradient descent to create an infinite ensemble.
result InfiniteBoost achieves high accuracy on regression, classification, and ranking tasks.
The paper discusses the importance of infinite-mean models in finance and risk management.
problem Classic statistical models assume finite mean or variance, which is not suitable for heavy-tailed data.
method Discussion and recent results on infinite-mean models in economics and finance.
result Classic statistical results for finite-mean models often fail or flip for infinite-mean models.
We present a nonparametric prior over reversible Markov chains. We use completely random measures, specifically gamma processes, to construct a countably infinite graph with weighted edges. By enforcing symmetry to make the edges undirected we define a prior over random walks on graphs that results in a reversible Mark…
Enhances clustering for functional data, robust to outliers.
problem Challenges of clustering infinite-dimensional functional data and outlier sensitivity.
method Extends OCLUST algorithm to handle functional data, trimming outliers.
result Strong performance in clustering and outlier identification on simulated and real-world datasets.
The paper studies a rebalanced dataset for imbalanced classification using Centered Random Forests.
problem Imbalanced classification where one class is underrepresented.
method Theoretical analysis of Centered Random Forests (CRF) with rebalanced datasets and debiasing techniques.
result Theoretical Central Limit Theorem (CLT) for the infinite CRF and debiased estimator IS-ICRF.
This research optimizes Andrews plots for better visual clarity in high-dimensional data.
problem Visualizing high-dimensional datasets with clarity and aesthetics.
method Developed a method to add spectral smoothing to Andrews plots to reduce visual clutter.
result Optimal spatial-spectral smoothing leads to more aesthetically pleasing and clutter-free visualizations.
New mechanism for pure differential privacy on functional summaries using Laplace-like process.
problem Challenges in achieving differential privacy for complex, structured functional summaries.
method Independent Component Laplace Process (ICLP) mechanism for infinite-dimensional Hilbert space.
result Effective enhancement of utility of private summaries through oversmoothing.
The study reveals a transition in neural network performance from infinite-width to variance-limited behavior as dataset size increases.
problem Understanding the transition from infinite-width to variance-limited behavior in neural networks.
method Empirical study of the transition from infinite-width to variance-limited behavior as a function of sample size and network width.
result The critical sample size \( P^* \) is approximately \( \sqrt{N} \) for polynomial regression with ReLU networks.
Paper evaluates deep generative models' ability to generalize geometric concepts.
problem Measuring deep generative models' ability to generalize across different geometric tasks.
method Raven's Progressive Matrices analogy, Infinite World dataset, Zero-Shot Intelligence Metric ZSI.
result Identifies bottlenecks and proposes optimization methods for few-shot and zero-shot learning.
Infinitely wide neural nets perform well on small datasets.
problem Performing well on small datasets with limited training samples.
method Using Neural Tangent Kernels (NTKs) for kernel regression.
result NTK SVM outperforms Random Forests and Convolutional NTK on CIFAR-10 with 10-640 training samples.
Infinitely wide GCNs perform as GPs for graph semi-supervised learning.
problem Graph-based semi-supervised classification with limited labeled data.
method Proposes GPGC model combining GCNs and GPs for semi-supervised learning.
result GPGC outperforms state-of-the-art methods on various datasets.
Extends metrics for SPD matrices to infinite dimensions.
problem Lack of generalized forms for Riemannian metrics.
method Unitized Hilbert-Schmidt operators and extended Mahalanobis norm.
result Improved performance in high-dimensional comparisons.
Adaptive kernels from neural networks improve model performance.
problem Improving neural network performance through adaptive kernels.
method Deriving adaptive kernels from infinite-width neural networks using feature learning and gradient flow training.
result Adaptive kernels achieve lower test loss compared to traditional kernels.
A new method for online multi-label stream classification.
problem Challenges in classifying continuous data streams with concept drift and delayed labels.
method Online unsupervised incremental method based on self-organizing maps.
result The method is highly competitive in both stationary and concept drift scenarios.
Bayesian linear networks reveal optimal depth and width trade-offs.
problem Understanding how depth, width, and dataset size affect model quality in linear networks.
method Zero noise Bayesian inference with Gaussian weight priors and mean squared error.
result Optimal predictions at infinite depth and maximized Bayesian model evidence at infinite depth.
Spectral methods learn flexible topic models with topic correlations.
problem Learning topic models with arbitrary topic correlations.
method Flexible topic model using Normalized Infinitely Divisible (NID) distributions, learned via spectral methods.
result Improved perplexity on real datasets compared to baseline.
The paper proposes a Gaussian mixture model for Hilbert-space-valued data.
problem Challenges in characterizing probability measures for infinite-dimensional random objects.
method Gaussian mixture framework based on kernel mean embeddings.
result The proposed algorithm yields a dense class of approximations in infinite-dimensional spaces.
Paper tackles deep learning confounding factors, learns unseen factors.
problem Learning from data with unknown and potentially infinite confounding factors.
method Combines deep generative models with Bayesian non-parametric factor models (Indian Buffet Process).
result Model can learn from data with unknown and potentially infinite confounding factors.
The paper analyzes the Rashomon ratio for infinite classifier families and shows its importance for choosing good classifiers.
problem Analyzing the Rashomon ratio for infinite classifier families.
method Quantifying the Rashomon ratio in two examples and providing guarantees for estimating it.
result A large Rashomon ratio guarantees choosing a classifier with good empirical accuracy will not significantly increase empirical loss.
KIP meta-learning compresses datasets significantly.
problem Training data size and quality issues in machine learning.
method Kernel Inducing Points (KIP) for dataset compression.
result Significant reduction in dataset size with similar model performance.
Infinite mixture models are commonly used for clustering. One can sample from the posterior of mixture assignments by Monte Carlo methods or find its maximum a posteriori solution by optimization. However, in some problems the posterior is diffuse and it is hard to interpret the sampled partitionings. In this paper, we…
Transformers with CoT don't enhance reasoning power across all tasks.
problem Does CoT enhance the reasoning power of transformers?
method Examined the memorization capabilities of fixed-precision transformers with and without CoT.
result Transformers with CoT cannot memorize all reasoning tasks, leading to a negative answer.
A novel training strategy speeds up learning of infinite RBM models.
problem Slow convergence of infinite RBM models due to dependency between hidden units.
method Randomly regrouping hidden units before each gradient descent step.
result Significant acceleration in learning and enhancement of generalization ability.
Optimal transport for functional data using Hilbert-Schmidt operators.
problem Optimal transport for distributions on function spaces with partially represented stochastic maps.
method Regularization technique to restrict transport maps to Hilbert-Schmidt operators, developing an efficient algorithm.
result Existence, uniqueness, and consistency of the Hilbert-Schmidt operator estimate for the transport map.
The study analyzes transfer learning in infinite-width neural networks, improving generalization on target tasks.
problem Improving generalization in neural networks when using pretraining on a source task.
method Developed a theory under gradient flow for infinitely wide networks, analyzing fine-tuning and joint pretraining.
result Summary statistics of randomly initialized networks after pretraining are adaptive kernels that depend on both source and target data.
The paper explains generalization in kernel regression and deep neural networks using spectral bias and task-model alignment.
problem Understanding generalization in machine learning models, especially deep neural networks.
method Analytical expression for generalization error derived from statistical mechanics, applied to various kernels and data distributions.
result Spectral bias and task-model alignment explain generalization in kernel regression and deep neural networks.
Empirical study compares wide neural networks to kernel methods, resolving open questions.
problem Understanding the relationship between wide neural networks and kernel methods.
method Large-scale empirical study using various neural network architectures and kernel methods.
result Wide neural networks outperform fully-connected finite-width networks in some cases, but underperform convolutional finite-width networks.
This work challenges the Neural Tangent Kernel's role in overparameterized neural networks, especially with large width and depth.
problem The Neural Tangent Kernel's behavior in overparameterized neural networks with large width and depth is unclear.
method Experimental and theoretical analysis of ReLU networks with large width and depth.
result The aggregate norm of hidden neuron deviations does not vanish in infinitely-wide ReLU networks, indicating non-trivial behavior.
RFAD uses random features to speed up dataset distillation.
problem Efficiently compress large datasets for reduced storage and computation.
method Random feature approximation of the Neural Network Gaussian Process kernel.
result At least 100-fold speedup over KIP with competitive accuracy.
The paper develops methods to estimate value functions in reinforcement learning for infinite horizon problems.
problem Constructing confidence intervals for value functions in infinite horizon reinforcement learning.
method Modeling the Q-function using series/sieve method and recursively updating the policy using SAVE method.
result The proposed CI achieves nominal coverage even when the optimal policy is not unique.
New framework combines simple machines into complex ones for better neural network performance.
problem Improving neural network performance with limited training data.
method Developed a framework using topology and functional analysis to combine simple machines into complex ones, and used kernel methods to find optimal architectures.
result Kernel-inspired networks can outperform classical neural networks when training data is small.
Infinite Tucker Decomposition (InfTucker) and random function prior models, as nonparametric Bayesian models on infinite exchangeable arrays, are more powerful models than widely-used multilinear factorization methods including Tucker and PARAFAC decomposition, (partly) due to their capability of modeling nonlinear rel…
The paper extends graph-based semi-supervised learning to infinite-dimensional Wasserstein space.
problem Graph-based semi-supervised learning in high-dimensional data.
method Laplace Learning in the Wasserstein space, proving variational convergence and characterizing the Laplace-Beltrami operator.
result Consistent classification performance in high-dimensional settings.
Infinite BART model selects number of trees and allows different functions for clusters.
problem Regression and classification analysis with automatic tree selection and cluster-specific functions.
method Incorporates an Indian Buffet process prior to select a subset of decision trees for each observation.
result Infinite BART model outperforms classic BART on simulated and real datasets.
New HDP-HMM model accurately segments Internet path delays.
problem Automating the statistical analysis of large-scale network delay measurements.
method Infinite Hidden Markov Model (HDP-HMM) for trace segmentation.
result The HDP-HMM model performs as well as human cognition in pattern recognition.
Metric evaluates symmetry-breaking in datasets, revealing severe biases.
problem Symmetry-breaking in datasets can hinder the performance of symmetry-aware methods.
method Developed a metric to quantify symmetry-breaking using a two-sample classifier test.
result Symmetry-breaking can prevent optimal performance of invariant methods, even when labels are invariant.
Infinite rank surface cluster algebras extend traditional concepts to surfaces with accumulation points.
problem Extending surface cluster algebras to infinite surfaces with accumulation points.
method Consider infinite mutation sequences and hyperbolic structures to define cluster variables as lambda lengths of arcs.
result Established transitivity of infinite mutation sequences on triangulations of infinite surfaces and provided expansion formulas for cluster variables.
New method synthesizes piano training data, improving transcription performance.
problem Lack of large piano datasets limits note onset transcription models.
method Synthesizes arbitrary training data, models piano dynamics, avoids disentanglement problem.
result Achieves good transcription performance on MAPS dataset and excellent generalization.