Reply to Ogburn et al. on their critique of Wang and Blei's work.
problem Critique of Wang and Blei's work on the blessings of multiple causes.
method Discussion and clarification of Wang and Blei's claims and findings.
result Wang and Blei's premise is correct and there are no foundational errors.
Enhanced TSFMs improve time series forecasting accuracy and reliability.
problem Variance, bias, and uncertainty in TSFMs' predictions on real data.
method Statistical and ensemble techniques including bagging, stacking, residual modeling, and prediction intervals.
result Hybrid models consistently outperform standalone TSFMs across multiple horizons.
This paper provides a mathematical foundation for deep neural networks solving PDEs.
problem Mathematical foundation for deep neural networks solving high-dimensional PDEs.
method Decomposed generalization error into approximation and training errors; derived gradient flow in the wide network limit.
result Generalization error tends to zero as the number of neurons and training time tend to infinity.
Establishes a microstructural foundation for a rough log-normal volatility model.
problem Developing a robust model for financial volatility under microstructural effects.
method Introduced a sequence of order-driven financial market models with Poisson process arrivals and analyzed their convergence to a log-normal rough volatility model.
result Weak convergence of price-volatility process to a log-normal rough volatility model with established weak error rates.
Develops PromptShift-CRC for drift-aware conformal risk control in foundation models under prompt and domain shift.
problem Fixed calibration risk in foundation models due to prompt and domain shift.
method Embeds prompts and responses, measures drift, gives more weight to recent examples, and updates risk online.
result Develops method to control risk up to terms for distribution mismatch and weighted quantile uncertainty.
Study identifies and analyzes three types of errors in learning Fourier operators.
problem Statistical, discretization, and truncation errors in learning Fourier operators.
method Analysis of a Discrete Fourier Transform (DFT) based least squares estimator.
result Established upper and lower bounds on statistical, discretization, and truncation errors.
VMoER improves uncertainty quantification in MoE layers for scalable foundation models.
problem Uncertainty quantification in large-scale models like MoE layers.
method Structured Bayesian approach with amortized variational inference over routing logits and temperature parameter inference.
result Improves routing stability, reduces calibration error, and increases AUROC by 12%.
This paper provides theoretical foundations for using quantized actions in behavior cloning.
problem Applying autoregressive models to continuous control requires discretizing actions through quantization, which is poorly understood.
method The paper analyzes quantization error propagation and statistical sample complexity, and proposes model-based augmentation.
result Behavior cloning with quantized actions achieves optimal sample complexity, matching existing lower bounds.
Paper presents a copula-based method to efficiently generate correlated sample paths from multi-step time series models.
problem Generating realistic correlation structures in multi-step forecast sample paths is expensive and time-consuming.
method Copula-based approach to generate correlated sample paths in one forward pass.
result Improved sample path quality and significant speedup over autoregressive sampling.
Unified framework HASSLE-free decomposes large model weights into sparse and low-rank components.
problem Efficiently compress large foundation models to reduce inference costs.
method Designs a unified framework for sparse plus low-rank matrix decomposition with a local layer-wise reconstruction error objective.
result HASSLE-free framework significantly outperforms state-of-the-art methods in compression and evaluation benchmarks.
The paper explores neural scaling laws for deep operator networks, offering a theoretical foundation.
problem Understanding neural scaling laws in deep operator networks.
method Theoretical analysis of approximation and generalization errors.
result Established a theoretical framework to quantify neural scaling laws for deep operator networks.
Cold-start PV forecasting uses synthetic histories to train time-series foundation models.
problem Cold-start PV forecasting
method Zero-shot pipeline with synthetic histories
result TabPFN-TS achieves the lowest error under Real Feedback strategy
Chronos models improve financial forecasting by integrating multivariate data.
problem Improving financial forecasting accuracy using multivariate data.
method Evaluation of Chronos-2 on multivariate and univariate financial forecasting models.
result Multivariate forecasts consistently outperform univariate forecasts, especially for interest rates.
Develops probabilistic safety regions for scalable classifiers.
problem Minimizing misclassification errors in supervised classification.
method Introduces probabilistic safety regions and scalable classifiers.
result Probabilistic certifications for classifier performance.
New framework assesses extreme errors in machine learning models.
problem Current validation methods fail to quantify extreme errors in high-stakes domains.
method Uses Extreme Value Theory (EVT) to estimate worst-case failures.
result Establishes EVT as a fundamental tool for assessing model reliability.
The paper analyzes error propagation in dynamic programming for stochastic control and option pricing.
problem Error propagation in dynamic programming for stochastic control and option pricing.
method Formulated a general dynamic programming framework, used RKHSs for nonparametric regression, and Monte Carlo subsampling for estimating continuation value.
result Proposed a rigorous error decomposition and control mechanism for error propagation in dynamic programming.
Artificial Intelligence (AI) systems sometimes make errors and will make errors in the future, from time to time. These errors are usually unexpected, and can lead to dramatic consequences. Intensive development of AI and its practical applications makes the problem of errors more important. Total re-engineering of the…
LPF provides formal guarantees for aggregating multi-evidence in probabilistic tasks.
problem Lack of formal guarantees for multi-evidence reasoning in AI.
method LPF uses variational autoencoders and Sum-Product Networks to aggregate evidence items.
result Proves multiple formal guarantees including calibration preservation and error decay.
Researchers improve spectrum reconstruction formula with proof.
problem Improving accuracy of spectrum estimation in PCA.
method Analytical derivation of approximation formula for PCA.
result Order of error for approximation formula is c-dependent. Simplifies VAE for anomaly detection using rate-distortion theory.
problem Anomaly detection in unsupervised learning systems.
method Revisit VAE from information theory, incorporate model uncertainty.
result Competitive performance on benchmark datasets.
We identify a trade-off between robustness and accuracy that serves as a guiding principle in the design of defenses against adversarial examples. Although this problem has been widely studied empirically, much remains unknown concerning the theory underlying this trade-off. In this work, we decompose the prediction er…
Derives error bounds for stochastic iterative algorithms using Stein's method.
problem Bounding errors in stochastic iterative algorithms like SGD and SGLD.
method Uses infinite-dimensional Stein's method of exchangeable pairs to derive functional approximation error bounds.
result Establishes non-asymptotic error bounds for algorithm sample paths and variance of iterate averages.
Group-invariant neural networks improve approximation accuracy for symmetric functions.
problem Improving approximation accuracy for symmetric functions using neural networks.
method Investigates the generalization error of group-invariant neural networks within the Barron framework.
result Group invariance introduces a factor δ that can significantly improve approximation accuracy when it is small.
Quantum codes on hyperbolic lattices outperform Euclidean ones with higher rates and lower overhead.
problem Improving quantum error correction performance with hyperbolic lattices.
method Unified framework using Hyperbolic Cycle Basis algorithm for CSS codes construction and benchmarking.
result Achieved higher encoding rates and lower qubit overhead in hyperbolic quantum error correction codes.
Anisotropic data structure affects learning dynamics and generalization error in linear networks.
problem Understanding the impact of data anisotropy on learning dynamics and generalization error in linear networks.
method Examined a spiked covariance structure as a model of anisotropy in a two-layer linear network in a linear regression setting.
result Learning dynamics proceed in two phases: initially driven by input-output correlation, then by other principal directions of the data structure. Derived an analytical expression for the generalization error.
We analyze the Hessian spectra of large models up to 100B parameters.
problem Accurate Hessian spectra of large foundation models are difficult to obtain.
method We use shard-local finite-difference Hessian vector products and stochastic Lanczos quadrature.
result We produce the first large-scale spectral density estimates of foundation models.
DM approximates submanifolds with error bounds.
problem Understanding the accuracy of Diffusion Maps in embedding submanifolds.
method Deriving geometric properties and deriving bounds on embedding errors.
result Error bounds for DM embeddings and tangent spaces.
Defines Learning Analytics' foundational structure and scope.
problem Lack of theoretical foundation in Learning Analytics.
method Proposes an axiomatic theory based on psychological learning and LA methodology.
result Clarifies the epistemological stance of Learning Analytics and its limitations.
Survey of financial foundation models for diverse applications.
problem Challenges in applying general-purpose FMs to financial tasks.
method Review of financial foundation models (FFMs) in three modalities.
result Emergence of FFMs designed specifically for finance.
Paper presents a method to reduce prediction variance of DNNs for unknown systems.
problem Uncertainty in DNN predictions due to high variance.
method Ensemble averaging of multiple DNN models trained independently.
result Reduction in variance of DNN predictions, improving reliability.
New insights into double descent phenomenon in neural networks.
problem Understanding the double descent behavior in deep learning models.
method Linear teacher-student setup and tools from statistical physics.
result Distinct features are learned at different scales, leading to epoch-wise double descent.
A model learns causal graphs from summary statistics of synthetic data.
problem Causal discovery algorithms are brittle with large sets of variables and limited data.
method A supervised model trained on synthetic data predicts causal graphs from summary statistics.
result The model generalizes well beyond its training set and runs on large graphs.
This paper uses SDEs to analyze GANs training and long-run behavior.
problem Understanding the training process and long-run behavior of GANs.
method Established SDE approximations for GANs training and analyzed long-run behavior via invariant measures.
result The long-run behavior of GANs training can be studied via the invariant measures of its SDE approximations.
Introduces foundation priors for using model-generated data in empirical research.
problem Using model-generated data as real observations in empirical research.
method Introduces foundation priors as an exponential-tilted, generalized Bayesian update of the user's primitive prior.
result Synthetic data reflects both model patterns and user's priors, enabling principled use in empirical work.
Revises Bayesian model averaging for foundation models.
problem Ensemble pre-trained and lightly-finetuned foundation models for improved classification performance.
method Introduces trainable linear classifiers and computationally cheaper model averaging scheme (OMA).
result Ensembled models can better predict on various datasets.
This paper establishes a theoretical foundation for super-models via domain adaptation.
problem Reducing computational and data costs in AI for small and medium-sized enterprises.
method Two-stage diffusion process modeling, including pre-training and fine-tuning stages, using the Uhlenbeck-Ornstein process.
result The generalization error of the fine-tuning stage is dominant in domain adaptation.
Time series foundation models are well-calibrated, improving over baseline models.
problem Calibration of time series foundation models for practical applications.
method Systematic evaluations of five time series foundation models and two baselines, assessing calibration, prediction heads, and long-term forecasting.
result Time series foundation models are consistently better calibrated than baseline models and do not show over- or under-confidence.
Combines foundation models with weak supervision to improve NLP and video tasks.
problem Leveraging weak supervision with foundation models without labeled data.
method Liger, a combination of foundation model embeddings and weak supervision techniques.
result Liger outperforms existing weak supervision methods by 14.1 points on benchmark NLP and video tasks.
Foundation models trained on collider data improve jet generation tasks.
problem Improving foundation models for jet generation tasks.
method Pre-training OmniJet-α model on AspenOpenJets dataset. result Pre-trained model improves performance on jet generation tasks with domain shift.
The study reveals a linear relationship between source and target domain classification errors based on disagreement.
problem Evaluating model performance under distribution shift with limited labeled data.
method Developed a theoretical foundation for analyzing disagreement in high-dimensional random features regression.
result The disagreement-on-the-line phenomenon occurs when classification error under the source domain is a linear function of the target domain.
New measure assesses time series pre-training data quality without labels.
problem Challenges in collecting diverse pre-training datasets for time series classification.
method Contrastive-learning-based foundation model and contrastive accuracy measure.
result Contrastive accuracy correlates with model performance on downstream tasks.
EControl improves fast distributed optimization with compression and error control.
problem Stable convergence issues in distributed training with compression.
method Proposes EControl to regulate error compensation and prove fast convergence.
result Proves fast convergence for EControl in various convex settings without additional assumptions.
SGNs use Hamiltonian mechanics for invertible deep generative modeling.
problem Efficient and exact likelihood evaluation for deep generative models.
method Symplectic structure in latent space, Hamiltonian dynamics for data generation.
result Exact likelihood evaluation without Jacobian calculations.
Paper explores foundation models for dynamical systems using synthetic data.
problem Lack of synthetic data for dynamical systems training.
method Pretrained transformer model on synthetic dynamics functions sampled from RKHS.
result Pretrained model generalizes across various dynamical systems in simulations and hardware.
GAMformer bridges tabular models and interpretability, offering a single-pass approach.
problem Lack of interpretability in tabular foundation models like TabPFN.
method In-context learning for GAM shape functions, training on synthetic data.
result GAMformer performs comparably to other leading GAMs across various classification benchmarks.
TradeFM learns market microstructure from trade events, improving financial model accuracy.
problem Lack of generalizable models for market microstructure.
method Generative Transformer model trained on billions of trade events, using scale-invariant features and universal tokenization.
result TradeFM generates rollouts that match key stylized facts of financial returns and outperforms existing models.
Paper speeds up large foundation models for time series data.
problem Resource-intensive foundation models limit accessibility.
method Dimensionality reduction techniques, including PCA and neural network adapters.
result Up to 10x speedup and 4.5x more datasets fit on a single GPU.
A permutation-based SW test achieves minimax-optimal power for two-sample testing.
problem Nonparametric two-sample testing using the sliced Wasserstein distance.
method Proposes a permutation-based SW test and analyzes its performance.
result Achieves minimax separation rate n−1/2 over multinomial and bounded-support alternatives.