New insights into tSNE for large datasets.
problem Limitations of tSNE in handling large datasets.
method Identified continuum limit of tSNE objective function, proposed rescaled model.
result Rescaled model has a consistent limit for large datasets.
The MBO scheme for data clustering is analyzed in the large data limit, proving convergence to optimal partition problems.
problem Analyzing the MBO scheme for data clustering in the large data limit.
method Implicit gradient descent on the thresholding energy of a similarity graph.
result The MBO scheme outcomes converge to minimizers of a weighted optimal partition problem.
Study of large mass limits of G2 and Calabi-Yau monopoles on specific manifolds.
problem Understanding the behavior of monopoles in the large mass limit on G2 and Calabi-Yau manifolds.
method Developed a structure theory for the limit of S U ( 2 ) SU(2) S U ( 2 ) G 2 G_2 G 2 -monopoles and Calabi-Yau monopoles, extracting singular abelian G2-monopoles with Dirac singularities. result Proved an energy identity for monopole bubbles in the large mass limit.
Bayesian Neural Networks are robust to gradient-based attacks in the large-data limit.
problem Vulnerability of deep learning models to adversarial attacks.
method Analysis of adversarial attacks in the large-data, overparametrized limit for Bayesian Neural Networks.
result BNN posteriors are robust to gradient-based adversarial attacks in the limit.
This paper evaluates LLMs on large graph property estimation tasks.
problem Limited context length of LLMs limits their evaluation on large graphs.
method Developed EstGraph dataset and introduced four tasks for LLMs to estimate large graph properties.
result LLMs perform better on graph property estimation tasks when provided with context-rich prompts based on random walks.
Theory explains neural network scaling with dataset and model size.
problem Neural network scaling laws with dataset and model size.
method Identified variance-limited and resolution-limited scaling behaviors.
result Four scaling regimes explained: infinite data, infinite width, resolution-limited, and large width.
Large corporate credit models may be adapted for small business risk assessment.
problem Limited data and lack of credit analysts for small businesses.
method Adapting large corporate credit risk models for small businesses.
result Adapted models can predict small business credit risk effectively.
Derives ideal train/test split for ridge regression in large data limit.
problem Finding optimal train/test split for ridge regression in large data scenarios.
method Mathematical derivation of optimal train/test split, considering ridge tuning parameter and asymptotic behavior.
result The optimal train/test split for ridge regression in the large data limit depends weakly on the ridge tuning parameter alpha.
Survey on DNNs for speech processing, focusing on limited data challenges.
problem Challenges in training DNNs for speech tasks with limited data.
method Overview of techniques for few-shot speech processing.
result Promising few-shot techniques for speech processing.
Scalings in which the graph Laplacian approaches a differential operator in the large graph limit are used to develop understanding of a number of algorithms for semi-supervised learning; in particular the extension, to this graph setting, of the probit algorithm, level set and kriging methods, are studied. Both optimi…
Graph Laplacians computed from weighted adjacency matrices are widely used to identify geometric structure in data, and clusters in particular; their spectral properties play a central role in a number of unsupervised and semi-supervised learning algorithms. When suitably scaled, graph Laplacians approach limiting cont…
Paper proposes Universum GANs to improve GANs with limited labeled data.
problem Limited labeled data makes supervised learning challenging.
method Proposes Universum GANs with evolving discriminator loss.
result Improved discriminator accuracy and high quality data generation.
We study the dynamics of the limit order book of liquid stocks after experiencing large intra-day price changes. In the data we find large variations in several microscopical measures, e.g., the volatility the bid-ask spread, the bid-ask imbalance, the number of queuing limit orders, the activity (number and volume) of…
Study shows algorithms benefit from limited target data with many source domains.
problem Adapting to new domains with scarce labeled target data.
method New family of model selection algorithms.
result Beneficial guarantees in scenarios with limited target data.
While deep learning has been incredibly successful in modeling tasks with large, carefully curated labeled datasets, its application to problems with limited labeled data remains a challenge. The aim of the present work is to improve the label efficiency of large neural networks operating on audio data through a combin…
The large-N limit of Segal-Bargmann transform on spheres is studied.
problem Understanding the behavior of Segal-Bargmann transform on spheres as dimension increases.
method Analyzing the large-N limit of the transform on S N − 1 ( N ) S^{N-1}(\sqrt N) S N − 1 ( N ) , describing geometric models, and showing the transform remains unitary. result The limiting transform is still a unitary map from the limiting domain onto the limiting range.
Framework uses synthetic data from pretrained models to improve predictive modeling.
problem Limited effectiveness of synthetic data from generative models for improving predictive performance.
method Proposes an end-to-end framework that generates and filters synthetic data through domain-specific statistical methods.
result Consistent improvements in predictive performance across various settings.
GNTK reveals convergence of GNNs on large graphs.
problem Understanding and optimizing GNNs on large graphs.
method Graph Neural Tangent Kernels (GNTK) and graphons.
result GNTKs converge to graphon NTKs on large graphs, enabling task inference.
This article provides an original understanding of the behavior of a class of graph-oriented semi-supervised learning algorithms in the limit of large and numerous data. It is demonstrated that the intuition at the root of these methods collapses in this limit and that, as a result, most of them become inconsistent. Co…
Given a data set and a subset of labels the problem of semi-supervised learning on point clouds is to extend the labels to the entire data set. In this paper we extend the labels by minimising the constrained discrete p p p -Dirichlet energy. Under suitable conditions the discrete problem can be connected, in the large da…
Bayesian data sketching speeds up inference for large functional data.
problem Slow posterior computations in Bayesian varying coefficient models for large data.
method Compress functional response and predictor matrix using random linear transformation.
result Fully model-based Bayesian inference on compressed data.
New framework analyzes SGD dynamics in large samples and dimensions.
problem Analyzing stochastic gradient descent in large-scale settings.
method Inspired by random matrix theory, new framework for fixed stepsize and finite sum settings.
result SGD dynamics become deterministic in the large sample and dimensional limit, governed by a Volterra integral equation.
The covariance matrix is formulated in the framework of a linear multivariate ARCH process with long memory, where the natural cross product structure of the covariance is generalized by adding two linear terms with their respective parameter. The residuals of the linear ARCH process are computed using historical data …
New method adds interactions to interpretable models for large-scale data.
problem Limited model complexity and lack of interactions in interpretable models.
method Factorization method to derive scalable higher-order tensor product spline models.
result Incorporates all higher-order interactions of non-linear feature effects without computational penalties.
New method recovers causal networks from short time-series data.
problem Inferring causal relationships from short time-series data in complex systems.
method Large-scale Nonlinear Granger Causality (lsNGC) approach.
result Captures meaningful interactions from limited observational data.
FaStR improves scalability for time-aware RS with varying coefficients.
problem Limited applicability of structured regression models to large-scale data with categorical effects and many interactions.
method Combines structured additive regression and factorization approaches in a neural network-based model implementation.
result FaStR scales better and performs competitively with other time-aware RS in prediction performance.
We study optimal estimation for sparse principal component analysis when the number of non-zero elements is small but on the same order as the dimension of the data. We employ approximate message passing (AMP) algorithm and its state evolution to analyze what is the information theoretically minimal mean-squared error …
BLoB fine-tunes LLMs with Bayesian methods to improve uncertainty estimation.
problem Overconfidence in LLMs during inference for domain-specific tasks.
method Continuous Bayesian low-rank adaptation during fine-tuning.
result BLoB improves generalization and uncertainty estimation in LLMs.
Dirichlet process mixture (DPM) models tend to produce many small clusters regardless of whether they are needed to accurately characterize the data - this is particularly true for large data sets. However, interpretability, parsimony, data storage and communication costs all are hampered by having overly many clusters…
Large learning rates work surprisingly well in standard parameterization, contrary to theory.
problem Theoretical limits of large learning rates do not match practical network behavior.
method Fine-grained analysis of learning rates and network behavior under cross-entropy loss.
result There are two distinct sub-regimes of unstable learning rates, with a controlled divergence regime where features continue to evolve.
FNO model predicts GCS pressure fields with 81% less data, even with limited high-fidelity data.
problem Accurate prediction of complex physical behaviors in large-scale 3D geological carbon storage problems with limited data.
method Multi-fidelity Fourier Neural Operator (FNO) for efficient training with multi-fidelity datasets.
result Multi-fidelity FNO model predicts pressure fields with reasonable accuracy even with limited high-fidelity data.
Infinitesimal boosting converges to a deterministic process in large sample limit.
problem Characterizing the asymptotic behavior of infinitesimal gradient boosting in large sample sizes.
method Proving convergence to a deterministic process using large sample theory and differential equations.
result The test error decreases over time in the population limit.
New optimal prior avoids bias in complex models with limited data.
problem Bias in inference from limited data using Jeffreys prior.
method Developed a principled choice of measure that avoids bias, dependent on data quantity.
result Optimal prior leads to unbiased inference in complex models.
We determine the critical batch size for large language models and find it scales with data size, not model size.
problem Determining the optimal batch size for large-scale model training.
method We propose a measure of critical batch size, pre-trained models, and systematic hyper-parameter sweeps.
result The critical batch size scales primarily with data size, not model size.
Simulates realistic execution and costs in limit order books.
problem Realistic simulation of limit order books for large-tick assets.
method Tractable representation of spread and volume imbalance; calibrated event timing; feedback mechanism for market impact.
result Simulator yields realistic behavior and sensitivity to execution parameters.
New research limits how well attackers can guess if data points were in a model's training set.
problem Revealing membership of data points in machine learning models.
method Theoretical analysis of statistical limits for membership inference attacks.
result The effectiveness of membership inference attacks is limited by a constant that quantifies data distribution diversity.
Analyzes SGD dynamics in two-layer networks, bridging different regimes.
problem Understanding SGD dynamics in high-dimensional and mean-field settings.
method Rigorous analysis via deterministic low-dimensional description of sufficient statistics.
result Infinite-width dynamics remains close to a low-dimensional subspace.
The paper analyzes Laplace learning for Gaussian measure data in infinite dimensions, proving convergence.
problem Analyzing Laplace learning for infinite-dimensional Gaussian measure data.
method Minimizes Dirichlet energy on a graph constructed from the full dataset.
result Proves pointwise convergence of the graph Dirichlet energy for Gaussian measure data.
LLMs help less-resourced researchers access costly data.
problem Unequal access to costly datasets limits research contributions.
method RAG framework with GPT-4o-mini for automated data collection.
result LLMs can collect CEO pay ratios and CAMs from corporate disclosures with high accuracy and low cost.
The paper identifies five extreme learning regimes for large linear autoencoders.
problem Understanding the learning dynamics of large weight-tied linear autoencoders.
method Formal loss-expansion hierarchy and analysis of gradient flow.
result Five extreme regimes associated with faces of a triangular prism.
MixKD improves large-scale language model compression and generalization.
problem Inefficient and resource-intensive large-scale language models.
method MixKD uses mixup data augmentation to enhance student model's generalization ability.
result MixKD leads to significant performance gains over standard KD and competitive baselines.
Study of limits of Einstein-Bogomol'nyi metrics on P^1 in two regimes.
problem Understanding limits of Einstein-Bogomol'nyi metrics on P^1.
method Analysis of two regimes: dissolving limit and large volume limit.
result Recovery of Einstein-Bogomol'nyi metrics on C with total string number N' for each N'.
Constructs foliations of critical surfaces for Hawking energy in asymptotically flat initial data sets.
problem Positivity and rigidity of Hawking quasi-local energy in asymptotically flat spacetimes.
method Lyapunov-Schmidt reduction within a Willmore-foliation framework.
result Existence and uniqueness of foliations by Hawking surfaces, positivity and large-sphere limit of Hawking energy.
The paper analyzes Bayesian neural networks trained with VI, proving a law of large numbers for different schemes.
problem Training Bayesian neural networks with variational inference.
method Analyzes three training schemes: exact estimation, Bayes by Backprop, and Minimal VI.
result All training schemes converge to the same mean-field limit.
The study connects K-stability and large complex structure limits in mirror symmetry.
problem Understanding K-stability and its relation to large complex structure limits in mirror symmetry.
method Analyzing Kähler test configurations and their mirror Landau-Ginzburg models, studying scaling behavior, and focusing on specific limiting cases.
result New formulae for the Donaldson-Futaki invariant are derived in terms of theta functions on the mirror in certain limiting cases.
This paper improves spectral clustering for large datasets using the Nystrom method.
problem Spectral clustering's scalability issues with large datasets.
method A principled spectral clustering algorithm exploiting Nystrom approximation's spectral properties.
result Improved spectral clustering efficiency and accuracy compared to existing methods.
Memory-efficient learning for large-scale imaging systems.
problem Memory limitations in GPUs for real-world large-scale inverse problems.
method Exploits reversibility of network layers to enable data-driven design.
result Demonstrated on small-scale and large-scale real-world systems.
Machine learning's predictive power is limited by sample size, as shown by the Limits-to-Learning Gap.
problem The limitations of machine learning in approximating true data-generating processes.
method Characterization of a universal lower bound (LLG) quantifying the discrepancy between empirical fit and population benchmark.
result Standard ML approaches can substantially understate true predictability in financial data.