We use learning curves to analyze deep networks and evaluate model design.
problem Evaluate design choices in deep networks.
method Propose a method to robustly estimate learning curves, abstract their parameters, and evaluate different parameterizations.
result Interesting observations on the effectiveness of different parameterizations.
Machine learning predicts arithmetic curve invariants with high accuracy.
problem Classifying arithmetic curves based on their invariants.
method Training machine learning algorithms on datasets of elliptic and genus 2 curves.
result High accuracy in classifying curves, including rank, torsion, and integral points.
New model predicts neural network performance from early training epochs, incorporating architecture impact.
problem Predicting neural network performance from early training epochs, neglecting architecture impact.
method Architecture-aware graph ordinary differential equation model.
result Model outperforms state-of-the-art methods for MLP and CNN learning curves.
Paper benchmarks machine learning for detecting process curve drifts.
problem Detecting drifts in multivariate manufacturing process data.
method Synthetic data generation and evaluation score introduction.
result Existing algorithms often fail with complex drift scenarios.
A new metric-based principal curve method learns 1D manifolds from spatial data.
problem Learning 1D manifolds from spatial data.
method Metric-based Principal Curve (MPC) approach.
result The method effectively learns the shape of 1D manifolds from synthetic and real datasets.
The paper explores using machine learning for yield curve calibration in multiple markets.
problem Calibration challenges in multiple yield curve markets.
method Gaussian process regression and Adam optimizer.
result Good results for single curve markets, but many challenges for multi curve markets.
This work studies learning curves for revenue maximization algorithms.
problem Understanding the performance of revenue-maximizing algorithms as they learn from more data.
method Initiates the study of learning curves for revenue maximization, providing a near-complete characterization of their rate of decay.
result Learning curves for revenue maximization can decay arbitrarily slowly or almost exponentially fast, depending on the distribution and optimal revenue.
Deep learning approximates geometric measures of planar curves.
problem Approximating differential invariants of planar curves.
method Utilizing deep neural networks to estimate geometric measures of planar curves.
result Deep neural networks can learn to overcome instabilities and sampling artifacts.
Learning curves show more data doesn't always improve performance.
problem The surprising finding that more data doesn't always lead to better generalization performance.
method A survey of learning curves, focusing on those that indicate more data doesn't necessarily improve performance.
result Learning curves can show that more data doesn't always lead to better generalization performance.
New method saves computational budget by ranking and transferring learning curves.
problem Expensive automated machine learning methods for hyperparameter and neural architecture optimization.
method Tackles as a ranking and transfer learning problem, optimizing a pairwise ranking loss and leveraging learning curves from other datasets.
result Accelerates neural architecture search by a factor of up to 100 without significant performance degradation.
Efficiently models learning curves using Gaussian processes with latent Kronecker structure.
problem Joint modeling of machine learning model performance across hyper-parameters and training progress.
method Imposes latent Kronecker structure to leverage efficient product kernels and handle missing values.
result Matches the performance of a Transformer on a learning curve prediction task.
Deep learning models forecast multiple yield curves with improved accuracy.
problem Globalization of financial markets affects yield curves.
method Combines self-attention mechanism and nonparametric quantile regression.
result Effective point and interval forecasts of future yields.
Learn2Evaluate uses learning curves to estimate high-dimensional prediction performance.
problem Estimating test performance in high-dimensional data settings is challenging.
method Learn2Evaluate uses learning curves to estimate test performance at the total sample size.
result Learn2Evaluate provides a lower confidence bound for performance estimation.
Framework for transferring discount curve estimates across fixed-income product classes.
problem Challenges in estimating discount curves from sparse or noisy data.
method Proposes a vector-valued kernel ridge regression (KR) framework with economic regularization.
result Transfer learning tightens confidence intervals and improves extrapolation performance.
Paper introduces k-DTW for robust curve comparison.
problem Robust dissimilarity measure for polygonal curves.
method Introduces k-Dynamic Time Warping (k-DTW) as a novel dissimilarity measure.
result k-DTW is more robust to outliers and has stronger metric properties than DTW.
ResNets learn the geodesic curve in Wasserstein space.
problem Characterize the dynamics of deep residual networks during training.
method Modeling ResNet dynamics using continuity equations and optimal transport.
result ResNets learn the geodesic curve in the Wasserstein space.
A meta-learning approach for efficient algorithm selection in budget-limited scenarios.
problem Efficiently selecting the best-performing machine learning algorithm with limited computational resources.
method A Markov Decision Process framework where an agent decides whether to train, wake up, or start new algorithms based on partial learning curves.
result Meta-learning from learning curves improves algorithm selection, especially when learning curves do not intersect frequently.
Machine learning predicts Shafarevich-Tate group orders of elliptic curves.
problem Predicting the order of the Shafarevich-Tate group of elliptic curves.
method Train feed-forward neural network and regression models on elliptic curve invariants.
result Models achieve high accuracy (>0.9) and predict orders not seen during training. Machine learning classifies complex geometric patterns with high accuracy.
problem Classifying extension degree of dessins d'enfants over the rationals.
method Deep feed-forward neural network trained on machine learning.
result 0.92 accuracy in classification with 0.03 standard error.
LC-PFN predicts learning curve performance more accurately and faster than MCMC.
problem Bayesian extrapolation of learning curves is computationally expensive and overly restrictive.
method Prior-Data Fitted Neural Networks (PFNs) for approximate Bayesian inference.
result LC-PFN outperforms MCMC in accuracy and is significantly faster.
The paper shows how the generalization curve can have multiple peaks, influenced by data and learning algorithm biases.
problem Understanding the generalization behavior of linear regression models under varying parameterizations.
method Analyzes generalization loss in linear regression models with varying parameterizations, both under- and over-parameterized.
result The generalization curve can have an arbitrary number of peaks, and their locations can be controlled.
Accelerates pulsar light curve inference with learned representations and optimization.
problem Computational expense of Markov chain Monte Carlo methods for posterior inference.
method Combining U-Net latent representations with local simulator-guided optimization.
result 120x reduction in inference time (24 hours to 12 minutes) with accuracy preserved.
Machine learning accurately distinguishes Sato-Tate groups for hyperelliptic curves.
problem Arithmetic of hyperelliptic curves and Sato-Tate conjecture.
method Bayesian classifier and machine learning techniques applied to L-functions of hyperelliptic curves.
result Machine learning can distinguish Sato-Tate groups with high accuracy and speed.
The paper clusters PK curves using ML, finding it useful for identifying similar patterns.
problem Improving drug development and patient outcomes through ML in pharmacogenomics.
method Unsupervised clustering of PK curves using various dissimilarity measures.
result Euclidean distance is most suitable for clustering PK curves, and clustering can validate pharmacogenomic results.
When confronted with massive data streams, summarizing data with dimension reduction methods such as PCA raises theoretical and algorithmic pitfalls. Principal curves act as a nonlinear generalization of PCA and the present paper proposes a novel algorithm to automatically and sequentially learn principal curves from d…
We propose probabilistic models that can extrapolate learning curves of iterative machine learning algorithms, such as stochastic gradient descent for training deep networks, based on training data with variable-length learning curves. We study instantiations of this framework based on random forests and Bayesian recur…
Study reveals learning curves and benign overfitting in spectral algorithms for large dimensions.
problem Understanding learning curves and benign overfitting in spectral algorithms for large-dimensional data.
method Analysis of learning curves and benign overfitting in spectral algorithms for inner-product kernels on the sphere and general domains.
result Characterization of three distinct regimes: over-regularized, under-regularized, and interpolation regimes, revealing benign overfitting across both under-regularized and interpolation regimes.
We introduce a new deep-learning based algorithm to evaluate options in affine rough stochastic volatility models. Viewing the pricing function as the solution to a curve-dependent PDE (CPDE), depending on forward curves rather than the whole path of the process, for which we develop a numerical scheme based on deep le…
New combinatorial dimension VCL refines learning curve theory.
problem Explaining the behavior of learning curves for specific distributions.
method Introducing combinatorial dimension VCL to characterize strong minimax lower bounds.
result Learning rate can be decomposed into linear and exponential components.
Transformer model removes noise from light curves efficiently.
problem Challenges in processing astrophysical light curves due to noise.
method Denoising Time Series Transformer (DTST) model trained with masked objective.
result DTST model excels at removing noise and outliers in time series datasets.
Exact risk and learning rate curves derived for adaptive SGD on high-dimensional problems.
problem Analyzing risk and learning rate dynamics in high-dimensional optimization problems.
method Developed a framework to give exact expressions for risk and learning rate curves using ODEs.
result Exact expressions for risk and learning rate curves, with detailed analysis of two adaptive learning rates.
The Receiver Operating Characteristic (ROC) curve is a representation of the statistical information discovered in binary classification problems and is a key concept in machine learning and data science. This paper studies the statistical properties of ROC curves and its implication on model selection. We analyze the …
Develops a simple model to understand learning curves for arbitrary power laws.
problem Lack of theoretical understanding of scaling laws in machine learning.
method Analyzes a toy model to determine if learning curves are universal or depend on data distribution.
result Determines that learning curves can exhibit n−β for arbitrary power β>0. A robust machine learning approach forecasts U.S. Treasury yields, reducing risk for investors.
problem Noisy and uncertain U.S. Treasury yields pose risk to forecast users.
method Formulates yield curve forecasting as a distributionally robust problem, combining factor models and machine learning.
result Robust forecast combinations improve out-of-sample performance across different maturity periods.
New optimization method improves AUC for binary classification and changepoint detection.
problem Non-convex AUC and sub-optimal points in ROC curves.
method AUM (Area Under Min(FP, FN)) surrogate loss function based on sorting and summing ROC curve points.
result AUM minimization learning algorithm improves AUC and speeds up compared to previous methods.
Estimating what would be an individual's potential response to varying levels of exposure to a treatment is of high practical relevance for several important fields, such as healthcare, economics and public policy. However, existing methods for learning to estimate counterfactual outcomes from observational data are ei…
Controlled interventions provide the most direct source of information for learning causal effects. In particular, a dose-response curve can be learned by varying the treatment level and observing the corresponding outcomes. However, interventions can be expensive and time-consuming. Observational data, where the treat…
Slice Tuner optimizes data acquisition for accurate and fair machine learning models.
problem Acquiring enough data for all slices of data to ensure accurate and fair models.
method Selective data acquisition using convex optimization and iterative learning curve updates.
result Significantly outperforms baselines in model accuracy and fairness.
JAXFit speeds up curve fitting on GPUs.
problem Nonlinear least squares curve fitting problems.
method Trust region method on GPU with automatic differentiation.
result Significantly faster than CPU and other GPU libraries.
Proposes FunNoL for better curve classification and reconstruction in multivariate functional data.
problem Linear methods fail to capture nonlinear structures in multivariate functional data.
method Functional nonlinear learning (FunNoL) method using nonlinear mapping.
result FunNoL outperforms FPCA in curve classification and reconstruction, especially in multivariate settings.
The paper analyzes learning curves for kernel ridge regression with dot-product kernels.
problem Understanding the learning curves for different scaling regimes of data and model.
method Precise formulas for mean test error, bias, and variance in the mo∞ with m/dr constant regime. result A peak in the learning curve at m≈dr/r! for any integer r. Unified theory for neural scaling laws in hierarchically compositional data.
problem Understanding neural scaling laws in hierarchically compositional data.
method Probabilistic context-free grammars and power-law distributed production rules.
result Unified learning curve behavior for classification and next-token prediction tasks.
Yield curve forecasting is an important problem in finance. In this work we explore the use of Gaussian Processes in conjunction with a dynamic modeling strategy, much like the Kalman Filter, to model the yield curve. Gaussian Processes have been successfully applied to model functional data in a variety of application…
We study the average case performance of multi-task Gaussian process (GP) regression as captured in the learning curve, i.e. the average Bayes error for a chosen task versus the total number of examples n for all tasks. For GP covariances that are the product of an input-dependent covariance function and a free-form …
Statistical approaches for Functional Data Analysis concern the paradigm for which the individuals are functions or curves rather than finite dimensional vectors. In this paper, we particularly focus on the modeling and the classification of functional data which are temporal curves presenting regime changes over time.…
Theory and method for reducing prediction variance in noisy feature-subsampled ridge ensembles.
problem Reduction of prediction variance in noisy data with feature bagging.
method Developed analytical learning curves for noisy ridge ensembles, introduced heterogeneous feature ensembling.
result Subsampling shifts the double-descent peak, leading to improved performance over a single linear predictor.
This paper introduces a novel monotone curve estimation framework based on convex duality.
problem Estimating smooth, continuous, and monotonic curves in data.
method Convex duality and optimal transport theories.
result Established statistical guarantees for monotone curve estimates.
Method learns symmetries in curves without augmentation.
problem Symmetries in datasets like rotations and scalings.
method Geometric learning using principal fiber bundles.
result 2-parameter family of canonical curve parameterizations.