Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

1223 · Sep 202019922001200920172026
48 results for Occam's Razor

Geometric Occam's Razor shapes deep learning solutions.

problem Understanding the regularization in over-parameterized neural networks.
method Analyzing the geometric model complexity and Dirichlet energy in neural networks.
result Over-parameterized neural networks are implicitly regularized by geometric model complexity.

The article applies Occam's Razor to non-parametric model building, minimizing the number of bits for data encoding.

problem Overlooking the role of model parameters in data encoding leads to inefficient probability density estimators.
method Extends bit counting to model parameters, providing a true measure of complexity for parametric models.
result Minimizing total bit requirement leads to smoother, more efficient probability density estimates and fewer relevant parameters.

Deep neural networks perform well due to a balance of architecture, training, and structured data.

problem Understanding why overparameterized DNNs perform well.
method Bayesian approach to analyze the interplay between network architecture, training algorithms, and data structure.
result Structured data and an inductive bias towards simple functions explain DNN success.

New approach to nonuniform learnability using measure theory.

problem Nonuniform learnability of hypotheses with varying sample sizes.
method Measure theoretic approach to redefine nonuniform learnability, introducing a new algorithm (Generalize Measure Learnability).
result Achieved statistical consistency in learning countable hypothesis classes.

Paper proposes a method to estimate scientific parameters in hybrid models without relying on model architecture.

problem Estimating unknown parameters in hybrid models combining machine learning and scientific models.
method Sharpness-aware minimization adapted for hybrid modeling, focusing on model simplicity.
result Demonstrates effectiveness of SAM-based hybrid model learning for scientific parameter estimation.

This study explores how feature graphs enhance GNNs' performance in modeling interactions.

problem Improving GNNs' ability to model feature interactions effectively.
method Investigates feature graphs and their importance in GNNs, using experiments and theoretical support.
result Edges between interacting features are crucial for GNNs, while non-interaction edges can degrade performance.

New model estimates sparse transport maps for high-dimensional data.

problem Estimating optimal transport maps in high-dimensional spaces.
method Proposes a new model using a family of translation invariant costs and sparsity-inducing norms.
result Sparse transport maps that apply Occam's razor to reduce complexity.

New method measures generalizability of deep neural networks based on decision boundary complexity.

problem Lack of generalization methods for deep neural networks.
method Created Decision Boundary Complexity (DBC) score to measure DNN complexity.
result Simpler decision boundaries lead to better generalizability, supporting Occam's Razor.

NLR models often perform worse than LR for outlying input data in environmental sciences.

problem NLR models often give poor predictions for input data outside the training domain.
method Screened input data for outliers, using linear extrapolation for outliers based on NLR within the non-outlier domain.
result NLROR_{\mathrm{OR}} approach reduces poor extrapolation and tends to outperform NLR and LR for outliers.

We exhibit a strong link between frequentist PAC-Bayesian risk bounds and the Bayesian marginal likelihood. That is, for the negative log-likelihood loss function, we show that the minimization of PAC-Bayesian generalization risk bounds maximizes the Bayesian marginal likelihood. This provides an alternative explanatio…

2016-05-27abs ↗pdf ↗

Bayesian nonparametric models, such as Gaussian processes, provide a compelling framework for automatic statistical modelling: these models have a high degree of flexibility, and automatically calibrated complexity. However, automating human expertise remains elusive; for example, Gaussian processes with standard kerne…

2015-10-26abs ↗pdf ↗

Study reveals how model volume affects learning curves in machine learning.

problem Understanding the double descent risk phenomenon in machine learning.
method Investigates the role of model volume using MDL, Occam's Razor, and information geometry.
result Model volume can explain the double descent risk, suggesting better generalization with increased dimensionality.

The study examines causal razors and their logical relations, highlighting a dilemma in causal discovery.

problem Selecting a reasonable scoring criterion for causal discovery algorithms.
method Review and logical comparison of numerous causal razors, focusing on parameter minimality in multinomial models.
result Parameter minimality poses a dilemma in selecting a reasonable scoring criterion for causal discovery algorithms.

Why do deep neural networks (DNNs) benefit from very high dimensional parameter spaces? Their huge parameter complexities vs stunning performance in practice is all the more intriguing and not explainable using the standard theory of model selection for regular models. In this work, we propose a geometrically flavored …

2019-05-27abs ↗pdf ↗

Transfer learning aims at transferring knowledge from a well-labeled domain to a similar but different domain with limited or no labels. Unfortunately, existing learning-based methods often involve intensive model selection and hyperparameter tuning to obtain good results. Moreover, cross-validation is not possible for…

2019-04-02abs ↗pdf ↗

Matrix completion and approximation are popular tools to capture a user's preferences for recommendation and to approximate missing data. Instead of using low-rank factorization we take a drastically different approach, based on the simple insight that an additive model of co-clusterings allows one to approximate matri…

2014-12-31abs ↗pdf ↗

We present an integer programming framework to build accurate and interpretable discrete linear classification models. Unlike existing approaches, our framework is designed to provide practitioners with the control and flexibility they need to tailor accurate and interpretable models for a domain of choice. To this end…

2014-05-16abs ↗pdf ↗

Study improves neural network performance in sequential learning for image classification.

problem Improving neural network performance in sequential learning for image classification.
method Evaluation of approaches for computing prequential description lengths, proposing forward-calibration and replay-streams.
result Improved description lengths for image classification datasets, outperforming previous results.

The paper explores how complex models can improve system identification beyond traditional limits.

problem Balancing model richness and spurious learning in system identification.
method Investigates the double-descent phenomenon in the context of dynamic systems.
result Complex models can improve system identification performance beyond the point of interpolation.

The recent empirical success of unsupervised cross-domain mapping algorithms, between two domains that share common characteristics, is not well-supported by theoretical justifications. This lacuna is especially troubling, given the clear ambiguity in such mappings. We work with adversarial training methods based on IP…

2018-07-23abs ↗pdf ↗

Bayesian framework detects symmetries in chaotic dynamical systems.

problem Detecting symmetries in chaotic attractors for insights into dynamical system structure.
method Bayesian framework using Gibbs posterior constructed from Wasserstein distances.
result Bayesian framework accurately recovers symmetries under high noise and small sample sizes.

We consider the problem of identifying the causal direction between two discrete random variables using observational data. Unlike previous work, we keep the most general functional model but make an assumption on the unobserved exogenous variable: Inspired by Occam's razor, we assume that the exogenous variable is sim…

2016-11-12abs ↗pdf ↗

Leo Breiman's Rashomon Effect and Occam Dilemma are re-evaluated in the context of modern machine learning.

problem The tradeoff between model complexity and accuracy in machine learning.
method Modern perspective on Breiman's arguments using current computational capabilities.
result Algorithmic models can be accurate without being complex, nullifying the Occam Dilemma.

This work characterizes optimal multiclass learning with regularization.

problem The empirical risk minimization (ERM) algorithm fails in multiclass learning settings.
method Using one-inclusion graphs (OIGs), the work introduces optimal learning algorithms that relax structural risk minimization and incorporate unsupervised learning.
result An optimal learner is introduced that uses a local regularization function and an unsupervised learning stage to learn the regularizer.

New estimate reduces overfitting risk in machine learning models.

problem Error rate on test data may not reflect true population error due to adaptive data analysis practices.
method Introduces Rip van Winkle's Razor, a simple estimate of overfit to test data based on information content.
result Shows non-vacuous estimate of deviation in many modern settings.

We introduce the problem of learning mixtures of kk subcubes over {0,1}n\{0,1\}^n, which contains many classic learning theory problems as a special case (and is itself a special case of others). We give a surprising nO(logk)n^{O(\log k)}-time learning algorithm based on higher-order multilinear moments. It is not possible to l…

2018-03-17abs ↗pdf ↗