Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

3517021,0521,403 · Jun 202019922001200920172026
48 results for model zoo

This paper improves model generalization by integrating diverse pretrained models.

problem Leveraging diverse pretrained models for robust out-of-distribution generalization.
method Characterize and integrate diverse pretrained models based on diversity and correlation shifts.
result Demonstrates state-of-the-art out-of-distribution generalization performance.

We study del Pezzo surfaces that are quasismooth and well-formed weighted hypersurfaces. In particular, we find all such surfaces whose alpha-invariant of Tian is greater than 2/3.

2009-04-01abs ↗pdf ↗

In the eyes of a rationalist like Descartes or Spinoza, human reasoning is flawless, marching toward uncovering ultimate truth. A few centuries later, however, culminating in the work of Kahneman and Tversky, human reasoning was portrayed as anything but flawless, filled with numerous misjudgments, biases, and cognitiv…

2018-10-15abs ↗pdf ↗

Since the discovery of differential calculus by Newton and Leibniz and the subsequent continuous growth of its applications to physics, mechanics, geometry, etc, it was observed that partial derivatives in the study of various natural problems are (self-)organized in certain structures usually called geometric. Tensors…

2015-11-21abs ↗pdf ↗

We consider the groups DiffB(Rn)\operatorname{Diff}_{\mathcal B}(\mathbb R^n), DiffH(Rn)\operatorname{Diff}_{H^\infty}(\mathbb R^n), and DiffS(Rn)\operatorname{Diff}_{\mathcal S}(\mathbb R^n) of smooth diffeomorphisms on Rn\mathbb R^n which differ from the identity by a function which is in either B\mathcal B (bounded in all derivatives),…

2012-11-24abs ↗pdf ↗

HotNAS reduces AI search time from hundreds of GPU hours to less than 3 GPU hours.

problem High time required for AI solution generation from scratch.
method HotNAS starts from a 'hot' state using pre-trained models and integrates compression during co-search.
result Reduces search time from 200 GPU hours to less than 3 GPU hours.

A zoo of deep nets is available these days for almost any given task, and it is increasingly unclear which net to start with when addressing a new task, or which net to use as an initialization for fine-tuning a new model. To address this issue, in this paper, we develop knowledge flow which moves 'knowledge' from mult…

2019-04-11abs ↗pdf ↗

Datasets are growing not just in size but in complexity, creating a demand for rich models and quantification of uncertainty. Bayesian methods are an excellent fit for this demand, but scaling Bayesian inference is a challenge. In response to this challenge, there has been considerable recent work based on varying assu…

2016-02-16abs ↗pdf ↗

The paper tackles model selection for unseen tasks by capturing relationships among checkpoints.

problem Deciding which model combinations are likely to be effective for a new task is difficult.
method The paper models the task space as a Gaussian process and identifies representative checkpoints using mutual information and a greedy algorithm.
result Representative checkpoints generalize to new tasks with superior performance.

Study classifies translating solitons in Minkowski 3-space, revealing singularities.

problem Classifying translating solitons in Minkowski 3-space.
method Introduced the concept of translating solitons on a light-like direction, classified them into graphical and invariant families.
result All time-like examples are incomplete and some have singularities.

Corporate bond factor research is flawed due to measurement errors and ex-post filtering.

problem Replication crisis in corporate bond factor research.
method Analysis of 108 signals across nine thematic clusters, correction of transaction prices and return filtering.
result Majority of previously documented factors do not produce statistically significant alphas after correction.

Suppose XX and YY are finite complexes, with YY simply connected. Gromov conjectured that the number of mapping classes in [X,Y][X,Y] which can be realized by LL-Lipschitz maps grows asymptotically as LαL^α, where αα is an integer determined by the rational homotopy type of YY and the rational cohomology of XX. Thi…

2018-05-31abs ↗pdf ↗

A study finds that only a few factors explain corporate bond risk, rendering extensive bond factor literature redundant.

problem The redundancy of extensive bond factor literature in explaining corporate bond risk premia.
method Bayesian Model Averaging Stochastic Discount Factor analysis of 18 quadrillion models.
result A Bayesian Model Averaging SDF explains risk premia better than low-dimensional models, with an out-of-sample Sharpe ratio of 1.5 to 1.8.

BOAT optimizes multiple antibody properties efficiently.

problem Balancing multiple drug-like properties in antibody design.
method Bayesian optimization framework coupling surrogate modeling and genetic algorithm.
result Competitive performance with state-of-the-art multi-objective protein optimization methods.

We provide a detailed description of solutions of Curve Shortening in Rn\R^n that are invariant under some one-parameter symmetry group of the equation, paying particular attention to geometric properties of the curves, and the asymptotic properties of their ends. We find generalized helices, and a connection with curv…

2012-07-17abs ↗pdf ↗

DiffEqFlux.jl is a library for fusing neural networks and differential equations. In this work we describe differential equations from the viewpoint of data science and discuss the complementary nature between machine learning models and differential equations. We demonstrate the ability to incorporate DifferentialEqua…

2019-02-06abs ↗pdf ↗

Automates galaxy morphology classification with less human labelling.

problem Insufficient human-labeled galaxy images for accurate classification.
method Developed a VAE with equivariant transformer layers and a classifier network.
result Improves accuracy with fewer labels and unlabelled data.

We explain SSL objectives as log-likelihoods in a data curation model.

problem Lack of understanding of SSL objectives as log-likelihoods.
method Formulate SSL objectives as a log-likelihood in a generative model of data curation.
result SSL methods can be understood as lower-bounds on a principled log-likelihood.

Let C[M]C^{[M]} be a (local) Denjoy-Carleman class of Beurling or Roumieu type, where the weight sequence M=(Mk)M=(M_k) is log-convex and has moderate growth. We prove that the groups DiffB[M](Rn){\operatorname{Diff}}\mathcal{B}^{[M]}(\mathbb{R}^n), DiffW[M],p(Rn){\operatorname{Diff}}W^{[M],p}(\mathbb{R}^n), ${\operatorname{Diff}}{\mathcal{S}}{}_…

2014-04-28abs ↗pdf ↗

Numerous methods for crafting adversarial examples were proposed recently with high success rate. Since most existing machine learning based classifiers normalize images into some continuous, real vector, domain firstly, attacks often craft adversarial examples in such domain. However, "adversarial" examples may become…

2019-05-19abs ↗pdf ↗

The paper introduces BCART models for aggregate claim amount, improving frequency-severity and joint modeling.

problem Modeling aggregate claim amount with frequency-severity and joint dependencies.
method Developed three types of BCART models: frequency-severity, sequential, and joint models. Used various distributions for claim severity data.
result Weibull distribution outperforms gamma and lognormal for right-skewed, heavy-tailed claim severity data.

The paper uses model-based trees to create interpretable surrogate models for complex machine learning models.

problem Interpreting complex machine learning models.
method Using model-based trees to partition feature space and create interpretable models.
result Model-based trees generate optimal surrogate models that balance interpretability and performance.

The study examines how model predictions hold up under model extensions.

problem Model predictions may not be robust under model extensions, limiting their applicability.
method The study uses causal ordering to assess robustness of qualitative model predictions and characterizes model extensions that preserve predictions.
result Conditions and techniques are provided to assess robustness of model predictions under model extensions.

Revises Bayesian model averaging for foundation models.

problem Ensemble pre-trained and lightly-finetuned foundation models for improved classification performance.
method Introduces trainable linear classifiers and computationally cheaper model averaging scheme (OMA).
result Ensembled models can better predict on various datasets.

Paper introduces symmetric divergence link models for probability distributions.

problem Symmetric divergence measures for probability distributions.
method Two general classes of link models: one for survival functions and another for cumulative probability distribution functions.
result Advantages of symmetric divergence measures over asymmetric measures for model averaging and feature assessment.

Researchers review challenges in interpreting additive models, especially neural additive models.

problem Challenges in interpreting additive models, particularly neural additive models.
method Review of generalized additive models and discussion of nonidentifiability.
result Challenges in claiming interpretability or suitability for safety-critical applications of additive models.

Sigma models linked to Gross-Neveu models via quiver varieties.

problem Understanding the relationship between sigma models and Gross-Neveu models.
method Exploring the mathematical correspondence between sigma models and Gross-Neveu models, including their geometric and trigonometric/elliptic deformations.
result Sigma models are mathematically equivalent to Gross-Neveu models under certain conditions.