Study compares different scoring rules for machine-learned weather forecasts, finding scale-awareness improves forecast realism.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.
Unified theory for neural scaling laws in hierarchically compositional data.
Unified quadrature framework for large-scale kernel machines.
This thesis examines the accuracy of scaling VaR estimates for longer holding periods.
This work provides a scaling rule for model EMA optimization across batch sizes.
New methods prune unpromising rules from KGs, improving scalability and runtime.
New Fourier features improve high-precision approximation in large-scale problems.
The lasso model has been widely used for model selection in data mining, machine learning, and high-dimensional statistical analysis. However, with the ultrahigh-dimensional, large-scale data sets now collected in many real-world applications, it is important to develop algorithms to solve the lasso that efficiently sc…
Lasso is a widely used regression technique to find sparse representations. When the dimension of the feature space and the number of samples are extremely large, solving the Lasso problem remains challenging. To improve the efficiency of solving large-scale Lasso problems, El Ghaoui and his colleagues have proposed th…
New tuning rules for Metropolis algorithms derived from Bayesian large-sample asymptotics.
MOSS optimizes decision rules for accuracy and stability.
Recently, to solve large-scale lasso and group lasso problems, screening rules have been developed, the goal of which is to reduce the problem size by efficiently discarding zero coefficients using simple rules independently of the others. However, screening for overlapping group lasso remains an open challenge because…
FIRE extracts interpretable rules from tree ensembles.
Abstraction and realization are bilateral processes that are key in deriving intelligence and creativity. In many domains, the two processes are approached through rules: high-level principles that reveal invariances within similar yet diverse examples. Under a probabilistic setting for discrete input spaces, we focus …
Model selection for time series forecasting can be biased by the distribution of scores.
Generative models learn rules at different timescales, revealing a 'innovation window'.
Neural network models of early sensory processing typically reduce the dimensionality of streaming input data. Such networks learn the principal subspace, in the sense of principal component analysis (PCA), by adjusting synaptic weights according to activity-dependent learning rules. When derived from a principled cost…
We give a description of local and global moves on a class of locally planar trivalent graphs and we show that it contains -Scale calculus, therefore in particular untyped lambda calculus. Surprisingly, the beta reduction rule comes from a local "sewing" transformation of trivalent locally planar graphs.
Coordinate descent methods employ random partial updates of decision variables in order to solve huge-scale convex optimization problems. In this work, we introduce new adaptive rules for the random selection of their updates. By adaptive, we mean that our selection rules are based on the dual residual or the primal-du…
Diffusion models learn hierarchical composition rules from data.
Optimal Volt/VAR control rules designed using deep learning.
ASTRA uses unlabeled data and weak rules to train deep models effectively.
We derive generalization and excess risk bounds for neural nets using a family of complexity measures based on a multilevel relative entropy. The bounds are obtained by introducing the notion of generated hierarchical coverings of neural nets and by using the technique of chaining mutual information introduced in Asadi…
Enhances sequence memory capacity in neural networks.
Two local learning rules are investigated to avoid weight transport in neural networks.
ScoreStop uses gradient tests to stop gradient boosting early.
New decision-theoretic characterization separates belief and decision posteriors.
Proposes a neural network for efficient imbalance electricity price forecasting.
FinReflectKG builds a comprehensive financial knowledge graph from SEC filings, improving extraction quality.
Boosting is a learning scheme that combines weak prediction rules to produce a strong composite estimator, with the underlying intuition that one can obtain accurate prediction rules by combining "rough" ones. Although boosting is proved to be consistent and overfitting-resistant, its numerical convergence rate is rela…
Identifies learning rules from neural network observables.
Learning an efficient update rule from data that promotes rapid learning of new tasks from the same distribution remains an open problem in meta-learning. Typically, previous works have approached this issue either by attempting to train a neural network that directly produces updates or by attempting to learn better i…
Using a model of wealth distribution where traders are characterized by quenched random saving propensities and trade among themselves by bipartite transactions, we mimic the enhanced rates of trading of the rich by introducing the preferential selection rule using a pair of continuously tunable parameters. The biparti…
We propose a probabilistic formulation that enables sequential detection of multiple change points in a network setting. We present a class of sequential detection rules for certain functionals of change points (minimum among a subset), and prove their asymptotic optimality properties in terms of expected detection del…
The paper studies scaling laws for associative memory mechanisms.
Approximate dynamic programming (ADP) has proven itself in a wide range of applications spanning large-scale transportation problems, health care, revenue management, and energy systems. The design of effective ADP algorithms has many dimensions, but one crucial factor is the stepsize rule used to update a value functi…
Develops deep jump learning for continuous treatment OPE.
Learning and memory in the brain are implemented by complex, time-varying changes in neural circuitry. The computational rules according to which synaptic weights change over time are the subject of much research, and are not precisely understood. Until recently, limitations in experimental methods have made it challen…
New research shows shrinkage methods re-scale portfolio efficient frontiers under distributional misspecification.
New priors can update posteriors without re-estimating likelihoods.
VB uses natural gradients in information geometry.
We empirically test predictability on asset price by using stock selection rules based on maximum drawdown and its consecutive recovery. In various equity markets, monthly momentum- and weekly contrarian-style portfolios constructed from these alternative selection criteria are superior not only in forecasting directio…
With a growing interest in using non-representative samples to train prediction models for numerous outcomes it is necessary to account for the sampling design that gives rise to the data in order to assess the generalized predictive utility of a proposed prediction rule. After learning a prediction rule based on a non…
Robust support vector machine (RSVM) has been shown to perform remarkably well to improve the generalization performance of support vector machine under the noisy environment. Unfortunately, in order to handle the non-convexity induced by ramp loss in RSVM, existing RSVM solvers often adopt the DC programming framework…
Bayesian neural networks are shown to be minimax and admissible under certain conditions.
In this paper, several modifications are introduced to the functional approximation method iterLap to reduce the approximation error, including stopping rule adjustment, proposal of new residual function, starting point selection for numerical optimisation, scaling of Hessian matrix. Illustrative examples are also prov…
What makes a task relatively more or less difficult for a machine compared to a human? Much AI/ML research has focused on expanding the range of tasks that machines can do, with a focus on whether machines can beat humans. Allowing for differences in scale, we can seek interesting (anomalous) pairs of tasks T, T'. We d…