Improves risk control in predictions using semi-supervised calibration.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New tuning rules for Metropolis algorithms derived from Bayesian large-sample asymptotics.
This work improves Bayesian Optimization for setting DNN hyper-parameters.
Two local learning rules are investigated to avoid weight transport in neural networks.
Generative networks minimize predictive scoring rules for probabilistic forecasting.
New rules reduce SLOPE model fitting time by screening out irrelevant variables.
A new principle for optimizer selection improves training speed and performance.
NDI enables high-quality QSM without parameter tuning.
Using a model of wealth distribution where traders are characterized by quenched random saving propensities and trade among themselves by bipartite transactions, we mimic the enhanced rates of trading of the rich by introducing the preferential selection rule using a pair of continuously tunable parameters. The biparti…
A prediscretisation of numerical attributes which is required by some rule learning algorithms is a source of inefficiencies. This paper describes new rule tuning steps that aim to recover lost information in the discretisation and new pruning techniques that may further reduce the size of rule models and improve their…
Unified quadrature framework for large-scale kernel machines.
Bayesian method infers local rules for collective animal movement.
Is cognition a collection of loosely connected functions tuned to different tasks, or can there be a general learning algorithm? If such an hypothetical general algorithm did exist, tuned to our world, could it adapt seamlessly to a world with different laws of nature? We consider the theory that predictive coding is s…
We consider the setting of sequential prediction of arbitrary sequences based on specialized experts. We first provide a review of the relevant literature and present two theoretical contributions: a general analysis of the specialist aggregation rule of Freund et al. (1997) and an adaptation of fixed-share rules of He…
Interpretable classifiers have recently witnessed an increase in attention from the data mining community because they are inherently easier to understand and explain than their more complex counterparts. Examples of interpretable classification models include decision trees, rule sets, and rule lists. Learning such mo…
We introduce a technique that can automatically tune the parameters of a rule-based computer vision system comprised of thresholds, combinational logic, and time constants. This lets us retain the flexibility and perspicacity of a conventionally structured system while allowing us to perform approximate gradient descen…
The random forest algorithm (RF) has several hyperparameters that have to be set by the user, e.g., the number of observations drawn randomly for each tree and whether they are drawn with or without replacement, the number of variables drawn randomly for each split, the splitting rule, the minimum number of samples tha…
Approximate dynamic programming (ADP) has proven itself in a wide range of applications spanning large-scale transportation problems, health care, revenue management, and energy systems. The design of effective ADP algorithms has many dimensions, but one crucial factor is the stepsize rule used to update a value functi…
Paper analyzes neural network distances and stability, leading to a new learning rule.
Here, we study different update rules in stochastic gradient descent (SGD) for online forecasting problems. The selection of the learning rate parameter is critical in SGD. However, it may not be feasible to tune this parameter in online learning. Therefore, it is necessary to have an update rule that is not sensitive …
The paper develops a stationary-distribution theory for Random Forest ensemble size selection.
The ability to learn and adapt in real time is a central feature of biological systems. Neuromorphic architectures demonstrating such versatility can greatly enhance our ability to efficiently process information at the edge. A key challenge, however, is to understand which learning rules are best suited for specific t…
Meta-analysis improves personalized treatment rules across multiple sites.
Proposes a cost-sensitive method to generate probabilistic SVM outputs.
A new method improves few-shot image classification by updating top layers.
A spherical topological manifold of dimension n-1 forms a prototile on its cover, the (n-1)-sphere. The tiling is generated by the fixpoint-free action of the group of deck transformations. By a general theorem, this group is isomorphic to the first homotopy group. Multiplicity and selection rules appear in the form of…
We consider the problem of efficient "on the fly" tuning of existing, or {\it legacy}, Artificial Intelligence (AI) systems. The legacy AI systems are allowed to be of arbitrary class, albeit the data they are using for computing interim or final decision responses should posses an underlying structure of a high-dimens…
A novel distributed adaptive NN classifier for large data sets.
Proposes RPG-RT for red-teaming T2I models without internal access.
EB-TCε identifies the best arm with ε confidence in stochastic bandits.
Stochastic Gradient Descent (SGD) methods are prominent for training machine learning and deep learning models. The performance of these techniques depends on their hyperparameter tuning over time and varies for different models and problems. Manual adjustment of hyperparameters is very costly and time-consuming, and e…
Proposes robust ITRs integrating multiple datasets to handle posterior shift.
Cryptonite tests NLP models with cryptic crossword clues.
Paper proposes a new framework for individualized treatment rules that generalize better across different distributions.
Improves robustness of high-dimensional regression with rank objective and group lasso regularization.
We revisit the classical Douglas-Rachford (DR) method for finding a zero of the sum of two maximal monotone operators. Since the practical performance of the DR method crucially depends on the stepsizes, we aim at developing an adaptive stepsize rule. To that end, we take a closer look at a linear case of the problem a…
A new type of distributional regression tree uses soft split rules for better predictive performance.
We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.
A cubing strategy identifies stable hyperparameter regions for uncertainty quantification in spatial deep learning.
Machine learning algorithms frequently require careful tuning of model hyperparameters, regularization terms, and optimization parameters. Unfortunately, this tuning is often a "black art" that requires expert experience, unwritten rules of thumb, or sometimes brute-force search. Much more appealing is the idea of deve…
Hybrid AI and rule-based framework de-identifies medical imaging data.
In order to interact intelligently with objects in the world, animals must first transform neural population responses into estimates of the dynamic, unknown stimuli which caused them. The Bayesian solution to this problem is known as a Bayes filter, which applies Bayes' rule to combine population responses with the pr…
Transformers learn to generalize unseen tasks by composing self-attention layers.
Statistical guarantees for hyperparameter selection
We can define a neural network that can learn to recognize objects in less than 100 lines of code. However, after training, it is characterized by millions of weights that contain the knowledge about many object types across visual scenes. Such networks are thus dramatically easier to understand in terms of the code th…
Proposes a low-cost method to set hyperparameters using optimized default values.
VORACE uses random classifiers to vote for the best class, saving time and expertise.
Supervised linear feature extraction can be achieved by fitting a reduced rank multivariate model. This paper studies rank penalized and rank constrained vector generalized linear models. From the perspective of thresholding rules, we build a framework for fitting singular value penalized models and use it for feature …