A new router uses attention-based reinforcement learning to solve detailed routing problems efficiently.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
ASTRA uses unlabeled data and weak rules to train deep models effectively.
Prior-weighted logistic regression has become a standard tool for calibration in speaker recognition. Logistic regression is the optimization of the expected value of the logarithmic scoring rule. We generalize this via a parametric family of proper scoring rules. Our theoretical analysis shows how different members of…
Transformers learn to generalize unseen tasks by composing self-attention layers.
Backpropagation and the chain rule of derivatives have been prominent; however, the total derivative rule has not enjoyed the same amount of attention. In this work we show how the total derivative rule leads to an intuitive visual framework for creating gradient estimators on graphical models. In particular, previous …
Centroid Transformers reduce memory and computation by summarizing inputs into centroids.
Estimates classification rules from partially classified data.
Factor graphs have recently gained increasing attention as a unified framework for representing and constructing algorithms for signal processing, estimation, and control. One capability that does not seem to be well explored within the factor graph tool kit is the ability to handle deterministic nonlinear transformati…
Interpretable classifiers have recently witnessed an increase in attention from the data mining community because they are inherently easier to understand and explain than their more complex counterparts. Examples of interpretable classification models include decision trees, rule sets, and rule lists. Learning such mo…
A central goal of meta-learning is to find a learning rule that enables fast adaptation across a set of tasks, by learning the appropriate inductive bias for that set. Most meta-learning algorithms try to find a \textit{global} learning rule that encodes this inductive bias. However, a global learning rule represented …
Transformers model contextual relations using probabilistic measures, revealing their expressive power.
Self-attention optimizers converge to optimal weights, revealing bias patterns.
Bayesian neural networks are shown to be minimax and admissible under certain conditions.
Recent exploration of optimal individualized decision rules (IDRs) for patients in precision medicine has attracted a lot of attention due to the heterogeneous responses of patients to different treatments. In the existing literature of precision medicine, an optimal IDR is defined as a decision function mapping from t…
A modern Hopfield network improves deep learning with better memory and attention.
The purpose of this paper relies on the study of long term affine yield curves modeling. It is inspired by the Ramsey rule of the economic literature, that links discount rate and marginal utility of aggregate optimal consumption. For such a long maturity modelization, the possibility of adjusting preferences to new ec…
The paper argues that machine learning is a falsificationist process.
New methods improve Byzantine robustness in distributed learning.
Given an -sample of random vectors whose joint law is unknown, the long-standing problem of supervised classification aims to \textit{optimally} predict the label of a given a new observation . In this context, the nearest neighbor rule is a popular flexible and intuitive method …
SDPA is shown to be an optimal transport problem in deep learning.
KATA improves associative recall by optimizing feature maps derived from nonnegative attention weights.
Estimates proper calibration errors and refinement terms in probabilistic predictions.
Paper proposes a math-inspired L2O model for better generalization.
A Nash game theory approach allocates capital requirements among financial institutions.
DeepRC uses Hopfield networks and attention to classify immune repertoires.
Proposes a method to create fair ITRs that balance value and fairness.
Gradient-based meta-learning has proven to be highly effective at learning model initializations, representations, and update rules that allow fast adaptation from a few samples. The core idea behind these approaches is to use fast adaptation and generalization -- two second-order metrics -- as training signals on a me…
GEAR uses auxiliary data to estimate optimal decisions in studies with limited primary outcomes.
A new method improves few-shot image classification by updating top layers.
Modelling and exploiting teammates' policies in cooperative multi-agent systems have long been an interest and also a big challenge for the reinforcement learning (RL) community. The interest lies in the fact that if the agent knows the teammates' policies, it can adjust its own policy accordingly to arrive at proper c…
Transformer-based multi-scale model outperforms traditional methods in solving PDEs on irregular domains.
Learning to remember long sequences remains a challenging task for recurrent neural networks. Register memory and attention mechanisms were both proposed to resolve the issue with either high computational cost to retain memory differentiability, or by discounting the RNN representation learning towards encoding shorte…
SemiGNN detects financial fraud using social relations and multi-view data.
Skein theory classifies UFCs with specific fusion rules.
Finding regions for which there is higher controversy among different classifiers is insightful with regards to the domain and our models. Such evaluation can falsify assumptions, assert some, or also, bring to the attention unknown phenomena. The present work describes an algorithm, which is based on the Exceptional M…
Pulli kolam is a ubiquitous art form in south India. It involves drawing a line looped around a collection of dots (pullis) place on a plane such that three mandatory rules are followed: all line orbits should be closed, all dots are encircled and no two lines can overlap over a finite length. The mathematical foundati…
The problem of frequent pattern mining has been studied quite extensively for various types of data, including sets, sequences, and graphs. Somewhat surprisingly, another important type of data, namely rank data, has received very little attention in data mining so far. In this paper, we therefore addresses the problem…
Molecular graph generation is a fundamental problem for drug discovery and has been attracting growing attention. The problem is challenging since it requires not only generating chemically valid molecular structures but also optimizing their chemical properties in the meantime. Inspired by the recent progress in deep …
Machine learning selects the best prediction rules from noisy data.
Nonconvex optimization problems arise in different research fields and arouse lots of attention in signal processing, statistics and machine learning. In this work, we explore the accelerated proximal gradient method and some of its variants which have been shown to converge under nonconvex context recently. We show th…
We discuss risk measures representing the minimum amount of capital a financial institution needs to raise and invest in a pre-specified eligible asset to ensure it is adequately capitalized. Most of the literature has focused on cash-additive risk measures, for which the eligible asset is a risk-free bond, on the grou…
A new framework for efficient sequence maps using Bayesian filtering and covariance.
Interpretable classification models are built with the purpose of providing a comprehensible description of the decision logic to an external oversight agent. When considered in isolation, a decision tree, a set of classification rules, or a linear model, are widely recognized as human-interpretable. However, such mode…
We present a memory augmented neural network for natural language understanding: Neural Semantic Encoders. NSE is equipped with a novel memory update rule and has a variable sized encoding memory that evolves over time and maintains the understanding of input sequences through read}, compose and write operations. NSE c…
Recent studies in the literature have paid much attention to the sparsity in linear classification tasks. One motivation of imposing sparsity assumption on the linear discriminant direction is to rule out the noninformative features, making hardly contribution to the classification problem. Most of those work were focu…
New model learns causal world dynamics from state space models.
Zero-Copy Architecture Detects Cross-Company Financial Signals Instantly.
Cosmos models scenes using neural encodings and symbolic attributes for compositional generalization.