The study proposes algorithms to minimize rating discordance in missing data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Framework integrates financial and annual report data for better corporate credit ratings.
A non-trivial probability structure is evident in the binary data extracted from the up/down price movements of very high frequency data such as tick-by-tick data for USD/JPY. In this paper, we analyze the Sony bank USD/JPY rates, ignoring the small deviations from the market price. We then show there is a similar non-…
Active data collection improves convergence rates in operator learning.
This research examines how the error rate of nearest neighbor classifiers varies with dataset size.
IUS framework predicts EUR/USD exchange rate with improved accuracy.
Most of the existing recommender systems use the ratings provided by users on individual items. An additional source of preference information is to use the ratings that users provide on sets of items. The advantages of using preferences on sets are two-fold. First, a rating provided on a set conveys some preference in…
This paper uses LLMs to improve equity stock ratings by ingesting diverse financial and news data.
CCR-CNN uses CNN to predict corporate credit ratings from financial data.
Paper studies CLT rates for dependent data in Wasserstein-p distance.
The paper models rating transitions and calibrates them to market data for XVA calculations.
We calibrate and test various variants of field theory models of the interest rate with data from eurodollars futures. A model based on a simple psychological factor are seen to provide the best fit to the market. We make a model independent determination of the volatility function of the forward rates from market data…
Continuous, ubiquitous monitoring through wearable sensors has the potential to collect useful information about users' context. Heart rate is an important physiologic measure used in a wide variety of applications, such as fitness tracking and health monitoring. However, wearable sensors that monitor heart rate, such …
Estimates rate-distortion function for large datasets using neural networks.
Rate-In dynamically adjusts dropout rates during inference to improve uncertainty estimation in neural networks.
Exact risk and learning rate curves derived for adaptive SGD on high-dimensional problems.
Paper analyzes deep learning models for credit rating prediction using text and numerical data.
We first show that there are in fact triangular arbitrage opportunities in the spot foreign exchange markets, analyzing the time dependence of the yen-dollar rate, the dollar-euro rate and the yen-euro rate. Next, we propose a model of foreign exchange rates with an interaction. The model includes effects of triangular…
Study shows deep linear networks can converge to flatter minima at large learning rates.
Rating prediction is an important application, and a popular research topic in collaborative filtering. However, both the validity of learning algorithms, and the validity of standard testing procedures rest on the assumption that missing ratings are missing at random (MAR). In this paper we present the results of a us…
A Hawkes process model with a time-varying background rate is developed for analyzing the high-frequency financial data. In our model, the logarithm of the background rate is modeled by a linear model with a relatively large number of variable-width basis functions, and the parameters are estimated by a Bayesian method…
Crowdsourcing is an effective tool for human-powered computation on many tasks challenging for computers. In this paper, we provide finite-sample exponential bounds on the error rate (in probability and in expectation) of hyperplane binary labeling rules under the Dawid-Skene crowdsourcing model. The bounds can be appl…
We introduce an autoregressive-type model with self-modulation effects for a foreign exchange rate by separating the foreign exchange rate into a moving average rate and an uncorrelated noise. From this model we indicate that traders are mainly using strategies with weighted feedbacks of the past rates in the exchange …
Analyzes Indian commercial dynamism using time series data.
We prove new fast learning rates for the one-vs-all multiclass plug-in classifiers trained either from exponentially strongly mixing data or from data generated by a converging drifting distribution. These are two typical scenarios where training data are not iid. The learning rates are obtained under a multiclass vers…
Study expands multiclass classification models with new rates and partial concept classes.
MOB-dS uses permutation to correct for dependency in discrete survival data.
Data of the form of event times arise in various applications. A simple model for such data is a non-homogeneous Poisson process (NHPP) which is specified by a rate function that depends on time. We consider the problem of having access to multiple independent observations of event time data, observed on a common inter…
New research shows unlabeled data is equally valuable as labeled data in certain semi-supervised learning scenarios.
Paper improves classification rates for private data.
Study examines how data augmentation impacts optimization in linear regression.
Paper analyzes faster convergence rates for reinforcement learning from offline data.
In this paper, we investigate the statistical convergence rate of a Bayesian low-rank tensor estimator. Our problem setting is the regression problem where a tensor structure underlying the data is estimated. This problem setting occurs in many practical applications, such as collaborative filtering, multi-task learnin…
Entropy rate of sequential data-streams naturally quantifies the complexity of the generative process. Thus entropy rate fluctuations could be used as a tool to recognize dynamical perturbations in signal sources, and could potentially be carried out without explicit background noise characterization. However, state of…
We introduce a new concept, data irrecoverability, and show that the well-studied concept of data privacy is sufficient but not necessary for data irrecoverability. We show that there are several regularized loss minimization problems that can use perturbed data with theoretical guarantees of generalization, i.e., loss…
Large learning rates cause oscillations in NN weights that improve generalization.
New guarantees for ERM with adaptively collected data.
In banking practice, rating transition matrices have become the standard approach of deriving multi-year probabilities of default (PDs) from one-year PDs, the latter normally being available from Basel ratings. Rating transition matrices have gained in importance with the newly adopted IFRS 9 accounting standard. Here,…
We present two methodologies on the estimation of rating transition probabilities within Markov and non-Markov frameworks. We first estimate a continuous-time Markov chain using discrete (missing) data and derive a simpler expression for the Fisher information matrix, reducing the computational time needed for the Wald…
Epidemiological model updates infection rates in China, US, Italy.
The paper uses machine learning and Lie groups to improve rating transitions and XVA calculations.
We apply multiple testing procedures to the validation of estimated default probabilities in credit rating systems. The goal is to identify rating classes for which the probability of default is estimated inaccurately, while still maintaining a predefined level of committing type I errors as measured by the familywise …
This paper reviews methods for distributed training of machine learning models from high-rate streams.
This paper reviews methods for constructing confidence intervals for error rates in 1:1 matching tasks.
Classifiers built with neural networks handle large-scale high dimensional data, such as facial images from computer vision, extremely well while traditional statistical methods often fail miserably. In this paper, we attempt to understand this empirical success in high dimensional classification by deriving the conver…
For environmental problems such as global warming future costs must be balanced against present costs. This is traditionally done using an exponential function with a constant discount rate, which reduces the present value of future costs. The result is highly sensitive to the choice of discount rate and has generated …
Optimal learning rate schedules for SGD in changing data distributions.
BEER accelerates decentralized nonconvex optimization to rate.