We have recently proposed a new information-based approach to model selection, the Frequentist Information Criterion (FIC), that reconciles information-based and frequentist inference. The purpose of this current paper is to provide a simple example of the application of this criterion and a demonstration of the natura…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In the information-based paradigm of inference, model selection is performed by selecting the candidate model with the best estimated predictive performance. The success of this approach depends on the accuracy of the estimate of the predictive complexity. In the large-sample-size limit of a regular model, the predicti…
There are three principle paradigms of statistical inference: (i) Bayesian, (ii) information-based and (iii) frequentist inference. We describe an objective prior (the weighting or -prior) which unifies objective Bayes and information-based inference. The -prior is chosen to make the marginal probability an unbia…
Change-point analysis is a flexible and computationally tractable tool for the analysis of times series data from systems that transition between discrete states and whose observables are corrupted by noise. The change-point algorithm is used to identify the time indices (change points) at which the system transitions …
We consider a class of generalized capital asset pricing models in continuous time with a finite number of agents and tradable securities. The securities may not be sufficient to span all sources of uncertainty. If the agents have exponential utility functions and the individual endowments are spanned by the securities…
We present simple and computationally efficient nonparametric estimators of Rényi entropy and mutual information based on an i.i.d. sample drawn from an unknown, absolutely continuous distribution over . The estimators are calculated as the sum of -th powers of the Euclidean lengths of the edges of the `genera…
New bound on machine learning model performance using Jensen-Shannon information.
Trustworthy AI is a critical issue in machine learning where, in addition to training a model that is accurate, one must consider both fair and robust training in the presence of data bias and poisoning. However, the existing model fairness techniques mistakenly view poisoned data as an additional bias to be fixed, res…
We analyze the adversarial examples problem in terms of a model's fault tolerance with respect to its input. Whereas previous work focuses on arbitrarily strict threat models, i.e., -perturbations, we consider arbitrary valid inputs and propose an information-based characteristic for evaluating tolerance to diverse …
Unknown constraints arise in many types of expensive black-box optimization problems. Several methods have been proposed recently for performing Bayesian optimization with constraints, based on the expected improvement (EI) heuristic. However, EI can lead to pathologies when used with constraints. For example, in the c…
In the information-based approach to asset pricing the market filtration is modelled explicitly as a superposition of signals concerning relevant market factors and independent noise. The rate at which the signal is revealed to the market then determines the overall magnitude of asset volatility. By letting this inform…
This paper provides sufficient conditions for the time of bankruptcy (of a company or a state) for being a totally inaccessible stopping time and provides the explicit computation of its compensator in a framework where the flow of market information on the default is modelled explicitly with a Brownian bridge between …
Proposes mutual information for regression without prior knowledge.
A fast feature selection method using OLS and SOCC for classification.
We are working to develop automated intelligent agents, which can act and react as learning machines with minimal human intervention. To accomplish this, an intelligent agent is viewed as a question-asking machine, which is designed by coupling the processes of inference and inquiry to form a model-based learning unit.…
In this paper we introduce a class of information-based models for the pricing of fixed-income securities. We consider a set of continuous- time information processes that describe the flow of information about market factors in a monetary economy. The nominal pricing kernel is at any given time assumed to be given by …
A pricing formula for discount bonds, based on the consideration of the market perception of future liquidity risk, is established. An information-based model for liquidity is then introduced, which is used to obtain an expression for the bond price. Analysis of the bond price dynamics shows that the bond volatility is…
AES uses α-divergence to select informative points for BO, improving optimization performance.
A new method improves uncertainty estimation in deep learning, especially for hard-to-label samples.
Bayesian Neural Networks improve high-dimensional level set estimation.
In a standard setting of Bayesian optimization (BO), the objective function evaluation is assumed to be highly expensive. Multi-fidelity Bayesian optimization (MFBO) accelerates BO by incorporating lower fidelity observations available with a lower sampling cost. In this paper, we focus on the information-based approac…
In this paper we examine inefficiencies and information disparity in the Japanese stock market. By carefully analysing information publicly available on the internet, an `outsider' to conventional statistical arbitrage strategies--which are based on market microstructure, company releases, or analyst reports--can never…
TES optimizes black-box functions efficiently with minimal approximations.
Multivariate pattern analyses approaches in neuroimaging are fundamentally concerned with investigating the quantity and type of information processed by various regions of the human brain; typically, estimates of classification accuracy are used to quantify information. While a extensive and powerful library of method…
A new framework for asset price dynamics is introduced in which the concept of noisy information about future cash flows is used to derive the price processes. In this framework an asset is defined by its cash-flow structure. Each cash flow is modelled by a random variable that can be expressed as a function of a colle…
We propose a betting strategy based on Bayesian logistic regression modeling for the probability forecasting game in the framework of game-theoretic probability by Shafer and Vovk (2001). We prove some results concerning the strong law of large numbers in the probability forecasting game with side information based on …
GIBBON unifies Bayesian optimization for various problem types.
Method quantifies uncertainties in complex MRF models.
New algorithms improve convergence rates for non-log-concave sampling and log-partition estimation.
A new framework for asset pricing based on modelling the information available to market participants is presented. Each asset is characterised by the cash flows it generates. Each cash flow is expressed as a function of one or more independent random variables called market factors or "X-factors". Each X-factor is ass…
This article attempts to place the emergence of probabilistic numerics as a mathematical-statistical research field within its historical context and to explore how its gradual development can be related both to applications and to a modern formal treatment. We highlight in particular the parallel contributions of Sul'…
In this paper we present a novel quasi-Newton algorithm for use in stochastic optimisation. Quasi-Newton methods have had an enormous impact on deterministic optimisation problems because they afford rapid convergence and computationally attractive algorithms. In essence, this is achieved by learning the second-order (…
Regularization is a big issue for training deep neural networks. In this paper, we propose a new information-theory-based regularization scheme named SHADE for SHAnnon DEcay. The originality of the approach is to define a prior based on conditional entropy, which explicitly decouples the learning of invariant represent…
Regularization is a big issue for training deep neural networks. In this paper, we propose a new information-theory-based regularization scheme named SHADE for SHAnnon DEcay. The originality of the approach is to define a prior based on conditional entropy, which explicitly decouples the learning of invariant represent…
Two new undersampling methods improve classification accuracy for imbalanced datasets.
RES improves robustness in Bayesian optimization.
Combines historical and market data for better portfolio selection.
We apply information-based complexity analysis to support vector machine (SVM) algorithms, with the goal of a comprehensive continuous algorithmic analysis of such algorithms. This involves complexity measures in which some higher order operations (e.g., certain optimizations) are considered primitive for the purposes …
In financial markets, the information that traders have about an asset is reflected in its price. The arrival of new information then leads to price changes. The `information-based framework' of Brody, Hughston and Macrina (BHM) isolates the emergence of information, and examines its role as a driver of price dynamics.…
Improved ARMA-GARCH model for illiquid assets like cryptocurrencies.
Complex dynamical systems driven by the unravelling of information can be modelled effectively by treating the underlying flow of information as the model input. Complicated dynamical behaviour of the system is then derived as an output. Such an information-based approach is in sharp contrast to the conventional mathem…
Estimating the causal effects of an intervention from high-dimensional observational data is difficult due to the presence of confounding. The task is often complicated by the fact that we may have a systematic missingness in our data at test time. Our approach uses the information bottleneck to perform a low-dimension…
We analyze random feature and two-layer neural networks using duality framework.
In reinforcement learning, an agent learns to reach a set of goals by means of an external reward signal. In the natural world, intelligent organisms learn from internal drives, bypassing the need for external signals, which is beneficial for a wide range of tasks. Motivated by this observation, we propose to formulate…
Typical neural networks with external memory do not effectively separate capacity for episodic and working memory as is required for reasoning in humans. Applying knowledge gained from psychological studies, we designed a new model called Differentiable Working Memory (DWM) in order to specifically emulate human workin…
New study shows exponential sample growth for ReQU neural networks.
New metric improves clustering in persistent homology.
Skewness dispersion predicts future stock market returns, especially in months with monetary policy announcements.