Robust risk minimisation has several advantages: it has been studied with regards to improving the generalisation properties of models and robustness to adversarial perturbation. We bound the distributionally robust risk for a model class rich enough to include deep neural networks by a regularised empirical risk invol…
In this paper we formulate in general terms an approach to prove strong consistency of the Empirical Risk Minimisation inductive principle applied to the prototype or distance based clustering. This approach was motivated by the Divisive Information-Theoretic Feature Clustering model in probabilistic space with Kullbac…
Deviation inequalities for stochastic approximation methods.
problem Establishing bounds on the deviation of stochastic approximation methods.
method Martingale approximation method for separately Lipschitz functions.
result Established various deviation inequalities for stochastic approximation by averaging and minimization.
The study analyzes multi-class teacher-student perceptron performance and generalization errors.
problem Analyzing multi-class classification with the teacher-student perceptron.
method Deriving asymptotic expressions for Bayes-optimal and empirical risk minimization (ERM) generalization errors.
result Regularised cross-entropy minimization yields close-to-optimal accuracy for multi-class classification.
In machine learning we often try to optimise a decision rule that would have worked well over a historical dataset; this is the so called empirical risk minimisation principle. In the context of learning from recommender system logs, applying this principle becomes a problem because we do not have available the reward …
Study robust linear regression with outliers, providing exact asymptotics for ERM performance.
problem Robust linear regression in high-dimension with outliers.
method Analyzes ℓ2, ℓ1, and Huber losses, providing asymptotic performance metrics. result Optimally-regularised ERM is asymptotically consistent with simple calibration, but Huber loss requires norm calibration.
Adaptive model learns from time series data with changing distributions.
problem Predicting time series data under distribution shift.
method Formulates distribution shift as weighted empirical risk minimization. Uses a gradient-based learning method for a forgetting mechanism.
result Proposes an efficient method for adaptive time series prediction.
The paper optimizes forecasting for risk-adjusted decisions under trading frictions.
problem Optimizing forecasting accuracy for investment decisions in the presence of transaction costs.
method Develops a utility-weighted calibration criterion to minimize decision loss net of costs.
result Utility-weighted calibration reduces decision loss by over 30% and improves Sharpe ratio.
The question of optimal portfolio is addressed. The conventional Markowitz portfolio optimisation is discussed and the shortcomings due to non-Gaussian security returns are outlined. A method is proposed to minimise the likelihood of extreme non-Gaussian drawdowns of the portfolio value. The theory is called Leptokurti…
Expectiles were defined using a minimisation principle. They form a special class of coherent risk measures. We will describe the scenario set and we will show that there is a most severe commonotonic risk measure that is smaller than the given expectile.
Gradient descent performs well on weakly convex losses, offering generalization guarantees.
problem Learning with weakly convex losses using gradient descent.
method Analyzing the stability of gradient descent through the smallest eigenvalue of the Hessian.
result Generalization error bounds hold under a wider range of step sizes.
Upper bounds and lower bounds show ERM outperforms DG methods in various settings.
problem Limitations of domain generalisation methods in various settings.
method Upper bounds and lower bounds on excess risk of ERM, and analysis of DG settings.
result It is not possible to significantly outperform ERM in DG settings.
Study characterizes learning from heavy-tailed data in high dimensions using superstatistical methods.
problem Characterizing learning from heavy-tailed data in high-dimensional settings.
method Empirical risk minimization with double-stochastic processes and superstatistical analysis.
result Analytical characterization of separability transition and generalization performance.
New method for PU learning with instance-dependent propensity scores.
problem Learning from positive and unlabeled data with instance-dependent labeling.
method Empirical risk minimization of joint risk function, alternating optimization of posterior probability and propensity score.
result The method achieves comparable or better performance than state-of-the-art methods.
In this paper, a new approach to computing the generalisation performance is presented that assumes the distribution of risks, ρ(r), for a learning scenario is known. From this, the expected error of a learning machine using empirical risk minimisation is computed for both classification and regression problems. A cr…
Study examines insider trading in short-selling restricted markets.
problem Analyzing insider trading opportunities in short-selling prohibited markets.
method Introducing minimal supermartingale measure and analyzing its properties in relation to minimal martingale measure.
result Conditions under which both measures fail to exist, indicating insider information affecting market perception.
Holdout set improves risk score accuracy without biasing predictions.
problem Updating risk scores can lead to biased estimates when directly applied.
method Use a holdout set of non-intervention population to update risk scores.
result Optimal holdout size reduces adverse outcomes to optimal level.
Adaptive reward models capture individual preferences from human feedback.
problem Learning a reward model that can be specialised to a user.
method Empirical risk minimisation and PAC bound analysis.
result Adaptive reward models benefit from the heterogeneity of user preferences.
New findings show second-order scoring rules can't accurately represent epistemic uncertainty.
problem Lack of epistemic uncertainty representation in second-order learners.
method Generalised second-order scoring rules introduced to prove theoretical limitations.
result No loss function incentivizes second-order learners to accurately represent epistemic uncertainty.
The paper addresses missing data imputation issues by correcting for distribution shift.
problem Missing data imputation and the resulting distribution shift between observed and full data.
method Formulates imputation as a risk minimization problem and proposes a novel algorithm to correct for distribution shift.
result The proposed algorithm consistently improves imputation accuracy, reducing RMSE and Wasserstein distance by 3% and 7%, respectively.
Randomised classifiers outperform deterministic ones in strategic classification.
problem Strategic modification of features by agents in classification tasks.
method Theoretical analysis of randomised classifiers in strategic classification.
result Randomised classifiers can achieve better accuracy than deterministic ones under certain conditions.
Portfolio optimisation typically aims to provide an optimal allocation that minimises risk, at a given return target, by diversifying over different investments. However, the potential scope of such risk diversification can be limited if investments are concentrated in only one country, or more specifically one currenc…
The study characterizes learning Gaussian mixtures using GLMs in high dimensions.
problem Learning Gaussian mixtures with generalised linear models in high-dimensional settings.
method Empirical risk minimization with convex loss and regularisation.
result Exact asymptotics of the ERM estimator for Gaussian mixtures in high dimensions.
New bounds improve generalization in machine learning with high probability.
problem Erratic behavior of KL divergence limits practical applications.
method Replaced KL divergence with Wasserstein distance for better bounds.
result Proved high probability generalization bounds for i.i.d. and non-i.i.d. data.
Study optimizes Bitcoin futures hedging to reduce liquidation risk.
problem Optimizing hedging strategies to minimize liquidation risk in Bitcoin futures.
method Derived a semi-closed form optimal hedging strategy considering spot and futures extreme returns, loss aversion, leverage, and collateral management.
result Optimal strategy reduces both hedged portfolio variance and liquidation probability.
Collider regression improves predictive performance in regression tasks.
problem Discarding prior causal knowledge in regression tasks.
method Collider regression framework incorporating probabilistic causal knowledge from collider structures.
result Proves positive generalization benefit and provides closed-form estimators.
Solves ambiguity in incomplete markets by minimizing price measure entropy.
problem Ambiguity in pricing incomplete markets.
method Minimizes the entropy of the price measure from the economic measure, subject to mark-to-market constraints.
result Resolves ambiguity and provides a consistent pricing measure.
In this article, we derive concentration inequalities for the cross-validation estimate of the generalization error for empirical risk minimizers. In the general setting, we prove sanity-check bounds in the spirit of \cite{KR99} \textquotedblleft\textit{bounds showing that the worst-case error of this estimate is not m…
Analyzes deep neural networks training errors with SGD and random init.
problem Lack of rigorous understanding of deep learning algorithms.
method Mathematical analysis of deep learning with SGD and random init.
result First full error analysis for deep learning with SGD and random init.
A method improves Cryo-EM 3D map refinement by regularizing rotation estimation.
problem Noise-robustness vs. data-consistency in Cryo-EM 3D map reconstruction.
method Ellipsoidal support lifting (ESL) for regularizing and approximating the global minimizer over Riemannian manifolds.
result The induced bias due to regularizing effect of ESL estimates better rotations than global optimisation.
Stochastic RNNs classify biological neural network paths with robust error bounds.
problem Classifying biological neural network paths.
method Modelled as a continuous-time stochastic recurrent neural network (RNN) with identity activation function, analysed in the robust regime.
result Generalisation error bound holds with high probability, showing the empirical risk minimiser is the best-in-class hypothesis.
We present a model of predatory traders interacting with each other in the presence of a central reserve (which dissipates their wealth through say, taxation), as well as inflation. This model is examined on a network for the purposes of correlating complexity of interactions with systemic risk. We suggest the use of s…
Study optimal consumption and investment for investors with Epstein-Zin preferences.
problem Optimal consumption and investment for investors with Epstein-Zin preferences in an incomplete market.
method Variational characterisation and direct method to prove existence of optimal policies.
result Existence and uniqueness of optimal consumption and investment policies.
New method uses kernel Stein discrepancy for measure transport without strict continuity constraints.
problem Minimizing Kullback-Leibler divergence for posterior approximation.
method Proposes minimizing kernel Stein discrepancy instead of Kullback-Leibler divergence.
result Demonstrates consistency and competitiveness of the new method.
A novel optimisation framework through quadratic nonlinear projection is introduced for credit portfolio when the portfolio risk is measured by Conditional Value-at-Risk (CVaR). The whole optimisation procedure to search toward the optimal portfolio state is conducted by a series of single-step optimisations under the …
A thesis submitted for the degree of Doctor of Philosophy of The Australian National University. In this work we introduce several new optimisation methods for problems in machine learning. Our algorithms broadly fall into two categories: optimisation of finite sums and of graph structured objectives. The finite sum pr…
We propose graph-dependent implicit regularisation strategies for distributed stochastic subgradient descent (Distributed SGD) for convex problems in multi-agent learning. Under the standard assumptions of convexity, Lipschitz continuity, and smoothness, we establish statistical learning rates that retain, up to logari…
Sharp stability result for maps near infinitely concentrated minimisers.
problem Stability of maps near minimisers with infinite concentration.
method Dynamic approach to deform maps into harmonic maps, controlling topology changes.
result Sharp quantitative estimates on map distance to infinitely concentrated minimisers.
Deep models can fit noisy labels, but robustness and reliability are still issues.
problem Training deep models with noisy labels leads to unreliable uncertainty quantification.
method Analysis of conditional distribution over noisy labels and evaluation of robust loss functions.
result Strictly proper and robust loss functions preserve accuracy but do not guarantee reliability.
New results on hypersurfaces show no branch points, improving smoothness.
problem Analyzing area minimising hypersurfaces mod p without branch points.
method General analysis of immersed stable minimal hypersurfaces with alternating orientation.
result Area minimising hypersurfaces mod p do not admit immersed branch points.
We study the problem of finding strain-minimising stream surfaces in a divergence-free vector field. These surfaces are generated by motions of seed curves that propagate through the field in a strain minimising manner, i.e., they move without stretching or shrinking, preserving the length of their arbitrary arc. In ge…
Loss minimisation fails to capture epistemic uncertainty in second-order predictors.
problem Capturing epistemic uncertainty in machine learning models.
method Analysis of a second-order learner approach using loss minimisation.
result Loss minimisation does not faithfully represent epistemic uncertainty in second-order predictors.
New method optimises worst-case risk under model uncertainty.
problem Minimizing expected risk under posterior beliefs leads to sub-optimal decisions due to model uncertainty.
method Distributionally Robust Optimisation with Bayesian Ambiguity Sets (DRO-BAS)
result Improved out-of-sample robustness in the Newsvendor problem.
Ensuring that classifiers are non-discriminatory or fair with respect to a sensitive feature (e.g., race or gender) is a topical problem. Progress in this task requires fixing a definition of fairness, and there have been several proposals in this regard over the past few years. Several of these, however, assume either…
Let Ω⊂R3 be a Lipschitz domain, and consider a harmonic map v:Ω→S2 with boundary data v∣∂Ω=φ which minimises the Dirichlet energy. For p≥2, we show that any energy minimiser u whose boundary map ψ has a small W1,p-distance to φ is close t…
We prove existence and regularity of minimisers for the Canham-Helfrich energy in the class of weak (possibly branched and bubbled) immersions of the 2-sphere. This solves (the spherical case) of the minimisation problem proposed by Helfrich in 1973, modelling lipid bilayer membranes. On the way to prove the main res…
Enhanced feature learning using neural networks and kernel methods with improved robustness.
problem Improving feature learning and function estimation in supervised learning.
method Regularised empirical risk minimisation with a new kernel approach.
result The proposed method, BKerNN, converges to the minimal risk with explicit high-probability rates.
An important task in computational statistics and machine learning is to approximate a posterior distribution p(x) with an empirical measure supported on a set of representative points {xi}i=1n. This paper focuses on methods where the selection of points is essentially deterministic, with an emphasis on achi…