Proposes a new framework for learning image augmentations to improve classification performance.
problem Improving classification performance with a given class of predictors.
method Transformed Risk Minimization (TRM) framework that optimizes both predictive models and data transformations.
result Performance of TRM with SCALE algorithm compares favorably to prior methods on CIFAR10/100.
We illustrate how to compute local risk minimization (LRM) of call options for exponential Lévy models. We have previously obtained a representation of LRM for call options; here we transform it into a form that allows use of the fast Fourier transform method suggested by Carr & Madan. In particular, we consider Merton…
Dual optimization connects ERM-fDR to normalization function.
problem Empirical risk minimization with f-divergence regularization.
method Dual formulation, Legendre-Fenchel transform, implicit function theorem, nonlinear ODE.
result Computational method to calculate normalization function efficiently.
Transformer model improves asset allocation by unifying forecasting and optimization.
problem Separation of forecasting and optimization leads to suboptimal portfolios.
method Signature Informed Transformer using path signatures and specialized attention.
result Direct minimization of Conditional Value at Risk improves performance.
New method estimates Schrödinger bridge potentials via empirical risk minimization.
problem Estimating Schrödinger bridge potentials from samples.
method Rewriting Schrödinger system as a fixed-point equation and estimating the potential via empirical risk minimization.
result Uniform concentration of empirical risk around population counterpart under sub-Gaussian assumptions.
Paper extends transfer learning for decision rules, improving treatment rule estimation.
problem Estimating optimal individualized treatment rules under changing conditions.
method Bayes decision rules and low-dimensional empirical risk minimization.
result Consistent estimators and risk bounds established under mild conditions.
Transformers recall from long distributions with statistical guarantees.
problem Designing Transformers that can recall from arbitrarily long, distributional contexts.
method Recast associative memory as probability measures, decomposing the task into recall and prediction.
result A shallow measure-theoretic Transformer learns the recall-and-predict map under spectral assumptions.
Study shows neural network parameters converge to ridgelet spectrum.
problem Characterization of local minima in over-parametrized neural networks.
method Developed ridgelet transform to analyze neural network parameters.
result Distribution of parameters converges to ridgelet spectrum.
In this paper we consider the problem of calculating the quantiles of a risky position, the dynamic of which is described as a continuous time regime-switching jump-diffusion, by using Fourier Transform methods. Furthermore, we study a classical option-based portfolio strategy which minimizes the Value-at-Risk of the h…
Over-parameterized CNNs show U-shaped test risk with depth increase.
problem Understanding the impact of depth on test risk in over-parameterized CNNs.
method Empirical image classification experiments and linear regression framework.
result Test risk is U-shaped with increasing depth in over-parameterized CNNs.
ERM with f-divergence regularization yields unique solution.
problem Optimizing empirical risk with f-divergence. method Mild conditions on f lead to unique optimal measure. result Equivalence of ERM-fDR to different f-divergence regularization. New regularization method reduces support of empirical risk minimization solutions.
problem Regularization in empirical risk minimization with relative entropy.
method Introduces Type-II regularization, characterizes solutions, analyzes properties of relative entropy.
result Type-II regularization collapses solution support into reference measure's support.
We consider the problem of minimizing a sum of clipped convex functions; applications include clipped empirical risk minimization and clipped control. While the problem of minimizing the sum of clipped convex functions is NP-hard, we present some heuristics for approximately solving instances of these problems. These h…
A generalization of expectiles for d-dimensional multivariate distribution functions is introduced. The resulting geometric expectiles are unique solutions to a convex risk minimization problem and are given by d-dimensional vectors. They are well behaved under common data transformations and the corresponding sample v…
Mixup improves model accuracy and calibration through data transformation and random perturbation.
problem Improving model accuracy and calibration in machine learning.
method Interprets Mixup as empirical risk minimization with data transformation and random perturbation.
result Mixup induces multiple known regularization schemes that prevent overfitting and overconfident predictions.
Enhances Transformers for better risk assessment in finance.
problem Transformer models lack sensitivity to extreme financial losses.
method Integrates Loss-at-Risk function with Value at Risk (VaR) and Conditional Value at Risk (CVaR).
result Improves risk prediction and management in financial datasets.
Paper presents ERM with f-divergence regularization and its properties.
problem Minimizing empirical risk with f-divergence constraints. method Introduces normalization function and solves ERM-fDR via ODE. result Characterizes difference between empirical risks and provides numerical algorithm.
The paper shows that benchmark-neutral pricing minimizes option prices.
problem Pricing extreme-maturity European put options on diversified indices.
method Benchmark-neutral pricing applied to a drifted time-transformed squared Bessel process.
result Benchmark-neutral price is the minimal possible price, risk-neutral price is more expensive.
Study excess risk in statistical inference with transformations.
problem Excess risk in estimating random variables from feature vectors and transformations.
method Characterize lossless transformations, develop test statistics, and information-theoretic bounds.
result Strongly consistent partitioning test statistic for lossless transformations.
Entropy asymmetry affects regularization in ERM, leading to biased solutions.
problem Analyzing the impact of relative entropy asymmetry in ERM regularization.
method Examined Type-I and Type-II ERM-RER, comparing their solutions and properties.
result Type-II ERM-RER regularization introduces a strong bias against training data.
Transform non-private e-values into differentially private ones.
problem Leaking sensitive data through non-private e-values.
method Developed a novel biased multiplicative noise mechanism.
result Differentially private e-values maintain strong statistical power and asymptotic equivalence to non-private ones.
Bayesian approach optimizes in-context learning for state space models.
problem Optimizing in-context learning for state space models.
method Bayesian optimal sequential prediction over latent sequence tasks.
result Bayesian optimal predictor converges to posterior predictive mean.
We derive bounds for a notion of adversarial risk, designed to characterize the robustness of linear and neural network classifiers to adversarial perturbations. Specifically, we introduce a new class of function transformations with the property that the risk of the transformed functions upper-bounds the adversarial r…
The standard risk minimization paradigm of machine learning is brittle when operating in environments whose test distributions are different from the training distribution due to spurious correlations. Training on data from many environments and finding invariant predictors reduces the effect of spurious features by co…
Transformer learns CoVaR from financial news, improving systemic risk forecasts.
problem Quantifying systemic financial risk using conditional Value-at-Risk (CoVaR).
method Transformer-based approach integrating financial news articles with market data.
result Transformer CoVaR improves out-of-sample forecasts and identifies market stress periods.
NKI integrates obfuscated datasets using nonlinear kernels for improved data collaboration.
problem Privacy-preserving data collaboration with reduced reconstruction risk.
method Formulates linear kernel integration, kernelizes it, and introduces graph regularization and centering constraints.
result NKI improves classification accuracy over existing linear integration methods under nonlinear dimensionality reduction.
Variance plays a crucial role in risk-sensitive reinforcement learning, and most risk measures can be analyzed via variance. In this paper, we consider two law-invariant risks as examples: mean-variance risk and exponential utility risk. With the aid of the state-augmentation transformation (SAT), we show that, the two…
Study finds formulas for minimal submanifolds using Möbius transformations.
problem Understanding minimal submanifolds in Euclidean space.
method Monotonicity formulas for minimal submanifolds involving Möbius transformations.
result Proved formulas for minimal submanifolds under Möbius transformations.
Linear Transformer Block combines MLP and linear attention for near-optimal ICL in linear regression.
problem Achieving near-optimal in-context learning (ICL) risk for linear regression with a Gaussian prior.
method Combines linear attention and MLP components in a Linear Transformer Block (LTB). Establishes correspondence with one-step gradient descent estimators (GDext−β). result LTB achieves nearly Bayes optimal ICL risk for linear regression with a Gaussian prior.
This study identifies financial risk paths in digital-transformed enterprises.
problem Identifying financial risks in digital-transformed enterprises.
method DEMATEL-ISM-MICMAC method.
result Political and economic environment affects enterprise's financial structure.
New dual formulation reduces generalization error for ERM-fDR.
problem Generalization error in constrained optimization problems.
method Introduces a dual formulation of ERM-fDR using Legendre-Fenchel transform and implicit function theorem.
result Explicit characterizations of generalization error for algorithms under mild conditions.
Different types of training data have led to numerous schemes for supervised classification. Current learning techniques are tailored to one specific scheme and cannot handle general ensembles of training data. This paper presents a unifying framework for supervised classification with general ensembles of training dat…
Study Transformer layers under cross-entropy training using mean field control.
problem Understanding the behavior of Transformer layers in cross-entropy training.
method Continuous-depth mean field control analysis, treating depth as time and layer parameters as controls.
result Derivation of a Pontryagin condition for the limiting population problem, involving the softmax residual.
Paper uses Time Series Transformer for bank stability prediction.
problem Predicting bank stability using complex financial data.
method Time Series Transformer model with self-attention mechanism.
result Time Series Transformer model outperforms other models in MSE and MAE.
Unified framework connects credit risk metrics with information theory.
problem Disconnection between industry-standard metrics and statistical theory.
method Unified information-theoretic framework, proving IV equals PSI, deriving standard errors, formalizing trade-off, automated binning with XGBoost.
result Unified framework connects IV and PSI, providing statistical foundation for metrics.
Algorithm finds optimal affine transformation to minimize overall distortion.
problem Minimizing distortion in affine transformations.
method Riemannian geometry approach to define and minimize distortion.
result Mean distorting transformation found for minimizing overall distortion.
In this paper we consider Fourier transform techniques to efficiently compute the Value-at-Risk and the Conditional Value-at-Risk of an arbitrary loss random variable, characterized by having a computable generalized characteristic function. We exploit the property of these risk measures of being the solution of an ele…
Paper proposes Multi-Transformer for more accurate stock volatility forecasts.
problem Accurate equity risk models needed for effective risk management.
method Introduces Multi-Transformer neural network architecture, adapted from Transformer models.
result Empirical results show Multi-Transformer leads to more accurate risk measures.
Quantum SVT reduces credit risk analysis costs.
problem Efficiently estimating credit risk metrics using quantum computing.
method Quantum Singular Value Transformation (QSVT) to reduce state preparation costs.
result Significant reduction in implementation costs for quantum credit risk analysis.
Transformers can cluster data from Gaussian mixtures without supervision.
problem Clustering data from Gaussian mixtures without labeled data.
method Theoretical analysis of attention-based layers, focusing on a simplified two-head attention layer and an identity matrix attention layer.
result Attention-based layers can align with true mixture centroids and adapt to input-specific distributions.
Enhanced Transformer models predict ETF portfolio performance by optimizing covariance and semi-covariance matrices.
problem Static covariance estimates fail to capture dynamic market fluctuations and non-linear correlations.
method Transformer-based models for real-time covariance and semi-covariance predictions.
result Portfolios optimized with semi-covariance matrix outperform those with standard covariance matrix, especially in volatile conditions.
New algorithm corrects bias in LDP-released data for better analysis.
problem Bias in data released under Local Differential Privacy (LDP).
method Inverse Weierstrass Private Stochastic Gradient Descent (IWP-SGD).
result Converges to true population risk minimizer at O(1/n) rate. The paper introduces a new method for risk measurement using weak optimal transport.
problem Risk measurement in insurance and financial contexts.
method Convex risk measures with weak optimal transport penalties, explicit representation via nonlinear transform, computational aspects, and approximations using neural networks.
result Explicit representation and computational methods for risk measures.
A new test evaluates risk estimation accuracy using probability integral transform.
problem Measuring the accuracy of financial market risk estimations.
method Probability Integral Transform (PIT) of ex post realized returns against ex ante probability distributions.
result The new test shows the importance of capturing the dynamic of financial markets.
Develops new methods to estimate treatment effects in survival data with competing risks.
problem Estimating treatment effects in survival data with competing risks.
method Censoring Unbiased Transformations (CUTs) for survival outcomes with and without competing risks.
result Consistent estimates of heterogeneous cumulative incidence effects and total effects using HTE learners.
This paper proves IRM minimizes o.o.d. risk under certain conditions.
problem Deep networks can fail to generalize to new domains with different distributions.
method Proves IRM minimizes o.o.d. risk through a bi-level optimization problem.
result IRM minimizes o.o.d. risk under specific conditions.
Paper classifies minimal graph transformations into new families of surfaces.
problem Classifying minimal graph transformations into new families of surfaces.
method Formulated and solved a coupled system of partial differential equations, reduced to solving an ordinary differential equation.
result Established rigorous equivalence to a modified problem for a harmonic function, yielding new families of minimal surfaces.
The paper analyzes the performance of empirical risk minimization for p-norm linear regression.
problem Empirical risk minimization on p-norm linear regression. method Analyzes performance under various conditions and moment assumptions.
result High probability excess risk bounds for empirical risk minimizer, matching asymptotic rates.