Separable losses are inconsistent for structured prediction models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this dissertation, we focus on several important problems in structured prediction. In structured prediction, the label has a rich intrinsic substructure, and the loss varies with respect to the predicted label and the true label pair. Structured SVM is an extension of binary SVM to adapt to such structured tasks. I…
A new framework learns differentiable structured losses from data.
Novel approach embeds loss tunnels in neural networks, revealing insights into their structure.
GCML preserves geometric structure in manifold clustering for diverse data types.
Structured entropy improves classification performance on structured targets.
Margin-based structured prediction commonly uses a maximum loss over all possible structured outputs \cite{Altun03,Collins04b,Taskar03}. In natural language processing, recent work \cite{Zhang14,Zhang15} has proposed the use of the maximum loss over random structured outputs sampled independently from some proposal dis…
We propose two structural models for stochastic losses given default which allow to model the credit losses of a portfolio of defaultable financial instruments. The credit losses are integrated into a structural model of default events accounting for correlations between the default events and the associated losses. We…
This work improves structured prediction by learning the balance between signal and random noise.
Linear-Core Surrogates combine fast optimization and statistical efficiency in classification and structured prediction.
Study predicts SGD test loss for structured features.
Study how data structure impacts classification performance in overparametrized models.
We consider the problem of training probabilistic conditional random fields (CRFs) in the context of a task where performance is measured using a specific loss function. While maximum likelihood is the most common approach to training CRFs, it ignores the inherent structure of the task's loss function. We describe alte…
We provide novel theoretical insights on structured prediction in the context of efficient convex surrogate loss minimization with consistency guarantees. For any task loss, we construct a convex surrogate that can be optimized via stochastic gradient descent and we prove tight bounds on the so-called "calibration func…
We present and study a partial-information model of online learning, where a decision maker repeatedly chooses from a finite set of actions, and observes some subset of the associated losses. This naturally models several situations where the losses of different actions are related, and knowing the loss of one action p…
The paper introduces a new loss function to prevent overfitting in semi-supervised graph networks.
We consider the problem of concurrent portfolio losses in two non-overlapping credit portfolios. In order to explore the full statistical dependence structure of such portfolio losses, we estimate their empirical pairwise copulas. Instead of a Gaussian dependence, we typically find a strong asymmetry in the copulas. Co…
Paper corrects Max-Margin loss for multi-label tasks.
A new framework for structured prediction on non-vectorial spaces.
Study compact Willmore surfaces without complex structure convergence, computing energy loss and geodesic lengths.
Extends LOSS invariant naturality to positive contact surgeries.
New tools reveal simple structure in complex hyperparameter loss surfaces near optima.
We present a fully-supervized method for learning to segment data structured by an adjacency graph. We introduce the graph-structured contrastive loss, a loss function structured by a ground truth segmentation. It promotes learning vertex embeddings which are homogeneous within desired segments, and have high contrast …
This work uses PAC-Bayes for structured prediction with ILE, yielding insights and algorithms.
Spectral images captured by satellites and radio-telescopes are analyzed to obtain information about geological compositions distributions, distant asters as well as undersea terrain. Spectral images usually contain tens to hundreds of continuous narrow spectral bands and are widely used in various fields. But the vast…
Forecastability measures predictive information across horizons.
New method uses FY loss for better inverse optimization.
This work proposes the Bregman-Tweedie classification model and analyzes the domain structure of the extended exponential function, an extension of the classic generalized exponential function with additional scaling parameter, and related high-level mathematical structures, such as the Bregman-Tweedie loss function an…
Operator-Valued Kernels (OVKs) and associated vector-valued Reproducing Kernel Hilbert Spaces provide an elegant way to extend scalar kernel methods when the output space is a Hilbert space. Although primarily used in finite dimension for problems like multi-task regression, the ability of this framework to deal with i…
Proves and tests methods for learning time-series with breaks.
We formalize and study the natural approach of designing convex surrogate loss functions via embeddings, for problems such as classification, ranking, or structured prediction. In this approach, one embeds each of the finitely many predictions (e.g.\ rankings) as a point in , assigns the original loss val…
We demonstrate that the gain/loss asymmetry observed for stock indices vanishes if the temporal dependence structure is destroyed by scrambling the time series. We also show that an artificial index constructed by a simple average of a number of individual stocks display gain/loss asymmetry - this allows us to explicit…
Unified framework for structured prediction with partial labelling.
We propose in this paper a general framework for deriving loss functions for structured prediction. In our framework, the user chooses a convex set including the output space and provides an oracle for projecting onto that set. Given that oracle, our framework automatically generates a corresponding convex and smooth l…
We solve a high-dimensional model where nonlinear autoencoders detect hidden structure missed by PCA.
We propose and analyze a regularization approach for structured prediction problems. We characterize a large class of loss functions that allows to naturally embed structured outputs in a linear space. We exploit this fact to design learning algorithms using a surrogate loss approach and regularization techniques. We p…
In this paper we study a class of insurance products where the policy holder has the option to insure of its annual Operational Risk losses in a horizon of years. This involves a choice of out of years in which to apply the insurance policy coverage by making claims against losses in the given year. The…
Study analyzes landscape complexity of empirical loss functions with correlated data.
In this work we consider the stochastic minimization of nonsmooth convex loss functions, a central problem in machine learning. We propose a novel algorithm called Accelerated Nonsmooth Stochastic Gradient Descent (ANSGD), which exploits the structure of common nonsmooth loss functions to achieve optimal convergence ra…
New framework for data-driven hyperparameter tuning with structured loss.
Over the past decades, numerous loss functions have been been proposed for a variety of supervised learning tasks, including regression, classification, ranking, and more generally structured prediction. Understanding the core principles and theoretical properties underpinning these losses is key to choose the right lo…
A new adversarial attack method using structured search and contextual bandits.
Framework for differentiating WFSTs for structured loss functions.
We propose SEARNN, a novel training algorithm for recurrent neural networks (RNNs) inspired by the "learning to search" (L2S) approach to structured prediction. RNNs have been widely successful in structured prediction applications such as machine translation or parsing, and are commonly trained using maximum likelihoo…
Neural network training relies on our ability to find "good" minimizers of highly non-convex loss functions. It is well-known that certain network architecture designs (e.g., skip connections) produce loss functions that train easier, and well-chosen training parameters (batch size, learning rate, optimizer) produce mi…
A robust loss for anomaly mitigation and unsupervised contamination classification
Deep neural networks for structured prediction using kernel-induced losses.
We develop a model for contagion in reinsurance networks by which primary insurers' losses are spread through the network. Our model handles general reinsurance contracts, such as typical excess of loss contracts. We show that simpler models existing in the literature--namely proportional reinsurance--greatly underesti…