The paper tackles online learning with two types of losses and shows it's impossible without certain assumptions.
problem Online learning with primary and secondary losses where the secondary loss is bounded by a linear threshold.
method Analyzes the feasibility of achieving low regret with respect to the primary loss while keeping the secondary loss within a linear threshold.
result Achieving the goal is impossible without bounded variance assumption on the secondary loss.
Fuses ITRs for primary and secondary outcomes to minimize harm.
problem Learn an ITR maximizing primary outcome while minimizing harm to secondary outcomes.
method Introduces fusion penalty to encourage similar recommendations for different outcomes. Two algorithms estimate the ITR using surrogate loss functions.
result Agreement rate between primary and secondary optimal ITRs converges faster than ignoring secondary outcomes.
Aux-NAS uses auxiliary labels to improve primary task performance without extra inference cost.
problem Improving primary task performance using auxiliary labels without increasing inference cost.
method Architecture-based approach with a flexible asymmetric structure for primary and auxiliary tasks, using Neural Architecture Search (NAS) to evolve networks with only primary-to-auxiliary connections.
result Achieves improved performance on multiple tasks without increasing inference cost.
DPC uses physics and neural nets to solve SDEs.
problem Solving stochastic differential equations with missing physics.
method Physics-data fusion with conditional maximum mean discrepancy (CMMD) loss.
result DPC achieves highly accurate solutions on benchmark examples.
Fine-tuning LLMs improves capability but harms safety, study finds.
problem Balancing capability and safety in LLM fine-tuning.
method Theoretical framework and numerical experiments for two safety-aware fine-tuning strategies.
result Characterization of fundamental limits of safety-capability trade-off in LLM fine-tuning.
Learning with auxiliary tasks can improve the ability of a primary task to generalise. However, this comes at the cost of manually labelling auxiliary data. We propose a new method which automatically learns appropriate labels for an auxiliary task, such that any supervised learning task can be improved without requiri…
Hill-ADAM optimizes loss landscapes by exploring state space deterministically.
problem Escaping local minima in loss landscapes.
method Hill-ADAM alternates between minimizing and maximizing error to explore the loss space.
result Hill-ADAM finds the global minimum state in loss landscapes.
This research analyzes the consistency of convex and nonconvex surrogate losses for adversarially robust classification.
problem Ensuring classifiers are robust to adversarial perturbations.
method Analysis of convex and nonconvex surrogate losses through the lens of calibration.
result No convex surrogate loss is calibrated with respect to the adversarial 0-1 loss for linear models, but nonconvex losses can be calibrated under certain conditions.
We develop a model for contagion in reinsurance networks by which primary insurers' losses are spread through the network. Our model handles general reinsurance contracts, such as typical excess of loss contracts. We show that simpler models existing in the literature--namely proportional reinsurance--greatly underesti…
Bayesian neural networks improved with scalable approximate inference.
problem Performing approximate Bayesian inference in complex models like neural networks.
method Two models: primary for prediction, secondary for posterior approximation; optimised via gradient descent on posterior predictive distribution.
result Approach scales better than MCMC and more expressive than VIs, without adversarial training.
Paper examines risk measure expansions under FGM dependence, improving accuracy at extreme levels.
problem Capturing higher-order tail behavior and dependence effects in risk measures.
method Second-order asymptotic expansions using extreme value theory and regular variation theory.
result Second-order approximations reduce approximation errors, especially at extreme confidence levels.
Informer model with GMADL loss outperforms benchmarks in high frequency Bitcoin trading.
problem Developing automated trading strategies for high frequency Bitcoin data.
method Informer architecture with RMSE, GMADL, and Quantile loss functions.
result Informer model with GMADL loss function outperforms benchmarks in trading outcomes.
Statistical inference is considered for variables of interest, called primary variables, when auxiliary variables are observed along with the primary variables. We consider the setting of incomplete data analysis, where some primary variables are not observed. Utilizing a parametric model of joint distribution of prima…
A significant advance in accelerating neural network training has been the development of normalization methods, permitting the training of deep models both faster and with better accuracy. These advances come with practical challenges: for instance, batch normalization ties the prediction of individual examples with o…
Improved similarity search in embeddings using InfoNCE loss.
problem Improving similarity search in embedding models trained by contrastive learning.
method Introduced a new continuity bound for InfoNCE loss via Gâteaux differentiation, preserving the averaging effect of negative samples.
result Demonstrated that the averaging effect of k negative samples in InfoNCE loss carries over to stabilisation of generalisation error as k grows. Proposes a probabilistic framework for smart contract risk quantification.
problem Quantifying financial risk of smart contract cyber attacks and failures.
method Probabilistic graph-theoretical framework using bond percolation models.
result Analytical results and numerical examples for aggregate loss distribution.
Over the past few years many research efforts have been devoted to the field of affect analysis. Various approaches have been proposed for: i) discrete emotion recognition in terms of the primary facial expressions; ii) emotion analysis in terms of facial Action Units (AUs), assuming a fixed expression intensity; iii) …
Study blockchain's impact on primary financial market challenges.
problem Challenges of blockchain in securities issuance and trading.
method Hybrid method combining interviews and surveys.
result Complex due diligence, mismatch, and difficult monitoring are significant challenges.
SGD updates align with a low-rank subspace but do not lead to further loss reduction.
problem Understanding the training dynamics of deep neural networks, particularly the role of the dominant subspace.
method Exploring whether neural networks can be trained within the dominant subspace of the loss Hessian.
result SGD updates, when projected onto the dominant subspace, do not decrease the training loss further, suggesting spurious alignment.
Study confirms conjectures for topologically slice knots' concordance group.
problem Primary decomposition conjectures for knot concordance groups.
method Use of amenable L2-signatures, Ozsváth-Szabó d-invariants, and Némethi's Heegaard Floer homology results. result Smooth concordance group of topologically slice knots has large subgroup with true primary decomposition.
A new geometric model for V1 hypercolumns combines symplectic and spherical models.
problem Understanding the structure of V1 hypercolumns in the visual cortex.
method A differential geometric model based on conformal geometry.
result Combines features of symplectic and spherical models of hypercolumns.
Augmenting a neural network with memory that can grow without growing the number of trained parameters is a recent powerful concept with many exciting applications. We propose a design of memory augmented neural networks (MANNs) called Labeled Memory Networks (LMNs) suited for tasks requiring online adaptation in class…
We introduce new forecast encompassing tests for the risk measure Expected Shortfall (ES). The ES currently receives much attention through its introduction into the Basel III Accords, which stipulate its use as the primary market risk measure for the international banking regulation. We utilize joint loss functions fo…
Classifies Real primary Hopf surfaces and their associated groups.
problem Classifying Real primary Hopf surfaces and their associated groups.
method Complete classification up to Real biholomorphisms and equivariant diffeomorphisms.
result Detailed description of groups associated with Real primary Hopf surfaces.
In most machine learning applications, classification accuracy is not the primary metric of interest. Binary classifiers which face class imbalance are often evaluated by the Fβ score, area under the precision-recall curve, Precision at K, and more. The maximization of many of these metrics can be expressed as a con…
Improved online learning with time-varying constraints for complex domains.
problem Constrained online convex optimization with time-varying constraints.
method Constructing a composite surrogate loss and using the online Frank-Wolfe method.
result Novel regret and cumulative constraint violation bounds for strongly convex losses.
Flaky performance found in GNN SSL on RDBs, leading to worse linear evaluation.
problem Downstream task performances of GNN SSL on RDBs are poor.
method Proposed InfoNode to maximize mutual information between initial and final node representations.
result InfoNode improves GNN SSL performance on RDBs, supporting conjecture of conflict between SSL and GNN message passing.
Super Learner combines dynamic predictions from various models to improve survival estimates.
problem Challenges in obtaining optimal survival estimates for liver failure risk.
method Super Learner framework combining machine learning and statistical procedures.
result Super Learner outperformed individual models in primary biliary cholangitis application.
A new method detects concept drift without true labels.
problem Detecting concept drift in unsupervised settings.
method Student-teacher learning paradigm for drift detection.
result The method outperforms state-of-the-art approaches in experiments.
New findings on knot concordance show limitations to primary decompositions.
problem Understanding the structure of knot concordance groups.
method Analyzing homomorphisms and polynomial factorizations of Alexander polynomials.
result Primary decompositions of topologically slice knots cannot exist.
Bayesian framework for policy learning in decision problems.
problem Maximizing expected welfare in decision-making problems.
method Loss-based Bayesian updating and squared-loss surrogate for welfare maximization.
result General Bayes posterior over decision rules with Gaussian pseudo-likelihood interpretation.
New loss function reduces outage probability in ML-assisted resource allocation.
problem Minimizing outage probability in ML-assisted resource allocation systems.
method Developed a novel loss function and trained an ML model to address the outage probability challenge.
result Exact and asymptotic expressions for the system's outage probability were established.
Translation-based embedding models have gained significant attention in link prediction tasks for knowledge graphs. TransE is the primary model among translation-based embeddings and is well-known for its low complexity and high efficiency. Therefore, most of the earlier works have modified the score function of the Tr…
Understanding the morphological changes of primary neuronal cells induced by chemical compounds is essential for drug discovery. Using the data from a single high-throughput imaging assay, a classification model for predicting the biological activity of candidate compounds was introduced. The image recognition model wh…
Proposes rCV to preserve serial correlations in time-series models.
problem Loss of serial correlations in cross-validation for time-series models.
method Form k folds, generate k new partial time-series, reconstruct using imputation/smoothing, build primary models, evaluate performance.
result Avoids loss of serial correlations and preserves non-stationarity in predictions.
Learning with a primary objective, such as softmax cross entropy for classification and sequence generation, has been the norm for training deep neural networks for years. Although being a widely-adopted approach, using cross entropy as the primary objective exploits mostly the information from the ground-truth class f…
We address feature interpretation and reproducibility issues in dense nets, proposing a modified loss function.
problem Feature interpretation and reproducibility issues in dense nets.
method Proposed a modified loss function to circumvent basis collapse.
result Substantially concise nets with 100x fewer parameters and lower MSE loss.
The paper introduces PD learning to improve deep learning theory.
problem Lack of theoretical understanding in deep learning model fitting and generalization.
method Proposes a PD learning framework to analyze optimization and generalization mechanisms of deep learning.
result Established theoretical guarantees on optimizability and derived generalization error bounds.
We study holomorphic locally homogeneous geometric structures modelled on line bundles over the projective line. We classify these structures on primary Hopf surfaces. We write out the developing map and holonomy morphism of each of these structures explicitly on each primary Hopf surface.
Paper proposes a method to speed up discrete diffusion models by distilling many steps into few.
problem Challenges in capturing dependencies between elements in discrete diffusion models.
method Proposes 'mixture' models and loss functions to distill many sampling steps into few.
result Effective in distilling pretrained discrete diffusion models across image and language domains.
Large data collections required for the training of neural networks often contain sensitive information such as the medical histories of patients, and the privacy of the training data must be preserved. In this paper, we introduce a dropout technique that provides an elegant Bayesian interpretation to dropout, and show…
In this study, we generalize double tangent bundles to double jet bundles. We present a secondary vector bundle structure on a 1-jet of a vector bundle. We show that 1-jet of a vector bundle carries two vector bundle structures, namely primary and secondary structures. We also show that the manifold charts induced by p…
Improved PAC-Bayesian bounds by considering example difficulty.
problem Improving generalization bounds in machine learning.
method Introducing a modified excess risk that leverages example difficulty to reduce variance and tighten PAC-Bayesian bounds.
result Tighter PAC-Bayesian generalization bounds for machine learning models.
Paper proposes a deep hedging method for Bermudan swaptions to manage residual profit and loss.
problem Real-world market conditions differ from ideal assumptions in traditional hedging methods, leading to residual profit and loss.
method Deep hedging framework applied to Bermudan swaptions, allowing flexible risk measures and hedge strategies.
result Effective residual profit and loss management demonstrated through numerical analysis.
Automorphisms of Kodaira surfaces are shown to be affine transformations.
problem Characterizing automorphisms of Kodaira surfaces.
method Analyzing lifts to the universal cover and conditions on affine transformations.
result Precise description of Kodaira surfaces' automorphism groups.
In several natural language tasks, labeled sequences are available in separate domains (say, languages), but the goal is to label sequences with mixed domain (such as code-switched text). Or, we may have available models for labeling whole passages (say, with sentiments), which we would like to exploit toward better po…
Randomly chosen primary hidden units and derived secondary units reduce neural network complexity.
problem Large number of hidden units in neural networks.
method Introducing primary and secondary hidden units with random weights for primary units and derived weights for secondary units.
result Significant reduction in the number of hidden units without compromising accuracy.
The Softmax function is used in the final layer of nearly all existing sequence-to-sequence models for language generation. However, it is usually the slowest layer to compute which limits the vocabulary size to a subset of most frequent types; and it has a large memory footprint. We propose a general technique for rep…