Some statistical models are specified via a data generating process for which the likelihood function cannot be computed in closed form. Standard likelihood-based inference is then not feasible but the model parameters can be inferred by finding the values which yield simulated data that resemble the observed data. Thi…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New ensemble models classify mouse movement trajectories to assess survey question difficulty.
A new restart criterion for k-means++ improves clustering quality and adapts to data difficulty.
Fast approximate nearest neighbor (NN) search in large databases is becoming popular. Several powerful learning-based formulations have been proposed recently. However, not much attention has been paid to a more fundamental question: how difficult is (approximate) nearest neighbor search in a given data set? And which …
Measures difficulty of predictions to improve deep learning models.
New measure quantifies task difficulty for machine learning models.
The difficulty of classification affects the weight matrices' heavy tail appearance in deep learning networks.
Entity Linking (EL) is the task of automatically identifying entity mentions in a piece of text and resolving them to a corresponding entity in a reference knowledge base like Wikipedia. There is a large number of EL tools available for different types of documents and domains, yet EL remains a challenging task where t…
One of the central difficulties of settling the -bounded curvature conjecture for the Einstein -Vacuum equations is to be able to control the causal structure of spacetimes with such limited regularity. In this paper we show how to circumvent this difficulty by showing that the geometry of null hypersurfaces of En…
Improved PAC-Bayesian bounds by considering example difficulty.
Bayesian model identifies skill difficulties and student subgroups in engineering education.
A new method estimates uncertainty without explicit prediction models.
Time-continuous emotion prediction has become an increasingly compelling task in machine learning. Considerable efforts have been made to advance the performance of these systems. Nonetheless, the main focus has been the development of more sophisticated models and the incorporation of different expressive modalities (…
Traditional statistical analysis requires that the analysis process and data are independent. By contrast, the new field of adaptive data analysis hopes to understand and provide algorithms and accuracy guarantees for research as it is commonly performed in practice, as an iterative process of interacting repeatedly wi…
Deep networks prioritize easier examples over harder ones, leading to faster training.
Structural health monitoring is a condition-based field of study utilised to monitor infrastructure, via sensing systems. It is therefore used in the field of aerospace engineering to assist in monitoring the health of aerospace structures. A difficulty however is that in structural health monitoring the data input is …
Gradient boosting with randomized trees reduces discontinuities and complexity.
Mathematical Reinforcement Learning faces a 'Two-Hump' problem due to sparse rewards and a scarcity of intermediate 'hard-but-solvable' instances.
Constructs metrics with Q-curvature on manifolds with singularities.
This paper applies secure multi-party computation to K-means clustering to protect private data.
Item Response Theory (IRT) aims to assess latent abilities of respondents based on the correctness of their answers in aptitude test items with different difficulty levels. In this paper, we propose the -IRT model, which models continuous responses and can generate a much enriched family of Item Characteristic Cur…
Increasingly complex generative models are being used across disciplines as they allow for realistic characterization of data, but a common difficulty with them is the prohibitively large computational cost to evaluate the likelihood function and thus to perform likelihood-based statistical inference. A likelihood-free…
Deep learning based medical image diagnosis has shown great potential in clinical medicine. However, it often suffers two major difficulties in practice: 1) only limited labeled samples are available due to expensive annotation costs over medical images; 2) labeled images may contain considerable label noises (e.g., mi…
Data imbalance remains one of the most widespread problems affecting contemporary machine learning. The negative effect data imbalance can have on the traditional learning algorithms is most severe in combination with other dataset difficulty factors, such as small disjuncts, presence of outliers and insufficient numbe…
Two-sample feature selection is the problem of finding features that describe a difference between two probability distributions, which is a ubiquitous problem in both scientific and engineering studies. However, existing methods have limited applicability because of their restrictive assumptions on data distributoins …
Curriculum Learning - the idea of teaching by gradually exposing the learner to examples in a meaningful order, from easy to hard, has been investigated in the context of machine learning long ago. Although methods based on this concept have been empirically shown to improve performance of several learning algorithms, …
Transformers learn to adapt to different task difficulties and resist distribution shifts.
The paper explores when to prioritize easy or hard samples in learning tasks.
CLOPS improves deep learning for continuous physiological data.
Many real-world applications reveal difficulties in learning classifiers from imbalanced data. The rising big data era has been witnessing more classification tasks with large-scale but extremely imbalance and low-quality datasets. Most of existing learning methods suffer from poor performance or low computation effici…
Proposes a simple framework to balance task difficulty in multi-task learning.
Symmetry in inverse problems leads to multiple solutions, but breaking symmetry helps deep learning.
We propose a multi-scale stochastic volatility model in which a fast mean-reverting factor of volatility is built on top of the Heston stochastic volatility model. A singular pertubative expansion is then used to obtain an approximation for European option prices. The resulting pricing formulas are semi-analytic, in th…
Salzmann's legacy in mathematics documented.
Paper shows training can improve GCN performance without changing architecture.
This work sets theoretical limits on meta-learning performance.
Proposes a context-aware approach to deep autoencoder novelty detection.
This study introduces balanced DRPS and OrderedLogitNN for better QDE of discrete-level questions.
This work builds the connection between the regularity theory of optimal transportation map, Monge-Ampère equation and GANs, which gives a theoretic understanding of the major drawbacks of GANs: convergence difficulty and mode collapse. According to the regularity theory of Monge-Ampère equation, if the support of the …
New method for Bayesian neural networks reduces inference difficulty.
Learning shrinks hard tail, improving inference performance.
Study predicts firm defaults using machine learning on Italian credit data.
Artificial neural networks are simple and efficient machine learning tools. Defined originally in the traditional setting of simple vector data, neural network models have evolved to address more and more difficulties of complex real world problems, ranging from time evolving data to sophisticated data structures such …
Score matching is a popular method for estimating unnormalized statistical models. However, it has been so far limited to simple, shallow models or low-dimensional data, due to the difficulty of computing the Hessian of log-density functions. We show this difficulty can be mitigated by projecting the scores onto random…
Despite the significant advances in recent years, Generative Adversarial Networks (GANs) are still notoriously hard to train. In this paper, we propose three novel curriculum learning strategies for training GANs. All strategies are first based on ranking the training images by their difficulty scores, which are estima…
The economic crisis in Argentina around year 2002 provides a unique opportunity for Econophysics studies. The available data on individual income are analyzed to show that they correspond to non stationary states. However, the rather restricted size of the data survey imposes difficulties that must be overcome through …
Study reveals differences in label shift problem difficulty in supervised vs. unsupervised settings.
We revisit the elegant observation of T. Cover '65 which, perhaps, is not as well-known to the broader community as it should be. The first goal of the tutorial is to explain---through the prism of this elementary result---how to solve certain sequence prediction problems by modeling sets of solutions rather than the u…