This paper uses reference priors to improve deep learning models with unlabeled and labeled data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Leveraging reference-only samples for two-sample testing under size asymmetry
Synthetic reference strings are as effective as real ones for training citation parsing models.
The paper introduces a method for forecasting corporate sales growth using multiple reference variables.
The paper analyzes how conformal prediction works with contaminated reference data.
Method learns neural network to overestimate reference function with guarantees.
New method detects if data points were used in training models with low cost and high power.
The exploitation of large-scale population data has the potential to improve healthcare by discovering and understanding patterns and trends within this data. To enable high throughput analysis of cardiac imaging data automatically, a pipeline should comprise quality monitoring of the input images, segmentation of the …
This paper describes a reference architecture for self-maintaining systems that can learn continually, as data arrives. In environments where data evolves, we need architectures that manage Machine Learning (ML) models in production, adapt to shifting data distributions, cope with outliers, retrain when necessary, and …
New method learns cell trajectories from multiple snapshots.
New method generates synthetic time series paths with more flexibility.
Enhances anomaly detection using multiple reference datasets.
Choosing a reference group in Oaxaca-Blinder decomposition can reverse conclusions.
Proposes a new framework to optimize portfolios with reduced estimation errors.
We discuss several uses of blockchain (and, more generally, distributed ledger) technologies outside of cryptocurrencies with a pragmatic view. We mostly focus on three areas: the role of coin economies for what we refer to as data malls (specialized data marketplaces); data provenance (a historical record of data and …
A new method detects distribution shifts faster than existing CTMs.
This paper improves model training by using a reference model to guide target model training.
In this paper, we bound the error induced by using a weighted skeletonization of two data sets for computing a two sample test with kernel maximum mean discrepancy. The error is quantified in terms of the speed in which heat diffuses from those points to the rest of the data, as well as how at the weights on the refere…
Self-supervised methods learn from noisy data alone, useful for imaging problems.
The Minimal Learning Machine (MLM) is a nonlinear supervised approach based on learning a linear mapping between distance matrices computed in the input and output data spaces, where distances are calculated using a subset of points called reference points. Its simple formulation has attracted several recent works on e…
We construct master spaces for oriented torsion free sheaves coupled with morphisms into a fixed reference sheaf. These spaces are projective varieties endowed with a natural $\C^*$-action. The fixed point set of this action contains the moduli space of semistable oriented torsion free sheaves and the quot scheme assoc…
A theoretical study is presented for a simple linear classifier called reference distance estimator (RDE), which assigns the weight of each feature j as P(r|j)-P(r), where r is a reference feature relevant to the target class y. The analysis shows that if r performs better than random guess in predicting y and is condi…
New methods improve LLM preference optimization by intelligently weighting multiple reference models.
This paper develops a method to select a reference contract for multi-contract quoting to minimize execution risk.
The paper proposes to analyze a data set of Finnish ranks of academic publication channels with Extreme Learning Machine (ELM). The purpose is to introduce and test recently proposed ELM-based mislabel detection approach with a rich set of features characterizing a publication channel. We will compare the architecture,…
Study replicates reference-dependent preferences impact on risk-return trade-off in Chinese stock market.
A salient approach to interpretable machine learning is to restrict modeling to simple models. In the Bayesian framework, this can be pursued by restricting the model structure and prior to favor interpretable models. Fundamentally, however, interpretability is about users' preferences, not the data generation mechanis…
Methods for learning to search for structured prediction typically imitate a reference policy, with existing theoretical guarantees demonstrating low regret compared to that reference. This is unsatisfactory in many applications where the reference policy is suboptimal and the goal of learning is to improve upon it. Ca…
Data of practical interest - such as personal records, transaction logs, and medical histories - are sequential collections of events relevant to a particular source entity. Recent studies have attempted to link sequences that represent a common entity across data sets to allow more comprehensive statistical analyses a…
This research identifies flaws in drift detection methods and creates adversarial data streams to exploit them.
Real-world datasets are often biased with respect to key demographic factors such as race and gender. Due to the latent nature of the underlying factors, detecting and mitigating bias is especially challenging for unsupervised machine learning. We present a weakly supervised algorithm for overcoming dataset bias for de…
This paper solves the multiple reference model problem in RLHF with exact solutions and sample complexity guarantees.
We present a weakly-supervised data augmentation approach to improve Named Entity Recognition (NER) in a challenging domain: extracting biomedical entities (e.g., proteins) from the scientific literature. First, we train a neural NER (NNER) model over a small seed of fully-labeled examples. Second, we use a reference s…
Reference metrics are used to define the differential structure on multicube representations of manifolds, i.e., they provide a simple and practical way to define what it means globally for tensor fields and their derivatives to be continuous. This paper introduces a general procedure for constructing reference metrics…
Conformal Alignment ensures trustworthy outputs from foundation models.
This article considers the quasi-local conserved quantities with respect to a reference spacetime with a cosmological constant. We follow the approach developed by the authors in [25,26,7] and define the quasi-local energy as differences of surface Hamiltonians. The ground state for the gravitational energy is taken to…
A new approach is presented to describe the change in the statistics of the log return distribution of financial data as a function of the timescale. To this purpose a measure is introduced, which quantifies the distance of a considered distribution to a reference distribution. The existence of a small timescale regime…
In this paper the introduction of notion of reference vector paves the way for a combination of classical and social approaches in the framework of referential preferences given by matrix groups. It is shown that individual demand issue from rational decision does not depend on that reference.
The paper proposes a method to improve sales forecasts by selecting optimal reference classes.
This work extends entropic optimal transport to non-product reference couplings, focusing on Gaussian cases.
Study validates ML-UQ calibration statistics using simulated reference values.
This work studies an explicit embedding of the set of probability measures into a Hilbert space, defined using optimal transport maps from a reference probability density. This embedding linearizes to some extent the 2-Wasserstein space, and enables the direct use of generic supervised and unsupervised learning algorit…
The purpose of this paper is to discuss how topology and geometry provide, in many instances, the connective tissue that enables logical comprehension. We illustrate this theme with many examples including Venn diagrams, knot diagrams, knot-logical diagrams and an arrow of reference that elucidates self-reference and G…
Probabilistic generative models provide a powerful framework for representing data that avoids the expense of manual annotation typically needed by discriminative approaches. Model selection in this generative setting can be challenging, however, particularly when likelihoods are not easily accessible. To address this …
Optimizes a portfolio for an investor preferring accepted securities over a reference security.
Unstructured data refers to information that does not have a predefined data model or is not organized in a pre-defined manner. Loosely speaking, unstructured data refers to text data that is generated by humans. In after-sales service businesses, there are two main sources of unstructured data: customer complaints, wh…
Optimal order execution strategies for brokers under reference benchmarks.
Active-GRPO improves molecular optimization by actively deciding when to imitate or self-improve.