Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

2885758631,150 · Jun 202019922001200920172026
48 results for reference data

This paper uses reference priors to improve deep learning models with unlabeled and labeled data.

problem Improving deep learning models with limited labeled data and unlabeled data from the same or related tasks.
method Develops and applies generalizations of reference priors for deep networks to exploit unlabeled and labeled data.
result Demonstrates new semi-supervised learning and pretraining methods for transfer learning.

Synthetic reference strings are as effective as real ones for training citation parsing models.

problem Lack of training data for citation parsing, especially with deep neural networks.
method Trained Grobid with human-labelled and synthetically created reference strings, and evaluated retraining and out-of-sample data impact.
result Synthetic and real reference strings are equally effective for training Grobid, with retraining improving performance.

The paper introduces a method for forecasting corporate sales growth using multiple reference variables.

problem Forecasting corporate sales growth with multiple reference variables.
method Reference class selection using rank-based algorithms and principal components analysis for data dimension reduction.
result Dimension reduced variables with past sales growth rates and operating margins perform well in forecasting.

The paper analyzes how conformal prediction works with contaminated reference data.

problem The impact of contamination on the validity and power of conformal prediction methods.
method The paper analyzes the impact of contamination on the validity of conformal methods and proposes a data-cleaning framework to enhance power.
result The proposed data-cleaning framework can effectively enhance power while maintaining type-I error control.

New method detects if data points were used in training models with low cost and high power.

problem Detecting if a particular data point was used in training a model.
method Fine-grained modeling of null hypothesis in likelihood ratio tests, leveraging reference models and population data.
result RMIA has superior test power compared to prior methods, even at extremely low false positive rates.

This paper describes a reference architecture for self-maintaining systems that can learn continually, as data arrives. In environments where data evolves, we need architectures that manage Machine Learning (ML) models in production, adapt to shifting data distributions, cope with outliers, retrain when necessary, and …

2019-03-12abs ↗pdf ↗

New method generates synthetic time series paths with more flexibility.

problem Restrictions in generating synthetic paths using Brownian reference.
method Introduces Triangular-Reference Schrödinger Bridges (TR-SBTS) for time series generation.
result Generates synthetic paths with more flexibility in stochastic volatility and correlated noise.

Choosing a reference group in Oaxaca-Blinder decomposition can reverse conclusions.

problem The choice of reference group in Oaxaca-Blinder decomposition can lead to different conclusions.
method The study uses the Oaxaca-Blinder decomposition to investigate how the choice of reference group affects the results.
result The Oaxaca-Blinder decomposition can yield different conclusions based on the choice of reference group.

We discuss several uses of blockchain (and, more generally, distributed ledger) technologies outside of cryptocurrencies with a pragmatic view. We mostly focus on three areas: the role of coin economies for what we refer to as data malls (specialized data marketplaces); data provenance (a historical record of data and …

2018-02-21abs ↗pdf ↗

This paper improves model training by using a reference model to guide target model training.

problem Improving generalization and data efficiency in model training.
method DRRho risk minimization framework based on Distributionally Robust Optimization (DRO).
result DRRho risk minimization improves generalization and data efficiency compared to training without a reference model.

In this paper, we bound the error induced by using a weighted skeletonization of two data sets for computing a two sample test with kernel maximum mean discrepancy. The error is quantified in terms of the speed in which heat diffuses from those points to the rest of the data, as well as how at the weights on the refere…

2018-12-11abs ↗pdf ↗

We construct master spaces for oriented torsion free sheaves coupled with morphisms into a fixed reference sheaf. These spaces are projective varieties endowed with a natural $\C^*$-action. The fixed point set of this action contains the moduli space of semistable oriented torsion free sheaves and the quot scheme assoc…

1996-07-17abs ↗pdf ↗

A theoretical study is presented for a simple linear classifier called reference distance estimator (RDE), which assigns the weight of each feature j as P(r|j)-P(r), where r is a reference feature relevant to the target class y. The analysis shows that if r performs better than random guess in predicting y and is condi…

2013-08-18abs ↗pdf ↗

New methods improve LLM preference optimization by intelligently weighting multiple reference models.

problem Improving LLM preference optimization with multiple reference models.
method Introducing four new weighting strategies for multiple-reference preference optimization.
result All four new weighting strategies outperform current methods on preference accuracy.

This paper develops a method to select a reference contract for multi-contract quoting to minimize execution risk.

problem Minimizing execution risk in multi-contract quoting sequences.
method Develops a diagnostic framework using order-flow Hawkes forecasts and CLF to select a stable reference contract.
result Event-history and LOB-state signals offer complementary views for reference-contract selection.

The paper proposes to analyze a data set of Finnish ranks of academic publication channels with Extreme Learning Machine (ELM). The purpose is to introduce and test recently proposed ELM-based mislabel detection approach with a rich set of features characterizing a publication channel. We will compare the architecture,…

2019-12-19abs ↗pdf ↗

Study replicates reference-dependent preferences impact on risk-return trade-off in Chinese stock market.

problem Impact of reference-dependent preferences on risk-return trade-off in Chinese stock market.
method Utilized CGO proxy, econometric techniques (Dependent Double Sorting, Fama-MacBeth regressions), and data from 1995-2024.
result Reference-dependent preferences have a weaker or absent positive risk-return relationship in the Chinese market.

Methods for learning to search for structured prediction typically imitate a reference policy, with existing theoretical guarantees demonstrating low regret compared to that reference. This is unsatisfactory in many applications where the reference policy is suboptimal and the goal of learning is to improve upon it. Ca…

2015-02-08abs ↗pdf ↗

This research identifies flaws in drift detection methods and creates adversarial data streams to exploit them.

problem The challenge of detecting data distribution changes (drift) in real-time systems.
method Developed adversarial data streams to show weaknesses in existing drift detection schemes.
result Demonstrated that common drift detection methods can be fooled by adversarial data streams.

Real-world datasets are often biased with respect to key demographic factors such as race and gender. Due to the latent nature of the underlying factors, detecting and mitigating bias is especially challenging for unsupervised machine learning. We present a weakly supervised algorithm for overcoming dataset bias for de…

2019-10-26abs ↗pdf ↗

This paper solves the multiple reference model problem in RLHF with exact solutions and sample complexity guarantees.

problem Limitations of single reference models in aligning LLMs with human feedback.
method Integrates multiple reference models into RLHF frameworks, addressing theoretical challenges with exact solutions and sample complexity guarantees.
result First exact solution to the multiple reference model problem in reverse KL-regularized RLHF.

Conformal Alignment ensures trustworthy outputs from foundation models.

problem Ensuring outputs from foundation models align with human values in high-stakes tasks.
method A framework that trains an alignment predictor using reference data to select trustworthy outputs.
result Conformal Alignment accurately identifies trustworthy outputs via lightweight training over moderate reference data.

This article considers the quasi-local conserved quantities with respect to a reference spacetime with a cosmological constant. We follow the approach developed by the authors in [25,26,7] and define the quasi-local energy as differences of surface Hamiltonians. The ground state for the gravitational energy is taken to…

2016-03-09abs ↗pdf ↗

A new approach is presented to describe the change in the statistics of the log return distribution of financial data as a function of the timescale. To this purpose a measure is introduced, which quantifies the distance of a considered distribution to a reference distribution. The existence of a small timescale regime…

2005-09-30abs ↗pdf ↗

In this paper the introduction of notion of reference vector paves the way for a combination of classical and social approaches in the framework of referential preferences given by matrix groups. It is shown that individual demand issue from rational decision does not depend on that reference.

2014-02-14abs ↗pdf ↗

The paper proposes a method to improve sales forecasts by selecting optimal reference classes.

problem Improving forecasts of sales growth exposed to behavioural bias.
method Finding optimal reference classes for each company based on specific predictors and matching forecast distributions to actual sales.
result The past operating margins are strong predictors for future sales distributions.

This work extends entropic optimal transport to non-product reference couplings, focusing on Gaussian cases.

problem Finding a diffuse coupling between two measures with non-product reference couplings.
method Reduction of the entropic optimal transport problem to a matrix optimization problem.
result Complete description of the solution for non-product reference couplings, including primal and dual variables.

Study validates ML-UQ calibration statistics using simulated reference values.

problem Validation of ML-UQ calibration statistics is lacking due to lack of predefined reference values.
method Proposed validation workflow using simulated reference values derived from synthetic datasets.
result Some statistics, like CC and ENCE, are overly sensitive to generative distribution choice.

The purpose of this paper is to discuss how topology and geometry provide, in many instances, the connective tissue that enables logical comprehension. We illustrate this theme with many examples including Venn diagrams, knot diagrams, knot-logical diagrams and an arrow of reference that elucidates self-reference and G…

2015-08-25abs ↗pdf ↗

Probabilistic generative models provide a powerful framework for representing data that avoids the expense of manual annotation typically needed by discriminative approaches. Model selection in this generative setting can be challenging, however, particularly when likelihoods are not easily accessible. To address this …

2015-11-14abs ↗pdf ↗

Optimizes a portfolio for an investor preferring accepted securities over a reference security.

problem Investor preference for a set of securities over a reference security with constraints.
method Mean-variance optimization with Sharpe Ratio performance measurement.
result Derives an optimal portfolio that maximizes returns while minimizing risk.

Unstructured data refers to information that does not have a predefined data model or is not organized in a pre-defined manner. Loosely speaking, unstructured data refers to text data that is generated by humans. In after-sales service businesses, there are two main sources of unstructured data: customer complaints, wh…

2016-07-26abs ↗pdf ↗

Optimal order execution strategies for brokers under reference benchmarks.

problem Maximizing broker's utility of excess profit-and-loss subject to reference strategies.
method Formulated as a utility maximization problem, optimal strategies derived in closed form.
result General reference strategies can be approximated by piece-wise linear combinations of IS and TC orders.

Active-GRPO improves molecular optimization by actively deciding when to imitate or self-improve.

problem Training robust and efficient molecular optimization models with large language models.
method Active-GRPO combines imitation and reinforcement learning, upgrading references and policies dynamically.
result Improves molecular optimization performance, achieving statistically significant gains.