Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

2995988971,196 · Jun 202019922001200920172026
48 results for indirect data labeling

Machine learning models trained on indirect data labels can fail on real-world examples.

problem Validity issues in machine learning when target labels are indirectly defined.
method Identification of problematic datasets and models using a general procedure.
result Machine learning models trained on indirect data labels will fail on real-world examples.

Weakly-supervised learning is a paradigm for alleviating the scarcity of labeled data by leveraging lower-quality but larger-scale supervision signals. While existing work mainly focuses on utilizing a certain type of weak supervision, we present a probabilistic framework, learning from indirect observations, for learn…

2019-10-10abs ↗pdf ↗

PLRM synthesizes labels from mismatched sources for better training sets.

problem Creating labeled training sets is a major challenge in machine learning.
method PLRM uses probabilistic modeling to synthesize labels from indirect supervision sources with different output spaces.
result PLRM outperforms baselines by 2%-9% on various tasks.

Paper tackles RUL prediction with scarce data using indirect supervision.

problem Predicting RUL with indirect supervision and scarce time series data.
method Unified framework called parameterized static regression, handling data scarcity without interpolation.
result Competitive performance in prediction accuracy with simulated data scarcity.

Unified framework for learning with indirect supervision signals.

problem Learning from indirect supervision signals when gold labels are missing or costly.
method Developed a unified theoretical framework for multi-class classification with variable supervision.
result Introduced the concept of separation to characterize learnability and generalization bounds.

It is important to learn various types of classifiers given training data with noisy labels. Noisy labels, in the most popular noise model hitherto, are corrupted from ground-truth labels by an unknown noise transition matrix. Thus, by estimating this matrix, classifiers can escape from overfitting those noisy labels. …

2018-05-21abs ↗pdf ↗

New methods resolve conflicting treatment effect estimates in health tech assessments.

problem Conflicting conclusions from different sponsors analyzing the same data.
method Arbitrated indirect treatment comparisons (ArMAIC) targeting a common target population.
result Estimates treatment effects in a common target population, resolving the MAIC paradox.

Machine learning has become pervasive in multiple domains, impacting a wide variety of applications, such as knowledge discovery and data mining, natural language processing, information retrieval, computer vision, social and health informatics, ubiquitous computing, etc. Two essential problems of machine learning are …

2017-05-08abs ↗pdf ↗

Although deep learning has been applied to successfully address many data mining problems, relatively limited work has been done on deep learning for anomaly detection. Existing deep anomaly detection methods, which focus on learning new feature representations to enable downstream anomaly detection methods, perform in…

2019-11-19abs ↗pdf ↗

DynaCor detects noisy labels by learning from corrupted training signals.

problem Label noise in real-world datasets hinders model generalization.
method DynaCor introduces label corruption to indirectly simulate noisy labels and learns to distinguish clean from noisy instances.
result DynaCor outperforms state-of-the-art competitors in noisy label detection.

Study stability of trading strategy under market perturbations.

problem Dynamic stability of trading strategy under market changes.
method Established reverse conjugacy characterizations, proved continuity and convergence of indirect utility process.
result Continuity and first-order convergence of indirect utility process under market perturbations.

Estimates causal effects using machine learning for binary treatment and mediator.

problem Estimating direct and indirect quantile treatment effects under selection-on-observables.
method Double/debiased machine learning estimators based on efficient score functions.
result Uniform consistency and asymptotic normality of effect estimators.

The relationship between international trade and foreign direct investment (FDI) is one of the main features of globalization. In this paper we investigate the effects of FDI on trade from a network perspective, since FDI takes not only direct but also indirect channels from origin to destination countries because of f…

2017-05-05abs ↗pdf ↗

New method estimates corporate default probabilities using indirect data.

problem Lack of direct default rate data for corporate companies.
method Modeling default probability dynamics using Bank of Russia overdue debt data.
result Validated method produces trustworthy default probability series.

In this paper, we propose a simple model referred as Contradistinguisher (CTDR) for unsupervised domain adaptation whose objective is to jointly learn to contradistinguish on unlabeled target domain in a fully unsupervised manner along with prior knowledge acquired by supervised learning on an entirely different domain…

2019-09-08abs ↗pdf ↗

Researchers show how to secretly train models with hidden data, detect usage with high confidence.

problem Protecting training data from traceability in large language models.
method Gradient-based optimization to learn secret sequences absent from training data.
result Secret sequences can be learned by models without performance degradation, detectable with high confidence.

We address the problem of gauging the influence exerted by a given country on the global trade market from the viewpoint of complex networks. In particular, we apply the PWP method for computing indirect influences on the world trade network.

2014-11-27abs ↗pdf ↗

Reinforcement learning (RL) algorithms have been successfully applied to a range of challenging sequential decision making and control tasks. In this paper, we classify RL into direct and indirect RL according to how they seek the optimal policy of the Markov decision process problem. The former solves the optimal poli…

2019-12-23abs ↗pdf ↗

End-to-end algorithm for controlling bilinear systems with probabilistic noise.

problem Controlling bilinear systems with noisy data.
method Proposes an end-to-end algorithm using statistical learning theory and robust controller design.
result Derived finite sample identification error bounds and structurally suitable for control.

Driven by the goal to enable sleep apnea monitoring and machine learning-based detection at home with small mobile devices, we investigate whether interpretation-based indirect knowledge transfer can be used to create classifiers with acceptable performance. Interpretation-based indirect knowledge transfer means that a…

2019-03-06abs ↗pdf ↗

Unified framework for estimating indirect effects in observational studies with unmeasured confounding.

problem Challenges in evaluating indirect effects due to unmeasured confounding and unethical exposures.
method Developed a unified identification and estimation framework using proximal causal inference.
result Unified identification and estimation of PIIE and causal effect of an intervening variable in settings with pervasive unmeasured confounding.

This study compares direct and indirect methods for estimating own funds in life insurance, finding indirect methods more effective under realistic asset-liability coupling.

problem Computing own funds for life insurers using direct and indirect methods in a risk-neutral pricing framework.
method Introduced a novel family of mixed estimators including both direct and indirect methods, integrated into a control variate framework for variance reduction.
result The indirect method is more effective under realistic asset-liability coupling, but neither method is universally superior.

This paper tackles structure learning in indirect observations of Gaussian and non-Gaussian random vectors.

problem Learning the graphical structure of random vectors indirectly observed through a sensing matrix and corrupted noise.
method Parametric and non-parametric approaches for Gaussian and non-Gaussian distributions, respectively.
result Correct graphical structure can be recovered under indefinite sensing systems with insufficient samples.

Develops exact and invariant study-based decompositions for network meta-analysis.

problem Lack of exact contribution decompositions in network meta-analysis.
method Contrast-space projection formulation of NMA, study-based definition of direct and indirect evidence.
result Exact covariance-aware decompositions of NMA estimator into direct and indirect contributions.

Paper develops models to forecast private equity fund cash flows.

problem Limited literature on illiquid alternative asset cash flow forecasting.
method Develops benchmark model and two novel approaches (direct vs. indirect) using LSTM/GRU models and macroeconomic indicators.
result Direct model performs better and aligns with actual cash flows, but indirect model's performance is less clear.

This paper studies communication efficiency in federated learning by optimizing the sum-rate-distortion function for indirect multiterminal source coding.

problem Indirect multiterminal source coding in federated learning where edge devices send noisy gradients to the server.
method Analyzes the rate region for the quadratic vector Gaussian CEO problem under unbiased estimator and derives an explicit formula for the sum-rate-distortion function.
result Derives an explicit formula for the sum-rate-distortion function in the special case of identical gradients over edge devices.

We introduce the anti-profile Support Vector Machine (apSVM) as a novel algorithm to address the anomaly classification problem, an extension of anomaly detection where the goal is to distinguish data samples from a number of anomalous and heterogeneous classes based on their pattern of deviation from a normal stable c…

2013-01-15abs ↗pdf ↗

In this paper, we consider the variational regularization of manifold-valued data in the inverse problems setting. In particular, we consider TV and TGV regularization for manifold-valued data with indirect measurement operators. We provide results on the well-posedness and present algorithms for a numerical realizatio…

2018-04-27abs ↗pdf ↗

Study relaxes identification assumptions for natural direct effects in non-randomized settings.

problem Identifying causal direct effects under unmeasured confounding.
method Developed relaxed conditions for identifying natural direct effects in non-randomized settings.
result Identified natural direct effect under unmeasured confounding conditions.

In structured prediction problems where we have indirect supervision of the output, maximum marginal likelihood faces two computational obstacles: non-convexity of the objective and intractability of even a single gradient computation. In this paper, we bypass both obstacles for a class of what we call linear indirectl…

2016-08-10abs ↗pdf ↗

Study short-term wind power and speed predictions using machine learning.

problem Accurate short-term wind power and speed predictions for energy systems.
method Combining numerical weather prediction models with local observations, using machine learning for variable selection and forecasting.
result Improved wind power and speed predictions for 4-hour ahead using machine learning.