Machine learning models trained on indirect data labels can fail on real-world examples.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Weakly-supervised learning is a paradigm for alleviating the scarcity of labeled data by leveraging lower-quality but larger-scale supervision signals. While existing work mainly focuses on utilizing a certain type of weak supervision, we present a probabilistic framework, learning from indirect observations, for learn…
PLRM synthesizes labels from mismatched sources for better training sets.
Paper tackles RUL prediction with scarce data using indirect supervision.
Unified framework for learning with indirect supervision signals.
It is important to learn various types of classifiers given training data with noisy labels. Noisy labels, in the most popular noise model hitherto, are corrupted from ground-truth labels by an unknown noise transition matrix. Thus, by estimating this matrix, classifiers can escape from overfitting those noisy labels. …
Deep learning based task systems normally rely on a large amount of manually labeled training data, which is expensive to obtain and subject to operator variations. Moreover, it does not always hold that the manually labeled data and the unlabeled data are sitting in the same distribution. In this paper, we alleviate t…
New methods resolve conflicting treatment effect estimates in health tech assessments.
Machine learning has become pervasive in multiple domains, impacting a wide variety of applications, such as knowledge discovery and data mining, natural language processing, information retrieval, computer vision, social and health informatics, ubiquitous computing, etc. Two essential problems of machine learning are …
A new method simplifies noisy data filtering for CNNs.
Although deep learning has been applied to successfully address many data mining problems, relatively limited work has been done on deep learning for anomaly detection. Existing deep anomaly detection methods, which focus on learning new feature representations to enable downstream anomaly detection methods, perform in…
DynaCor detects noisy labels by learning from corrupted training signals.
This paper learns prior models from indirect data efficiently.
Study stability of trading strategy under market perturbations.
Existing methods for CWS usually rely on a large number of labeled sentences to train word segmentation models, which are expensive and time-consuming to annotate. Luckily, the unlabeled data is usually easy to collect and many high-quality Chinese lexicons are off-the-shelf, both of which can provide useful informatio…
Despite alarm over the reliance of machine learning systems on so-called spurious patterns, the term lacks coherent meaning in standard statistical frameworks. However, the language of causality offers clarity: spurious associations are due to confounding (e.g., a common cause), but not direct or indirect causal effect…
Estimates causal effects using machine learning for binary treatment and mediator.
The relationship between international trade and foreign direct investment (FDI) is one of the main features of globalization. In this paper we investigate the effects of FDI on trade from a network perspective, since FDI takes not only direct but also indirect channels from origin to destination countries because of f…
AI detects 38% NFT trades likely manipulated, improving on indirect methods.
New method estimates corporate default probabilities using indirect data.
In this paper, we propose a simple model referred as Contradistinguisher (CTDR) for unsupervised domain adaptation whose objective is to jointly learn to contradistinguish on unlabeled target domain in a fully unsupervised manner along with prior knowledge acquired by supervised learning on an entirely different domain…
Financial markets are exposed to systemic risk, the risk that a substantial fraction of the system ceases to function and collapses. Systemic risk can propagate through different mechanisms and channels of contagion. One important form of financial contagion arises from indirect interconnections between financial insti…
Researchers show how to secretly train models with hidden data, detect usage with high confidence.
We address the problem of gauging the influence exerted by a given country on the global trade market from the viewpoint of complex networks. In particular, we apply the PWP method for computing indirect influences on the world trade network.
Reinforcement learning (RL) algorithms have been successfully applied to a range of challenging sequential decision making and control tasks. In this paper, we classify RL into direct and indirect RL according to how they seek the optimal policy of the Markov decision process problem. The former solves the optimal poli…
Study shows cooperation can improve everyone's market efficiency.
End-to-end algorithm for controlling bilinear systems with probabilistic noise.
Driven by the goal to enable sleep apnea monitoring and machine learning-based detection at home with small mobile devices, we investigate whether interpretation-based indirect knowledge transfer can be used to create classifiers with acceptable performance. Interpretation-based indirect knowledge transfer means that a…
Unified framework for estimating indirect effects in observational studies with unmeasured confounding.
This study compares direct and indirect methods for estimating own funds in life insurance, finding indirect methods more effective under realistic asset-liability coupling.
New method for learning indirectly through control variables.
Indirect attacks can fool graph classifiers even with poisoned neighbors.
This paper tackles structure learning in indirect observations of Gaussian and non-Gaussian random vectors.
The objective optimization of medical imaging systems requires full characterization of all sources of randomness in the measured data, which includes the variability within the ensemble of objects to-be-imaged. This can be accomplished by establishing a stochastic object model (SOM) that describes the variability in t…
Develops exact and invariant study-based decompositions for network meta-analysis.
Paper develops models to forecast private equity fund cash flows.
The study tackles indirect discrimination in insurance pricing models.
We study the ever more integrated and ever more unbalanced trade relationships between European countries. To better capture the complexity of economic networks, we propose two global measures that assess the trade integration and the trade imbalances of the European countries. These measures are the network (or indire…
Paper proposes an algorithm to learn DAGs with indirect dependencies.
New benchmark PVR tests neural network reasoning about indirection.
This paper studies communication efficiency in federated learning by optimizing the sum-rate-distortion function for indirect multiterminal source coding.
We introduce the anti-profile Support Vector Machine (apSVM) as a novel algorithm to address the anomaly classification problem, an extension of anomaly detection where the goal is to distinguish data samples from a number of anomalous and heterogeneous classes based on their pattern of deviation from a normal stable c…
In this paper, we consider the variational regularization of manifold-valued data in the inverse problems setting. In particular, we consider TV and TGV regularization for manifold-valued data with indirect measurement operators. We provide results on the well-posedness and present algorithms for a numerical realizatio…
Study relaxes identification assumptions for natural direct effects in non-randomized settings.
Motivated by the need to audit complex and black box models, there has been extensive research on quantifying how data features influence model predictions. Feature influence can be direct (a direct influence on model outcomes) and indirect (model outcomes are influenced via proxy features). Feature influence can also …
In structured prediction problems where we have indirect supervision of the output, maximum marginal likelihood faces two computational obstacles: non-convexity of the objective and intractability of even a single gradient computation. In this paper, we bypass both obstacles for a class of what we call linear indirectl…
We present Vision-based Navigation with Language-based Assistance (VNLA), a grounded vision-language task where an agent with visual perception is guided via language to find objects in photorealistic indoor environments. The task emulates a real-world scenario in that (a) the requester may not know how to navigate to …
Extends causal inference to hidden mediators with proxies.