Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

2.7%5.5%8.2%10.9% · Feb 202619922001200920182026
48 results for Source Reliability

Improves neural network performance by dynamically adjusting model weights based on source reliability.

problem Training neural networks on data from unreliable sources leads to poor performance.
method Dynamic re-weighting strategy using likelihood tempering to adjust model weights based on estimated source reliability.
result Significant improvement in model performance when trained on mixtures of reliable and unreliable data sources.

Paper tackles multi-source transfer learning with diverse labeling volume and reliability.

problem Challenges in multi-source transfer learning with diverse labeling volume and reliability.
method Combines domain similarity and source reliability through a new transfer learning method, and integrates distribution matching and uncertainty sampling in pool-based active learning.
result Demonstrates superior performance over state-of-the-art transfer learning methods.

rMFBO improves MFBO by making it robust to unreliable low-fidelity sources.

problem Optimizing expensive functions with unreliable low-fidelity approximations.
method rMFBO (robust MFBO) integrates a theoretical guarantee to make GP-based MFBO robust to unreliable sources.
result rMFBO outperforms earlier MFBO methods on unreliable sources.

Proposes incorporating noise sources in machine learning evaluation for more reliable conclusions.

problem Inadequate handling of nondeterminism in machine learning research leads to unreliable results.
method Uses linear mixed effects models (LMEMs) and generalized likelihood ratio tests (GLRT) to analyze performance evaluation scores and assess performance differences.
result Demonstrates how to incorporate various sources of noise and data properties into statistical significance testing and reliability analysis.

Study on reliability of latent reuse in diffusion models under distribution shift.

problem When can latent spaces from a source dataset be reused for a target dataset with different distributions?
method Considered a source-target setting with approximately low-dimensional datasets near different subspaces. Analyzed the target-domain score error due to principal-angle misalignment and target ambient noise.
result Latent reuse is reliable only if the source and target subspaces are close and the target ambient noise is not too amplified.

Bayesian approach for aggregating unreliable data sources to create accurate heatmaps.

problem Classifying regions with sparse, unreliable data from multiple sources.
method Bayesian Gaussian Process classifier that models reliability and bias of each data source.
result Reduces crowdsourced data needed and improves accuracy of heatmaps.

Method quantifies sensitivity of reliability analysis to uncertainty sources.

problem Computational expense in reliability analysis of complex models.
method Gaussian process surrogate model, active learning, sensitivity analysis.
result Reduces main source of error in estimating rare event probabilities.

New method for reliability analysis using multi-fidelity models.

problem Reliability analysis of complex systems with high computational costs.
method Adaptive Multi-fidelity Gaussian Process for Reliability Analysis (AMGPRA) with collective learning function (CLF).
result AMGPRA achieves similar or higher accuracy with reduced computational costs compared to state-of-the-art methods.

Neural networks combining multiple data sources can reverse preferences, affecting decision reliability.

problem Preference reversals in neural networks under pooled data.
method Formalized through Case-Based Decision Theory, analyzed Gram geometry, introduced regularization, and developed auditing methods.
result Pooled refitting can reverse shared preferences, and conditions for preserving preferences are derived.

Study evaluates cross-validation methods for clinical ECG classification, finding leave-source-out more reliable.

problem Overoptimistic cross-validation estimates for new patient sources.
method Empirical evaluation of K-fold and leave-source-out cross-validation methods.
result Leave-source-out cross-validation provides more reliable performance estimates.

Theoretical validation of linear PCA and ICA for accurate nonlinear BSS.

problem Blind source separation for high-dimensional nonlinear source mixtures.
method Theoretical validation of a cascade of linear PCA and ICA.
result Zero-element-wise-error nonlinear BSS is achieved under certain conditions.

Open-source Vizier optimizes complex systems for Google and beyond.

problem Optimizing large-scale systems with multiple objectives and constraints.
method Distributed, fault-tolerant, flexible API for blackbox optimization.
result OSS Vizier supports a wide range of optimization problems and is available as open-source.

This paper examines risks and uncertainties of changing data sources in machine learning for official statistics.

problem Risks and uncertainties associated with changing data sources in machine learning for official statistics.
method An overview of risks, causes, and repercussions of changing data sources, with a checklist of measures.
result Maintaining integrity, reliability, consistency, and relevance in official statistics.

New algorithm reduces age of information in wireless networks with unknown channel reliability.

problem Learning optimal source-channel pairs to minimize age of information in wireless networks.
method Introduces AoI regret, novel learning algorithm with bounded AoI regret.
result Developed a learning algorithm with O(1)O(1) AoI regret, improving upon Θ(logT)Θ(\log T).

A novel deep learning technique combines multiple modalities, improving performance.

problem Challenges in leveraging different modalities due to noise and conflicts.
method Proposes a deep neural network that multiplicatively combines information from different modalities.
result Consistent accuracy improvements on three multimodal classification tasks.

An important preprocessing step in most data analysis pipelines aims to extract a small set of sources that explain most of the data. Currently used algorithms for blind source separation (BSS), however, often fail to extract the desired sources and need extensive cross-validation. In contrast, their rarely used probab…

2018-03-23abs ↗pdf ↗

Change detection (CD) in time series data is a critical problem as it reveal changes in the underlying generative processes driving the time series. Despite having received significant attention, one important unexplored aspect is how to efficiently utilize additional correlated information to improve the detection and…

2016-03-31abs ↗pdf ↗

Unstructured data refers to information that does not have a predefined data model or is not organized in a pre-defined manner. Loosely speaking, unstructured data refers to text data that is generated by humans. In after-sales service businesses, there are two main sources of unstructured data: customer complaints, wh…

2016-07-26abs ↗pdf ↗

Motivation: Modelling methods that find structure in data are necessary with the current large volumes of genomic data, and there have been various efforts to find subsets of genes exhibiting consistent patterns over subsets of treatments. These biclustering techniques have focused on one data source, often gene expres…

2015-12-29abs ↗pdf ↗

This dissertation tackles challenges in reliable machine learning measurement.

problem Challenges in reproducibility, scalability, and uncertainty quantification in machine learning.
method Develops criteria for meaningful metrics and methodologies for scalable, reliable measurement.
result Provides methods for evaluating generative-AI systems and quantifying memorization.

A new method reduces energy consumption in machine learning by using multiple, less costly data sources.

problem High computational and energy costs in machine learning model training.
method Augmented Gaussian Process (AGP-MISO) with multi-source optimization.
result The AGP-MISO method reduces computational time and energy consumption compared to traditional approaches.

Algorithm finds function contours using multiple approximations.

problem Locating contours of expensive-to-evaluate functions.
method Uses multiple biased and noisy approximations to locate contours efficiently by maximizing entropy reduction.
result Maximizes reduction of contour entropy per unit cost.

The paper proposes a method to adapt models from source to target domains by calibrating their predictive uncertainties.

problem Inferring class labels for unlabeled target domain given a related labeled source dataset.
method The approach involves calibrating predictive uncertainties quantified as Renyi entropy, using variational Bayes learning and sample variance regularization.
result The proposed method effectively adapts models across three domain-adaptation tasks.

New framework improves model reliability under distribution shifts.

problem Lack of formal guarantees connecting shift magnitude to prediction reliability in TTA methods.
method Develops a PAC-Bayesian framework interpreting MMD-balls as credal sets.
result Establishes generalization bounds and provides epistemic uncertainty quantification.

Quantum classification robustness improved via quantum hypothesis testing.

problem Vulnerability of quantum classification algorithms to input perturbations.
method Formalized link between quantum hypothesis testing and robustness, developed practical protocols.
result Tight robustness condition independent of noise source (natural or adversarial).

Estimates calibration error under label shift without labels.

problem Ensuring model reliability in the face of dataset shift without access to labels.
method Importance re-weighting of the labeled source distribution to estimate calibration error under label shift.
result Effective and reliable CE estimation with respect to the shifted target distribution.

Non-availability of reliable and sustainable electric power is a major problem in the developing world. Renewable energy sources like solar are not very lucrative in the current stage due to various uncertainties like weather, storage, land use among others. There also exists various other issues like mis-commitment of…

2017-11-08abs ↗pdf ↗

New method uses neural networks to identify sources from limited data in complex systems.

problem Identifying sources from noisy and limited data in high-dimensional systems.
method Calibrating deep neural network surrogates to ensemble simulations and using Bayesian optimization for source identification.
result Reliable source identification with uncertainty quantification using limited data and auxiliary processes.

mfEGRA uses active learning to efficiently locate failure boundaries in reliability analysis.

problem Prohibitive cost of reliability analysis using Monte Carlo sampling for high-fidelity models.
method Develops a multifidelity active learning method using data-driven adaptively refined surrogates.
result Significant computational savings (46-48%) compared to single-fidelity EGRA.

CAMul forecasts with calibrated and accurate multi-view time-series data.

problem Combining diverse data sources for reliable time-series forecasting.
method CAMul integrates multi-modal data views dynamically, assigning importance based on context.
result CAMul outperforms state-of-the-art models by 25% in accuracy and calibration.

DeepTrust uses NLP to quickly identify and verify financial anomalies on Twitter.

problem Unreliable information in financial markets leading to unexpected price changes.
method Machine learning for anomaly detection, NLP for information retrieval and reliability assessment.
result DeepTrust outperforms baseline classifiers in identifying financial anomalies.

In the modern era, abundant information is easily accessible from various sources, however only a few of these sources are reliable as they mostly contain unverified contents. We develop a system to validate the truthfulness of a given statement together with underlying evidence. The proposed system provides supporting…

2018-02-15abs ↗pdf ↗

Study finds machine learning interpretations are often unstable and unreliable.

problem Reliability of machine learning interpretations in high-stakes domains.
method Stability study on global interpretations using tabular data.
result Popular interpretation methods are frequently unstable, less stable than predictions, and not associated with prediction accuracy.