Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

4182123164 · Jun 202019922001200920182026
48 results for evidence extraction

System automates identification of cancer drug repurposing from PubMed.

problem Manual extraction of cancer drug repurposing evidence from scientific publications is infeasible.
method NLP pipeline including querying, filtering, entity extraction, classification, and study type classification.
result Automated system extracts cancer drug repurposing evidence from PubMed abstracts.

This article introduces a framework to estimate the value of evidence-based decision making.

problem Lack of empirical tools to assess the value of evidence-based decision making and optimize statistical precision.
method Empirical framework using parametric and nonparametric empirical Bayes methods.
result The value of statistical evidence depends on how organizations translate it into policy decisions.

GEAR uses graphs to integrate and reason over multiple evidence for fact verification.

problem Fact verification requires integrating and reasoning over multiple pieces of evidence.
method GEAR employs a graph-based framework to transfer information among evidence and uses BERT for improved performance.
result GEAR achieves a promising test FEVER score of 67.10% on a large-scale benchmark dataset.

The paper extracts structured data from physician-patient conversations, reducing clerical burden.

problem Mining insights from physician-patient conversations for electronic health record documentation.
method Created a dataset of transcripts and summaries, extracted noteworthy utterances, and improved model performance.
result Extracting noteworthy utterances significantly boosts model performance for recognizing diagnoses and RoS abnormalities.

Gradient-based method extracts slow features from high-dimensional data.

problem Extracting meaningful low-dimensional features from high-dimensional, temporally varying data.
method Power Slow Feature Analysis (PowerSFA) using gradient-based training of differentiable architectures.
result PowerSFA effectively extracts meaningful low-dimensional features in various data types.

Proposes OpenKI for better web-scale knowledge extraction and alignment.

problem Combining OpenIE and KB for web-scale knowledge extraction and alignment.
method Instance-level inference using neighborhood information from KB and OpenIE extractions, with attention mechanisms.
result Significantly improves performance on OpenIE extractions and semi-structured data.

Proposes a Bayesian approach to explain, justify, and quantify uncertainty in DNNs.

problem Lack of transparency and confidence in DNNs for critical applications.
method Bayesian approach to extract explanations, justifications, and uncertainty estimates from black box DNNs.
result Improves interpretability and reliability of DNNs, validated on CIFAR-10.

The paper explores methods to explain brain tumor segmentation models and extract visualizations.

problem Improving interpretability of deep learning models for brain tumor segmentation.
method Exploring techniques to explain brain tumor segmentation models and extract visualizations.
result Brain tumor segmentation networks learn human-understandable disentangled concepts and use a top-down approach.

A new method using mean shift clustering speeds up Bayesian evidence calculation.

problem Difficulty in Nested Sampling algorithm convergence and systematic errors.
method Mean shift cluster recognition method integrated into NestedFit.
result Significant reduction in computation time and uncertainty of Bayesian evidence.

This paper uses feature preprocessing and RRL to automate profitable financial trading.

problem Automating profitable financial trading strategies.
method Feature preprocessing (PCA, DWT) followed by Recurrent Reinforcement Learning (RRL).
result The proposed strategy is effective, robust, and mitigates RRL's drawbacks.

Automates U.S. visa petition document classification and RFE response generation.

problem Manual effort in organizing visa petition documents and responding to RFEs.
method Ensemble of image and text classifiers for document categorization and text classifier for RFE evidence identification.
result Achieves considerable accuracy in automated responses while reducing processing time.

Survey of large language models in financial prediction and trading.

problem Improving predictability and robustness of financial predictions and trading decisions.
method Task-centered taxonomy, review of empirical evidence, design patterns, benchmarks, and challenges analysis.
result Improved predictability and robustness of financial predictions and trading decisions through large language models.

Two models incorporate market microstructure noise into asset pricing and option valuation.

problem Effect of market microstructure noise on asset pricing and option valuation.
method Developed two models: a continuous-time Black-Scholes-Merton model and a discrete binomial tree model.
result Extracted coefficients to quantify noise impact on volatility and drift.

This research examines how data transformations affect adversarial robustness in recurrent neural networks.

problem Adversarial examples reduce machine learning accuracy, especially in high-dimensional datasets.
method Analysis of feature selection, dimensionality reduction, and trend extraction techniques on recurrent neural networks.
result Data transformations may increase vulnerability to adversarial samples, but only if they approximate intrinsic dimensionality and maintain manifold coverage.

Paper introduces methods to automatically generate SOAP notes from patient-physician conversations.

problem Burden of creating digital SOAP notes by physicians.
method Cluster2Sent algorithm for summarizing patient-physician conversations.
result Cluster2Sent algorithm outperforms existing methods by 8 ROUGE-1 points.

KM-GPT automates IPD reconstruction from KM plots with high accuracy and scalability.

problem Manual digitization of IPD from KM plots is error-prone and lacks scalability.
method KM-GPT integrates advanced image preprocessing, multi-modal reasoning, and iterative reconstruction algorithms.
result KM-GPT generates high-quality IPD without manual input or intervention, achieving superior accuracy.

Manifold regularization improves DNN performance in speech recognition.

problem Training deep neural networks for speech recognition is challenging.
method Manifold learning is used to regularize DNNs, preserving speech feature relationships.
result Manifold regularized DNNs reduce word error rate by up to 37%.

This paper automates mining of COVID-19 scholarly articles using machine learning.

problem Time-consuming and impractical manual extraction of relevant COVID-19 research articles.
method Used machine learning approaches, specifically clustering and parallel one-class support vector machines (OCSVMs), on the CORD-19 dataset.
result Parallel OCSVMs outperform other methods for both original and reduced feature space.

FIBS extracts relevant features from IBTSs for classification.

problem Classifying interval-based temporal sequences (IBTSs) using common algorithms is challenging.
method FIBS extracts features from IBTSs based on relative frequency and temporal relations, incorporating a filter-based selection strategy to avoid irrelevant features.
result FIBS effectively represents IBTSs for classification algorithms, providing similar or better accuracy compared to state-of-the-art competitors.

We make a precision test of a recently proposed conjecture relating Chern-Simons gauge theory to topological string theory on the resolution of the conifold. First, we develop a systematic procedure to extract string amplitudes from vacuum expectation values (vevs) of Wilson loops in Chern-Simons gauge theory, and then…

2000-04-27abs ↗pdf ↗

RB-Modulation trains free diffusion models without external adapters.

problem Training-free personalization of diffusion models with style and content control.
method Stochastic optimal control with a style descriptor and cross-attention aggregation.
result Precise content and style extraction and control without external adapters.

Improved price bounds for multi-asset derivatives using market option data.

problem Creating robust price bounds for multi-asset derivatives under market-implied dependence.
method Extracting inter-asset dependence information from market option prices and applying modified martingale optimal transport.
result Improved price bounds for multi-asset derivatives, demonstrating relevance and tractability.

We construct a financial "Turing test" to determine whether human subjects can differentiate between actual vs. randomized financial returns. The experiment consists of an online video-game (http://arora.ccs.neu.edu) where players are challenged to distinguish actual financial market returns from random temporal permut…

2010-02-24abs ↗pdf ↗

BERT captures linguistic features in separate semantic and syntactic subspaces.

problem Understanding how transformer models like BERT represent linguistic features internally.
method Qualitative and quantitative investigations of BERT's internal representations.
result Evidence of a fine-grained geometric representation of word senses and syntactic representations.

Study shows mutual information can reward structure learning agents without expert systems.

problem Designing rewards for structure learning agents in natural language environments.
method Revisited Information Theory of unsupervised induction of phrase-structure grammars, using random sets of linguistic samples.
result Empirical evidence that simulated semantic structures can be distinguished from random ones by mutual information among their constituents.

Study reveals a hidden cost in derivatives markets through option-implied discount factors.

problem The hidden cost in derivatives markets, not visible in price space.
method Minute-level NBBO data on options, reduced-form specification linking carry gap to implementation risk, trading frictions, and financial conditions.
result An annualized carry gap exists, linked to implementation risk and financial conditions.

A non-trivial probability structure is evident in the binary data extracted from the up/down price movements of very high frequency data such as tick-by-tick data for USD/JPY. In this paper, we analyze the Sony bank USD/JPY rates, ignoring the small deviations from the market price. We then show there is a similar non-…

2005-09-30abs ↗pdf ↗

The Efficient Market Hypothesis (EMH) is widely accepted to hold true under certain assumptions. One of its implications is that the prediction of stock prices at least in the short run cannot outperform the random walk model. Yet, recently many studies stressing the psychological and social dimension of financial beha…

2013-10-20abs ↗pdf ↗

The paper explores how to handle uncertain evidence in probabilistic models.

problem Handling uncertain evidence in probabilistic models and stochastic simulators.
method The paper considers distributional evidence, Jeffrey's rule, and virtual evidence as methods for interpreting uncertain evidence.
result The paper provides guidelines on how to account for uncertain evidence and highlights the importance of careful consideration.

FOCA method prevents co-adaptation between feature extractor and classifier.

problem Co-adaptation between feature extractor and classifier degrades neural network performance.
method FOCA method uses randomly-generated, weak classifiers to optimize feature extractor without explicit co-adaptation.
result FOCA features form a point-like distribution within the same class under special conditions.

Generative models improve image probability estimation but lack interpretability.

problem Lack of interpretability in generative models for natural image distributions.
method Extracted explicit probability density estimates from GANs and analyzed latent representations.
result Natural image density functions are difficult to interpret.

Complex systems are typically represented by large ensembles of observations. Correlation matrices provide an efficient formal framework to extract information from such multivariate ensembles and identify in a quantifiable way patterns of activity that are reproducible with statistically significant frequency compared…

2011-06-02abs ↗pdf ↗