Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

1.6%3.3%4.9%6.5% · Jun 202019922001200920182026
48 results for code readability

Context2Name predicts natural names from minified code, improving code readability.

problem Minified code makes it hard to understand natural names in JavaScript.
method Combines static analysis and neural networks to predict natural names.
result Successfully predicts 47.5% of all minified identifiers in real-world code.

The simulator is an R package that streamlines the process of performing simulations by creating a common infrastructure that can be easily used and reused across projects. Methodological statisticians routinely write simulations to compare their methods to preexisting ones. While developing ideas, there is a temptatio…

2016-06-30abs ↗pdf ↗

SPX optimizes multiple graph drawing metrics for better readability.

problem Graph drawing algorithms often optimize one metric at a time, leading to suboptimal layouts.
method Introduces Stress-Plus-X (SPX) framework that optimizes stress, crossings, angles, and upwardness simultaneously.
result SPX achieves results close to state-of-the-art algorithms that optimize metrics individually.

SigD2 reduces noisy rules in rule-based classifiers for better accuracy and readability.

problem Redundant and noisy rules in rule-based classifiers reduce model accuracy and readability.
method Two-stage pruning strategy and ensemble methods (bagging and boosting) to reduce noise and improve model performance.
result SigD2 and ACboost ensemble models outperform state-of-the-art classifiers in terms of accuracy and rule count.

A novel memory mechanism for reinforcement learning agents that stores past events in human-readable language.

problem Lack of interpretability in reinforcement learning agent's memory mechanisms.
method Uses CLIP to associate visual inputs with language tokens, then feeds these tokens to a pretrained language model.
result Significantly faster convergence on challenging continuous recognition tasks.

IdBench benchmarks semantic representations of identifiers, revealing strengths and weaknesses.

problem Evaluating semantic representations of identifiers in source code.
method Created a benchmark using developer ratings, evaluated natural language and source code embeddings, and compared lexical string distance functions.
result No single technique provides a satisfactory representation of semantic similarities, but ensemble models can improve performance.

Shai-am simplifies ML for finance, solving code structure and scalability issues.

problem Challenges in integrating ML for investment strategies, including code structure and scalability.
method Integrates a Python framework with modern open-source technologies to manage containerized pipelines and unified interfaces.
result Facilitates collaborative work in quantitative finance by enhancing reusability and readability.

Torch-Struct simplifies structured prediction for deep learning.

problem Difficulty in integrating structured prediction algorithms with deep learning frameworks.
method Develops a library (Torch-Struct) that integrates structured prediction with vectorized, auto-differentiation-based frameworks.
result Significant performance gains over fast baselines and cross-algorithm efficiency.

Team QCRI-MIT detects hyperpartisan news with 72.9% accuracy.

problem Detecting hyperpartisan news from biased political content.
method Logistic regression model using engineered features from propaganda detection.
result Significant performance improvements with better feature pre-processing.

Solves the challenge of retrieving item-specific financial information from Form 10-Q filings.

problem Retrieving item-specific information from Form 10-Q filings with varying formats and machine-readable hierarchy.
method Complements a rule-based algorithm with a Convolutional Neural Network (CNN) image classifier to itemize 10-Q files.
result Demonstrates a generalized pipeline for rapid data retrieval from a large volume of textual data.

Generative Map learns interpretable neural network maps for camera localization.

problem Creating interpretable maps for neural network-based camera localization.
method Combining generative models with Kalman filters and incorporating additional sensor information.
result Generative Map predicts images closely resembling the true scene and achieves comparable localization performance.

Given a Morse function f on a closed manifold M with distinct critical values, and given a field F, there is a canonical complex, called the Morse-Barannikov complex, which is equivalent to any Morse complex associated with f and whose form is simple. In particular, the homology of M with coefficients in F is immediate…

2015-09-11abs ↗pdf ↗

We present the mathematical background of a software package that computes triangulations of mapping tori of surface homeomorphisms, suitable for Jeff Weeks's program SnapPea. It consists of two programs. jmt computes triangulations and prints them in a human-readable format. jsnap converts this format into SnapPea's t…

2000-12-01abs ↗pdf ↗

This article demonstrates that convolutional operation can be converted to matrix multiplication, which has the same calculation way with fully connected layer. The article is helpful for the beginners of the neural network to understand how fully connected layer and the convolutional layer work in the backend. To be c…

2017-12-04abs ↗pdf ↗

This paper proposes a framework to learn explainable rules from knowledge graphs for better recommendation.

problem Combining side information with explainability in recommendation systems.
method Joint learning framework integrating rule induction from knowledge graphs with a rule-guided neural recommendation model.
result Significant improvements in item recommendation performance over baselines.

The paper explains deep neural network predictions using logical proxies.

problem Creating understandable explanations for deep neural network predictions.
method Randomized propositionalization and Bayes-like approach to identify logical proxies.
result Models in first-order logic can approximate DRM's predictions in local regions.

We consider the problem of finding the minimizer of a function f:RdRf: \mathbb{R}^d \rightarrow \mathbb{R} of the finite-sum form minf(w)=1/ninfi(w)\min f(w) = 1/n\sum_{i}^n f_i(w). This problem has been studied intensively in recent years in the field of machine learning (ML). One promising approach for large-scale data is to use a stoc…

2017-10-27abs ↗pdf ↗

BTZSC benchmarks zero-shot text classification across diverse models.

problem Systematically comparing zero-shot text classification across various models.
method Comprehensive benchmark of 22 datasets, comparing NLI cross-encoders, embedding models, rerankers, and instruction-tuned LLMs.
result Rerankers and instruction-tuned LLMs outperform NLI cross-encoders, with rerankers setting a new state-of-the-art.

Tracr compiles programs into transformer models for interpretability.

problem Uncertainty in understanding transformer model outputs due to unknown learned programs.
method Tracr compiles human-readable programs into known structure transformer models.
result Known structure of Tracr-compiled models serves as ground-truth for interpretability.

This is an introduction to the subject of the differential topology of the space of smooth loops in a finite dimensional manifold. It began as the background notes to a series of seminars given at NTNU and subsequently at Sheffield. I am posting them in the hope that they will be useful to people wishing to know a litt…

2005-10-05abs ↗pdf ↗

LLMs can help explain credit risk models but not autonomously.

problem Leveraging LLMs for post-hoc explainability in credit risk models.
method Comparison of LLM outputs with SHAP and coefficient-based attributions on three LMs.
result LLMs reliably preserve feature-importance rankings but poorly align with autonomous explanations.

A substantial progress in development of new and efficient tensor factorization techniques has led to an extensive research of their applicability in recommender systems field. Tensor-based recommender models push the boundaries of traditional collaborative filtering techniques by taking into account a multifaceted nat…

2016-03-19abs ↗pdf ↗

VALC provides concept-level interpretations of FLMs, overcoming word-level limitations.

problem Lack of higher-level structure interpretation in FLMs' attention weights.
method Formal definition of conceptual interpretation, variational Bayesian framework (VALC).
result VALC finds optimal language concepts for FLM predictions, providing concept-level interpretations.

Researchers often summarize their work in the form of posters. Posters provide a coherent and efficient way to convey core ideas from scientific papers. Generating a good scientific poster, however, is a complex and time consuming cognitive task, since such posters need to be readable, informative, and visually aesthet…

2016-04-05abs ↗pdf ↗

Benchmark for math reasoning models from human proofs.

problem Measuring and accelerating machine learning models in high-level mathematical reasoning.
method Built a non-synthetic dataset from theorem prover proofs, defined a task for model to fill in missing propositions, used hierarchical transformer to improve performance.
result Neural models can capture non-trivial mathematical reasoning, hierarchical transformer outperforms baseline.

Text clustering method replaces centroids with summaries for interpretability and scalability.

problem Efficiently clustering text data while maintaining interpretability and scalability.
method k-NLPmeans and k-LLMmeans, which periodically replace numeric centroids with textual summaries.
result Consistently outperforms classical baselines and recent LLM-based clustering methods.

Discover governing equations from data without specifying terms.

problem Discovering differential equations from data without predefined terms.
method Data-driven approach using genetic programming and automatic differentiation.
result Calibrated differential equations from various solutions of a differential equation.