Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

2.6%5.1%7.7%10.2% · Feb 201919922001200920182026
48 results for Experimental Flaws

Bayesian Deep Learning experiments often use weak baselines, leading to misleading conclusions.

problem Misleading conclusions in Bayesian Deep Learning due to weak baselines in experiments.
method Used a fixed number of iterations for baselines and compared them with models trained to convergence.
result Monte Carlo dropout baseline outperforms or performs competitively with superior methods.

Flawed groups are shown to include all finitely generated groups isomorphic to free products of nilpotent groups.

problem Characterizing flawed groups and understanding their topological properties.
method Analyzing finitely presented groups and their deformation retracts onto subspaces of character varieties.
result All finitely generated groups isomorphic to free products of nilpotent groups are flawed.

Improved software flaw detection using NAS on multimodal DL models.

problem Software flaw detection in multimodal deep learning models.
method Adapted NAS framework for multimodal learning, combined with multimodal deep learning models.
result Improved performance on the Juliet Test Suite.

PMI-Masking improves MLM pretraining by masking correlated spans efficiently.

problem Uniform token masking leads to inefficient and suboptimal performance in MLMs.
method PMI-Masking uses Pointwise Mutual Information to mask n-grams with high collocation.
result PMI-Masking reaches half the training time and improves performance.

New method identifies flawed internal models of the world in animals.

problem How animals make decisions with partial sensory information.
method Generalizes Inverse Rational Control to continuous nonlinear dynamics and noise.
result Identifies the best internal model explaining an agent's actions.

The abstract warns against flawed empirical research in machine learning.

problem Flawed empirical research in machine learning leading to unreliable results.
method Call for more awareness of experimental knowledge plurality and epistemic limitations.
result Current empirical machine learning research should be exploratory, not confirmatory.

Study shows over-sampling biases prediction results on imbalanced datasets.

problem Over-optimistic prediction results on imbalanced data.
method Applying over-sampling before partitioning training and testing sets.
result Over-sampling causes biased results and reduces predictive performance.

This article focuses on the work of O. Chanel and G. Chichilnisky (2013) on the flaws of expected utility theory while assessing the value of life. Expected utility is a fundamental tool in decision theory. However, it does not fit with the experimental results when it comes to catastrophic outcomes ---see, for example…

2015-08-25abs ↗pdf ↗

Recent anomaly detection benchmarks are flawed, potentially misleading progress.

problem Flawed benchmark datasets create misleading progress reports.
method Identified four flaws in benchmark datasets and introduced a new archive.
result Published comparisons may be unreliable due to flaws in benchmark datasets.

A multi-agent simulator evaluates trading strategies using Market Replay and Interactive Agent-Based Simulation.

problem Evaluate trading strategies using Market Replay and Interactive Agent-Based Simulation.
method Multi-agent simulator for Market Replay and Interactive Agent-Based Simulation.
result IABS provides a more realistic market environment for evaluating trading strategies.

This paper critiques flawed MVTS anomaly detection evaluation methods and proposes a simple baseline.

problem Flawed evaluation methods in MVTS anomaly detection research.
method Robust evaluation protocols, including PCA-based baseline.
result Simple PCA-based baseline outperforms many DL approaches.

Study examines flaws in probing LLMs' knowledge and introduces a new method.

problem Flaws in existing methods for probing the veracity of LLMs' internal knowledge.
method sAwMIL (Sparse-Aware Multiple-Instance Learning) combining multiple-instance learning with conformal prediction.
result LLMs encode a third type of signal distinct from true and false.

New research highlights flaws in evaluating clustering algorithms using classification datasets.

problem Flaws in evaluating clustering algorithms using classification datasets.
method Advanced visualization and dimension reduction techniques to expose flaws.
result Current practice of evaluating clustering algorithms may produce misleading results.

The paper tackles auction market design flaws by randomizing closing times and optimizing transaction fees.

problem Strategic traders exploit accumulated information to delay their orders, distorting auction efficiency.
method Randomizing auction closing times and designing optimal transaction fees policies.
result Policies encourage strategic traders to send orders earlier, improving auction market efficiency.

Paper introduces a benchmark for predicting bankruptcy from text data.

problem Lack of a common benchmark dataset and evaluation strategy for unstructured data in bankruptcy prediction.
method Describes and evaluates several baseline models, including a bag-of-words model.
result A lightweight bag-of-words model performs surprisingly well, especially when considering data from multiple years.

New methods needed to evaluate uncertainty estimates in neural networks.

problem Evaluating uncertainty estimates in neural networks is flawed and inconsistent.
method Proposes a simulation-based testing approach to address flaws in current methods.
result Current methods for evaluating uncertainty estimates have significant flaws and cannot accurately compare different methods.

Quantized neural networks are vulnerable to adversarial attacks.

problem Adversarial robustness of quantized neural networks.
method Investigated adversarial robustness of quantized neural networks under different threat models.
result Quantization does not offer robust protection and results in gradient masking.

We solve Hilbert's fifth problem for local groups: every locally euclidean local group is locally isomorphic to a Lie group. Jacoby claimed a proof of this in 1957, but this proof is seriously flawed. We use methods from nonstandard analysis and model our solution after a treatment of Hilbert's fifth problem for global…

2007-08-28abs ↗pdf ↗

Study flaws in generative model evaluation metrics, especially for diffusion models.

problem Flaws in existing metrics for evaluating generative models, particularly for diffusion models.
method Systematic study of generative models, human perception experiments, and analysis of feature extractors.
result State-of-the-art perceptual realism of diffusion models is not reflected in commonly reported metrics.

New bounds for unsupervised domain adaptation account for non-invertibility and support coverage.

problem Theoretical arguments for domain-invariant representations are flawed and do not account for non-invertibility and support coverage.
method Generalization bounds for any representation function acknowledging the cost of non-invertibility and penalizing distance between densities.
result Proposed bounds based on support coverage provide better generalization than current standard practice.

New methods reduce extrapolation errors in feature importance.

problem Flawed feature importance methods using unrestricted permutations lead to extrapolation errors.
method Three new approaches: conditional model reliance, Knockoffs with Gaussian transformation, and restricted ALE plot designs.
result Theoretical and numerical results show our strategies reduce/eliminate extrapolation.

This paper analyzes stablecoins to address cryptocurrency volatility.

problem Cryptocurrency price volatility hinders adoption.
method Survey of 24 stablecoin projects, combined with monetary policy insights.
result Stablecoin designs often prioritize 1-to-1 stabilization, but literature suggests smoothing volatility is more sustainable.

Neural networks and linear systems linked, revealing training loss and kernel limitations.

problem Exploring the training loss and limitations of neural networks and their kernels.
method Drawing connections between neural networks and under-determined linear systems, providing lower bounds, and analyzing gradient descent.
result Zero training loss achievable for neural networks under certain conditions, but not for ReLU kernels.