Recent anomaly detection benchmarks are flawed, potentially misleading progress.
problem Flawed benchmark datasets create misleading progress reports.
method Identified four flaws in benchmark datasets and introduced a new archive.
result Published comparisons may be unreliable due to flaws in benchmark datasets.
Flawed groups are shown to include all finitely generated groups isomorphic to free products of nilpotent groups.
problem Characterizing flawed groups and understanding their topological properties.
method Analyzing finitely presented groups and their deformation retracts onto subspaces of character varieties.
result All finitely generated groups isomorphic to free products of nilpotent groups are flawed.
Vision and language tasks often fail to test AI comprehensively.
problem Current vision and language tasks are flawed due to dataset and evaluation issues.
method Review of current state and proposal for improvement.
result State-of-the-art systems perform well due to dataset and evaluation flaws.
Study shows over-sampling biases prediction results on imbalanced datasets.
problem Over-optimistic prediction results on imbalanced data.
method Applying over-sampling before partitioning training and testing sets.
result Over-sampling causes biased results and reduces predictive performance.
Multimodal deep learning improves flaw detection in software programs.
problem Current flaw detection relies on single software representations.
method Adapted multimodal deep learning models for flaw detection.
result Multimodal models outperform traditional deep learning models.
This paper critiques flawed MVTS anomaly detection evaluation methods and proposes a simple baseline.
problem Flawed evaluation methods in MVTS anomaly detection research.
method Robust evaluation protocols, including PCA-based baseline.
result Simple PCA-based baseline outperforms many DL approaches.
Improved software flaw detection using NAS on multimodal DL models.
problem Software flaw detection in multimodal deep learning models.
method Adapted NAS framework for multimodal learning, combined with multimodal deep learning models.
result Improved performance on the Juliet Test Suite.
Study examines flaws in probing LLMs' knowledge and introduces a new method.
problem Flaws in existing methods for probing the veracity of LLMs' internal knowledge.
method sAwMIL (Sparse-Aware Multiple-Instance Learning) combining multiple-instance learning with conformal prediction.
result LLMs encode a third type of signal distinct from true and false.
Study flaws in generative model evaluation metrics, especially for diffusion models.
problem Flaws in existing metrics for evaluating generative models, particularly for diffusion models.
method Systematic study of generative models, human perception experiments, and analysis of feature extractors.
result State-of-the-art perceptual realism of diffusion models is not reflected in commonly reported metrics.
The paper highlights issues with fixed point claims in digital images.
problem Flaws in published assertions about fixed points in digital images.
method Continues a series of studies examining digital topology.
result Identifies and discusses problems with fixed point claims.
Detects unusual inputs to neural networks to prevent flawed predictions.
problem Erratic predictions from neural networks on unexpected inputs.
method Evaluates input unusualness by comparing its content to learned parameters.
result Simple, effective method for comparing input metrics across different scales.
The paper highlights issues in fixed point claims in digital topology.
problem Flaws in published assertions about fixed points in digital metric spaces.
method Continues a series of studies examining these flaws.
result Identifies and discusses problems in fixed point claims.
The F-measure or F-score is one of the most commonly used single number measures in Information Retrieval, Natural Language Processing and Machine Learning, but it is based on a mistake, and the flawed assumptions render it unsuitable for use in most contexts! Fortunately, there are better alternatives.
Like all sub-fields of machine learning Bayesian Deep Learning is driven by empirical validation of its theoretical proposals. Given the many aspects of an experiment it is always possible that minor or even major experimental flaws can slip by both authors and reviewers. One of the most popular experiments used to eva…
Experiments used in current continual learning research do not faithfully assess fundamental challenges of learning continually. Instead of assessing performance on challenging and representative experiment designs, recent research has focused on increased dataset difficulty, while still using flawed experiment set-ups…
Paper discusses flaws in traditional RL for lifelong learning.
problem Traditional RL fails to model lifelong learning systems.
method Simplified prototype of lifelong RL system.
result Insights into lifelong RL, showing traditional RL's limitations.
This paper is being withdrawn by the author due a serious flaw.
New flaw found in SAP defense, reducing its effectiveness to 0.1%.
problem Weakness in Stochastic Activation Pruning defense against adversarial attacks.
method Re-examined the implementation of SAP and introduced a new BPDA attack.
result SAP's effectiveness reduced to 0.1% when properly applied.
In general, recommendation can be viewed as a matching problem, i.e., match proper items for proper users. However, due to the huge semantic gap between users and items, it's almost impossible to directly match users and items in their initial representation spaces. To solve this problem, many methods have been studied…
uTSGAN improves on TSGAN for generating time series data.
problem Challenges in generating time-dependent data.
method Unified training of independent networks in TSGAN.
result uTSGAN outperforms TSGAN in 80% of benchmark datasets.
The paper proves deep learning can be robust with certain loss functions.
problem The robustness of deep learning models under flawed data.
method Empirical-risk minimization with unbounded, Lipschitz-continuous loss functions.
result These loss functions provide efficient prediction under minimal data assumptions.
This paper has been withdrawn by the author due to a serious flaw that needs to be fixed. That is in progress by the author.
Neural networks for stock price prediction often misrepresent model performance due to flawed error metrics.
problem Flawed prediction error metrics lead to unreliable model evaluations in the securities market.
method Used data from 20 stock datasets across multiple markets and evaluated with four prediction error measures.
result Prediction error value only partially reflects model accuracy and fails to represent stock price direction.
The paper tackles auction market design flaws by randomizing closing times and optimizing transaction fees.
problem Strategic traders exploit accumulated information to delay their orders, distorting auction efficiency.
method Randomizing auction closing times and designing optimal transaction fees policies.
result Policies encourage strategic traders to send orders earlier, improving auction market efficiency.
Research shows bias in machine learning can be due to algorithmic flaws, not just data.
problem Underestimation bias in machine learning algorithms.
method Initial research to understand factors contributing to bias in classification algorithms.
result Regularization methods to address overfitting can also accentuate bias.
New research highlights flaws in evaluating clustering algorithms using classification datasets.
problem Flaws in evaluating clustering algorithms using classification datasets.
method Advanced visualization and dimension reduction techniques to expose flaws.
result Current practice of evaluating clustering algorithms may produce misleading results.
The paper addresses flaws in fixed point assertions for digital images.
problem Deficiencies in previously published works on fixed point assertions for digital images.
method Continues a series of studies to identify and rectify issues in fixed point assertions.
result Identifies and corrects flaws in fixed point assertions for digital images.
We derive a new proof to show that the incremental resparsification algorithm proposed by Kelner and Levin (2013) produces a spectral sparsifier in high probability. We rigorously take into account the dependencies across subsequent resparsifications using martingale inequalities, fixing a flaw in the original analysis…
Coupled entropy corrects flaws in Tsallis entropy for complex systems.
problem Misinterpretation of generalized temperature and entropy.
method Derived from generalized Pareto and Student's t distributions.
result Provides balanced measure of uncertainty for complex systems.
New methods needed to evaluate uncertainty estimates in neural networks.
problem Evaluating uncertainty estimates in neural networks is flawed and inconsistent.
method Proposes a simulation-based testing approach to address flaws in current methods.
result Current methods for evaluating uncertainty estimates have significant flaws and cannot accurately compare different methods.
We analyze GANs using neural tangent kernels, revealing flaws and advancing understanding.
problem Flaws in previous GAN analysis models.
method Neural Tangent Kernel framework for infinite-width discriminator.
result New insights into GAN convergence and generated distribution.
We prove that the twisted Reidemeister torsion of a 3-manifold corresponding to a fibered class is monic and we show that it gives lower bounds on the Thurston norm. The former fixes a flawed proof in [FV10], the latter gives a quick alternative argument for the main theorem of [FK06].
New method identifies flawed internal models of the world in animals.
problem How animals make decisions with partial sensory information.
method Generalizes Inverse Rational Control to continuous nonlinear dynamics and noise.
result Identifies the best internal model explaining an agent's actions.
Smart contracts are a digital technology with potential but also flaws.
problem Understanding the potential and limitations of smart contracts.
method Exploratory study combining statistics, IT, and law.
result Smart contracts have both idealistic promises and practical challenges.
PMI-Masking improves MLM pretraining by masking correlated spans efficiently.
problem Uniform token masking leads to inefficient and suboptimal performance in MLMs.
method PMI-Masking uses Pointwise Mutual Information to mask n-grams with high collocation.
result PMI-Masking reaches half the training time and improves performance.
Paper introduces a benchmark for predicting bankruptcy from text data.
problem Lack of a common benchmark dataset and evaluation strategy for unstructured data in bankruptcy prediction.
method Describes and evaluates several baseline models, including a bag-of-words model.
result A lightweight bag-of-words model performs surprisingly well, especially when considering data from multiple years.
Study improves model robustness in noisy datasets.
problem Instance-specific label noise in robust classification tasks.
method Coordinated Sparse Recovery (CSR) method introduces a collaboration matrix and confidence weights to reduce generalization error.
result CSR and CSR+ significantly reduce generalization error compared to existing methods.
Notes a flaw in a proof about embedding graphs.
problem A flaw in proving Sachs' conjecture about graph embeddings.
method Analyzing Stanfield's proof for gaps.
result Identifies a significant error in the proof.
We solve Hilbert's fifth problem for local groups: every locally euclidean local group is locally isomorphic to a Lie group. Jacoby claimed a proof of this in 1957, but this proof is seriously flawed. We use methods from nonstandard analysis and model our solution after a treatment of Hilbert's fifth problem for global…
New bounds for unsupervised domain adaptation account for non-invertibility and support coverage.
problem Theoretical arguments for domain-invariant representations are flawed and do not account for non-invertibility and support coverage.
method Generalization bounds for any representation function acknowledging the cost of non-invertibility and penalizing distance between densities.
result Proposed bounds based on support coverage provide better generalization than current standard practice.
Corrects errors in Hans' pseudocovering spaces paper.
problem Mathematical errors and citation issues in Hans' pseudocovering spaces paper.
method Identifies and corrects errors in Hans' previous work.
result Addresses mathematical and citation errors in Hans' pseudocovering spaces paper.
AI bias arises from human-defined goals, not algorithmic flaws.
problem AI bias due to human-defined goals in LLMs.
method Purpose-conditioned cognition and revealing downstream use of LLM outputs.
result AI bias can be reduced by purpose-aware prompting but not fully by regularization.
New method corrects bias in feature importance measures of GBM.
problem Bias in feature importance measures of GBM.
method Cross-validated unbiased base learners.
result Significant improvement in feature importance measures with minimal computational cost.
Deep RL algorithms can overfit to early experiences, leading to poor performance.
problem Overfitting to early interactions in deep reinforcement learning.
method Proposed a mechanism to periodically reset part of the agent to mitigate overfitting.
result Periodic resetting improves performance in both discrete and continuous action domains.
In recent years, neural networks have demonstrated outstanding effectiveness in a large amount of applications.However, recent works have shown that neural networks are susceptible to adversarial examples, indicating possible flaws intrinsic to the network structures. To address this problem and improve the robustness …
New methods reduce extrapolation errors in feature importance.
problem Flawed feature importance methods using unrestricted permutations lead to extrapolation errors.
method Three new approaches: conditional model reliance, Knockoffs with Gaussian transformation, and restricted ALE plot designs.
result Theoretical and numerical results show our strategies reduce/eliminate extrapolation.
Intact-VAE estimates treatment effects with latent confounders.
problem Estimating treatment effects under unobserved confounding.
method Intact-VAE, a VAE variant, models latent confounders to identify treatment effects.
result Intact-VAE is a consistent estimator of treatment effects under certain settings.
Machine learning experiments show IID assumption is flawed for bathymetry editing.
problem Flawed IID assumption in machine learning for bathymetry editing.
method Real-world computer-assisted labeling task, IID assumption analysis.
result Common random split leads to poor performance in machine learning.