Graph-based approach repairs programs from diagnostic feedback.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Due to its potential to improve programmer productivity and software quality, automated program repair has been an active topic of research. Newer techniques harness neural networks to learn directly from examples of buggy programs and their fixes. In this work, we consider a recently identified class of bugs called va…
This paper presents a novel end-to-end approach to program repair based on sequence-to-sequence learning. We devise, implement, and evaluate a system, called SequenceR, for fixing bugs based on sequence-to-sequence learning on source code. This approach uses the copy mechanism to overcome the unlimited vocabulary probl…
Deep reinforcement learning has led to several recent breakthroughs, though the learned policies are often based on black-box neural networks. This makes them difficult to interpret and to impose desired specification constraints during learning. We present an iterative framework, MORL, for improving the learned polici…
Automatic program repair holds the potential of dramatically improving the productivity of programmers during the software development process and correctness of software in general. Recent advances in machine learning, deep learning, and NLP have rekindled the hope to eventually fully automate the process of repairing…
Proposes a method to repair arbitrage in option prices data.
MACER accelerates error repair by modularly identifying and applying fixes.
SED integrates synthesis, execution, and debugging for neural program synthesis.
The recent use of `Big Code' with state-of-the-art deep learning methods offers promising avenues to ease program source code writing and correction. As a first step towards automatic code repair, we implemented a graph neural network model that predicts token types for Javascript programs. The predictions achieve an a…
In this study we model the warranty claims process and evaluate the warranty servicing costs under non-renewing and renewing free repair warranties. We assume that the repair time for rectifying the claims is non-zero and the repair cost is a function of the length of the repair time. To accommodate the ageing of the p…
Deep learning had been used in program analysis for the prediction of hidden software defects using software defect datasets, security vulnerabilities using generative adversarial networks as well as identifying syntax errors by learning a trained neural machine translation on program codes. However, all these approach…
New method to recover over-parameterized models corrupted during estimation.
Motivated by the problem of automated repair of software vulnerabilities, we propose an adversarial learning approach that maps from one discrete source domain to another target domain without requiring paired labeled examples or source and target domains to be bijections. We demonstrate that the proposed adversarial l…
CLSVAE repairs systematic errors in images with minimal labeled data.
REPAIR mitigates variance collapse to enable linear interpolation between SGD solutions.
We focus on the problem of unsupervised cell outlier detection and repair in mixed-type tabular data. Traditional methods are concerned only with detecting which rows in the dataset are outliers. However, identifying which cells are corrupted in a specific row is an important problem in practice, and the very first ste…
SnareNet adds repair layers to neural networks to ensure outputs meet physical constraints.
LLMs excel at summarizing and repairing complex models without needing full models.
CROP verifies clean prefixes in reasoning traces, improving downstream repair accuracy.
CASP improves portfolio optimization by considering asset covariance.
We present a new application and covering number bound for the framework of "Machine Learning with Operational Costs (MLOC)," which is an exploratory form of decision theory. The MLOC framework incorporates knowledge about how a predictive model will be used for a subsequent task, thus combining machine learning with t…
Optimal pre-processing reduces disparate impact by minimizing total variation distance.
The above named paper has been withdrawn. A colleague has observed a gap in the proof of isotopy invariance, which can be repaired by reducing the coefficients (which lie in (1/6)Z) of the antisymmetric kanji with chords incident with more than one component modulo 8Z. An analogous issue arises in considering the effec…
Many modern data-intensive computational problems either require, or benefit from distance or similarity data that adhere to a metric. The algorithms run faster or have better performance guarantees. Unfortunately, in real applications, the data are messy and values are noisy. The distances between the data points are …
Withdrawn May 2005. There is an error in the even-dimensional case of the proof in the April 2005 version. The hoped-for 4-dimensional applications are unlikely to survive the repairs.
Professional software developers spend a significant amount of time fixing builds, but this has received little attention as a problem in automatic program repair. We present a new deep learning architecture, called Graph2Diff, for automatically localizing and fixing build errors. We represent source code, build config…
Suppose is a compact Riemannian manifold and an arbitrary point. We employ estimates on the volume growth around to prove that the only conformal compactification of is itself.
We study graphs of (generalized) joins and intersections of finitely generated subgroups of a free group. We show how to disprove a lemma of Imrich and Müller on these graphs and how to repair this lemma.
When the performance of a machine learning model varies over groups defined by sensitive attributes (e.g., gender or ethnicity), the performance disparity can be expressed in terms of the probability distributions of the input and output variables over each group. In this paper, we exploit this fact to reduce the dispa…
A complete error analysis of variational integrators is obtained, by blowing up the discrete variational principles, all of which have a singularity at zero time-step. Divisions by the time step lead to an order that is one less than observed in simulations, a deficit that is repaired with the help of a new past-future…
Algorithms learned from data are increasingly used for deciding many aspects in our life: from movies we see, to prices we pay, or medicine we get. Yet there is growing evidence that decision making by inappropriately trained algorithms may unintentionally discriminate people. For example, in automated matching of cand…
This paper addresses the problem of predicting duration of unplanned power outages, using historical outage records to train a series of neural network predictors. The initial duration prediction is made based on environmental factors, and it is updated based on incoming field reports using natural language processing …
Clients are increasingly looking for fast and effective means to quickly and frequently survey and communicate the condition of their buildings so that essential repairs and maintenance work can be done in a proactive and timely manner before it becomes too dangerous and expensive. Traditional methods for this type of …
The Cayley hyperbolic space minimizes volume entropy among finite-volume metrics.
The theory of quantum computation can be constructed from the abstract study of anyonic systems. In mathematical terms, these are unitary topological modular functors. They underlie the Jones polynomial and arise in Witten-Chern-Simons theory. The braiding and fusion of anyonic excitations in quantum Hall electron liqu…
New methods for equity fund selection and portfolio construction using mutual fund top holdings.
Many software analysis methods have come to rely on machine learning approaches. Code segmentation - the process of decomposing source code into meaningful blocks - can augment these methods by featurizing code, reducing noise, and limiting the problem space. Traditionally, code segmentation has been done using syntact…
The complement of an arrangement A of a finite number of affine hyperplanes in complex n-space has the structure of a poset of spaces indexed by the intersection poset, L(A). The space corresponding to G in L(A) is homotopy equivalent to the complement of the hyperplanes in the central arrangement A_G normal to G. This…
Machine learning (ML) can automate decision-making by learning to predict decisions from historical data. However, these predictors may inherit discriminatory policies from past decisions and reproduce unfair decisions. In this paper, we propose two algorithms that adjust fitted ML predictors to make them fair. We focu…
Improves probabilistic programming by analyzing program structure.
The Pearson distance between a pair of random variables with correlation , namely, 1-, has gained widespread use, particularly for clustering, in areas such as gene expression analysis, brain imaging and cyber security. In all these applications it is implicitly assumed/required that the distance …
We propose design guidelines for a probabilistic programming facility suitable for deployment as a part of a production software system. As a reference implementation, we introduce Infergo, a probabilistic programming facility for Go, a modern programming language of choice for server-side software development. We argu…
We introduce the notion of a stochastic probabilistic program and present a reference implementation of a probabilistic programming facility supporting specification of stochastic probabilistic programs and inference in them. Stochastic probabilistic programs allow straightforward specification and efficient inference …
We define a class of representations of the fundamental group of a closed surface of genus to : the pentagon representations. We show that they are exactly the non-elementary -representations of surface groups that do not admit a Schottky decomposition, i.e. a…
A neural program synthesis method with iterative fix operations.
A key feature of inductive logic programming (ILP) is its ability to learn first-order programs, which are intrinsically more expressive than propositional programs. In this paper, we introduce techniques to learn higher-order programs. Specifically, we extend meta-interpretive learning (MIL) to support learning higher…
We present a new algorithm for approximate inference in probabilistic programs, based on a stochastic gradient for variational programs. This method is efficient without restrictions on the probabilistic program; it is particularly practical for distributions which are not analytically tractable, including highly struc…
This paper corrects errors in UMAP's derivation and explains its properties.