Improves code2vec for Java classes by obfuscating variable names.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Most of the JavaScript code deployed in the wild has been minified, a process in which identifier names are replaced with short, arbitrary and meaningless names. Minified code occupies less space, but also makes the code extremely difficult to manually inspect and understand. This paper presents Context2Name, a deep le…
Graph-Structured Cache improves code completion and variable naming tasks.
Proposes a new variable grouping approach to improve Bayesian additive regression tree (BART) performance.
Harmonic functions of two variables are exactly those that admit a conjugate, namely a function whose gradient has the same length and is everywhere orthogonal to the gradient of the original function. We show that there are also partial differential equations controlling the functions of three variables that admit a c…
We present and implement two algorithms for analytic asymptotic evaluation of the marginal likelihood of data given a Bayesian network with hidden nodes. As shown by previous work, this evaluation is particularly hard for latent Bayesian network models, namely networks that include hidden variables, where asymptotic ap…
In this paper, we propose novel strategies for neutral vector variable decorrelation. Two fundamental invertible transformations, namely serial nonlinear transformation and parallel nonlinear transformation, are proposed to carry out the decorrelation. For a neutral vector variable, which is not multivariate Gaussian d…
Knoop enhances variable selection with over-parameterization and knockoffs.
SCORE improves tree-based predictions with boosted residual extraTrees.
A hybrid model for Bayesian optimization handles mixed variables using MCTS for categorical and GP for continuous.
In this paper, a class of statistics named ART (the alternant recursive topology statistics) is proposed to measure the properties of correlation between two variables. A wide range of bi-variable correlations both linear and nonlinear can be evaluated by ART efficiently and equitably even if nothing is known about the…
Method identifies variable matches from data values.
This paper considers the problem of minimizing an expectation function over a closed convex set, coupled with a {\color{black} functional or expectation} constraint on either decision variables or problem parameters. We first present a new stochastic approximation (SA) type algorithm, namely the cooperative SA (CSA), t…
In this note, we would like to find the laws of electrodynamics in simple economic systems. In this direction, we identify the chief economic variables and parameters, scalar and vector, which are amenable to be put directly into the crouch of the laws of electrodynamics, namely Maxwell's equations. Moreover, we obtain…
Within a supervised classification framework, labeled data are used to learn classifier parameters. Prior to that, it is generally required to perform dimensionality reduction via feature extraction. These preprocessing steps have motivated numerous research works aiming at recovering latent variables in an unsupervise…
Study tackles causal structure learning in linear models with unobserved variables and measurement error.
OutlierTree detects outliers using decision trees and provides explanations.
The author proposes a finance trading strategy named Entropy Oriented Trading and apply thermodynamics on the strategy. The state variables are chosen so that the strategy satisfies the second law of thermodynamics. Using the law, the author proves that the rate of investment (ROI) of the strategy is equal to or more t…
New measures quantify dependence between variables without distribution estimation.
New framework models uncertainty with uncertainty variables.
We show a connection between the Fourier spectrum of Boolean functions and the REINFORCE gradient estimator for binary latent variable models. We show that REINFORCE estimates (up to a factor) the degree-1 Fourier coefficients of a Boolean function. Using this connection we offer a new perspective on variance reduction…
Graphical models are commonly used tools for modeling multivariate random variables. While there exist many convenient multivariate distributions such as Gaussian distribution for continuous data, mixed data with the presence of discrete variables or a combination of both continuous and discrete variables poses new cha…
In this manuscript we analyse the leading statistical properties of fluctuations of (log) 3-month US Treasury bill quotation in the secondary market, namely: probability density function, autocorrelation, absolute values autocorrelation, and absolute values persistency. We verify that this financial instrument, in spit…
Efficient algorithm learns direct causes and effects from data.
Proposes ABC method for discrete data, improving likelihood-free inference.
A new method clusters mixed-type data efficiently.
Using the one dimensional free particle symmetries, the quantum finance symmetries are obtained. Namely, it is shown that Black-Scholes equation is invariant under Schrödinger group. In order to do this, the one dimensional free non-relativistic particle and its symmetries are revisited. To get the Black-Scholes equati…
We establish a new version of the first Noether Theorem, according to which the (equivalence classes of) first integrals of given Euler-Lagrange equations in one independent variable are in exact one-to-one correspondence with the (equivalence classes of) vector fields satisfying two simple geometric conditions, namely…
Langevin autoencoders improve deep latent variable models with efficient posterior sampling.
SlideVaR balances risk and prudence by considering variable investor attitudes.
This paper enhances stability selection by evaluating overall results robustness and identifying optimal regularization values.
A new method quickly identifies key variables and interactions.
Proposes a new model for high-dimensional data analysis with unknown link function.
We generalise surface cluster algebras to the case of infinite surfaces where the surface contains finitely many accumulation points of boundary marked points. To connect different triangulations of an infinite surface, we consider infinite mutation sequences. We show transitivity of infinite mutation sequences on tria…
The paper classifies symplectic invariants of specific singularities in integrable Hamiltonian systems.
New assumptions help identify causal relationships in data.
Proposes a Gaussian process for Koopman mode decomposition.
A deep learning approach for clustering time series of varying lengths.
mGENRE improves multilingual entity linking with autoregressive sequence prediction.
Three results in p-convex geometry are established. First is the analogue of the Levi problem in several complex variables, namely: local p-convexity implies global p-convexity. The second asserts that the support of a minimal p-dimensional current is contained in the p-hull of the boundary union with the "core" of the…
Estimates and infers multi-stage stationary treatment policies with variable selection.
Extends knockoff filter for composite null hypotheses in variable selection.
Paper proves causal direction can be inferred from data with limited randomness.
New methods encode high-cardinality string variables efficiently.
CLOUD method detects causal relationships in various data types without latent variable assumptions.
We introduce Thurstonian Boltzmann Machines (TBM), a unified architecture that can naturally incorporate a wide range of data inputs at the same time. Our motivation rests in the Thurstonian view that many discrete data types can be considered as being generated from a subset of underlying latent continuous variables, …
The paper proves a new method to improve generalization in covariate-shift scenarios.
ABM automates feature engineering and variable selection for loss-based models.