Reply to Ogburn et al. on their critique of Wang and Blei's work.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method selects facts in proofs using stateful recurrent neural networks.
ATPboost is a system for solving sets of large-theory problems by interleaving ATP runs with state-of-the-art machine learning of premise selection from the proofs. Unlike many previous approaches that use multi-label setting, the learning is implemented as binary classification that estimates the pairwise-relevance of…
In this paper, we demonstrate how to do automated theorem proving in the presence of a large knowledge base of potential premises without learning from human proofs. We suggest an exploration mechanism that mixes in additional premises selected by a tf-idf (term frequency-inverse document frequency) based lookup in a d…
LeanDojo removes barriers to theorem proving with open-source tools and data.
Neoclassical economics has two theories of competition between profit-maximizing firms (Marshallian and Cournot-Nash) that start from different premises about the degree of strategic interaction between firms, yet reach the same result, that market price falls as the number of firms in an industry increases. The Marsha…
Comment refutes the deconfounder method's premise about ignorability.
REALFIN benchmarks financial reasoning by removing implicit assumptions, revealing model weaknesses.
Any surface can be foliated into equipotential hypersurfaces of the level sets. A current result is that the contours are the progressing wave fronts of a certain hyperbolic partial differential equation, a wave equation. It is connected with the gradient lines, as well as with a corresponding eikonal equation. The lev…
PSC classifier improves HDLSS classification on class-imbalanced data.
The main points of the first section of the article written by S.I. Chernyshov, A.V. Voronin and S.A. Razumovsky arXiv:1003.4382), which deals with the fundamental bases of the macroeconomic theory, have been analyzed. An incorrectness of the Harrod's model of the economical growth in its generally accepted interpretat…
Fed-FEARE model extracts rules from multiple agencies' data securely.
It is known that the impact of transactions on stock price (market impact) is a concave function of the size of the order, but there exists little quantitative theory that suggests why this is so. I develop a quantitative theory for the market impact of hidden orders (orders that reflect the true intention of buying an…
This review compares various deep generative models.
Performing supervised learning from the data synthesized by using Generative Adversarial Networks (GANs), dubbed GAN-synthetic data, has two important applications. First, GANs may generate more labeled training data, which may help improve classification accuracy. Second, in scenarios where real data cannot be release…
Econophysics, is based on the premise that some ideas and methods from physics can be applied to economic situations. We intend to show in this paper how a physics concept such as entropy can be applied to an economic problem. In so doing, we demonstrate how information in the form of observable data and moment constra…
The discrete-time multifactor Vasiček model is a tractable Gaussian spot rate model. Typically, two- or three-factor versions allow one to capture the dependence structure between yields with different times to maturity in an appropriate way. In practice, re-calibration of the model to the prevailing market conditions …
Recent work establishes dataset difficulty and removes annotation artifacts via partial-input baselines (e.g., hypothesis-only models for SNLI or question-only models for VQA). When a partial-input baseline gets high accuracy, a dataset is cheatable. However, the converse is not necessarily true: the failure of a parti…
Two oppositely charged droplets of (say) water in e.g. oil or air will tend to drift together under the influence of their charges. As they make contact, one might expect them to coalesce and form one large droplet, and this indeed happens when the charge difference is sufficiently small. However, Ristenpart et al disc…
Designed to compete with fiat currencies, bitcoin proposes it is a crypto-currency alternative. Bitcoin makes a number of false claims, including: solving the double-spending problem is a good thing; bitcoin can be a reserve currency for banking; hoarding equals saving, and that we should believe bitcoin can expand by …
Many Machine Reading and Natural Language Understanding tasks require reading supporting text in order to answer questions. For example, in Question Answering, the supporting text can be newswire or Wikipedia articles; in Natural Language Inference, premises can be seen as the supporting text and hypotheses as question…
Standard acquisition functions are sufficient for asynchronous Bayesian optimization.
We show that conically smooth stratified spaces embed fully faithfully into -categories. This articulates a stratified generalization of the homotopy hypothesis proposed by Grothendieck. As such, each -category defines a stack on conically smooth stratified spaces, and we identify the descent conditions…
Paper develops a new model for predicting volatility surface.
Recursive prediction of graph signals with new nodes added.
Trading strategy uses Hoeffding's Inequality to predict financial regime change.
Much recent machine learning research has been directed towards leveraging shared statistics among labels, instances and data views, commonly referred to as multi-label, multi-instance and multi-view learning. The underlying premises are that there exist correlations among input parts and among output targets, and the …
New method infers co-expression networks robustly from multiple studies.
Ensemble learning is a statistical paradigm built on the premise that many weak learners can perform exceptionally well when deployed collectively. The BART method of Chipman et al. (2010) is a prominent example of Bayesian ensemble learning, where each learner is a tree. Due to its impressive performance, BART has rec…
This paper introduces C-DSL to improve data mining outcomes by considering context.
Our understanding of supercooled liquids and glasses has lagged significantly behind that of simple liquids and crystalline solids. This is in part due to the many possibly relevant degrees of freedom that are present due to the disorder inherent to these systems and in part to non-equilibrium effects which are difficu…
Detecting emergence of a low-rank signal from high-dimensional data is an important problem arising from many applications such as camera surveillance and swarm monitoring using sensors. We consider a procedure based on the largest eigenvalue of the sample covariance matrix over a sliding window to detect the change. T…
A new method combines federated learning and logistic regression for better credit scoring.
Polynomial inequalities lie at the heart of many mathematical disciplines. In this paper, we consider the fundamental computational task of automatically searching for proofs of polynomial inequalities. We adopt the framework of semi-algebraic proof systems that manipulate polynomial inequalities via elementary inferen…
The paper values reinsurance contracts for dynamic catastrophe claims without arbitrage.
IDVAE combines VAE and GAN without explicit discriminator.
In topic modeling, many algorithms that guarantee identifiability of the topics have been developed under the premise that there exist anchor words -- i.e., words that only appear (with positive probability) in one topic. Follow-up work has resorted to three or higher-order statistics of the data corpus to relax the an…
New method reduces bias in NLI models using ensemble adversarial training.
Trust-aware MAB improves learning performance by accounting for human deviation.
FLFE improves machine learning by efficiently and securely transforming features.
Seismic inversion method uses GAN to improve efficiency and accuracy.
New metric captures individual neuron tuning across neural networks.
Improved forecasting for irregularly-sampled time series using kernel flows.
FedCM measures contributions in real-time for federated learning.
Extracting and detecting spike activities from the fluorescence observations is an important step in understanding how neuron systems work. The main challenge lies in that the combination of the ambient noise with dynamic baseline fluctuation, often contaminates the observations, thereby deteriorating the reliability o…
ALMANACS benchmarks explainability methods on simulatability.
FNOs learn solution operators of dissipative equations efficiently via spectral methods.
The paper proposes a method to evaluate superhuman models by checking for logical inconsistencies.