Develops a dynamic latent-factor model for high-dimensional asset characteristics.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Survey on ML advances for personalized prediction considering entity characteristics.
Recent studies have shown that information disclosed on social network sites (such as Facebook) can be used to predict personal characteristics with surprisingly high accuracy. In this paper we examine a method to give online users transparency into why certain inferences are made about them by statistical models, and …
CCs learn high-dimensional distributions from heterogeneous data.
For assistive robots and virtual agents to achieve ubiquity, machines will need to anticipate the needs of their human counterparts. The field of Learning from Demonstration (LfD) has sought to enable machines to infer predictive models of human behavior for autonomous robot control. However, humans exhibit heterogenei…
A multi-task network avoids indirect discrimination in insurance pricing.
Learning and inference movement is a very challenging problem due to its high dimensionality and dependency to varied environments or tasks. In this paper, we propose an effective probabilistic method for learning and inference of basic movements. The motion planning problem is formulated as learning on a directed grap…
New method for Bayesian neural networks reduces inference difficulty.
We develop Riemannian Stein Variational Gradient Descent (RSVGD), a Bayesian inference method that generalizes Stein Variational Gradient Descent (SVGD) to Riemann manifold. The benefits are two-folds: (i) for inference tasks in Euclidean spaces, RSVGD has the advantage over SVGD of utilizing information geometry, and …
Proposes a probabilistic model to improve hydrology predictions and trust.
In this paper, we consider the sigmoid Gaussian Hawkes process model: the baseline intensity and triggering kernel of Hawkes process are both modeled as the sigmoid transformation of random trajectories drawn from Gaussian processes (GP). By introducing auxiliary latent random variables (branching structure, Pólya-Gamm…
Dynamic treatment effects estimated over time using covariate balancing.
The ever-increasing demand from mobile Machine Learning (ML) applications calls for evermore powerful on-chip computing resources. Mobile devices are empowered with heterogeneous multi-processor Systems-on-Chips (SoCs) to process ML workloads such as Convolutional Neural Network (CNN) inference. Mobile SoCs house sever…
The application of deep learning techniques resulted in remarkable improvement of machine learning models. In this paper provides detailed characterizations of deep learning models used in many Facebook social network services. We present computational characteristics of our models, describe high performance optimizati…
Proposes Infomax and Domain-Independent Representations for robust causal inference.
We pose causal inference as the problem of learning to classify probability distributions. In particular, we assume access to a collection , where each is a sample drawn from the probability distribution of , and is a binary label indicating whether "" or …
Improved neural models for diverse user event sequences.
Real Estate Investment Trusts (REITs) are the only truly liquid assets related to real estate investments. We study the behavior of U.S. REITs over the past three decades and document their return characteristics. REITs have somewhat less market risk than equity; their betas against a broad market index average about .…
Mobility datasets are fundamental for evaluating algorithms pertaining to geographic information systems and facilitating experimental reproducibility. But privacy implications restrict sharing such datasets, as even aggregated location-data is vulnerable to membership inference attacks. Current synthetic mobility data…
We present a new, fully generative model for constructing astronomical catalogs from optical telescope image sets. Each pixel intensity is treated as a random variable with parameters that depend on the latent properties of stars and galaxies. These latent properties are themselves modeled as random. We compare two pro…
We prove a formality theorem for the Fukaya categories of the symplectic manifolds underlying symplectic Khovanov cohomology, over fields of characteristic zero. The key ingredient is the construction of a degree one Hochschild cohomology class on a Floer A-infinity algebra associated to the (k,k)-nilpotent slice Y, ob…
Several methods exist to infer causal networks from massive volumes of observational data. However, almost all existing methods require a considerable length of time series data to capture cause and effect relationships. In contrast, memory-less transition networks or Markov Chain data, which refers to one-step transit…
Geometric analysis improves convergence of variational inference.
Paper proves Shapley value convergence in Bayesian learning games.
Paper defends diffusion models from membership inference attacks using Langevin dynamics.
Network-assisted regression uses conformal prediction for valid inference.
Estimates volatility of volatility and leverage effect using high-frequency options data.
This paper analyzes consumer choices over lunchtime restaurants using data from a sample of several thousand anonymous mobile phone users in the San Francisco Bay Area. The data is used to identify users' approximate typical morning location, as well as their choices of lunchtime restaurants. We build a model where res…
Proposes a new privacy notion for membership inference attacks on machine learning models.
Proposes a non-parametric method for deep discrete latent variable models.
New method corrects biased predictions and uncertainty estimates in classification with nuisance parameters.
Many biological characteristics of evolutionary interest are not scalar variables but continuous functions. Here we use phylogenetic Gaussian process regression to model the evolution of simulated function-valued traits. Given function-valued data only from the tips of an evolutionary tree and utilising independent pri…
Upper bound derived for informed traders' gains in a model, akin to thermodynamics.
Recent advances in high-throughput cDNA sequencing (RNA-Seq) technology have revolutionized transcriptome studies. A major motivation for RNA-Seq is to map the structure of expressed transcripts at nucleotide resolution. With accurate computational tools for transcript reconstruction, this technology may also become us…
Paper develops fast, flexible Hawkes process inference for space-time data.
In many contexts, we have access to aggregate data, but individual level data is unavailable. For example, medical studies sometimes report only aggregate statistics about disease prevalence because of privacy concerns. Even so, many a time it is desirable, and in fact could be necessary to infer individual level chara…
Paper presents a fast method for estimating hidden states in Bayesian models.
Study proposes a multi-agent framework to mitigate bias in sentiment analysis.
We present a novel approximate graph matching algorithm that incorporates seeded data into the graph matching paradigm. Our Joint Optimization of Fidelity and Commensurability (JOFC) algorithm embeds two graphs into a common Euclidean space where the matching inference task can be performed. Through real and simulated …
Simulation-based inference speeds up gravitational wave data analysis.
The paper studies causal effects of multiple treatments in healthcare databases with rare outcomes.
Proposes Causal-Batle for estimating treatment effects in small high-dimensional datasets.
AI-enhanced product embeddings boost demand analysis accuracy.
Ranking a set of objects involves establishing an order allowing for comparisons between any pair of objects in the set. Oftentimes, due to the unavailability of a ground truth of ranked orders, researchers resort to obtaining judgments from multiple annotators followed by inferring the ground truth based on the collec…
New algorithm speeds up IRT model fitting for large datasets.
Disease phenotyping algorithms process observational clinical data to identify patients with specific diseases. Supervised phenotyping methods require significant quantities of expert-labeled data, while unsupervised methods may learn non-disease phenotypes. To address these limitations, we propose the Semi-Supervised …
In this paper we present a statistical analysis about the characteristics that we intend to influence in the performance of the neural networks in terms of assertiveness in the prediction of Brazilian stock returns. We created a population of architectures for analysis and extracted the sample that had the best asserti…
Proposes a method for interpreting time-varying causal effect moderation in high-dimensional data.