This review discusses challenges and solutions for AI in chemical engineering.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proves stability of Minkowski space for specific initial data.
Survey on ML advances for personalized prediction considering entity characteristics.
Develops a dynamic latent-factor model for high-dimensional asset characteristics.
CCs learn high-dimensional distributions from heterogeneous data.
Study on 3D spacetimes, focusing on vacuum data and energy bounds.
Imbalanced data sets containing much more background than signal instances are very common in particle physics, and will also be characteristic for the upcoming analyses of LHC data. Following up the work presented at ACAT 2008, we use the multivariate technique presented there (a rule growing algorithm with the meta-m…
Generates synthetic data for benchmarking unsupervised outlier detection.
We show here that the Nielsen core of the bumping set of the domain of discontinuity of a Kleinian group is the boundary of the characteristic submanifold of the associated 3-manifold with boundary. Some examples of interesting characteristic submanifolds are given. We also give a construction of the characteristic…
A neural network method estimates densities from characteristic functions.
The study assesses machine learning generalization using various data set characteristics.
We prove a formula that relates the Euler-Poincaré characteristic of a closed semi-algebraic set to its Lipschitz-Killing curvatures
The paper analyzes heat trace asymptotics for de Rham and Dolbeault complexes in both real and complex settings.
Solves characteristic problem in general relativity for null data.
Path signatures adapted for Lie groups improve action recognition in computer vision.
Private business schools in India face a common problem of selecting quality students for their MBA programs to achieve the desired placement percentage. Generally, such data sets are biased towards one class, i.e., imbalanced in nature. And learning from the imbalanced dataset is a difficult proposition. This paper pr…
Representation learning (RL) plays an important role in extracting proper representations from complex medical data for various analyzing tasks, such as patient grouping, clinical endpoint prediction and medication recommendation. Medical data can be divided into two typical categories, outpatient and inpatient, that h…
Solves Einstein vacuum equations gluing problem for close Minkowski data.
We determine the expected curvature polynomial of random real projective varieties given as the zero set of independent random polynomials with Gaussian distribution, whose distribution is invariant under the action of the orthogonal group. In particular, the expected Euler characteristic of such random real projective…
The paper explores conditions for homothetic Killing vectors on spacetime hypersurfaces.
Paper glues characteristic data to Kerr spacetime, proving spacelike gluing.
The paper proposes a method to assess surrogate heterogeneity in non-randomized data.
The paper extends Chern-Weil-Lecomte map to -algebras.
In this paper we address the problem of discovering a small set of frequent serial episodes from sequential data so as to adequately characterize or summarize the data. We discuss an algorithm based on the Minimum Description Length (MDL) principle and the algorithm is a slight modification of an earlier method, called…
A subset of a group is characteristic if it is invariant under every automorphism of the group. We study word length in fundamental groups of closed hyperbolic surfaces with respect to characteristic generating sets consisting of a finite union of orbits of the automorphism group, and show that the translation length o…
This paper presents a novel data-driven technique based on the spatiotemporal pattern network (STPN) for energy/power prediction for complex dynamical systems. Built on symbolic dynamic filtering, the STPN framework is used to capture not only the individual system characteristics but also the pair-wise causal dependen…
In this work we introduce a mixture of GPs to address the data association problem, i.e. to label a group of observations according to the sources that generated them. Unlike several previously proposed GP mixtures, the novel mixture has the distinct characteristic of using no gating function to determine the associati…
The explosion of time series data in recent years has brought a flourish of new time series analysis methods, for forecasting, clustering, classification and other tasks. The evaluation of these new methods requires either collecting or simulating a diverse set of time series benchmarking data to enable reliable compar…
We study a Laplacian operator related to the characteristic cohomology of a smooth manifold endowed with a distribution. We prove that this Laplacian does not behave very well: it is not hypoelliptic in general and does not respect the bigrading on forms in a complex setting. We also discuss the consequences of these n…
The effectiveness of biosignal generation and data augmentation with biosignal generative models based on generative adversarial networks (GANs), which are a type of deep learning technique, was demonstrated in our previous paper. GAN-based generative models only learn the projection between a random distribution as in…
Study Euler characteristic of manifolds with almost nonnegative curvature operator, showing nonnegativity under certain conditions.
New framework for interpretable firm characteristics factors.
What are called secondary characteristic classes in Chern-Weil theory are a refinement of ordinary characteristic classes of principal bundles from cohomology to differential cohomology. We consider the problem of refining the construction of secondary characteristic classes from cohomology sets to cocycle spaces; and …
New framework for dense weighted networks with community-specific patterns.
In this paper we present a technique for using the bootstrap to estimate the operating characteristics and their variability for certain types of ensemble methods. Bootstrapping a model can require a huge amount of work if the training data set is large. Fortunately in many cases the technique lets us determine the eff…
Dynamic treatment effects estimated over time using covariate balancing.
Real estate appraisal is a complex and important task, that can be made more precise and faster with the help of automated valuation tools. Usually the value of some property is determined by taking into account both structural and geographical characteristics. However, while geographical information is easily found, o…
Optimal smooth subspaces approximate large data sets efficiently.
Many physical systems are described by partial differential equations (PDEs). Determinism then requires the Cauchy problem to be well-posed. Even when the Cauchy problem is well-posed for generic Cauchy data, there may exist characteristic Cauchy data. Characteristics of PDEs play an important role both in Mathematics …
Paper solves Einstein vacuum equations gluing problem with applications.
In this work a robust clustering algorithm for stationary time series is proposed. The algorithm is based on the use of estimated spectral densities, which are considered as functional data, as the basic characteristic of stationary time series for clustering purposes. A robust algorithm for functional data is then app…
Higher bootstrap rates than 1.0 improve random forest performance.
Paper defines and proves geometric uniqueness of Einstein field equations.
Given a finite simplicial complex L and a collection of pairs of spaces indexed by its vertex set, one can define their polyhedral product. We record a simple formula for its Euler characteristic. In special cases the formula simplifies further to one involving the h-polynomial of L.
Decision forests are widely used for classification and regression tasks. A lesser known property of tree-based methods is that one can construct a proximity matrix from the tree(s), and these proximity matrices are induced kernels. While there has been extensive research on the applications and properties of kernels, …
Current deep learning models are mostly build upon neural networks, i.e., multiple layers of parameterized differentiable nonlinear modules that can be trained by backpropagation. In this paper, we explore the possibility of building deep models based on non-differentiable modules. We conjecture that the mystery behind…
Testing the implementation of deep learning systems and their training routines is crucial to maintain a reliable code base. Modern software development employs processes, such as Continuous Integration, in which changes to the software are frequently integrated and tested. However, testing the training routines requir…
This paper improves fund net value prediction using ARIMA-LSTM hybrid model.