The main task in oil and gas exploration is to gain an understanding of the distribution and nature of rocks and fluids in the subsurface. Well logs are records of petro-physical data acquired along a borehole, providing direct information about what is in the subsurface. The data collected by logging wells can have si…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proves well-posedness for Einstein equations with specific boundary data.
Study proves well-posedness and scattering for wave equations on hyperbolic spaces with singular data.
We establish both local and global well-posedness for the heat flow of polyharmonic maps from to a compact Riemannian manifold without boundary for initial data with small BMO norms.
New methods predict neural network quality without access to training data.
We investigate the well-posedness of (i) the heat flow of harmonic maps from to a compact Riemannian manifold without boundary for initial data in BMO; and (ii) the hydrodynamic flow of nematic liquid crystals on for initial data in .
We find necessary and sufficient conditions ensuring that the vacuum development of an initial data set of the Einstein's field equations admits a conformal Killing vector. We refer to these conditions as conformal Killing initial data (CKID) and they extend the well-known Killing initial data (KID) that have been know…
Proves well-posedness for Einstein equations with specific boundary conditions.
New findings on robust learning with well-separated data.
Efficient approach improves prediction calibration for domain shifts.
Measuring domain relevance of data and identifying or selecting well-fit domain data for machine translation (MT) is a well-studied topic, but denoising is not yet. Denoising is concerned with a different type of data quality and tries to reduce the negative impact of data noise on MT training, in particular, neural MT…
Using tax and census data, we demonstrate that the distribution of individual income in the USA is exponential. Our calculated Lorenz curve without fitting parameters and Gini coefficient 1/2 agree well with the data. From the individual income distribution, we derive the distribution function of income for families wi…
Industry datasets used for text classification are rarely created for that purpose. In most cases, the data and target predictions are a by-product of accumulated historical data, typically fraught with noise, present in both the text-based document, as well as in the targeted labels. In this work, we address the quest…
We survey recent work on local well-posedness results for parabolic equations and systems with rough initial data.
For data sets populated by a very well modeled process and by another process of unknown probability density function (PDF), a desired feature when manipulating the fraction of the unknown process (either for enhancing it or suppressing it) consists in avoiding to modify the kinematic distributions of the well modeled …
Bayesian neural networks improve uncertainty in data-driven VFMs for oil and gas wells.
Recent studies have shown that online portfolio selection strategies that exploit the mean reversion property can achieve excess return from equity markets. This paper empirically investigates the performance of state-of-the-art mean reversion strategies on real market data. The aims of the study are twofold. The first…
CAMul forecasts with calibrated and accurate multi-view time-series data.
Deep neural networks (DNNs) are incredibly brittle due to adversarial examples. To robustify DNNs, adversarial training was proposed, which requires large-scale but well-labeled data. However, it is quite expensive to annotate large-scale data well. To compensate for this shortage, several seminal works are utilizing l…
Maximum likelihood estimation fails to be well-posed in Gaussian process regression.
A new metric-based principal curve method learns 1D manifolds from spatial data.
This paper establish the local (or global, resp.) well-posedness of the heat flow of biharmonic maps from to a compact Riemannian manifold without boundary with small local BMO (or BMO, resp.) norms.
t-SNE fails to reveal clusters even in well-clusterable data.
The paper explores how linear neural networks can overfit without bias when data is well-behaved.
With the advent of Big Data era, data reduction methods are highly demanded given its ability to simplify huge data, and ease complex learning processes. Concretely, algorithms that are able to filter relevant dimensions from a set of millions are of huge importance. Although effective, these techniques suffer from the…
Develops an online nonparametric classifier for massive data.
Deep learning calibrates CO2 storage formations from seismic and well data.
Machine learning predicts liquid water properties from cluster data.
Neural networks can interpolate noisy data and still generalize well.
Modeling data with linear combinations of a few elements from a learned dictionary has been the focus of much recent research in machine learning, neuroscience and signal processing. For signals such as natural images that admit such sparse representations, it is now well established that these models are well suited t…
The prediction of the gas production from mature gas wells, due to their complex end-of-life behavior, is challenging and crucial for operational decision making. In this paper, we apply a modified deep LSTM model for prediction of the gas flow rates in mature gas wells, including the uncertainties in input parameters.…
This paper tackles few-shot classification by improving GAN-based data augmentation.
Deep Gaussian processes on manifolds improve performance on complex data.
SrvfNet aligns multiple functional data to templates without supervision.
Principal components analysis (PCA) is a well-known technique for approximating a tabular data set by a low rank matrix. Here, we extend the idea of PCA to handle arbitrary data sets consisting of numerical, Boolean, categorical, ordinal, and other data types. This framework encompasses many well known techniques in da…
Study well-posedness of fast diffusion equation on noncompact manifolds.
Study uses LLMs to generate investor briefs from company reports and SEC filings.
Cellwise outliers challenge traditional methods in statistics and machine learning.
Accident detection is a vital part of traffic safety. Many road users suffer from traffic accidents, as well as their consequences such as delay, congestion, air pollution, and so on. In this study, we utilize two advanced deep learning techniques, Long Short-Term Memory (LSTM) and Gated Recurrent Units (GRUs), to dete…
A new kNN imputation method improves classification performance on datasets with missing data.
Framework generates fair synthetic data to avoid biases.
Machine learning portfolios perform well with simple imputation of missing data.
Finding the diameter of a dataset in multidimensional Euclidean space is a well-established problem, with well-known algorithms. However, most of the algorithms found in the literature do not scale well with large values of data dimension, so the time complexity grows exponentially in most cases, which makes these algo…
While adversarial training can improve robust accuracy (against an adversary), it sometimes hurts standard accuracy (when there is no adversary). Previous work has studied this tradeoff between standard and robust accuracy, but only in the setting where no predictor performs well on both objectives in the infinite data…
We study properties of Sobolev-type metrics on the space of immersed plane curves. We show that the geodesic equation for Sobolev-type metrics with constant coefficients of order 2 and higher is globally well-posed for smooth initial data as well as initial data in certain Sobolev spaces. Thus the space of closed plane…
Over the years data has become increasingly higher dimensional, which has prompted an increased need for dimension reduction techniques. This is perhaps especially true for clustering (unsupervised classification) as well as semi-supervised and supervised classification. Although dimension reduction in the area of clus…
Extracting latent low-dimensional structure from high-dimensional data is of paramount importance in timely inference tasks encountered with `Big Data' analytics. However, increasingly noisy, heterogeneous, and incomplete datasets as well as the need for {\em real-time} processing of streaming data pose major challenge…
Learning data representations that reflect the customers' creditworthiness can improve marketing campaigns, customer relationship management, data and process management or the credit risk assessment in retail banks. In this research, we adopt the Variational Autoencoder (VAE), which has the ability to learn latent rep…