Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

3416821,0231,364 · Jun 202019922001200920172026
48 results for well data

The main task in oil and gas exploration is to gain an understanding of the distribution and nature of rocks and fluids in the subsurface. Well logs are records of petro-physical data acquired along a borehole, providing direct information about what is in the subsurface. The data collected by logging wells can have si…

2017-05-10abs ↗pdf ↗

Proves well-posedness for Einstein equations with specific boundary data.

problem Proving well-posedness for Einstein equations with Dirichlet boundary data.
method Local-in-time well-posedness proof for vacuum Einstein equations with specific boundary conditions.
result Proves well-posedness for Einstein equations with Dirichlet boundary data under convexity-type assumptions.

Study proves well-posedness and scattering for wave equations on hyperbolic spaces with singular data.

problem Proving well-posedness and scattering for wave equations on hyperbolic spaces with singular initial data.
method Using weak-LpL^{p} spaces and dispersive estimates on Lorentz spaces, the study establishes global well-posedness and exponential asymptotic stability.
result Developed a scattering theory and constructed wave operators in a singular framework.

New methods predict neural network quality without access to training data.

problem Predicting neural network quality without access to training or testing data.
method Meta-analysis of pretrained models using norm and power law based metrics.
result Power law based metrics can better distinguish well-trained from poorly-trained models.

We find necessary and sufficient conditions ensuring that the vacuum development of an initial data set of the Einstein's field equations admits a conformal Killing vector. We refer to these conditions as conformal Killing initial data (CKID) and they extend the well-known Killing initial data (KID) that have been know…

2019-05-03abs ↗pdf ↗

Proves well-posedness for Einstein equations with specific boundary conditions.

problem Well-posedness of vacuum Einstein equations with twisted Dirichlet boundary conditions.
method Proves local-in-time well-posedness for the IBVP of the Einstein equations with specified conformal class and scalar densities.
result Proves well-posedness for the Einstein equations with twisted Dirichlet boundary conditions.

Using tax and census data, we demonstrate that the distribution of individual income in the USA is exponential. Our calculated Lorenz curve without fitting parameters and Gini coefficient 1/2 agree well with the data. From the individual income distribution, we derive the distribution function of income for families wi…

2000-08-21abs ↗pdf ↗

Bayesian neural networks improve uncertainty in data-driven VFMs for oil and gas wells.

problem Uncertainty and robustness in data-driven VFMs for oil and gas wells.
method Bayesian neural networks with variational inference for uncertainty quantification.
result Variational inference provides more robust predictions on future data.

CAMul forecasts with calibrated and accurate multi-view time-series data.

problem Combining diverse data sources for reliable time-series forecasting.
method CAMul integrates multi-modal data views dynamically, assigning importance based on context.
result CAMul outperforms state-of-the-art models by 25% in accuracy and calibration.

Deep neural networks (DNNs) are incredibly brittle due to adversarial examples. To robustify DNNs, adversarial training was proposed, which requires large-scale but well-labeled data. However, it is quite expensive to annotate large-scale data well. To compensate for this shortage, several seminal works are utilizing l…

2019-11-20abs ↗pdf ↗

Maximum likelihood estimation fails to be well-posed in Gaussian process regression.

problem Establishing well-posedness of maximum likelihood estimation in Gaussian process regression.
method Analyzing the conditions under which maximum likelihood estimation is not Lipschitz in the data with respect to the Hellinger distance.
result Maximum likelihood estimation is not well-posed in the noiseless data setting for any Gaussian process with a stationary covariance function whose lengthscale parameter is estimated using maximum likelihood.

The paper explores how linear neural networks can overfit without bias when data is well-behaved.

problem Understanding why linear neural networks can generalize well despite fitting noisy data.
method Analyzing two-layer linear neural networks trained with gradient flow, deriving bounds on excess risk.
result The excess risk depends on initialization quality and data covariance matrix properties.

Develops an online nonparametric classifier for massive data.

problem Challenges of batch kernel-based nonparametric classifiers in massive data.
method Online principle components analysis to reduce dimensionality, followed by stochastic approximation algorithm for real-time calculation.
result Online classifier provides the best trade-off between accuracy and computation cost.

Deep learning calibrates CO2 storage formations from seismic and well data.

problem Uncertainty in CO2 storage formation properties.
method Two deep learning models for well and seismic data, integrated into MCMC history matching.
result Significant uncertainty reduction in key parameters and accurate CO2 plume predictions.

Machine learning predicts liquid water properties from cluster data.

problem Accuracy of bulk properties from machine-learned potentials is limited by training data.
method Local, atom-centred descriptors enable prediction of bulk properties from cluster data.
result Excellent agreement with experimental and theoretical counterparts of liquid water properties.

Neural networks can interpolate noisy data and still generalize well.

problem Generalization of neural networks trained on noisy data.
method Two-layer neural networks trained to interpolation by gradient descent on corrupted labels.
result Neural networks can achieve zero training error and optimal test error.

Modeling data with linear combinations of a few elements from a learned dictionary has been the focus of much recent research in machine learning, neuroscience and signal processing. For signals such as natural images that admit such sparse representations, it is now well established that these models are well suited t…

2010-09-27abs ↗pdf ↗

This paper tackles few-shot classification by improving GAN-based data augmentation.

problem Improving few-shot classification performance using GANs with limited data.
method Fine-tuning GANs for few-shot classification, addressing training and evaluation challenges.
result Semi-supervised fine-tuning is a more effective approach for few-shot classification with limited data.

SrvfNet aligns multiple functional data to templates without supervision.

problem Aligning large collections of functional data to templates without labeled data.
method Generative deep learning framework using SRVF and fully-connected layers.
result Framework achieves alignment and optimal template prediction without supervision.

Principal components analysis (PCA) is a well-known technique for approximating a tabular data set by a low rank matrix. Here, we extend the idea of PCA to handle arbitrary data sets consisting of numerical, Boolean, categorical, ordinal, and other data types. This framework encompasses many well known techniques in da…

2014-10-01abs ↗pdf ↗

Study well-posedness of fast diffusion equation on noncompact manifolds.

problem Investigate well-posedness of fast diffusion equation in noncompact Riemannian manifolds.
method Establish existence and uniqueness of solutions for globally integrable initial data.
result Global solutions exist for initial data in Lloc1L^1_{\mathrm{loc}} on general Riemannian manifolds.

Study uses LLMs to generate investor briefs from company reports and SEC filings.

problem Improving data analysis for individual investors.
method Preprocessed data, used gpt-4o model in RAG regime, evaluated by investors.
result LLMs can generate useful investor briefs from company reports and SEC filings.

Cellwise outliers challenge traditional methods in statistics and machine learning.

problem Cellwise outliers contaminate over half cases in data matrices, complicating existing methods.
method Requires techniques different from casewise methods, focusing on non-intuitive equivariance properties.
result Substantial progress in estimating location and covariance matrices, regression methods, and tensor data.

Machine learning portfolios perform well with simple imputation of missing data.

problem Handling missing values in machine learning portfolios constructed from cross-sectional return predictors.
method Simple imputation with cross-sectional means compared to rigorous expectation-maximization methods.
result Simple imputation performs well due to the structure of missing data.

While adversarial training can improve robust accuracy (against an adversary), it sometimes hurts standard accuracy (when there is no adversary). Previous work has studied this tradeoff between standard and robust accuracy, but only in the setting where no predictor performs well on both objectives in the infinite data…

2019-06-14abs ↗pdf ↗

Over the years data has become increasingly higher dimensional, which has prompted an increased need for dimension reduction techniques. This is perhaps especially true for clustering (unsupervised classification) as well as semi-supervised and supervised classification. Although dimension reduction in the area of clus…

2017-12-22abs ↗pdf ↗