Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

12.6%25.1%37.7%50.3% · Jun 202019922001200920172026
48 results for data properties

Machine learning predicts liquid water properties from cluster data.

problem Accuracy of bulk properties from machine-learned potentials is limited by training data.
method Local, atom-centred descriptors enable prediction of bulk properties from cluster data.
result Excellent agreement with experimental and theoretical counterparts of liquid water properties.

Hybridizes physical and data-driven methods for predicting physicochemical properties.

problem Predicting physicochemical properties accurately using limited data.
method Distills physical method predictions into a prior model and combines with sparse experimental data using Bayesian inference.
result Significant improvements in predicting activity coefficients at infinite dilution compared to baselines and ensemble methods.

Multi-party machine learning leaks global dataset properties even with black-box access.

problem Leakage of global dataset properties in multi-party machine learning.
method Demonstrated leakage of sensitive attribute distributions in pooled data.
result A curious party can infer sensitive attribute distributions in other parties' data with high accuracy.

This paper analyzes how data and model properties affect membership inference attacks.

problem Understanding and mitigating the vulnerability of machine learning models to membership inference attacks.
method Empirical analysis of data and model properties on MIA success.
result Data and model properties, not just model overfitting, influence MIA success.

The paper discusses regularization properties of artificial data for deep learning. Artificial datasets allow to train neural networks in the case of a real data shortage. It is demonstrated that the artificial data generation process, described as injecting noise to high-level features, bears several similarities to e…

2019-08-19abs ↗pdf ↗

Paper tackles robust prediction of nuclear reactor materials under scarce data.

problem Challenges of data scarcity and uncertainty in nuclear reactor design.
method Meta-learning approach informed by uncertainty and prior knowledge.
result Achieves superior performance in rupture life prediction.

Two ML frameworks predict antibody properties using structural data.

problem Predicting antibody properties using sequence and structural data.
method ANTIPASTI and INFUSSE models using graph representations and neural networks.
result ANTIPASTI predicts binding affinity; INFUSSE predicts residue flexibility.

The study investigates how data variability impacts the generalization of neural networks.

problem Understanding the impact of data variability on neural network generalization.
method Developed a field-theoretic formalism to compute generalization properties of neural networks, focusing on data variability.
result Data variability leads to non-Gaussian action, affecting the learning curve and generalization properties of neural networks.

Study builds ML models to predict fuel properties accurately.

problem Accurate prediction of liquid fuel properties over a wide range of conditions.
method Used Gaussian Processes and probabilistic conditional generative learning to train ML models on fuel density data.
result ML models can predict fuel properties accurately across various pressure and temperature conditions.

Enhances machine learning models by preserving data structure, addressing statistical distortions.

problem Statistical distortions in synthetic data generated by Mixup.
method Proposes a generalized mixup method with a flexible weighting scheme to preserve data structure.
result Preserves statistical properties of original data while maintaining model performance.

Multitask Gaussian process regression reduces data generation costs for molecular property prediction.

problem Data bottleneck in training surrogate models for molecular properties.
method Multitask Gaussian process regression over heterogeneous data sources (CC and DFT).
result Predicts at CC-level accuracy with over an order of magnitude reduction in data generation cost.

New method filters large networks from financial data to reveal key subnetworks.

problem Filtering large dimensional networks to isolate key constituents.
method Exploits spectral properties of high-dimensional data networks, tuning for sparsity and consistency.
result Shows method can interpolate between zero and maximal filtering, preserving spectral properties.

Deep models generate geometric objects with global properties.

problem Comparing neural models' global properties from generated samples.
method Training on datasets of reflexive polytopes, comparing different representations.
result Models learn non-trivial global properties of geometric objects.

A new model uses Toeplitz matrices to analyze time-series data transitions.

problem Analyzing transitions in time-series data from nonautonomous systems.
method Deep Koopman-layered models with learnable Toeplitz matrices, leveraging Toeplitz matrices' universal property.
result The model demonstrates universality and generalization, outperforming existing methods.

Kernelized PCovR reveals structure-property relations in chemistry and materials.

problem Understanding structure-property relations in complex systems.
method Kernel Principal Covariates Regression (kernel PCovR) with sparsification.
result Kernelized PCovR effectively reveals and predicts structure-property relations.

Proposes a new model to predict polymer properties by integrating various data types.

problem Inaccurate polymer property prediction due to separate modeling of different data types.
method Multi-modal cascade feature transfer using GCN for chemical structure and molecular descriptors.
result Empirically evaluated model shows higher predictive performance than single-feature approaches.

Jointly estimates flow fields and particle properties from Lagrangian data.

problem Estimating flow fields and particle properties from sparse, noisy Lagrangian data.
method Data assimilation framework coupling Eulerian and Lagrangian models.
result Joint estimation of flow fields and particle properties in various flow regimes.

Study shows explanation disparities in machine learning models are influenced by data and model properties.

problem Disparities in post-hoc machine learning explanation methods across race and gender.
method Simulations and experiments on a real-world dataset to assess challenges to explanation disparities.
result Increased covariate shift, concept shift, and omission of covariates increase explanation disparities, especially for neural network models.

XIMP improves molecular property prediction by integrating multiple graph representations.

problem Graph neural networks struggle in data-scarce regimes and fail to surpass traditional methods.
method Cross-graph inter-message passing with multiple graph abstractions.
result XIMP outperforms state-of-the-art baselines across diverse molecular property tasks.

Paper introduces a method to predict molecule properties from diverse data sources.

problem Limited ability to accommodate scarce or fragmented training data.
method Adaptive Invariance using invariant risk minimization to generalize beyond heterogeneous data.
result Predictor outperforms state-of-the-art transfer learning methods by significant margin.

Automates GNN design for molecular property prediction.

problem Designing and tuning GNN architectures for molecular property prediction is labor-intensive.
method Developed a NAS approach to automatically discover high-performing GNN architectures for MPNNs.
result Automatically discovered MPNNs outperform manually designed GNNs in molecular property prediction.

This study compares transfer learning and self-supervised learning for better model performance.

problem Choosing between transfer learning and self-supervised learning for optimal model performance.
method Comprehensive comparative study of transfer learning and self-supervised learning under various data and task properties.
result Self-supervised learning outperforms transfer learning in certain applications, and vice versa.

Paper tackles multi-task learning for molecular property prediction with limited data.

problem Limited labeled data for each molecular property task in drug discovery.
method Proposes SGNN-EBM method to utilize relation graph between tasks and improve multi-task learning performance.
result Empirical results show the effectiveness of SGNN-EBM.

Deep learning model improves seismic rock property estimation.

problem Estimating reservoir rock properties from seismic reflection data.
method Proposes a deep learning-based seismic inversion workflow that models seismic traces spatiotemporally.
result Achieves best performance on SEAM dataset with r2r^{2} coefficient of 79.77\%

In this paper, a novel approach for coding nominal data is proposed. For the given nominal data, a rank in a form of complex number is assigned. The proposed method does not lose any information about the attribute and brings other properties previously unknown. The approach based on these knew properties can been used…

2016-01-08abs ↗pdf ↗

Scaling properties in financial fluctuations are reviewed from the standpoint of statistical physics. We firstly show theoretically that the balance of demand and supply enhances fluctuations due to the underlying phase transition mechanism. By analyzing tick data of yen-dollar exchange rates we confirm two fractal pro…

2000-08-03abs ↗pdf ↗

With the proliferation of mobile devices and the internet of things, developing principled solutions for privacy in time series applications has become increasingly important. While differential privacy is the gold standard for database privacy, many time series applications require a different kind of guarantee, and a…

2017-07-10abs ↗pdf ↗

Enhances scenario approach for certifying design properties post-design.

problem Certifying additional useful properties in designs not considered during the design phase.
method Two-level framework of appropriateness: baseline and post-design. Distribution-free upper bounds on risk derived.
result Distribution-free upper bounds on the risk of failing to meet post-design appropriateness.

The article proposes a deep learning method to test and infer the Markov property in time series data.

problem Testing and inferring the Markov property in high-dimensional time series data.
method Deep conditional generative learning to estimate conditional density functions and derive a doubly robust test statistic.
result The test controls the type-I error asymptotically and has power approaching one.

New method calibrates uncertainty in molecular property predictions.

problem Uncalibrated uncertainty estimates in molecular property predictions.
method Message Passing Neural Networks with calibrated probabilistic predictive distribution.
result Accurate molecular formation energy predictions with well-calibrated uncertainty.

ADS explains object differences by quantifying and removing underlying properties.

problem Explaining differences between two object images.
method Align-Deform-Subtract (ADS) framework that uses semantic alignments and iterative quantification/removal of differences.
result ADS provides disentangled error measures explaining object differences in terms of underlying properties.

We introduce generative adversarial models in which the discriminator is replaced by a calibrated (non-differentiable) classifier repeatedly enhanced by domain relevant features. The role of the classifier is to prove that the actual and generated data differ over a controlled semantic space. We demonstrate that such m…

2019-10-07abs ↗pdf ↗

Paper optimizes material microstructures with limited data using probabilistic methods.

problem Optimizing material properties with uncertain process-structure-property links.
method Flexible probabilistic formulation, data-driven surrogate, active learning.
result Significant improvement in accuracy with small training data.