Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

11.0%22.0%33.0%44.0% · Jun 202019922001200920182026
48 results for data handling

Paper proposes a probabilistic method to handle missing data in decision trees.

problem Handling missing data in decision trees.
method At deployment time, use density estimators to compute expected predictions. At learning time, fine-tune tree parameters to minimize expected prediction loss.
result Effective compared to baselines in experiments.

Paper develops a model for handling outliers and missing data.

problem Handling outliers and missing data in data with scale mixture of normal distributions.
method Developed a scale mixture of Normal distributions model with latent variables. Inference through Variational Bayesian Approximation.
result Model effectively handles outliers and missing data.

New method improves prediction accuracy in business process mining by handling concept drift.

problem Improving prediction quality in business process mining affected by concept drift.
method Systematically analyzed and compared different data selection strategies for retraining machine learning models.
result Improved accuracy from 0.5400 to 0.7010 with concept drift handling.

Trinary decision tree improves handling of missing data in machine learning.

problem Improving accuracy in decision tree algorithms when dealing with missing data.
method Introduces Trinary decision tree, which does not assume missing values contain information about the response.
result Trinary decision tree outperforms other algorithms in Missing Completely at Random settings, especially when data is only missing out-of-sample.

Study on missing data mechanisms and simple imputation methods in fairness of machine learning algorithms.

problem Impact of missing data mechanisms and simple imputation methods on fairness of machine learning algorithms.
method Three popular datasets for classification fairness were used. Missing values were generated using three missing data mechanisms. Various missing data handling techniques (listwise deletion, mean imputation, mode imputation, multiple imputation) were applied to the datasets. Fairness was assessed using classification algorithms (random forests).
result Missing data mechanism does not significantly impact fairness; listwise deletion gives highest fairness on average.

PyTorch Frame simplifies multi-modal tabular learning with modular data and model handling.

problem Handling complex multi-modal tabular data in deep learning.
method A PyTorch-based framework that provides a data structure, model abstraction, and integration with external models.
result Demonstrated the effectiveness of PyTorch Frame in implementing and applying diverse tabular models to complex multi-modal tabular data.

A new approach switches between simple and complex models to handle concept drifts in regression tasks.

problem Handling concept drifts in regression models to maintain accurate predictions over time.
method Error Intersection Approach: switches between simple and complex models based on drift detection.
result The Error Intersection Approach significantly outperforms baselines in handling concept drifts in a real-world taxi demand dataset.

A method for classification using pairwise similarities and unlabeled data.

problem Handling pairwise similarities and unlabeled data for classification.
method Empirical risk minimization approach to create an unbiased risk estimator.
result Derives an unbiased risk estimator for handling both similarities and unlabeled data.

The paper develops methods to handle missing data using regularized M-estimation in reproducing kernel Hilbert space.

problem Handling missing data in statistical analysis.
method Kernel ridge regression for imputation and maximum entropy method for propensity score estimation.
result The proposed methods achieve statistical consistency and asymptotic equivalence.

Improved online classification for manual material handling using wearable sensors.

problem Online monitoring of manual material handling activities using wearable sensors.
method Optimizes dictionary learning to improve sparse representation classification (SRC) accuracy and computational efficiency.
result Proposed method outperforms benchmark methods in accuracy and computational time for online monitoring.

New algorithms handle missing data to improve fairness in machine learning.

problem Missing values in data can lead to unfair outcomes in machine learning models.
method Developed scalable and adaptive algorithms to handle missing values while preserving predictive information.
result Our adaptive algorithms consistently achieve higher fairness and accuracy than standard impute-then-classify methods.

Combines pseudo-point and state space approximations for scalable GPs.

problem Handling large numbers of off-the-grid spatial data-points and long time-series.
method Combines pseudo-point approximations for spatial data with state space GP approximations for temporal data.
result Combined approach is more scalable and applicable to a greater range of spatio-temporal problems.

New method handles missing data and multiple data types in time series models.

problem Handling missing data and multiple data modalities in time series models.
method Factorized inference method for Multimodal Deep Markov Models (MDMMs).
result Method performs well even with high levels of missing data and outperforms existing approaches.

Develops new algorithms for QRF to handle mixed-frequency and longitudinal data.

problem Handling mixed-frequency and longitudinal data in quantile regression.
method Mixed-Frequency Quantile Regression Forest (MIDAS-QRF) and Finite Mixture Quantile Regression Forest (FM-QRF).
result Valid and flexible models for complex empirical settings in financial risk management and climate-change impact evaluation.

Harer-Kas-Kirby conjectured that every handle decomposition of the elliptic surface E(1)_{2,3} requires both 1- and 3-handles. We prove that the elliptic surface E(n)_{p,q} has a handle decomposition without 1-handles for n1n\geq 1 and (p,q)=(2,3),(2,5),(3,4),(4,5).

2008-02-22abs ↗pdf ↗

Study examines how network architecture handles increasing data complexity.

problem Understanding how network architecture affects performance with complex data.
method Empirical study comparing various network architectures on an image classification task with increasing class numbers.
result Modern architectures show better generalization performance with increasing data complexity.

We review Giroux's contact handles and contact handle attachments in dimension three and show that a bypass attachment consists of a pair of contact 1 and 2-handles. As an application we describe explicit contact handle decompositions of infinitely many pairwise non-isotopic overtwisted 3-spheres. We also give an alter…

2009-01-29abs ↗pdf ↗

Paper introduces methods to handle missing data in probabilistic regression trees.

problem Handling missing data in probabilistic regression trees.
method Three approaches: uniform probability, partial observation, and dimension-reduced smoothing.
result Preserves interpretability while extending applicability to incomplete datasets.

Topological Bayesian Optimization finds optimal structures using topological data.

problem Optimizing complex structured data like material or neural network structures.
method Extract topological information from structures using persistent homology, apply Bayesian optimization with kernels for persistence diagrams.
result Topological information improves search efficiency for optimal structures.

D.Nash defined a family of homotopy 4-spheres in [11]. Proving that his manifolds Sm,n,m,n{\mathcal S}_{m,n,m',n'} are all real S4S^4, we find that they have handle decomposition with no 1-handles, two 2-handles and two 3-handles. The handle structures give new potential counterexamples of Property 2R conjecture.

2011-02-15abs ↗pdf ↗

New methods handle both data and network heterogeneity in federated learning.

problem Challenges in federated learning due to data and network heterogeneity.
method Two novel client selection schemes that minimize theoretical runtime to convergence.
result Our methods are at least competitive to and up to 20 times better than existing baselines.

Kernel ridge regression imputation with consistent variance estimation for handling missing data.

problem Handling missing data in statistical analysis.
method Kernel ridge regression imputation combined with entropy method for variance estimation.
result Root-n consistency of the imputation estimator in a Sobolev space setting.

A 2-dimensional braid over an oriented surface-knot FF is presented by a graph called a chart on a surface diagram of FF. We consider 2-dimensional braids obtained by an addition of 1-handles equipped with chart loops. We introduce moves of 1-handles with chart loops, called 1-handle moves, and we investigate how muc…

2015-03-02abs ↗pdf ↗

The paper constructs contractible manifolds with knotted spheres.

problem Creating contractible manifolds with non-standard boundaries.
method Using handles and constructing knotted spheres in SnimesS2S^n imes S^2.
result Contractible (n+3)(n+3)-manifolds with non-standard boundaries are constructed.

This paper uses ML to identify prey handling in seals.

problem Automatically classify prey handling activity in seals for monitoring.
method Developed and compared three ML algorithms: Input Delay Neural Networks, Support Vector Machines, and Echo State Networks.
result Echo State Networks outperformed other algorithms in terms of accuracy and F1score.