Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

112224335447 · Jun 202019922001200920182026
48 results for automatic evaluation

This paper proposes automatic tuning of Bayesian Optimization's acquisition function.

problem Optimizing black-box functions with noisy, expensive evaluations and hyperparameter tuning.
method Exploring heuristics to automatically tune acquisition functions in Bayesian Optimization.
result Demonstrates effectiveness of heuristics in automatic Bayesian Optimization.

Optimizes crowdsourced preference-based subjective evaluation with online learning.

problem Large-scale evaluation of generative media using crowdsourcing due to combinatorial explosion.
method Automatic optimization of pair combination selections and evaluation volumes with online learning.
result Optimizes evaluation by reducing pair combinations and allocating optimal evaluation volumes.

NGE uses neural graphs to efficiently design robots.

problem Designing robots is hard due to combinatorial search space and evaluation costs.
method Formulated as graph search, NGE uses neural networks for policy parameterization and graph mutation with uncertainty.
result NGE significantly outperforms previous methods, discovering kinematically preferred structures.

New deep learning method validated across multiple sleep staging databases.

problem Improving automatic sleep scoring accuracy across different datasets.
method Ensemble of local models using deep learning for automatic sleep staging.
result Good general performance compared to human experts and state-of-the-art methods.

This paper describes the data collection effort that is part of the project Sprekend Nederland (The Netherlands Talking), and discusses its potential use in Automatic Accent Location. We define Automatic Accent Location as the task to describe the accent of a speaker in terms of the location of the speaker and its hist…

2016-02-08abs ↗pdf ↗

Deep learning predicts pharmaceutical formulations with high accuracy.

problem Laborious, time-consuming and costly traditional trial-and-error approach in pharmaceutical formulation development.
method Used deep learning for automatic feature extraction, developed automatic dataset selection algorithm, compared with six machine learning methods.
result Deep neural networks achieved accuracies above 80% in predicting pharmaceutical formulations.

AVATAR uses a surrogate model to quickly evaluate ML pipelines, saving time and resources.

problem Time-consuming evaluation of ML pipelines limits exploration of complex models.
method AVATAR employs a surrogate model to assess pipeline validity without execution.
result AVATAR accelerates ML pipeline evaluation, improving efficiency in complex scenarios.

Deep Reinforcement Learning (DRL) is a trending field of research, showing great promise in many challenging problems such as playing Atari, solving Go and controlling robots. While DRL agents perform well in practice we are still missing the tools to analayze their performance and visualize the temporal abstractions t…

2016-06-22abs ↗pdf ↗

D-Adaptation automatically sets optimal learning rates without manual tuning.

problem Optimizing learning rates for efficient convergence in machine learning.
method D-Adaptation, which asymptotically achieves optimal learning rates without back-tracking or additional evaluations.
result D-Adaptation automatically matches hand-tuned learning rates across diverse problems.

BayesAME automatically determines coreset size for efficient model evaluation.

problem Time-consuming and computationally expensive evaluation of large generative models.
method Bayesian active model evaluation (BayesAME) that automatically selects coreset size.
result BayesAME consistently outperforms existing methods in diverse benchmarks.

RobustSleepNet automates sleep stage classification for any PSG montage.

problem Manual sleep staging is tedious and expensive; automatic methods are limited by PSG montage and demographic differences.
method RobustSleepNet is a deep learning model that handles arbitrary PSG montages and is robust to demographic changes.
result RobustSleepNet achieves 97% F1 score on unseen data, outperforming specific training datasets by 2%.

Paper proposes efficient method for automatic renal segmentation in DCE-MRI.

problem Automatic segmentation of renal parenchyma in DCE-MRI images.
method Cascaded application of two 3D CNNs for localization and segmentation.
result Achieved high segmentation accuracy with mean dice coefficients of 91.4 and 83.6 for normal and abnormal kidneys, respectively.

EGO optimizes neural network architectures without manual tuning.

problem Designing optimal neural network architectures is difficult and time-consuming.
method Adapted EGO algorithm for efficient optimization of neural network architectures.
result Automatically optimized neural networks achieve competitive performance compared to hand-crafted ones.

TOPNet integrates task-based evaluation into machine learning models.

problem Non-differentiable task-based evaluation criteria in real-world applications.
method Task-Oriented Prediction Network (TOPNet) with learnable surrogate loss function.
result TOPNet significantly outperforms traditional and heuristic models in financial prediction tasks.

Paper proposes a method to automatically detect drift in machine learning models.

problem Detecting changes in class-label data distributions that affect model predictions.
method Self-evaluating predictive model degradation to detect concept drift.
result Effectiveness in automatically detecting and describing concept drift.

Efficient neural networks compute various differential operators cheaply.

problem Efficient computation of higher time complexity differential operators.
method Restricted neural network architectures with diagonal and hollow Jacobian matrices, allowing efficient extraction of dimension-wise derivatives.
result Demonstrated efficient computation of differential operators for various applications.

Convolutional neural networks improve KL grade prediction from Indian knee radiographs.

problem Improving accuracy of knee osteoarthritis grading from Indian radiographs.
method Two-stage approach: object detection followed by regression.
result Fine-tuning model on private hospital data reduces mean absolute error from 1.09 to 0.28.

Paper proposes a deep learning method for automatic seizure detection.

problem Manual seizure identification is time-consuming, labor-intensive, and error-prone.
method Leverages attention mechanism and BiLSTM to capture spatial and temporal features.
result Average sensitivity, specificity, and precision of 87.00%, 88.60%, and 88.63% respectively.

Paper develops a BERT-based classifier to reduce pathology report annotation workload.

problem Manual annotation of pathology reports is labor-intensive and time-consuming.
method Developed an automatic text classifier using BERT and introduced a human-centric metric to identify low-confidence cases.
result The model reduces manual annotation workload by 80% to 98%.

This work evaluates machine learning-based hotspot detectors on synthesized layout patterns.

problem Evaluating model robustness and generality of machine learning-based hotspot detectors.
method Developed an automatic layout generation tool to synthesize various layout patterns and tested machine learning-based detectors on these synthesized layouts.
result Machine learning-based detectors need continuous study for robustness and generality in DFM flows.

This paper improves auto-augment efficiency by sharing augmentation weights.

problem Efficient evaluation of augmentation policies for model training.
method Augmentation-Wise Weight Sharing (AWS) to create a fast yet accurate proxy task.
result Augmentation policies found achieve superior accuracies compared to existing methods.

A framework assesses the quality of crowdsourced weather data.

problem Quality control and assessment of crowdsourced weather data from third-party stations.
method Proposes a simple, scalable, and interpretable AI/Stats/ML framework to assess TPAWS data.
result Demonstrates the performance of the framework using synthetic and real data.

StratPPI improves prediction-powered inference with stratified sampling.

problem Improving statistical estimates with limited human-labeled data.
method Combining small human-labeled data with large automatic-labeled data, stratifying data for tighter confidence intervals.
result StratPPI provides substantially tighter confidence intervals than unstratified approaches.

Automatically finds efficient multi-task models with less data.

problem Training separate models requires more data, parameters, and time.
method Compact search space for multi-task architectures, feature distillation for quick evaluation.
result Automatically identifies multi-task architectures that balance resource requirements and performance.

Study improves machine learning models for GI tract disease detection using comprehensive evaluations and cross-dataset testing.

problem Incomplete or incorrect evaluation of machine learning models for GI tract diseases.
method Comprehensive evaluations of five machine learning models using Global Features and Deep Neural Networks, introducing performance hexagons and cross-dataset testing.
result Demonstrates the need for more sophisticated performance metrics and evaluation methods to build generalizable models.

New method improves hyperparameter tuning efficiency across similar tasks.

problem Mismatch between evaluations in current and previous tasks.
method Nested drop-out and auto-relevance determination for learning basis functions of increasing complexity.
result Improves sample efficiency in hyperparameter tuning across different data regimes.

A popular tool for unsupervised modelling and mining multi-aspect data is tensor decomposition. In an exploratory setting, where and no labels or ground truth are available how can we automatically decide how many components to extract? How can we assess the quality of our results, so that a domain expert can factor th…

2015-03-11abs ↗pdf ↗

DeepLine automates ML pipeline generation using reinforcement learning.

problem Automatic generation of end-to-end ML pipelines combining multiple algorithms.
method Deep Reinforcement Learning with hierarchical actions filtering.
result DeepLine outperforms state-of-the-art approaches in accuracy and computational cost.

Growing interest in automatic speaker verification (ASV)systems has lead to significant quality improvement of spoofing attackson them. Many research works confirm that despite the low equal er-ror rate (EER) ASV systems are still vulnerable to spoofing attacks. Inthis work we overview different acoustic feature spaces…

2017-05-24abs ↗pdf ↗

Automatic debiasing for causal and policy effects using Neural Nets and Random Forests.

problem Estimating causal and policy effects from high-dimensional or non-parametric regression functions.
method Automatic learning of Riesz representation using Neural Nets and Random Forests.
result Automatic debiasing method performs well compared to state-of-the-art algorithms.