Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

55110165220 · Jun 202019922001200920172026
48 results for research guidelines

Current advances in research, development and application of artificial intelligence (AI) systems have yielded a far-reaching discourse on AI ethics. In consequence, a number of ethics guidelines have been released in recent years. These guidelines comprise normative principles and recommendations aimed to harness the …

2019-02-28abs ↗pdf ↗

This chapter introduces reproducibility in machine learning for medical imaging.

problem Lack of reproducibility in machine learning for medical imaging.
method Distinguishes and defines types of reproducibility, outlines requirements, and discusses utility.
result Discussion on benefits and a plea for a non-dogmatic approach to reproducibility.

This research proposes methods to model and assess liability liquidity risk in asset management.

problem Lack of standardized models for liability liquidity risk in asset management.
method Statistical models, zero-inflated models, aggregate and individual-based approaches, and factor models.
result Developed mathematical and statistical approaches to estimate and assess redemption shocks.

Survival analysis models predict economic convergence across Americas.

problem Analyzing GDP per capita trajectories and convergence across the Americas.
method Survival analysis, machine learning, economic interpretation.
result DeepSurv captures non-linear interactions in GDP per capita trajectories.

This manuscript addresses the problem of the automatic lesion boundary detection in dermoscopy, using deep neural networks. An approach is based on the adaptation of the U-net convolutional neural network with skip connections for lesion boundary segmentation task. I hope this paper could serve, to some extent, as an e…

2018-11-23abs ↗pdf ↗

Mixed data comprises both numeric and categorical features, and mixed datasets occur frequently in many domains, such as health, finance, and marketing. Clustering is often applied to mixed datasets to find structures and to group similar objects for further analysis. However, clustering mixed data is challenging becau…

2018-11-11abs ↗pdf ↗

Sherpa.ai framework combines federated learning and differential privacy for edge AI services.

problem Protecting data privacy in edge AI services.
method Holistic federated learning and differential privacy approach with methodological guidelines.
result Demonstrated through classification and regression use cases.

The study examines how experimental design choices affect machine learning model performance.

problem Lack of guidelines on choosing experimental designs and machine learning models.
method 12 experimental designs, 7 families of predictive models, 7 test functions, 8 noise settings.
result Guidelines for practical applications of DOE and ML are provided.

We conduct an extensive evaluation of price jump tests based on high-frequency financial data. After providing a concise review of multiple alternative tests, we document the size and power of all tests in a range of empirically relevant scenarios. Particular focus is given to the robustness of test performance to the …

2017-08-31abs ↗pdf ↗

The paper examines the reliability of limit order book representations in the face of data perturbation.

problem The reliability of limit order book representations under data perturbation.
method Experimental analysis of existing representations and guidelines for future research.
result Existing representations of limit order book data are vulnerable to data perturbation.

Scikit-multiflow is a multi-output/multi-label and stream data mining framework for the Python programming language. Conceived to serve as a platform to encourage democratization of stream learning research, it provides multiple state of the art methods for stream learning, stream generators and evaluators. scikit-mult…

2018-07-12abs ↗pdf ↗

The paper provides guidelines for choosing between SBI methods in complex biological models.

problem Choosing appropriate SBI methods for real-world biological data.
method Comprehensive guidelines and application to agent-based models.
result Statistical SBI methods outperform neural SBI methods with sufficient computational resources.

Third part of a study on liquidity risk in asset management, focusing on managing the asset-liability liquidity risk.

problem Managing the asset-liability liquidity risk in asset management.
method Develops a methodological and practical framework for liquidity stress testing programs.
result Proposes measurement, management, and monitoring tools for controlling the liquidity gap.

Data quality issues have attracted widespread attention due to the negative impacts of dirty data on data mining and machine learning results. The relationship between data quality and the accuracy of results could be applied on the selection of the appropriate algorithm with the consideration of data quality and the d…

2018-03-16abs ↗pdf ↗

This study evaluates different normalizing flow architectures for MCMC.

problem Lack of systematic comparison of normalizing flow architectures in MCMC.
method Extensive evaluation of various normalizing flow architectures on different MCMC methods and target distributions.
result Contractive residual flows are the best general-purpose models for MCMC.

Framework for applying GPs to real-world data with scalability guidelines.

problem Deployment of Gaussian Processes (GPs) is hindered by computational costs and lack of guidelines.
method Proposed a framework for identifying GP suitability and setting up robust models, formalizing decisions of experienced practitioners.
result More accurate results at test time for glacier elevation change case study.

Reinforcement learning is a promising approach to developing hard-to-engineer adaptive solutions for complex and diverse robotic tasks. However, learning with real-world robots is often unreliable and difficult, which resulted in their low adoption in reinforcement learning research. This difficulty is worsened by the …

2018-03-19abs ↗pdf ↗

This review synthesizes uncertainty modeling in probabilistic image segmentation.

problem Relaxed Bayesian assumptions lead to missing uncertainty information in deep models.
method Standardizes theory, notation, and terminology for feature- and parameter-distribution modeling.
result Establishes a common framework for robust decision-making in segmentation tasks.

Unified survey of treatment effect heterogeneity and uplift modeling methods.

problem Estimating heterogeneous treatment effects and uplift modeling.
method Unified survey of treatment effect heterogeneity and uplift modeling approaches.
result Unified notations for comparing methods and applications in personalized marketing, medicine, and social studies.

The research proposes a stopping rule for reinforcement learning algorithms based on instance-dependent confidence.

problem Dramatic variation in convergence rates of reinforcement learning algorithms due to problem structure.
method Develops instance-dependent confidence regions and a data-dependent stopping rule for MDP policy evaluation and optimal value estimation.
result Proposes a stopping rule that adapts to the instance-specific difficulty of the problem, allowing for early termination.

Improving cancer treatment decisions requires considering causal effects, not just model accuracy.

problem Cancer outcome prediction models may cause harm when used for treatment decisions.
method Explains the importance of considering causal effects in model validation and provides guidelines.
result Building and validating models that are useful for decision making requires considering causal effects.

Paper derives policy rules from observational data for hepatitis C treatment.

problem Improving treatment guidelines for HIV/HCV co-infected patients.
method Weighted K-means algorithm for estimating CATEs, decision tree implementation.
result Identifies a subgroup with high spontaneous HCV clearance rate.

Research proposes a model to estimate transaction costs and assess asset liquidity risk.

problem Lack of standardized models for asset liquidity risk in asset management.
method Develops a market impact model and a two-regime model based on power-law property.
result Defines liquidity measures and applies model to stocks and bonds.

This paper asks, "Do classics exist in megaproject management?" We identify three types of classic texts: conventional, Kuhnian, and citation classics. We find that the answer to our question depends on the definition of "classic" employed. First, "citation classics" do exist in megaproject management, and they perform…

2017-09-06abs ↗pdf ↗

Scale of data and scale of computation infrastructures together enable the current deep learning renaissance. However, training large-scale deep architectures demands both algorithmic improvement and careful system configuration. In this paper, we focus on employing the system approach to speed up large-scale training.…

2017-08-10abs ↗pdf ↗

This paper uses deep learning to analyze sentiment in financial forums and improve stock market prediction.

problem Improving stock market prediction accuracy through sentiment analysis.
method Crawling financial forum data, training BERT model on financial corpus, and using maximum information coefficient.
result Sentiment features from financial text can reflect stock market fluctuations and improve prediction accuracy.

This study evaluates methods to measure traffic forecasting model confidence.

problem Lack of consensus on uncertainty types and techniques for traffic forecasting models.
method Reviews and compares different uncertainty estimation techniques using real traffic data.
result Empirical evidence shows benefits and caveats of various techniques.