Automatically assesses the quality of online health articles.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
IGGP learns game rules from varying quality game play, finding no overall trend.
SpecGrad improves neural vocoder sound quality by adapting diffusion noise to log-mel spectrogram.
We study the effect of the quality and quantity of side information on the recovery of a hidden community of size in a graph of size . Side information for each node in the graph is modeled by a random vector with the following features: either the dimension of the vector is allowed to vary with , while …
DeepSubQE estimates translation quality for subtitles, improving on existing methods.
ALPCAHUS clusters data from multiple subspaces with varying noise.
There is growing interest in applying machine learning methods to Electronic Medical Records (EMR). Across different institutions, however, EMR quality can vary widely. This work investigated the impact of this disparity on the performance of three advanced machine learning algorithms: logistic regression, multilayer p…
Empirical median performs well in estimating location with varying scales.
With super-resolution optical microscopy, it is now possible to observe molecular interactions in living cells. The obtained images have a very high spatial precision but their overall quality can vary a lot depending on the structure of interest and the imaging parameters. Moreover, evaluating this quality is often di…
We seek to better understand the difference in quality of the several publicly released embeddings. We propose several tasks that help to distinguish the characteristics of different embeddings. Our evaluation of sentiment polarity and synonym/antonym relations shows that embeddings are able to capture surprisingly nua…
We propose an alternative generator architecture for generative adversarial networks, borrowing from style transfer literature. The new architecture leads to an automatically learned, unsupervised separation of high-level attributes (e.g., pose and identity when trained on human faces) and stochastic variation in the g…
A new method generates counterfactual treatment outcomes for time-varying treatments.
Local EGOP learns functions varying along a few directions.
The presence of discrete dividends complicates the derivation and form of pricing formulas even for vanilla options. Existing analytic, numerical, and theoretical approximations provide results of varying quality and performance. Here, we compare the analytic approach, developed and effective for European puts and call…
Study analyzes neural network models to understand generalization performance.
Domain adaptation provides a powerful set of model training techniques given domain-specific training data and supplemental data with unknown relevance. The techniques are useful when users need to develop models with data from varying sources, of varying quality, or from different time ranges. We build CrossTrainer, a…
The paper investigates how dataset quality and heterogeneity affect model confidence in machine learning.
New methods for handling time-varying label noise in time series classification.
The performance of a reinforcement learning algorithm can vary drastically during learning because of exploration. Existing algorithms provide little information about the quality of their current policy before executing it, and thus have limited use in high-stakes applications like healthcare. We address this lack of …
Machine learning models have demonstrated vulnerability to adversarial attacks, more specifically misclassification of adversarial examples. In this paper, we propose a one-off and attack-agnostic Feature Manipulation (FM)-Defense to detect and purify adversarial examples in an interpretable and efficient manner. The i…
Many graph clustering quality functions suffer from a resolution limit, the inability to find small clusters in large graphs. So called resolution-limit-free quality functions do not have this limit. This property was previously introduced for hard clustering, that is, graph partitioning. We investigate the resolution-…
Develops ML-DQA for healthcare data quality assurance.
The paper proposes a novel Kernelized image segmentation scheme for noisy images that utilizes the concept of Smallest Univalue Segment Assimilating Nucleus (SUSAN) and incorporates spatial constraints by computing circular colour map induced weights. Fuzzy damping coefficients are obtained for each nucleus or center p…
CDC-FM improves generative model quality-generalization tradeoff by regularizing with geometry-aware noise.
The paper uses interpretable ML to secure data quality in IoT edge computing.
Model explains how stablecoin runs are influenced by large sales and reserve quality.
We demonstrate the use of conditional autoregressive generative models (van den Oord et al., 2016a) over a discrete latent space (van den Oord et al., 2017b) for forward planning with MCTS. In order to test this method, we introduce a new environment featuring varying difficulty levels, along with moving goals and obst…
This paper presents a new approach for filter design based on stochastic distances and tests between distributions. A window is defined around each pixel, samples are compared and only those which pass a goodness-of-fit test are used to compute the filtered value. The technique is applied to intensity Synthetic Apertur…
This paper formed part of a preliminary research report for a risk consultancy and academic research. Stochastic Programming models provide a powerful paradigm for decision making under uncertainty. In these models the uncertainties are represented by a discrete scenario tree and the quality of the solutions obtained i…
AdaPID optimizes diffusion-based samplers by dynamically adjusting schedules.
A new restart criterion for k-means++ improves clustering quality and adapts to data difficulty.
A simple method treats heteroscedastic variance variatively, improving model calibration and sample quality.
National statistical systems are the enterprises tasked with collecting, validating and reporting societal attributes. These data serve many purposes - they allow governments to improve services, economic actors to traverse markets, and academics to assess social theories. National statistical systems vary in quality, …
This research introduces a new strategy in cluster ensemble selection by using Independency and Diversity metrics. In recent years, Diversity and Quality, which are two metrics in evaluation procedure, have been used for selecting basic clustering results in the cluster ensemble selection. Although quality can improve …
New algorithm reduces regret and constraint violation in online convex optimization with predictions.
A popular approach of achieving fairness in optimization problems is by constraining the solution space to "fair" solutions, which unfortunately typically reduces solution quality. In practice, the ultimate goal is often an aggregate of sub-goals without a unique or best way of combining them or which is otherwise only…
Improves magnetic field mapping using an array of magnetometers with noisy input.
Novel Orlicz regrets consistently bound environmental variable statistics.
QC methods improve reliability of machine learning-based image segmentation.
Study finds visual explanations do not significantly improve human accuracy or trust in model predictions.
Our goal is to extract meaningful transformations from raw images, such as varying the thickness of lines in handwriting or the lighting in a portrait. We propose an unsupervised approach to learn such transformations by attempting to reconstruct an image from a linear combination of transformations of its nearest neig…
Increasing availability of vehicle GPS data has created potentially transformative opportunities for traffic management, route planning and other location-based services. Critical to the utility of the data is their accuracy. Map-matching is the process of improving the accuracy by aligning GPS data with the road netwo…
Bayesian methods improve text annotation quality.
This study analyzes how colonial rice trade in prewar Japan affected its rice market, considering several government interventions in the two rice futures exchanges in Tokyo and Osaka. We explore the interventions in the futures markets using two procedures. First, we measure the joint degree of efficiency in the marke…
P3BO optimizes biological sequence design by combining multiple methods.
The study forecasts water quality from satellite data using machine learning.
This study examines how DMMs affect market liquidity and competition.
Ultrasound (US) imaging is based on the time-reversal principle, in which individual channel RF measurements are back-propagated and accumulated to form an image after applying specific delays. While this time reversal is usually implemented as a delay-and-sum (DAS) beamformer, the image quality quickly degrades as the…