Novel framework for ML-assisted inference valid for any statistical task.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
FEDHC learns Bayesian networks efficiently for continuous data.
Software development effort estimation is considered a fundamental task for software development life cycle as well as for managing project cost, time and quality. Therefore, accurate estimation is a substantial factor in projects success and reducing the risks. In recent years, software effort estimation has received …
Software helps teach latent variable methods in multivariate data analytics.
Method controls extrapolation in prediction profiles for statistical and machine learning models.
FOSS is an acronym for Free and Open Source Software. The FOSS 2013 survey primarily targets FOSS contributors and relevant anonymized dataset is publicly available under CC by SA license. In this study, the dataset is analyzed from a critical perspective using statistical and clustering techniques (especially multiple…
Method learns software resource usage from snapshots.
A new fuzzy clustering method using hyperbolic smoothing for large datasets.
Models of complex systems are often formalized as sequential software simulators: computationally intensive programs that iteratively build up probable system configurations given parameters and initial conditions. These simulators enable modelers to capture effects that are difficult to characterize analytically or su…
Software package assesses spherical data distributions and clusters.
New GPU algorithm speeds up Gaussian Process analysis.
Linear cost method approximates Gaussian Matérn processes with exponentially convergent accuracy.
This paper introduces the R package sgmcmc; which can be used for Bayesian inference on problems with large datasets using stochastic gradient Markov chain Monte Carlo (SGMCMC). Traditional Markov chain Monte Carlo (MCMC) methods, such as Metropolis-Hastings, are known to run prohibitively slowly as the dataset size in…
Recent advances in cryptography promise to enable secure statistical computation on encrypted data, whereby a limited set of operations can be carried out without the need to first decrypt. We review these homomorphic encryption schemes in a manner accessible to statisticians and machine learners, focusing on pertinent…
DRIFT uses RL to automate functional software testing efficiently.
Multimodal deep learning improves flaw detection in software programs.
KnotPlot helps beginners and veterans use software for visualizing knots.
The task of clustering a set of objects based on multiple sources of data arises in several modern applications. We propose an integrative statistical model that permits a separate clustering of the objects for each data source. These separate clusterings adhere loosely to an overall consensus clustering, and hence the…
The modern data analyst must cope with data encoded in various forms, vectors, matrices, strings, graphs, or more. Consequently, statistical and machine learning models tailored to different data encodings are important. We focus on data encoded as normalized vectors, so that their "direction" is more important than th…
Research proposes an ensemble learning model for efficient software defect prediction.
Improving software quality through effective organizational learning.
Improved software flaw detection using NAS on multimodal DL models.
Method predicts hardware resource usage by control software with guaranteed linear convergence.
Existing language models such as n-grams for software code often fail to capture a long context where dependent code elements scatter far apart. In this paper, we propose a novel approach to build a language model for software code to address this particular issue. Our language model, partly inspired by human memory, i…
This paper tackles co-design of neural hardware and software to improve efficiency.
The public package registry npm is one of the biggest software registry. With its 216 911 software packages, it forms a big network of software dependencies. In this paper we evaluate various methods for finding similar packages in the npm network, using only the structure of the graph. Namely, we want to find a way of…
Although software analytics has experienced rapid growth as a research area, it has not yet reached its full potential for wide industrial adoption. Most of the existing work in software analytics still relies heavily on costly manual feature engineering processes, and they mainly address the traditional classification…
The purpose of this study is to introduce new design-criteria for next-generation hyperparameter optimization software. The criteria we propose include (1) define-by-run API that allows users to construct the parameter search space dynamically, (2) efficient implementation of both searching and pruning strategies, and …
Underlying cause of death coding from death certificates is a process that is nowadays undertaken mostly by humans with a potential assistance from expert systems such as the Iris software. It is as a consequence an expensive process that can in addition suffer from geospatial discrepancies, thus severely impairing the…
Python tools for 3D shape analysis on Kendall's space.
Survey of software developers' experience with Github Copilot tool.
Paper develops multilingual job classification for ISCO and KZiS.
We introduce a new graphical model for tracking radio-tagged animals and learning their movement patterns. The model provides a principled way to combine radio telemetry data with an arbitrary set of userdefined, spatial features. We describe an efficient stochastic gradient algorithm for fitting model parameters to da…
Develops regression trees for estimating cumulative incidence curves in competing risks.
RobPy offers robust statistical methods in Python.
Existing malware detectors on safety-critical devices have difficulties in runtime detection due to the performance overhead. In this paper, we introduce PROPEDEUTICA, a framework for efficient and effective real-time malware detection, leveraging the best of conventional machine learning (ML) and deep learning (DL) te…
Flexible framework for deep distributional regression models.
Neural networks speed up statistical inference.
A deep-learning inference accelerator is synthesized from a C-language software program parallelized with Pthreads. The software implementation uses the well-known producer/consumer model with parallel threads interconnected by FIFO queues. The LegUp high-level synthesis (HLS) tool synthesizes threads into parallel FPG…
Unified platform for statistical and machine learning in bioinformatics.
RAMANMETRIX simplifies Raman spectroscopy data analysis.
HarDNN detects and protects CNNs from hardware errors.
Software estimates inequality in random systems with changing communities.
Targeted Learning uses robust statistics for reproducible research.
AEC Games model represents software MARL environments better than POSGs.
The question of how best to estimate a continuous probability density from finite data is an intriguing open problem at the interface of statistics and physics. Previous work has argued that this problem can be addressed in a natural way using methods from statistical field theory. Here I describe new results that allo…
Forest management relies on the evaluation of silviculture practices. The increase in natural risk due to climate change makes it necessary to consider evaluation criteria that take natural risk into account. Risk integration in existing software requires advanced programming skills.We propose a user-friendly software …
With the rise of social media like Twitter and of software distribution platforms like app stores, users got various ways to express their opinion about software products. Popular software vendors get user feedback thousandfold per day. Research has shown that such feedback contains valuable information for software de…