Advocates for Marr's levels of analysis to unify machine learning debates.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The Frame Problem (FP) is a puzzle in philosophy of mind and epistemology, articulated by the Stanford Encyclopedia of Philosophy as follows: "How do we account for our apparent ability to make decisions on the basis only of what is relevant to an ongoing situation without having explicitly to consider all that is not …
The Availability bias, manifested in the over-representation of extreme eventualities in decision-making, is a well-known cognitive bias, and is generally taken as evidence of human irrationality. In this work, we present the first rational, metacognitive account of the Availability bias, formally articulated at Marr's…
A new algorithm improves Wasserstein discriminant analysis for better data classification.
The level crossing and inverse statistics analysis of DAX and oil price time series are given. We determine the average frequency of positive-slope crossings, , where is the average waiting time for observing the level again. We estimate the probability , which provides us the probab…
A Python tool generates synthetic data for cluster analysis from high-level descriptions.
New algorithm stabilizes bi-level hyperparameter optimization.
As mobile devices become more and more popular, mobile gaming has emerged as a promising market with billion-dollar revenues. A variety of mobile game platforms and services have been developed around the world. A critical challenge for these platforms and services is to understand the churn behavior in mobile games, w…
This paper introduces compositional data analysis for financial ratios, improving industry-level analysis.
DeepSupp detects financial support levels using attention mechanisms.
Probe-level models have led to improved performance in microarray studies but the various sources of probe-level contamination are still poorly understood. Data-driven analysis of probe performance can be used to quantify the uncertainty in individual probes and to highlight the relative contribution of different noise…
Factor analysis provides linear factors that describe relationships between individual variables of a data set. We extend this classical formulation into linear factors that describe relationships between groups of variables, where each group represents either a set of related variables or a data set. The model also na…
Symbolic data analysis (SDA) is an emerging area of statistics concerned with understanding and modelling data that takes distributional form (i.e. symbols), such as random lists, intervals and histograms. It was developed under the premise that the statistical unit of interest is the symbol, and that inference is requ…
Novel approach models life events using causal discovery and survival analysis.
Study finds environmental liability insurance reduces industrial carbon emissions.
Local Neural Operators enable efficient system-level analysis of complex PDEs.
BSD is a Bayesian framework for analyzing neural spectral data.
Introduces Motion Programs for better video analysis of human motion.
The so-called level crossing analysis has been used to investigate the empirical data set. But there is a lack of interpretation for what is reflected by the level crossing results. The fractional Gaussian noise as a well-defined stochastic series could be a suitable benchmark to make the level crossing findings more s…
The level set tree approach of Hartigan (1975) provides a probabilistically based and highly interpretable encoding of the clustering behavior of a dataset. By representing the hierarchy of data modes as a dendrogram of the level sets of a density estimator, this approach offers many advantages for exploratory analysis…
Bayesian CNN improves MRI stroke diagnosis accuracy and uncertainty quantification.
Quantitative Investment, built on the solid foundation of robust financial theories, is at the center stage in investment industry today. The essence of quantitative investment is the multi-factor model, which explains the relationship between the risk and return of equities. However, the multi-factor model generates e…
We applied Generative Adversarial Networks (GANs) to learn a model of DOOM levels from human-designed content. Initially, we analysed the levels and extracted several topological features. Then, for each level, we extracted a set of images identifying the occupied area, the height map, the walls, and the position of ga…
Linear Discriminant Analysis (LDA) is a well-known method for dimensionality reduction and classification. Previous studies have also extended the binary-class case into multi-classes. However, many applications, such as object detection and keyframe extraction cannot provide consistent instance-label pairs, while LDA …
Short-term load forecasting (STLF) is essential for the reliable and economic operation of power systems. Though many STLF methods were proposed over the past decades, most of them focused on loads at high aggregation levels only. Thus, low-aggregation load forecast still requires further research and development. Comp…
Study uses image analysis to predict MSI status in tumors.
This paper presents an analysis of the study variables such as gdp, employment levels, the level of R & D and technology that will serve as the basis for stochastic modeling of production possibilities frontier in the goodness of fractal dimensions Ex Ante and Ex Post a priori to determine the levels of causality immed…
Hybrid models improve groundwater level prediction and uncertainty analysis.
In this paper, we use several techniques with conventional vocal feature extraction (MFCC, STFT), along with deep-learning approaches such as CNN, and also context-level analysis, by providing the textual data, and combining different approaches for improved emotion-level classification. We explore models that have not…
Paper uses LLMs for sector allocation, showing better returns.
Assessment of risk levels for existing credit accounts is important to the implementation of bank policies and offering financial products. This paper uses cluster analysis of behaviour of credit card accounts to help assess credit risk level. Account behaviour is modelled parametrically and we then implement the behav…
This paper detects function-level obfuscation in binary code using graph-based methods.
We study heterogeneity in the effect of a mindset intervention on student-level performance through an observational dataset from the National Study of Learning Mindsets (NSLM). Our analysis uses machine learning (ML) to address the following associated problems: assessing treatment group overlap and covariate balance,…
We propose a diffusion process to describe the global dynamic evolution of credit operations at a national level given observed operations at a subnational level in a sovereign country. Empirical analysis with a unique dataset from Brazilian federate constituents supports the conclusions. Despite the heterogeneity obse…
New method for high-fidelity shape representations from raw data.
This paper analyzes user-level local differential privacy in distributed systems.
With the increasing popularity of video sharing websites such as YouTube and Facebook, multimodal sentiment analysis has received increasing attention from the scientific community. Contrary to previous works in multimodal sentiment analysis which focus on holistic information in speech segments such as bag of words re…
InstanceFlow visualizes classifier confusion over training epochs.
The paper studies spectral analysis on complex spaces and finds explicit formulas for eigensections.
The paper breaks down AUC into cluster-level components for better model diagnostics.
Study robust mean estimation under coordinate-level corruptions using Hamming distance.
We introduce in this paper a new way of optimizing the natural extension of the quantization error using in k-means clustering to dissimilarity data. The proposed method is based on hierarchical clustering analysis combined with multi-level heuristic refinement. The method is computationally efficient and achieves bett…
In this article we review several techniques to extract information from stock market data. We discuss recurrence analysis of time series, decomposition of aggregate correlation matrices to study co-movements in financial data, stock level partial correlations with market indices, multidimensional scaling and minimum s…
The statistical properties of a stochastic process may be described (1)by the expectation values of the observables, (2)by the probability distribution functions or (3)by probability measures on path space. Here an analysis of level (3) is carried out for market fluctuation processes. Gibbs measures and chains with com…
The clusters of a distribution are often defined by the connected components of a density level set. However, this definition depends on the user-specified level. We address this issue by proposing a simple, generic algorithm, which uses an almost arbitrary level set estimator to estimate the smallest level at which th…
We present five methods to the problem of network anomaly detection. These methods cover most of the common techniques in the anomaly detection field, including Statistical Hypothesis Tests (SHT), Support Vector Machines (SVM) and clustering analysis. We evaluate all methods in a simulated network that consists of nomi…
Improved penalty-based methods for bilevel optimization with reduced complexity.
Dataset analyzes tweets' impact on stock returns.