Morpheo is a transparent and secure machine learning platform collecting and analysing large datasets. It aims at building state-of-the art prediction models in various fields where data are sensitive. Indeed, it offers strong privacy of data and algorithm, by preventing anyone to read the data, apart from the owner an…
cGAP visualizes high-dimensional categorical data with interpretable geometric structure.
problem Lack of visualization tools for high-dimensional categorical data.
method cGAP uses Homogeneity Analysis (HOMALS) to embed data in a 3D space and maps it to colors.
result cGAP reveals coherent clusters, outliers, and local-to-global structure in categorical data.
cGAP visualizes high-dimensional categorical data with interpretable geometric structure.
problem Lack of visualization tools for high-dimensional categorical data.
method cGAP uses Homogeneity Analysis (HOMALS) to embed data in a 3D space and maps it to colors for visualization.
result cGAP reveals coherent clusters, outliers, and local-to-global structure in categorical data.
A new framework uses uncertainty to learn from raw data without explicit models.
problem Limitations of traditional machine learning models and lack of interpretability.
method Introduces a model-free framework using surprisal (information theoretic uncertainty) to analyze and infer from raw data.
result Achieves at or near state-of-the-art performance across various machine learning tasks.
The traditional offline approaches are no longer sufficient for building modern recommender systems in domains such as online news services, mainly due to the high dynamics of environment changes and necessity to operate on a large scale with high data sparsity. The ability to balance exploration with exploitation make…
This communication is based on an original approach linking economical factors to technical and methodological ones. This work is applied to the decision process for mix production. This approach is relevant for costing driving systems. The main interesting point is that the quotation factors (linked to time indicators…
In Alain Connes noncommutative geometry, the question of the existence of a non-trivial integral can be described in terms of the singular traceability of the compact operator |D|^(-d), D being the Dirac operator, namely of the existence of a finite non-trivial singular trace on the ideal generated by |D|^(-d). A condi…
A trace on the C^*-algebra A of quasi-local operators on an open manifold is described, based on the results in \cite{RoeOpen}. It allows a description `a la Novikov-Shubin \cite{NS2} of the low frequency behavior of the Laplace-Beltrami operator. The 0-th Novikov-Shubin invariant defined in terms of such a trace is pr…
Approach to verify neural network training integrity.
problem Poisoning attacks during neural network training.
method Use of cryptographic mechanisms to verify training integrity.
result Provable verification of neural network training integrity.
The paper proposes using non-isotropic distances for more accurate trace link recommendation.
problem Time-consuming and error-prone creation and maintenance of trace links.
method Geometric viewpoint on semantic similarity using non-linear similarity measures.
result Non-isotropic distances improve trace link recommendation accuracy.
New framework for reinforcement learning with sporadic state observations.
problem Partial observability in reinforcement learning.
method Action-Triggered Sporadically Traceable Markov Decision Processes (ATST-MDPs).
result Optimistic algorithm achieving regret bound for episodic learning.
Unified AI system for data quality control and governance in regulated environments.
problem Isolated data quality control steps in existing systems.
method AI-driven framework integrating rule-based, statistical, and AI methods.
result Empirical gains in anomaly detection, reduced manual remediation, improved auditability.
New framework improves text watermark detection under imperfect pseudorandomness.
problem Structured dependence in generated text from language models causes Type I error control issues.
method Hierarchical two-layer partition, minimal units, non-asymptotic efficiency measure, minimax hypothesis testing.
result Closed-form optimal rules for watermark detection under imperfect pseudorandomness.
Prototype model improves model auditing and understanding.
problem Auditing and understanding modern language models is expensive and approximate.
method Introduced a sparse, non-negative mixture of learned prototypes trained with clustering objectives.
result Prototype models either surpass or remain within 2.5 percentage points of dense baselines on downstream tasks.
Developing an AI economist agent using RAG, knowledge graphs, and LLMs for economic scenario analysis.
problem Economic scenario analysis using large language models and knowledge graphs.
method Proposing an RAG-based AI economist framework that utilizes knowledge graphs and LLMs.
result Improves economic coherence and traceability in generated reports.
Estimates proportions of LLM-generated text in mixed documents.
problem Estimating the proportion of text generated by a pre-specified LLM in mixed documents.
method Developed estimators for two observation regimes: full observation and pivotal reduction, and established sample complexity bounds.
result Full observation estimators require fewer samples than pivotal reduction estimators.
Review of uncertainty representation methods in risk management.
problem Inadequate consideration of uncertainty in risk management.
method Systematic literature review of 370 publications.
result Probabilistic methods are predominant, but fuzzy and evidence-based approaches are also useful.
New method reduces computational cost for selective inference.
problem Over-conditioning in selective inference.
method Parametric programming-based selective inference (PP-based SI) with bounded p-values.
result Reduced computational cost while maintaining desired precision.
E-LMC improves spatial field prediction accuracy by linearizing complex fields.
problem Predicting complex spatial fields with high accuracy and efficiency.
method Introducing an invertible neural network to linearize nonlinear spatial fields, enabling the use of LMC for nonlinear problems.
result Maximum improvement of about 40% over original LMC, outperforming other models.
LargeMvC-Net improves scalability of multi-view clustering.
problem Scalability issues in multi-view clustering.
method Deep unfolding of multi-view clustering into a network architecture with three modules.
result LargeMvC-Net consistently outperforms state-of-the-art methods in scalability and effectiveness.
Bayesian framework improves ML classification models' uncertainty estimates.
problem Ensuring trustworthy AI predictions with explicit uncertainty quantification.
method Proposes a Bayesian framework for generative ML classification models that accounts for input measurement uncertainty.
result The BQDA model outperforms other models in terms of interpretability, explicit uncertainty modeling, and computational efficiency.
This work presents MeKDDaM-SAGA, computer-aided automation software for implementing a novel knowledge discovery and data mining process model that was designed for performing justifiable, traceable and reproducible metabolomics data analysis. The process model focuses on achieving metabolomics analytical objectives an…
One of the key requirements for incorporating machine learning into the drug discovery process is complete reproducibility and traceability of the model building and evaluation process. With this in mind, we have developed an end-to-end modular and extensible software pipeline for building and sharing machine learning …
Paper studies S-rectangular DR-RL models for robust reinforcement learning with near-optimal sample complexity.
problem Addressing distributional discrepancies in reinforcement learning environments.
method Empirical value iteration algorithm for divergence-based S-rectangular DR-RL models.
result Near-optimal sample complexity bound of O(∣S∣∣A∣(1−γ)−4ε−2). Exactly solvable model reveals how data geometry influences ML bias.
problem How data geometry affects machine learning bias.
method High-dimensional data imbalance model, statistical physics tools.
result Exact predictions for fairness metrics and mitigation strategies.
LOT framework embeds high-dimensional cell data into interpretable Euclidean space.
problem Lack of interpretable methods for high-dimensional cell data.
method Adapts Linear Optimal Transport (LOT) to irregular point clouds.
result Accurate and interpretable classification and synthetic data generation.
SHARC explains machine learning risk models for regulatory capital, linking outputs to scenarios.
problem Inability to explain machine learning model outputs to regulatory bodies.
method SHAP-based explainability framework for Hybrid GPR-HS architecture and SVaR stress-testing.
result SHARC links SVaR outputs to scenario inputs, providing auditable traceability.
Machine learning has recently been widely adopted to address the managerial decision making problems, in which the decision maker needs to be able to interpret the contributions of individual attributes in an explicit form. However, there is a trade-off between performance and interpretability. Full complexity models a…
We show how different approaches to developing marketing strategies depending on the type of environment a firm faces, where environments are distinguished in terms of their systems properties not their context. Particular emphasis is given to turbulent environments in which outcomes are not a priori predictable and ar…
A multi-agent system improves crypto portfolio management by processing diverse data types.
problem Managing cryptocurrency portfolios requires processing various data types under high volatility.
method A multi-agent system with three specialized agents for market dynamics, news sentiment, and signal fusion.
result The best configuration, Hierarchical (Skill), achieved a 133.52% cumulative return and 1.502 Sharpe ratio.
Artificial intelligence (AI) generally and machine learning (ML) specifically demonstrate impressive practical success in many different application domains, e.g. in autonomous driving, speech recognition, or recommender systems. Deep learning approaches, trained on extremely large data sets or using reinforcement lear…