Improved knowledge graph embedding using taxonomic information.
problem Learning about domains in knowledge graphs using embedding models.
method Minimal modifications to existing knowledge graph completion methods to incorporate taxonomic information.
result Our model is fully expressive, respecting subclass and subproperty information.
META2 improves taxonomic classification and abundance estimation in metagenomics with deep learning and memory efficiency.
problem Memory constraints and inefficiencies in taxonomic classification and abundance estimation for metagenomics.
method Developed a novel memory-efficient read classification technique combining deep learning and locality-sensitive hashing, and formulated abundance estimation as a Multiple Instance Learning problem.
result Our approach outperforms conventional methods in both single-read taxonomic classification and abundance estimation, especially when memory is limited.
The step of expert taxa recognition currently slows down the response time of many bioassessments. Shifting to quicker and cheaper state-of-the-art machine learning approaches is still met with expert scepticism towards the ability and logic of machines. In our study, we investigate both the differences in accuracy and…
Metagenomics characterizes the taxonomic diversity of microbial communities by sequencing DNA directly from an environmental sample. One of the main challenges in metagenomics data analysis is the binning step, where each sequenced read is assigned to a taxonomic clade. Due to the large volume of metagenomics datasets,…
New method interprets object representations from human behavior.
problem Understanding how mental object representations relate to human behavior.
method Sparse, non-negative representations of objects estimated from behavioral judgments.
result Representations predict latent object similarity and are interpretable.
Traditional disease surveillance can be augmented with a wide variety of real-time sources such as, news and social media. However, these sources are in general unstructured and, construction of surveillance tools such as taxonomical correlations and trace mapping involves considerable human supervision. In this paper,…
Generative model for high-dimensional categorical data using Gaussian-Dirichlet fields.
problem Efficiently modeling and predicting high-dimensional categorical data.
method Combines Dirichlet and Gaussian processes for spatio-temporal modeling.
result Model accurately approximates categorical data in unobserved locations.
Task2Vec creates task embeddings for meta-learning.
problem Creating a framework for selecting feature extractors for new tasks.
method Process images through a probe network to compute task embeddings based on Fisher information matrix.
result Task embeddings predict task similarities and feature extractor performance.
ProGen models protein sequences for synthetic biology.
problem Generating proteins without structural annotations.
method Trained a 1.2B-parameter language model on 280M protein sequences.
result ProGen generates proteins with fine-grained control and accuracy.
Adaptive and sequential experiment design is a well-studied area in numerous domains. We survey and synthesize the work of the online statistical learning paradigm referred to as multi-armed bandits integrating the existing research as a resource for a certain class of online experiments. We first explore the tradition…
New statistical models for predicting ranked preferences from partial orders.
problem Statistical models overlook information in list length.
method Composite and augmented ranking models for joint modeling of partial orders and list lengths.
result Augmented ranking models best predict both length and preferences.
Machine learning predicts plant phenotypes from soil microbiome data.
problem Predicting plant phenotypes from soil microbiome data.
method Two models (random forest and Bayesian neural network) were used to predict plant phenotypes from soil properties and microbial population density.
result Human decisions and normalization strategies significantly impact model performance.
Image classification system identifies bumble bee species from images.
problem Manual identification of bumble bee species is time-consuming and requires expert knowledge.
method Transfer learning using Inception, VGG16, VGG19, and ResNet models.
result Inception and VGG classifiers achieved up to 23% accuracy for single species identification.
Semantic segmentation was seen as a challenging computer vision problem few years ago. Due to recent advancements in deep learning, relatively accurate solutions are now possible for its use in automated driving. In this paper, the semantic segmentation problem is explored from the perspective of automated driving. Mos…
Study re-evaluates MIMIC-III codes, finding many are under-coded.
problem Validity of MIMIC-III clinical codes is questionable.
method Open-source, reproducible methodology for assessing codes.
result Most frequently assigned codes are under-coded up to 35%
AI enhances microbiology and microbiome research through machine learning.
problem Understanding microbial life and its impact on health and the environment.
method AI-driven approaches including machine learning and deep learning.
result Transformative role in enhancing microbial life understanding.
This review establishes a taxonomy for modular neural networks.
problem Scaling ANNs for complex and multi-disciplinary problems.
method Systematic analysis of modularization techniques in MNNs.
result A universal framework for studying MNNs.
DeepMaxent uses neural networks to improve species distribution models.
problem Sampling biases and lack of absence data in presence-only observations.
method DeepMaxent employs neural networks to learn shared features among species using the maximum entropy principle.
result DeepMaxent outperforms traditional methods in predicting species distributions, especially in unevenly sampled regions.
New method extracts high-accuracy, high-fidelity neural networks.
problem Extracting accurate and functionally-equivalent neural networks.
method Learning-based attack with practical functionally-equivalent extraction.
result First practical functionally-equivalent extraction attack.
The paper introduces a method to control false splits in tree-based data aggregation.
problem Identifying the correct subgroups to treat as a single entity in tree-based data.
method Introduces the 'false split rate' and proposes a multiple hypothesis testing algorithm for tree-based aggregation.
result The proposed algorithm controls the false split rate, demonstrating its effectiveness on stock volatility and taxi fare data.
A taxonomy of saliency metrics helps in pruning deep neural networks.
problem Difficulty in separating saliency metric effectiveness from pruning algorithms.
method Proposed a taxonomy based on four orthogonal components.
result Constructed metrics can outperform existing state-of-the-art metrics.
A model predicts how hyperparameters affect pruning performance.
problem Predicting the impact of hyperparameters on pruning performance.
method Phenomenological model using temperature-like and load-like parameters.
result A sharp transition phenomenon in pruning performance.
Model interprets image classification using hierarchical prototypes.
problem Lack of hierarchical interpretation in vision models.
method Uses hierarchically organized prototypes to classify objects at each level of a taxonomy.
result Model interprets image classification at each level of the taxonomy.
Framework uncovers symmetric and asymmetric species associations from data.
problem Retrieving bidirectional species associations from co-occurrence data.
method Machine learning framework modeling latent embeddings and joint generative model.
result Framework successfully recovers known symmetric and asymmetric associations.
This paper shows excessive invariance in adversarial robust models can make them more vulnerable to certain types of attacks.
problem Excessive invariance in adversarial robust models can make them more vulnerable to certain types of attacks.
method Analytical constructions and empirical studies of vision classifiers with state-of-the-art robustness to perturbation-based adversaries constrained by an ℓp norm. result Robustness to perturbation-based adversarial examples does not guarantee general robustness and can increase vulnerability to invariance-based adversarial examples.
Text2Node maps medical phrases to a taxonomy, overcoming coding standard limitations.
problem Limited data interchangeability between EHR systems due to different coding standards.
method Text2Node uses word and node embeddings, along with mapping functions, to generalize from limited training data.
result Text2Node achieves high accuracy in mapping phrases to a taxonomy, even for unseen concepts.
We simplify information measure computation using learned features.
problem Computing information measures from raw data is computationally expensive.
method Developed a separable design for computing information measures from learned feature representations.
result A variety of information measures can be computed efficiently through learned feature representations.
A new framework for information theory considers computational constraints.
problem Understanding information in complex systems with computational limitations.
method Variational extension of Shannon's information theory with computational constraints.
result Predictive V-information can be created through computation and reliably estimated from data. An asymmetric information model is introduced for the situation in which there is a small agent who is more susceptible to the flow of information in the market than the general market participant, and who tries to implement strategies based on the additional information. In this model market participants have access t…
We study a simple model of an asset market with informed and non-informed agents. In the absence of non-informed agents, the market becomes information efficient when the number of traders with different private information is large enough. Upon introducing non-informed agents, we find that the latter contribute signif…
Model shows too much information can make markets inefficient.
problem Information overload in financial markets.
method Original model using information overload to show inefficiency.
result Efficient market hypothesis ceases to be true with infinite information.
New method quantifies redundant information using information bottleneck.
problem Quantifying redundant information among multiple sources.
method Formulated as an information bottleneck problem, termed redundancy bottleneck.
result Extracts information that best predicts the target without revealing source identity.
New method uses mutual info and network science to explain deep learning models.
problem Interpreting deep neural networks for understanding their decision-making process.
method Coupling mutual information with network science to quantify information flow in deep learning models.
result Proposed NIF technique for codifying information flow in deep learning models.
This paper reviews information theory in open-world machine learning.
problem Lack of a unified theoretical foundation for open-world machine learning.
method Synthesis of information theoretic approaches.
result Established a pathway toward provable and trustworthy open world intelligence.
Generalizes information theory to evolving belief.
problem Measuring change in belief over time.
method Derives a general theory of information from first principles.
result Recover all information measures and interprets entropy as expected gain.
Paper proposes a framework to identify and obfuscate sensitive features via information density estimation.
problem Identifying and protecting sensitive attributes from leakage in obfuscation mechanisms.
method Information density estimation to identify leaking features, followed by a targeted obfuscation mechanism.
result Proven leakage guarantee in terms of Eγ-divergence for the obfuscation mechanism. Review of information plane analyses in neural networks, highlighting mixed results and methodological challenges.
problem Understanding the relationship between information-theoretic compression and neural network performance.
method Literature review and detailed analysis of information quantity estimation methods.
result Information plane compression is not necessarily information-theoretic but compatible with geometric compression.
Information geometry offers new tools for statistical analysis.
problem Statistical analysis of probability distributions.
method Geometric perspective on statistical manifolds.
result New applications in radar sensing, signal processing, etc.
In financial markets valuable information is rarely circulated homogeneously, because of time required for information to spread. However, advances in communication technology means that the 'lifetime' of important information is typically short. Hence, viewed as a tradable asset, information shares the characteristics…
Introduces relative information gain for improving Gaussian process regression rates.
problem Improving the sample complexity of estimating or maximizing unknown functions.
method Introduces relative information gain, interpolates between effective dimension and information gain, and proves PAC-Bayesian bounds.
result Obtains minimax-optimal rates of convergence through the relative information gain.
This paper uses information theory to improve risk modeling in big data.
problem Insufficient application of information theory in actuarial science.
method Explores information theory to uncover performance limits of insurance big data systems.
result Guidance for risk modeling and actuarial pricing systems.
In information theory, Fisher information and Shannon information (entropy) are respectively used to quantify the uncertainty associated with the distribution modeling and the uncertainty in specifying the outcome of given variables. These two quantities are complementary and are jointly applied to information behavior…
Investors pay for additional asset information based on utility maximization.
problem Determining the optimal price for additional asset information.
method Solving a stochastic control problem with partial information and utility maximization.
result Investors choose to purchase information at a deterministic time.
New method detects information leakage using approximate Bayes predictor.
problem Unintentional exposure of sensitive information via observable data.
method Statistical learning theory and information theory framework, approximating Bayes predictor's log-loss and accuracy.
result MI can be accurately estimated to detect ILs, outperforming state-of-the-art baselines.
Market strategies minimize Fisher information to minimize risk.
problem Applying minimum Fisher information principle to market dynamics.
method Analytical extension to quantum harmonic oscillator eigenstates and Gibbs distribution.
result Minimizing Fisher information reduces information and risk.
New approach quantifies overfitting in high-dimensional regression.
problem Quantifying and avoiding overfitting in large neural networks.
method Information bottleneck theory to minimize residual information while maximizing relevant bits.
result Characterized the relative information efficiency of randomized regression compared to optimal algorithms.
Proposes a new framework for decomposing information in multivariate settings.
problem Decomposing information from multiple sources about a target variable.
method Formal analogy with set theory and Blackwell order to define PID.
result Framework can be generalized to various information theories.
Unified notation simplifies information-theoretic concepts in machine learning.
problem Opaque notation for information-theoretic quantities in machine learning.
method Proposed a practical and unified notation for information-theoretic quantities.
result Unified notation facilitates new intuitions and rederivations in machine learning.