Proposes a new method to estimate causal effects of unstructured treatments.
problem Estimating causal effects of unstructured treatments like text or images.
method Maximally Influential Feature (MIF) method to identify key features influencing outcomes.
result Developed algorithms to estimate and apply the MIF to improve outcomes in various contexts.
Proposes a method to infer causal effects from unstructured outcomes like text and images.
problem Traditional causal inference methods are inadequate for outcomes like text and images.
method Identifies the maximally contrasting feature (MCF) and learns a feature-scoring function to estimate causal effects.
result The method recovers salient aspects of outcomes changed by treatments in text and image settings.
Paper provides conditions for reliable use of pre-trained embeddings in econometrics.
problem Uncertainty in using pre-trained embeddings for econometric tasks.
method Derives sufficient conditions and convergence rates for machine learning models with pre-trained embeddings.
result Establishes theoretical foundations for reliable use of pre-trained embeddings in econometrics.
Paper uses text mining to predict reliability issues from customer complaints.
problem Detecting emerging reliability issues in after-sales service businesses.
method Essential text mining concepts applied to analyze customer complaints.
result Proactive detection of reliability problems through customer feedback.
Functional Retrofitting improves embedding of unstructured data into knowledge graphs.
problem Combining unstructured data with knowledge graphs that have diverse entities and relations.
method Explicitly models pairwise relations with a variety of penalty functions and allows encoding of relation semantics.
result Significantly outperforms existing retrofitting methods on complex knowledge graphs.
Satellite imagery helps adjust for unobserved confounders in observational studies.
problem Adjusting for confounding factors in observational studies with non-tabular data like satellite imagery.
method Formalizing conditions for causal effect identification, estimation, and sensitivity analysis.
result Demonstrated the use of satellite imagery as a proxy for unobserved confounders in anti-poverty aid programs.
Study uses LLMs to create personalized treatment plans for rare gynecological tumors.
problem Suboptimal management and poor prognosis due to low incidence and heterogeneity of rare gynecological tumors.
method Developed a digital twin system using LLMs to integrate clinical and biomarker data.
result LLM-enabled digital twins efficiently model individual patient trajectories and identify potential treatment options.
UQE uses LLMs to analyze unstructured data efficiently.
problem Efficient analytics on unstructured data.
method Proposes UQE, a query engine that uses LLMs to interpret UQL queries.
result Demonstrates efficient analytics on various unstructured data types.
Paper uses RL and DCAE to classify large unstructured data with fewer features.
problem Classifying large unstructured data with high precision using fewer features.
method Deep Convolutional Autoencoder (DCAE) for feature learning and Double DQN/Retrace RL algorithms for policy optimization.
result The approach achieves high classification precision with fewer features than traditional methods.
Combines deep learning with instrumental variables for causal effect prediction.
problem Causal effect estimation in the presence of unobserved confounders.
method Deep instrumental variables networks (D-IV) combining first and second stage neural nets.
result Flexible D-IV framework resolves causal estimation into manageable prediction tasks.
This study shows unstructured clinical notes can improve mortality prediction.
problem Lack of effective use of unstructured clinical notes in mortality prediction.
method Used a hierarchical architecture with convolutional and recurrent layers to predict in-hospital mortality from unprocessed clinical notes.
result Achieved higher metrics in mortality prediction compared to structured data approaches.
GMLS-Nets extend CNNs to unstructured data points.
problem Learning from irregularly spaced data points in science and engineering.
method Introducing GMLS for non-parametric estimation and parameterizing it for learning operators with unstructured stencils.
result GMLS-Nets provide a framework for functional regression and quantity prediction from unstructured data.
This paper explains how transformers learn from unstructured data in ICL.
problem Understanding how transformers learn from unstructured data in in-context learning.
method A simple transformer model with one or two attention layers and positional encoding is used to study the role of each component in ICL.
result A transformer with two attention layers and a look-ahead attention mask can learn from unstructured data.
New scalable multi-class SVM for structured and unstructured data.
problem Classification of structured and unstructured data.
method Bayesian multi-class support vector machine with pseudo-likelihood, variational inference, and inducing point approximation.
result Outperforms competitor methods in training time and accuracy.
SparseRT accelerates sparse computations on GPUs for deep learning inference.
problem Efficiently handling unstructured sparsity patterns on GPUs for deep learning.
method SparseRT, a code generator that leverages unstructured sparsity for accelerating sparse linear algebra operations.
result Geometric mean speedups of 3.4x at 90% sparsity and 5.4x at 95% sparsity for 1x1 convolutions and fully connected layers.
Study introduces a benchmark suite for evaluating neural MI estimators on real-world unstructured datasets.
problem Lack of comprehensive evaluation methods for neural MI estimators on real-world unstructured datasets.
method Developed a benchmark suite using same-class sampling and a binary symmetric channel trick.
result Showed accurate manipulation of true MI values of real-world datasets.
LLMs learn new tasks from unstructured data, but it depends on word co-occurrence and positional information.
problem Understanding how LLMs can learn new tasks from unstructured data without explicit training.
method Examined the capabilities of LLMs trained on unstructured data, focusing on sequence model requirements and training data structure.
result Many ICL capabilities can emerge from word co-occurrence in unstructured data, but positional information is crucial for certain tasks.
Develops Φ-DVAE for assimilating unstructured data into physical models.
problem Challenges in incorporating unstructured data into physical models.
method Physics-informed dynamical variational autoencoder (Φ-DVAE) combining latent state-space model and VAE. result Demonstrates data-efficient dynamics encoding with competitive performance and uncertainty quantification.
Word2vec model captures disease attributes from unstructured text.
problem Augmenting disease surveillance with unstructured data requires accurate taxonomical correlations and trace mapping.
method Developed a disease vocabulary driven word2vec model (Dis2Vec) to model diseases and their attributes.
result Dis2Vec outperforms traditional word2vec methods in capturing taxonomical attributes across different disease classes.
LLMs help automate extraction of actuarial variables from unstructured claims data.
problem Manual processing of unstructured claims data is time-consuming and inconsistent.
method Two-stage processing architecture using LLMs, modular Python pipeline.
result LLM-based extraction achieved high accuracy and practical actuarial value.
Proposes neural network for causal inference with multimodal data.
problem Estimating causal effects with text and image data as confounders.
method Double machine learning framework adapted to partially linear models, semi-synthetic dataset generation.
result Improved performance in causal effect estimation with multimodal data.
Prototype learns automotive industry ontology from unstructured data.
problem Automatic learning of domain-specific ontologies from unstructured text data.
method Two-stage classification system: first classifier for concepts and irrelevant collocates, second classifier for concept types.
result Prototype validated with automotive industry complaint and repair data.
GPI uses GenAI models to infer causal and predictive effects from unstructured data.
problem Estimating causal and predictive effects from unstructured data like text and images.
method Leverages open-source GenAI models to generate and represent unstructured data, applying machine learning to these representations.
result GPI efficiently estimates causal and predictive effects with quantified uncertainty, without fine-tuning.
Enhanced regime shifts detection using unstructured text and financial data.
problem Detecting regime shifts in financial markets is challenging due to noisy and multicollinear data.
method Combines LLM reasoning on unstructured text and statistical validation on financial time series.
result Framework achieves F1 score of 0.82, outperforming pure data-driven methods.
Paper tackles product categorization with structured and unstructured attributes for large-scale eCommerce.
problem Challenges in categorizing products with thousands of classes and millions of products.
method Compares hierarchical and flat models, uses Deep Learning for feature extraction, combines structured and unstructured attributes.
result Flat models perform better in specific cases, and the proposed approach handles faulty attribute names and values.
Develops methods to improve demand counterfactuals from imperfect proxies.
problem Imperfect proxies in demand models lead to biased counterfactuals and invalid inference.
method Practical toolkit for market-level and individual data, requiring minimal computation.
result Improves substitution prediction and counterfactual performance.
Population-based learning improves representation of unstructured data.
problem Improving representation of unstructured data.
method Instantiating Lewis signaling games within a population of agents.
result Population-based learning produces better representations than single-agent learning.
A new method monitors unstructured 3D shapes without registration.
problem Error-prone registration and mesh reconstruction steps in PCD monitoring.
method Intrinsic geometric properties of shapes, using Laplacian and geodesic distances.
result Effective monitoring of defects without registration and mesh reconstruction.
Neural network solves BVPs with unstructured data.
problem Solving Boundary Value Problems (BVPs) with numerical methods.
method Neural Network based numerical method for solving BVPs.
result Validated the method for Laplace and Poisson equations.
Deep neural networks predict prostate motion from MR images.
problem Predicting prostate motion during ultrasound-guided interventions.
method Biomechanically-trained deep neural networks on unstructured nodes.
result Trained networks yield near real-time inference with 0.017 mm error.
This paper compares unstructured and structured EM-based semi-supervised learning methods.
problem Semi-supervised learning with EM algorithm for structured prediction.
method Comparative study between unstructured and structured EM-based semi-supervised learning methods.
result Structured EM is more robust to class confusion in flood mapping datasets.
New insights on pruning deep networks by preserving function locality.
problem Designing effective pruning methods for deep neural networks.
method Revisited loss modeling using first and second order Taylor expansions, emphasizing locality.
result Both first and second order Taylor expansions can achieve similar performance in pruning.
Deep learning models detect and classify log anomalies.
problem Anomaly detection in unstructured log data.
method Auto-LSTM, Auto-BLSTM, and Auto-GRU models for feature extraction.
result Models outperform other algorithms on various log data sets.
PODNet discovers plannable options from unstructured demonstrations.
problem Learning from unstructured, multi-objective demonstrations.
method Custom categorical variational autoencoder, recurrent option inference network, option-conditioned policy network, and option dynamics model.
result PODNet enables learning from demonstration for multiple tasks and planning.
ASOS improves fashion product recommendations by learning from unstructured data.
problem Lack of consistent product information in e-commerce.
method Developed a hybrid recommender system to learn product attributes from unstructured data.
result Quantitative understanding of products improves recommendation accuracy.
Automated tests detect interactions in unstructured data.
problem Detecting interactions between latent variables in low-dimensional systems.
method Derive two interaction tests based on pairwise interventions and integrate them into an active learning pipeline.
result Tests can identify more known biological interactions than random search and standard active learning baselines.
Quantum algorithm finds extrema in discrete optimisation problems.
problem Finding extrema in discrete optimisation functions.
method Quantum unstructured search algorithm (QSERA) to map and find extrema.
result Quadratic speed-up over classical algorithms for discrete optimisation.
This paper analyzes text in financial disclosures to improve financial analysis.
problem Insufficient analysis of unstructured text in financial disclosures.
method Reviews and explores methods in computational linguistics and NLP.
result Highlights limitations of sentiment metrics and suggests future research areas.
Study copyright's impact on creative industries using AI-generated fonts.
problem Estimating supply and demand in creative industries with AI-generated content.
method Neural network embeddings, spatial regression, event-study analyses, structural model of supply and demand.
result Copyright can raise consumer welfare by encouraging product relocation.
Study identifies three sub-phenotypes of AKI with different severity.
problem Tackles the heterogeneity of AKI to improve targeted interventions.
method Used a memory network-based deep learning approach on EHR data.
result Identified three distinct sub-phenotypes of AKI with varying severity.
Multi-cell cooperative processing with limited backhaul traffic is studied for cellular uplinks. Aiming at reduced backhaul overhead, a sparsity-regularized multi-cell receive-filter design problem is formulated. Both unstructured distributed cooperation as well as clustered cooperation, in which base station groups ar…
This study creates a new conceptual framework for news aggregation.
problem Aggregating and presenting diverse news sources in a unified and accessible way.
method Developed a mobile app that analyzes unstructured data patterns to create a conceptual framework for news.
result Users can easily find and navigate through updated news using a new conceptual multilevel structure.
New algorithm identifies outliers in PCA without needing parameters.
problem Robust PCA with unknown outlier fraction and subspace dimension.
method Non-iterative, parameter-free method for structured and unstructured outliers.
result Analytical guarantees and performance comparison with existing methods.
Optimizes data labeling for causal effect estimation with missing outcomes.
problem Estimating causal effects with missing outcome data and budget constraints.
method Optimizes batch sampling probability to minimize variance of causal inference estimator.
result Achieves lower mean-squared error with fewer labeled data points.
Framework extracts symptoms from EHRs for rapid disease outbreak detection.
problem Extracting relevant data from unstructured medical texts.
method Conformal active learning for efficient data mining.
result Framework achieves strong performance with minimal manual labeling.
BCCNet combines biased crowd labels to train classifiers for disaster response.
problem Improper labels from citizen scientists limit machine learning applications.
method Bayesian classifier combination neural network (BCCNet) aggregates and trains classifiers from imperfect labels.
result BCCNet effectively processes large unstructured data for disaster prevention and response.
To date, there have been massive Semi-Structured Documents (SSDs) during the evolution of the Internet. These SSDs contain both unstructured features (e.g., plain text) and metadata (e.g., tags). Most previous works focused on modeling the unstructured text, and recently, some other methods have been proposed to model …
Bayesian method for semi-structured models accounts for both types of uncertainty.
problem Lack of work on epistemic uncertainty in semi-structured regression models.
method Bayesian approximation with subspace inference for joint posterior sampling.
result Validated approach recovers structured effect posteriors and approaches full-space posterior.