Mixed labyrinth fractals can have finite or infinite arc lengths.
problem Characterizing the length of arcs in mixed labyrinth fractals.
method Analyzing sequences of labyrinth patterns to determine arc lengths.
result Arc lengths can be finite or infinite depending on pattern choice.
New method trains Boltzmann machines without supervision.
problem Training unsupervised learning models.
method Mixed binary quadratic feasibility problem formulation.
result Theory validated on XOR patterns.
Labyrinth fractals are dendrites in the unit square. They were introduced and studied in the last decade first in the self-similar case [Cristea & Steinsky (2009,2011)], then in the mixed case [Cristea & Steinsky (2017), Cristea & Leobacher (2017)]. Supermixed fractals constitute a significant generalisation of mixed l…
The study uses Hidden Markov Models to analyze student enrollment patterns and academic performance.
problem Limited understanding of how enrollment patterns affect academic performance.
method Applied Hidden Markov Models to categorize enrollment strategies and compare academic outcomes.
result Mixed enrollment strategies lead to better academic performance, especially during part-time semesters.
New model learns relative importance of multiple seasonal patterns in time series data.
problem Complex seasonal patterns in business time series data.
method Mixed hierarchical seasonality (MHS) model using Stan.
result Significant improvements in prediction error and predictive density compared to existing models.
The majority of real-world networks are dynamic and extremely large (e.g., Internet Traffic, Twitter, Facebook, ...). To understand the structural behavior of nodes in these large dynamic networks, it may be necessary to model the dynamics of behavioral roles representing the main connectivity patterns over time. In th…
MC-GMENN improves neural networks for clustered data using Monte Carlo methods.
problem Improving neural network performance on clustered data with correlations.
method MC-GMENN employs Monte Carlo methods to train generalized mixed effects neural networks.
result MC-GMENN outperforms existing models in generalization and quantifying inter-cluster variance.
FAMDAD detects anomalies in mixed data using kurtosis-weighted Factor Analysis.
problem Detecting anomalies in high-dimensional mixed data.
method kurtosis-weighted Factor Analysis of Mixed Data (FAMDAD).
result Anomalies are highly separable in the first and last few dimensions of the FAMDAD embedding.
DeepCausalMMM models marketing impacts using deep learning and causal inference.
problem Traditional MMM approaches struggle with non-linear dynamics and temporal patterns.
method Combines deep learning, causal inference, and marketing science. Uses GRUs for temporal patterns and DAG structure for channel dependencies.
result Captures non-linear dynamics and temporal patterns in marketing impacts.
Paper proposes a method to identify wind hazard types and predict extreme wind speeds.
problem Difficulty in identifying wind hazard types from meteorological data records.
method Numerical pattern recognition method with feature extraction and generalization.
result Algorithm performance validated using K-fold cross-validation and real-world data.
MPTE uses Transformer attention to estimate mixed-frequency factor models.
problem Estimating factor models in panel datasets with mixed frequencies and nonlinear signals.
method Mixed-Panels-Transformer Encoder (MPTE) with attention mechanisms.
result MPTE achieves competitive performance in nonlinear forecasting environments.
Identifying important components or factors in large amounts of noisy data is a key problem in machine learning and data mining. Motivated by a pattern decomposition problem in materials discovery, aimed at discovering new materials for renewable energy, e.g. for fuel and solar cells, we introduce CombiFD, a framework …
HyDaP clusters mixed-type data efficiently.
problem Clustering data with mixed types (continuous and categorical).
method Two-step hybrid approach: density-based and partition-based for continuous variables, partition-based for mixed data.
result HyDaP outperforms existing methods in clustering electronic health records.
Develops M2 model for next-basket recommendation considering user preferences, item popularity, and transition patterns.
problem Next-basket recommendation problem considering user preferences, item popularity, and transition patterns.
method Mixed model with preferences, popularities, and transitions (M2) using ed-Trans for transition patterns among items.
result Significantly outperforms state-of-the-art methods on all datasets in all tasks, with up to 22.1% improvement.
Deep-SITAR uses autoencoders to predict growth patterns.
problem Predicting individual growth trajectories from population data.
method Deep learning framework integrating autoencoders and B-spline models.
result Deep-SITAR predicts individual growth without full model re-estimation.
Neural networks struggle with abstract patterns, new RBP structures improve performance.
problem Neural networks fail to learn abstract patterns based on identity rules.
method Proposed Relation Based Pattern (RBP) extensions to neural network structures.
result Neural networks with RBP structures achieve perfect performance on synthetic and real-world sequence prediction tasks.
Paper uses deep learning to improve thermal-hydraulic simulations.
problem Limited credibility of thermal-hydraulic codes in real plant conditions.
method Feature Similarity Measurement (FSM) and deep learning.
result Deep learning constructs relationships between local physical features and simulation errors.
Word embedding maps words into a low-dimensional continuous embedding space by exploiting the local word collocation patterns in a small context window. On the other hand, topic modeling maps documents onto a low-dimensional topic space, by utilizing the global word collocation patterns in the same document. These two …
New model predicts energy prices under different scenarios.
problem Complex causal relationships in energy markets with continuous regime changes.
method Augmented Time Series Structural Causal Models (ATSCM) integrating neural causal discovery.
result Enables novel counterfactual queries in energy markets.
A new method improves network modeling by mixing multiple models.
problem Learning network connections from noisy data.
method Mixing multiple models to improve individual model performances.
result The method outperforms existing approaches even when models are misspecified.
We introduce a mixed-effects model to learn spatiotempo-ral patterns on a network by considering longitudinal measures distributed on a fixed graph. The data come from repeated observations of subjects at different time points which take the form of measurement maps distributed on a graph such as an image or a mesh. Th…
PAIN network improves imputation for mixed datasets.
problem Missing data in diverse scientific domains.
method Dynamic adaptive imputation using statistical methods, random forests, and autoencoders.
result PAIN outperforms traditional imputation methods in preserving data distributions.
Recently, l2,1 matrix norm has been widely applied to many areas such as computer vision, pattern recognition, biological study and etc. As an extension of l1 vector norm, the mixed l2,1 matrix norm is often used to find jointly sparse solutions. Moreover, an efficient iterative algorithm has been designed…
Models of bags of words typically assume topic mixing so that the words in a single bag come from a limited number of topics. We show here that many sets of bag of words exhibit a very different pattern of variation than the patterns that are efficiently captured by topic mixing. In many cases, from one bag of words to…
Two algorithms learn Gaussian graphical models from Glauber dynamics trajectories, achieving optimal performance.
problem Learning Gaussian graphical models from a single trajectory of a dependent stochastic process.
method Two algorithms based on dueling-neighborhood search and local statistics built from the update sequence of Glauber dynamics.
result Achieve κ−2 dependence of the information-theoretic lower bounds, mixing-free and signal-optimal. Blind Source Separation (BSS) has proven to be a powerful tool for the analysis of composite patterns in engineering and science. We introduce Convex Analysis of Mixtures (CAM) for separating non-negative well-grounded sources, which learns the mixing matrix by identifying the lateral edges of the convex data scatter p…
We present an efficient algorithm for the inference of stochastic block models in large networks. The algorithm can be used as an optimized Markov chain Monte Carlo (MCMC) method, with a fast mixing time and a much reduced susceptibility to getting trapped in metastable states, or as a greedy agglomerative heuristic, w…
This study compares transfer learning and multi-agent learning for AI-driven traffic agents.
problem Improving traffic flow in mixed-intelligence highway scenarios.
method Online MIT DeepTraffic simulation, deep reinforcement learning, elitist evolutionary algorithm, hyperparameter search, transfer learning, multi-agent learning.
result Transfer learning and multi-agent learning yield different average speeds for AI-driven traffic agents.
Tissue heterogeneity is a major confounding factor in studying individual populations that cannot be resolved directly by global profiling. Experimental solutions to mitigate tissue heterogeneity are expensive, time consuming, inapplicable to existing data, and may alter the original gene expression patterns. Here we a…
New method identifies causal variables from partially observed data.
problem Learning from unpaired observations with instance-dependent partial observability.
method Proposes two methods enforcing sparsity in the inferred representation.
result Establishes two identifiability results for linear and piecewise linear mixing functions.
Guided warping augments time series data by aligning features with a teacher.
problem Small time series datasets limit neural network performance.
method Guided warping with a discriminative teacher to augment data deterministically.
result Significant improvement in performance on various time series datasets.
SNI framework for mixed-type data imputation interprets and explains missing values.
problem Missing data in mixed-type databases skew analysis results.
method SNI couples statistical priors with neural attention to impute and explain missing values.
result SNI provides interpretable feature dependency diagnostics and soft regularization of attention.
The study introduces a holdout-based framework to assess synthetic data fidelity and privacy.
problem Evaluating the quality and privacy of synthetic data solutions for mixed-type tabular data.
method Holdout-based empirical assessment framework measuring fidelity and privacy risk.
result Synthetic data samples are as close to the training as to the holdout data, indicating generalization and independence from individual records.
GLMM trees identify subgroups with different growth patterns in longitudinal data.
problem Identifying subgroups with distinct growth trajectories in longitudinal studies.
method Extended GLMM trees for longitudinal data.
result Extended GLMM trees outperform other methods in accuracy and speed.
EFA extends self-attention to handle mixed data types and dynamic relevance.
problem Handling high-dimensional, mixed data types with dynamic relevance.
method Probabilistic generative model using self-attention and latent factor model.
result EFA consistently outperforms existing models in complex latent structure capture and reconstruction.
Survey of data augmentation techniques for time series classification with neural networks.
problem Small datasets in time series recognition.
method Four families of data augmentation: transformation-based, pattern mixing, generative models, and decomposition methods.
result Empirical evaluation of 12 data augmentation methods on 128 datasets.
By generalizing the measurements on the game experiments of mixed strategy Nash equilibrium, we study the dynamical pattern in a representative dynamic stochastic general equilibrium (DSGE). The DSGE model describes the entanglements of the three variables (output gap [y], inflation [π] and nominal interest rate [$…
New method for MTL with varying sparsity patterns across tasks.
problem Jointly training multiple linear models with differing sparsity patterns.
method Mixed-integer programming formulation and scalable algorithms.
result Our methods leverage shared support information to improve variable selection.
OMERF extends random forest for hierarchical data and ordinal responses.
problem Analyzing hierarchical data and ordinal responses using tree-based methods.
method Ordinal Mixed-Effects Random Forest (OMERF) that preserves flexibility and hierarchical structure.
result OMERF identifies discriminating student characteristics and estimates school effects.
CoCoAFusE fuses expert predictions to model complex patterns with interpretability and uncertainty.
problem Modeling complex patterns with interpretability and uncertainty quantification.
method Competitive/Collaborative Fusion of Experts (CoCoAFusE) that fuses expert distributions in addition to mixing.
result CoCoAFusE avoids multimodality artifacts and provides tighter credible bounds on the response variable.
We present a new technique called contrastive principal component analysis (cPCA) that is designed to discover low-dimensional structure that is unique to a dataset, or enriched in one dataset relative to other data. The technique is a generalization of standard PCA, for the setting where multiple datasets are availabl…
Paper proposes a novel optimization method for disaggregating smart meter data.
problem Energy disaggregation, inferring appliance-specific energy consumption from aggregate meter data.
method Two-stage optimization approach: first phase uses mixed integer programming, second phase binary quadratic optimization with penalty terms and appliance constraints.
result Proposed method successfully reconstructs appliance signatures, overcoming previous optimization-based methods' limitations.
Proposes ESCA model to analyze mixed data types in multiple sets of measurements.
problem Separating common and distinct information in mixed data types from multiple sources.
method Exponential Family Simultaneous Component Analysis (ESCA) model with structured sparse loading matrix.
result The proposed method effectively disentangles global, local common and distinct information.
TimeMixer predicts global financial asset volatility, excelling in short-term forecasts.
problem Predicting volatility in global financial markets is challenging due to complexity and non-linear dynamics.
method Uses TimeMixer, a multiscale-mixing model for forecasting across different scales.
result TimeMixer performs exceptionally well in short-term volatility forecasting but less so in longer-term predictions.
Enhances TCK for missing data and incomplete labels in time series.
problem Missing data and incomplete labels in time series analysis.
method Ensemble learning with Bayesian mixture models, representation of missing patterns, semi-supervised learning.
result Improved accuracy in similarity learning for time series with missing and incomplete labels.
Proposes a proportional masking strategy for better tabular data imputation.
problem Heterogeneity of tabular data disrupts uniform random masking in MAEs.
method Computes missingness statistics, generates proportional masks, uses MLP token mixing.
result Proportional masking preserves missingness distribution, improves imputation performance.
Theoretical analysis of deep neural networks for time series data.
problem Theoretical development for deep neural networks on temporally dependent observations is lacking.
method Established non-asymptotic bounds for prediction error of deep neural networks under mixing-type assumptions.
result Deep neural networks can model non-linear time series data with additional logarithmic factors due to dependence.
Paper uses RL to optimize branching strategy in B&B algorithms.
problem Optimizing Branch and Bound algorithms for mixed integer linear programs.
method FMSTS, a Reinforcement Learning approach for variable selection.
result FMSTS outperforms commercial solvers in efficiency and generalization.