Flexics samples patterns with guarantees, addressing flexibility and accuracy issues.
problem Pattern explosion and limited sampling accuracy with existing methods.
method Leverages SAT sampling and pattern mining algorithms to support flexible quality measures and constraints.
result Flexics provides strong guarantees on sampling accuracy while being flexible and efficient.
Pattern sampling reduces time series classification complexity.
problem High computational complexity of exhaustive search for shapelets.
method Pattern sampling using a weighted trie to extract discriminative patterns.
result Significant reduction in computational and memory resources.
LOUPE optimizes MRI sub-sampling patterns using machine learning.
problem Optimizing sub-sampling patterns for MRI scans to improve reconstruction accuracy.
method End-to-end learning strategy combining sub-sampling pattern optimization and reconstruction model training.
result LOUPE yields more accurate reconstructions compared to standard under-sampling schemes.
LetSIP learns relevant patterns for user interests in data mining.
problem Redundancy in pattern mining makes it hard for analysts to identify relevant patterns.
method Combines pattern sampling with interactive data mining, using user feedback to learn sampling distribution.
result Favourable trade-offs in quality-diversity and exploitation-exploration compared to existing methods.
LOUPE optimizes MRI under-sampling patterns for faster scans.
problem Accelerating MRI scans while maintaining image quality.
method End-to-end learning framework that trains on full-resolution scans.
result LOUPE-optimized masks yield superior reconstructions with 8x faster scans.
New matrix completion method for arbitrary sampling patterns using network flows.
problem Matrix completion under arbitrary sampling patterns.
method Network flow approach to matrix completion.
result Minimax optimal estimation for individual entries.
MCRapper efficiently computes patterns in data using Monte-Carlo Rademacher Averages.
problem Finding statistically significant patterns in data with limited samples.
method Monte-Carlo Empirical Rademacher Averages (MCERA) for poset families.
result MCRapper provides upper bounds to the discrepancy of functions, enabling efficient pattern mining.
New method optimizes MRI sampling patterns for faster scans.
problem Accelerate MRI scans without sacrificing image quality.
method Joint learning of adaptive sampling patterns and model-based recovery.
result Improved MR image quality compared to other methods.
R-GPM enables efficient graph pattern mining through user-defined relations.
problem Efficient graph pattern mining through user-defined relations.
method Parallel computing framework with MCMC sampling algorithm and optimizations.
result Efficient estimators for graph pattern statistics with up to 3-orders-of-magnitude computational cost reduction.
The paper analyzes conditions for low-rank tensor completion using TT decomposition.
problem Conditions for finite completability of low-rank tensors.
method Algebraic geometric analysis on the TT manifold, focusing on the independence of polynomials defined by sampling patterns and TT decompositions.
result Deterministic and probabilistic conditions for finite completability of tensors with high probability.
Two supervised methods classify single-molecule patterns from X-ray imaging.
problem Classifying high-quality patterns from noisy, stochastic XFEL data.
method Supervised template-based learning methods: Eigen-Image and Log-Likelihood classifiers.
result Classifiers can find best-matched templates within milliseconds and parallelize for XFEL repetition rate.
Spatio-temporal data compression method reduces memory usage.
problem Efficiently storing and analyzing large spatio-temporal datasets.
method Adaptive sampling of tensor slices to compress and preserve structure.
result SkeTenSmooth outperforms other sampling methods in retaining patterns.
PbP strategy improves logistic model prediction with missing values.
problem Predicting with missing inputs in logistic models.
method Pattern-by-Pattern (PbP) strategy for logistic models with missing values.
result PbP accurately approximates Bayes probabilities under GPMM across various missing data scenarios.
The paper analyzes how sampling design affects machine learning model generalization.
problem The impact of sampling properties on machine learning model generalization.
method Spectral analysis of the generalization error in Euclidean space using Fourier analysis.
result Estimation of expected error bounds and convergence rates for various sampling patterns.
Moon phases added to stock market analysis for better pattern recognition.
problem Finding meaningful patterns in stock market data using irregular time sampling.
method Incorporating Moon phases into the Gregorian calendar time sampling methods for stock market analysis.
result Moon phases provide unique, irregular sampling features for stock market pattern recognition.
MPPN network improves long-term time series forecasting accuracy.
problem Inaccurate long-term time series forecasting due to noise and lack of interpretability.
method MPPN network constructs context-aware multi-resolution semantic units and employs multi-periodic pattern mining and channel adaptive module.
result MPPN significantly outperforms state-of-the-art methods on nine real-world benchmarks.
New method for tensor completion with reduced sampling requirements.
problem Tensors with low CP rank completion under sampling.
method Analyzing manifold structure and defining polynomials based on sampling pattern and CP decomposition.
result Deterministic and probabilistic conditions for finite and unique completability with reduced sampling requirements.
CAG method predicts nonlinear solid mechanics responses in real-time with high accuracy and efficiency.
problem Real-time prediction of nonlinear solid mechanics responses.
method Clustering adaptive Gaussian process regression (CAG) method.
result Offers predictions within a second with high precision using only 20 samples.
Generative Adversarial Networks create realistic geology from sparse measurements.
problem Building models of subsurface geology from sparse physical measurements.
method Semantic inpainting with Generative Adversarial Networks.
result Generated samples mimic a distribution of geological patterns, not a single image.
CEDA analyzes large categorical datasets using tree geometry and binary codes.
problem Analyzing large categorical datasets with extreme-K samples. method CEDA uses tree geometry and binary codes to analyze categorical data.
result CEDA discovers patterns and evaluates their reliability in large categorical datasets.
Detects audio adversarial examples using anomalous pattern detection.
problem Identifies adversarial audio attacks in deep neural networks.
method Applies anomalous pattern detection in activation space of audio models.
result Can detect adversarial examples with up to 0.98 AUC, no degradation on benign samples.
A hybrid K-NN and SVM technique improves classification accuracy.
problem Improving classification accuracy in pattern recognition.
method Discriminative nearest neighbour classification combined with SVM.
result The hybrid technique outperforms state-of-the-art methods.
New guarantees for matrix completion from any deterministic sampling patterns.
problem Proving guarantees for low-rank matrix completion from non-random sampling schemes.
method Introduced a graph with observed entries as edges to analyze the performance of constrained nuclear norm minimization algorithm.
result The algorithm can successfully complete the matrix if the observation graph is well-connected and has similar node degrees.
CDPA identifies common and distinctive patterns in high-dimensional datasets.
problem Existing methods fail to capture the common pattern between coefficient matrices of shared latent factors.
method Proposes CDPA, an unsupervised learning method that incorporates both common and distinctive patterns of coefficient matrices.
result CDPA provides better characterization of common and distinctive patterns in high-dimensional datasets.
WTM reduces clause usage and computation time for pattern recognition.
problem High computation time and memory usage in Tsetlin Machine.
method Weighting clauses and using binomial sampling to reduce complexity.
result WTM achieves similar accuracy with fewer clauses and faster training.
New framework for unbiased sampling of temporal networks.
problem Challenges in analyzing and modeling large, continuous temporal networks.
method General framework for unbiased temporal network sampling with online, single-pass algorithms and unbiased estimators.
result Effective algorithms for fast, accurate, and memory-efficient statistical estimation of temporal network patterns and properties.
Pattern ensembling fills in missing or inaccurate trajectory data.
problem Incompleteness, missing information, and inaccuracies in geolocation data.
method Probabilistically ensemble similar trajectory patterns from the vicinity.
result Reconstructs missing or unreliable trajectory segments effectively.
Paper proposes efficient algorithm for recovering sparsity pattern from deterministic missing data.
problem Recovering sparsity pattern from datasets with deterministic missing structure.
method Proposes an efficient algorithm for missing value imputation using topological property of censorship filter.
result Consistently recovers the sparsity pattern with high probability in polynomial time and logarithmic sample complexity.
The standard approach to compressive sampling considers recovering an unknown deterministic signal with certain known structure, and designing the sub-sampling pattern and recovery algorithm based on the known structure. This approach requires looking for a good representation that reveals the signal structure, and sol…
Randomization helps verify if data mining results are due to inherent patterns.
problem Verify if data mining results are due to inherent patterns or coincidental findings.
method Metropolis sampling based on local swaps to randomize data while preserving discovered patterns.
result Randomized data often reveals that clustering results imply frequent pattern discovery.
We consider the problem of inferring the interactions between a set of N binary variables from the knowledge of their frequencies and pairwise correlations. The inference framework is based on the Hopfield model, a special case of the Ising model where the interaction matrix is defined through a set of patterns in the …
In the panoply of pattern classification techniques, few enjoy the intuitive appeal and simplicity of the nearest neighbor rule: given a set of samples in some metric domain space whose value under some function is known, we estimate the function anywhere in the domain by giving the value of the nearest sample per the …
Paper proposes active learning for hotspot detection in VLSI design.
problem Hotspot detection in VLSI design is computationally expensive and relies on costly reference libraries.
method Active learning-based layout pattern sampling and hotspot detection flow.
result Significantly reduces lithography simulation overhead with satisfactory detection accuracy.
A method to learn from noisy labels without a clean validation set.
problem Learning from samples with noisy labels.
method Limited Gradient Descent: modifying labels to estimate optimal stopping timing.
result Comparable and sometimes superior generalization performance compared to methods relying on clean validation sets.
Approach detects outliers in large data sets by consistent data points.
problem Lack of training data for supervised learning models.
method Two-phase approach: first phase identifies non-outliers, second phase uses one-class classifier.
result Quickly retrieves samples for consistent and non-outlier data sets.
A new MDS method uses derivative-free optimization for manifold learning.
problem Learning the intrinsic geometry of high-dimensional manifolds.
method Pattern Search Multidimensional Scaling (PS-MDS) using General Pattern Search (GPS) framework.
result PS-MDS accurately infers manifold geometry in clean and noisy synthetic datasets.
Discovering statistically significant patterns from databases is an important challenging problem. The main obstacle of this problem is in the difficulty of taking into account the selection bias, i.e., the bias arising from the fact that patterns are selected from extremely large number of candidates in databases. In …
FSR efficiently discovers significant patterns with few resampled datasets.
problem Mining significant patterns in transactional data, especially subgroups.
method FSR uses resampling to bound the supremum deviation of quality statistics, providing rigorous guarantees on false discoveries.
result FSR effectively discovers significant subgroups with a small number of resampled datasets.
New statistical measures assess group separability in low-dimensional geometrical spaces.
problem Lack of statistical measures to evaluate group separability in low-dimensional geometrical spaces.
method Proposed three statistical measures (PSI-ROC, PSI-PR, PSI-P) based on Projection Separability rationale.
result Statistical-based measures outperform traditional cluster validity indices in evaluating group separability.
Study finds conformity bias drives music sampling traditions.
problem How frequency-based bias drives cultural diversity in music sampling.
method Agent-based simulations in approximate Bayesian computation framework.
result Sampling patterns at population-level consistent with conformity bias.
Improved acoustic scene classification with factorized CNN.
problem Acoustic scene classification in varying environments.
method Large-margin factorized CNN with triplet loss.
result Improved performance and better generalization on unseen data.
Paper proposes TRA to learn multiple stock trading patterns.
problem Inconsistent i.i.d. assumption limits stock prediction performance.
method TRA architecture with Optimal Transport for pattern assignment.
result Improves information coefficient (IC) by 0.04-0.06 compared to baselines.
Designs a classifier for malware detection using MDL principle.
problem Static malware detection in PE files.
method Clustering, closed frequent pattern mining, MDL principle for compression.
result Our classifier performs close to deep neural networks.
A distributed algorithm learns patterns in large images and signals.
problem High-dimensional optimization in large images and signals.
method Distributed asynchronous algorithm with locally greedy coordinate descent.
result Patterns can be learned on large scales images from the Hubble Space Telescope.
This paper aims at the problem of link pattern prediction in collections of objects connected by multiple relation types, where each type may play a distinct role. While common link analysis models are limited to single-type link prediction, we attempt here to capture the correlations among different relation types and…
Guided warping augments time series data by aligning features with a teacher.
problem Small time series datasets limit neural network performance.
method Guided warping with a discriminative teacher to augment data deterministically.
result Significant improvement in performance on various time series datasets.
Proposes PENNs for deep learning with missing data.
problem Deep learning with missing covariates in multivariate nonparametric regression.
method Pattern Embedded Neural Networks (PENNs) combining imputation and neural networks.
result PENNs achieve minimax rate of convergence for typical cases, improving standard neural networks.
RotatE embeds knowledge graphs using rotations in complex space for better link prediction.
problem Predicting missing links in knowledge graphs.
method RotatE models relations as rotations in complex vector space, using self-adversarial negative sampling.
result RotatE models and infers various relation patterns, outperforming existing models.