Joint peak detection is a central problem when comparing samples in genomic data analysis, but current algorithms for this task are unsupervised and limited to at most 2 sample types. We propose PeakSegJoint, a new constrained maximum likelihood segmentation model for any number of sample types. To select the number of…
FLOPART solves peak detection by creating accurate train and test set predictions.
problem Correctly detecting peaks in sequential data.
method Dynamic programming changepoint algorithm with zero train label errors.
result FLOPART provides highly accurate predictions on both train and test sets.
A nonparametric method for time series analysis extracts envelopes, detects peaks, and clusters data.
problem Extracting envelopes, detecting peaks, and clustering in time series data.
method Iterative procedure that minimizes L1 drift to create upper and lower bounding signals, using Viterbi-like path tracking and optimal elimination rules. result Efficiently calculated solution with near-linear time complexities for various applications.
Novel graph-based method detects R-peaks in noisy ECG signals without preprocessing.
problem Detecting R-peaks in noisy ECG signals for real-time analysis.
method Graph-constrained Changepoint Detection (GCCD) approach.
result GCCD achieves high sensitivity, positive predictivity, and low detection error rate.
Improved peak detection in ChIP-seq data reduces over-dispersion.
problem Over-dispersion in ChIP-seq data reduces peak detection accuracy.
method Supervised segmentation models with alternative noise assumptions.
result Improved peak detection accuracy compared to natural assumptions.
Paper proposes an extension of Peak criterion for selecting kernel bandwidth in SVDD for large datasets.
problem Selecting optimal kernel bandwidth parameter for SVDD in large datasets.
method Extend Peak criterion method for large datasets, modifying existing methods for comparison.
result Proposed method gives good results and demonstrates advantage over existing methods.
A Fourier transform approach optimizes clustering algorithms.
problem Optimizing clustering algorithms for accuracy and reliability.
method Fourier transform and Gaussian filtering to smooth density functions, detecting peaks as cluster centroids.
result Remarkable accuracy in finding cluster centroids, overcoming initialization problems.
New algorithm detects changes in genomic data faster and more accurately.
problem Detecting changes in genomic data with constraints.
method Adapting a functional pruning technique to solve constrained changepoint detection problems.
result Log-linear time complexity algorithm achieves state-of-the-art accuracy.
HyPV-LEAD detects cryptocurrency anomalies proactively, improving financial security.
problem Cryptocurrency anomalies like mixing, fraud, and pump-and-dump operations are hard to detect due to class imbalance and temporal volatility.
method HyPV-LEAD integrates lead time into anomaly detection through window-horizon modeling, Peak-Valley sampling, and hyperbolic embedding.
result HyPV-LEAD achieves a PR-AUC of 0.9624 on Bitcoin transaction data, significantly outperforming state-of-the-art methods.
Deep learning detects arrhythmia from RR-interval ECG data.
problem Diagnosing arrhythmia using ECG data.
method Convolutional neural network (CNN) on time-sliced RR-interval data.
result Compact system achieves accurate arrhythmia detection.
New clustering algorithm for mixed data improves applicability and efficiency.
problem Clustering large, mixed data with improved accuracy and efficiency.
method Developed a new clustering algorithm using peak-finding technique, reducing computational complexity.
result Algorithm detects outliers, clusters of lower density, and determines correct number of clusters.
Proposes a new hierarchical clustering method combining DP and DBSCAN strengths.
problem Combining strengths of DP and DBSCAN for arbitrary shape clusters.
method DC-HDP: Combines Density Peak and Density-Connectivity approaches.
result Produces best clustering results on 14 datasets.
Self-supervised model detects phoneme boundaries without annotations.
problem Unsupervised phoneme segmentation without manual annotations.
method Convolutional neural network trained with Noise-Contrastive Estimation.
result Model outperforms baselines on TIMIT and Buckeye corpora.
Study analyzes Bitcoin price dynamics from 2012 to 2018, identifying major price peaks.
problem Understanding Bitcoin price fluctuations and predicting market crashes.
method Automatic peak detection, Lagrange Regularisation Method, LPPLS model for predicting crashes.
result Identification of 3 major and 10 smaller price peaks over the analyzed period.
Paper proposes bypassing implicit assumption in GM-based AD methods.
problem Lack of anomalous data and implicit assumption in GM-based AD methods.
method Integrating Discriminative idea to GMM for AD tasks (DiGMM).
result Establishes a connection between generative and discriminative models for AD.
Quantile gradient boosted trees outperform other models in predicting NO2 concentration distributions.
problem Forecasting high NO2 concentration episodes for effective air quality management.
method Compared 10 probabilistic forecasting models for NO2 concentration prediction.
result Quantile gradient boosted trees model outperformed others in predicting NO2 concentration distributions.
Method identifies financial rogue waves close to their onset.
problem Identifying extreme financial events close to their onset.
method Analogy between rogue waves in optics and financial volatility, using Schrödinger equation with potential shaped by Kerr nonlinearity.
result Numerical gradient spikes at the onset of extreme financial events.
Automates key steps in NMR protein structure elucidation.
problem Laborious data analysis limits NMR spectroscopy potential.
method Combination of deep learning, non-parametric models, and combinatorial optimization.
result Automated detection and assignment of signals in NMR data.
Paper presents a method for identifying isotope envelopes in MALDI-ToF data.
problem Deisotoping of isotopic peaks in MALDI-ToF molecular imaging data.
method Uses Mamdani-Assilan fuzzy system and spatial maps of molecular distribution to identify isotope envelopes.
result Proposed method detects overlapping envelopes and analyzes large data sets.
Study quantifies pedestrian traffic patterns in NYC.
problem Understanding pedestrian traffic dynamics in urban areas.
method Publicly available traffic camera data in NYC, time series analysis.
result Pedestrian traffic exhibits diurnal patterns with weekday peaks and no peak on weekends.
We consider a mean-reverting stochastic volatility model which satisfies some relevant stylized facts of financial markets. We introduce an algorithm for the detection of peaks in the volatility profile, that we apply to the time series of Dow Jones Industrial Average and Financial Times Stock Exchange 100 in the perio…
Bayesian model predicts emotion from fitness tracker heartbeat data.
problem Predicting emotional valence from consumer fitness tracker heartbeat data.
method End-to-end Bayesian deep learning model using PPG data.
result Peak F1 score of 0.7 for emotional valence classification.
A robust method for decomposing spectral peaks robust to distortion and interference.
problem Decomposing spectral peaks in the presence of distortion and interference.
method Optimizing a nonparametric approach using pseudo-symmetric functions with nonincreasing behavior.
result Decomposed spectral peaks show pseudo-orthogonal behavior and power preserving equality.
SPADE improves demand forecasting accuracy by 4.5% for post-promotion periods.
problem Overreacting to peak events in demand forecasting leads to biased forecasts.
method SPADE splits forecasting into two tasks: one for peak events and another for post-peak events, using masked convolution filters and a specialized Peak Attention module.
result Overall PPE improvement of 4.5%, 30% improvement for most affected forecasts after promotions and holidays, and 3.9% improvement in PE accuracy.
DADC algorithm improves clustering for data with varying density.
problem Sparse cluster loss and cluster fragmentation in density peak clustering.
method Domain-adaptive density measurement, cluster center self-identification, and cluster self-ensemble.
result DADC achieves more reasonable clustering results on data with varying density.
We find empirically a characteristic sharp peak-flat trough pattern in a large set of commodity prices. We argue that the sharp peak structure reflects an endogenous inter-market organization, and that peaks may be seen as local ``singularities'' resulting from imitation and herding. These findings impose a novel strin…
In this work, we propose a simple yet effective solution to the problem of connectome inference in calcium imaging data. The proposed algorithm consists of two steps. First, processing the raw signals to detect neural peak activities. Second, inferring the degree of association between neurons from partial correlation …
Mass spectrometry (MS) is an important technique for chemical profiling which calculates for a sample a high dimensional histogram-like spectrum. A crucial step of MS data processing is the peak picking which selects peaks containing information about molecules with high concentrations which are of interest in an MS in…
A new method for analyzing high-dimensional data using density peaks.
problem Analyzing complex, high-dimensional data sets.
method Density Peak clustering combined with a non-parametric density estimator.
result Automatic identification of density peaks and valleys in data.
Study on-chain peak shaving to reduce Ethereum transaction costs.
problem Reducing transaction costs in blockchain networks, especially during congested periods.
method Analyzing transaction-level data from multiple firms across various industries to understand scheduling responses and cost management strategies.
result Firms' scheduling responses to congestion vary, leading to different fee savings and residual costs.
The paper explains two distinct peaks in generalization error for neural networks and simpler models, each governed by different factors.
problem Understanding the peaks in generalization error for neural networks and simpler models.
method Analysis of random feature models and comparison with numerical experiments involving deep neural networks.
result The peaks at N=P and N=D are distinct and governed by different factors (noise sensitivity vs. initialization noise). The heuristic identification of peaks from noisy complex spectra often leads to misunderstanding of the physical and chemical properties of matter. In this paper, we propose a framework based on Bayesian inference, which enables us to separate multipeak spectra into single peaks statistically and consists of two steps.…
Deep learning detects atrial fibrillation from wearable sensor data.
problem Detecting atrial fibrillation from raw sensor data.
method Convolutional-recurrent neural network with long short-term memory, end-to-end learning.
result State-of-the-art AFib detection with high accuracy.
Bayesian framework integrates spectral deconvolution with expert reasoning for robust peak estimation.
problem Challenges in extracting meaningful peaks from noisy or complex spectra.
method Bayesian spectral deconvolution coupled with a physical-property regression layer.
result Recovery of weak peaks in poly(lactic acid) IR spectra related to degradation rates.
The paper shows how the generalization curve can have multiple peaks, influenced by data and learning algorithm biases.
problem Understanding the generalization behavior of linear regression models under varying parameterizations.
method Analyzes generalization loss in linear regression models with varying parameterizations, both under- and over-parameterized.
result The generalization curve can have an arbitrary number of peaks, and their locations can be controlled.
PEAKS selects key training examples incrementally based on prediction error and kernel similarity.
problem Dynamic data selection in deep learning models.
method Prediction Error Anchored by Kernel Similarity (PEAKS) for incremental data selection.
result PEAKS outperforms existing selection strategies and yields better performance returns as training data size grows.
PEAK tests means of multiple data streams with sequential betting.
problem Testing means of multiple data streams with nonparametric methods.
method Sequential, nonparametric testing using a betting scheme.
result PEAK provides up to 85% reduction in samples for stopping.
Study analyzes cryptocurrency pump-and-dump dynamics using minute-level data.
problem Identifying and quantifying insider trading in cryptocurrency markets.
method Algorithmic identification of insider volume spikes, conservative profit bounds calculation, social-media verification.
result Median returns above 100%, upper-quartile returns exceeding 2000% for insider profits.
Finite-time queue peaks in stochastic networks have logarithmic scaling after geometric thresholds.
problem Queue peak laws in stochastic networks with geometric thresholds.
method Self-normalization mechanism
result Logarithmic scaling of queue peaks after geometric thresholds.
LLMs learn peaked distributions slowly due to power-law losses.
problem Slow convergence of loss in training large language models.
method Systematic analysis of toy models and empirical evaluation of LLMs.
result Power-law time scaling with an exponent of 1/3 for learning peaked distributions.
Peaking phenomenon in semi-supervised learning observed and explained.
problem The peaking phenomenon in semi-supervised learning.
method Simulation studies and approximation of the learning curve.
result The learning curve in semi-supervised learning has a steeper incline and a more gradual decline.
Bayesian Quadrature improves ensembling for neural networks with dispersed likelihood peaks.
problem Ensembling neural networks struggles with dispersed, narrow peaks in likelihood surfaces.
method Uses Bayesian Quadrature to construct weighted ensembles of architectures.
result Empirically outperforms state-of-the-art baselines in test likelihood, accuracy, and expected calibration error.
Paper finds linear laws in Bitcoin price changes, aiding anomaly detection.
problem Detecting anomalies in Bitcoin price changes.
method Time embedding of autocorrelation function, binary series generation, stepped time windows.
result Linear laws became more complex before major market events, suggesting price manipulation.
By combining (i) the economic theory of rational expectation bubbles, (ii) behavioral finance on imitation and herding of investors and traders and (iii) the mathematical and statistical physics of bifurcations and phase transitions, the log-periodic power law model has been developed as a flexible tool to detect bubbl…
During a stock market peak the price of a given stock (i) jumps from an initial level p1(i) to a peak level p2(i) before falling back to a bottom level p3(i). The ratios A(i)=p2(i)/p1(i) and B(i)=p3(i)/p1(i) are referred to as the peak- and bottom-amplitude respectively. The paper show…
New insights into overfitting peaks in generalization error for l2 and l1 penalized interpolation.
problem Understanding the phenomenon of overfitting peaks in generalization error for modern machine learning models.
method Introducing a generative and fitting model pair (MiSpaR) and deriving analytical risk curves for l2 and l1 penalties. result The overfitting peak can be dissociated from the point of model flexibility, complicating the interpretation of overfitting as a boundary between classical and modern regimes.
In this paper, the fractional order curvature equation (−Δ)γu=(1+εK(x))uN−2γN+2γ in RN is considered. Assuming K(x) has two critical points satisfying certain local conditions, we prove the existence of two-peak solutions.
Populations of species in ecosystems are often constrained by availability of resources within their environment. In effect this means that a growth of one population, needs to be balanced by comparable reduction in populations of others. In neutral models of biodiversity all populations are assumed to change increment…