A new clustering algorithm GDT improves on HDBSCAN for uneven data.
problem Data clustering with uneven distribution and high noise.
method GDT combines local and global structures, forming local clusters and estimating a global topological graph based on connectivity between clusters.
result GDT achieves SOTA performance on various datasets with low time complexity.
Study investigates one-shot semi-supervised learning for image classification.
problem Training deep networks requires many labeled samples, limiting adoption.
method Empirical investigation of FixMatch method for one-shot semi-supervised learning.
result Uneven class accuracy is a barrier to high performance in one-shot semi-supervised learning.
New algorithm for robust circular coordinates in recurrent time series data.
problem Inefficient and sensitive methods for finding circular coordinates on recurrent data.
method Subsampling, aligning, and averaging to correct uneven sampling density.
result More robust and efficient circular coordinates for neuronal recordings.
New findings reveal discount regularization can be seen as a strong prior, leading to poor performance in unevenly sampled data.
problem Discount regularization leads to poor performance in unevenly sampled data.
method Equivalence theorem showing discount regularization as a strong prior, setting regularization parameters locally for individual state-action pairs.
result Discount regularization can be seen as a strong prior, leading to poor performance in unevenly sampled data.
We present surrogate regret bounds for arbitrary surrogate losses in the context of binary classification with label-dependent costs. Such bounds relate a classifier's risk, assessed with respect to a surrogate loss, to its cost-sensitive classification risk. Two approaches to surrogate regret bounds are developed. The…
Background elimination for noisy character images or character images from real scene is still a challenging problem, due to the bewildering backgrounds, uneven illumination, low resolution and different distortions. We propose a stroke-based character reconstruction(SCR) method that use a weighted quadratic Bezier cur…
We study historical calibration of one- and two-factor models that are known to describe relatively well the dynamics of energy underlyings such as spot and index natural gas or oil prices at different physical locations or regional power prices. We take into account uneven frequency of data due to weekends, holidays, …
Traditionally, practitioners initialize the {\tt k-means} algorithm with centers chosen uniformly at random. Randomized initialization with uneven weights ({\tt k-means++}) has recently been used to improve the performance over this strategy in cost and run-time. We consider the k-means problem with semi-supervised inf…
Deep learning predicts AMD progression from longitudinal fundus images.
problem Predicting future stages of age-related macular degeneration (AMD).
method InceptionV3 feature vectors, interval scaling, Recurrent Neural Network.
result 0.878 sensitivity, 0.887 specificity, 0.950 AUC.
GrateTile optimizes CNN feature map storage for efficient data access.
problem Efficient storage and access of sparse CNN feature maps.
method Divides feature maps into uneven-sized subtensors, compresses and stores them in a compressed yet accessible format.
result Average 55% DRAM bandwidth reduction with minimal indexing overhead.
New spectral clustering method for graphs with uneven node degrees.
problem Challenges in community detection for graphs with heterogeneous degree distributions.
method Spectral clustering on spherical coordinates with degree correction.
result Improved performance in representing computer networks.
We develop an approach to learn an interpretable semi-parametric model of a latent continuous-time stochastic dynamical system, assuming noisy high-dimensional outputs sampled at uneven times. The dynamics are described by a nonlinear stochastic differential equation (SDE) driven by a Wiener process, with a drift evolu…
Regularizes decision trees to reduce inference time by up to 4x with minimal accuracy loss.
problem Optimizing decision tree execution time on resource-constrained devices.
method Regularizes impurity computation during CART algorithm training to favor highly asymmetric distributions.
result Reduces inference time by up to 4x with minimal accuracy loss.
Joint diffusion models improve data representation for both generation and prediction.
problem Inconsistent performance between generation and classification tasks in joint models.
method Extended vanilla diffusion model with a classifier for joint end-to-end training.
result Joint diffusion model outperforms state-of-the-art hybrid methods in classification and generation.
The paper develops a method to learn robust decision policies from observational data, reducing high-cost outcomes.
problem Learning safe decision policies from observational data with high-risk outcomes.
method Develops a method to learn policies that reduce high-cost outcomes, valid under finite samples and uneven feature overlap.
result Validates the method with real and synthetic data, providing statistical bounds on decision costs.
Suppose you look at today's stock prices and bet on the value of the first digit. One could guess that a fair bet should correspond to the frequency of 1/9=11.11 for each digit from 1 to 9. This is by no means the case, and one can easily observe a strong prevalence of the small values over the large ones. The fir…
New neural network improves MRI reconstruction for non-Cartesian data.
problem Improving MRI reconstruction for non-Cartesian data acquisitions.
method Density-compensated unrolled neural networks.
result Density-compensated unrolled neural networks outperform baselines.
This research debiases machine unlearning by using counterfactual examples.
problem Machine unlearning processes can be biased, leading to inaccurate results.
method Intervention-based approach using counterfactual examples to mitigate biases.
result The method outperforms existing baselines on evaluation metrics.
Distributed machine learning is an approach allowing different parties to learn a model over all data sets without disclosing their own data. In this paper, we propose a weighted distributed differential privacy (WD-DP) empirical risk minimization (ERM) method to train a model in distributed setting, considering differ…
Improved nuclear cross section fitting with weighted Levenberg-Marquardt method.
problem Challenging optimization in multichannel nuclear cross section data.
method Weighted Levenberg-Marquardt algorithm with Fisher Information Metric.
result More physically consistent fits for raw and smoothed datasets.
Study on droplet flow on uneven surfaces, proving existence and properties.
problem Understanding droplet movement on irregular surfaces.
method Existence of smooth flow and 1/2-Hölder continuous minimizing movement solutions.
result Properties of minimizing movements including comparison principles and uniform boundedness.
MMCGAN uses explicit manifold learning to improve GAN performance.
problem GAN mode collapse and unstable training.
method Introduces Minimum Manifold Coding (MMC) as a prior to guide GAN training.
result MMCGAN effectively alleviates mode collapse and stabilizes GAN training.
Sparse random projection (RP) is a popular tool for dimensionality reduction that shows promising performance with low computational complexity. However, in the existing sparse RP matrices, the positions of non-zero entries are usually randomly selected. Although they adopt uniform sampling with replacement, due to lar…
WamOL uses PINNs to efficiently calibrate IVS from sparse data.
problem Calibrating time-dependent IVS from sparse market data.
method Physics-Informed Neural Networks (PINNs) with adaptive reweighting.
result WamOL outperforms in calibrating intraday IVS from uneven data.
Study shows non-systematic bias in customer satisfaction surveys limits data value.
problem Non-systematic bias in customer satisfaction surveys limits data value.
method Used real customer satisfaction survey data of a large retail bank to show the irreducible error and suggest thoughtful survey design methods.
result A thoughtful survey design can reduce non-systematic error in customer satisfaction surveys.
Bubbles are essential in certain economic models with high growth and low interest rates.
problem Asset price bubbles exceeding fundamental values.
method Developed the Bubble Necessity Theorem in economic models with specific growth and interest rate conditions.
result Bubbles are inevitable in certain economic scenarios with high growth and low interest rates.
Using results by Donaldson and Auroux on pseudo-holomorphic curves as well as Duval's rational convexity construction, the paper investigates the existence of smooth Lagrangian surfaces representing 2-dimensional homology classes in complex projective surfaces. We prove that if the projective surface X is minimal, of g…
A new DRL system using LSTM improves stock trading performance.
problem Adapting DRL to financial data with low signal-to-noise ratios.
method Cascaded LSTM networks for feature extraction and reinforcement learning.
result Our model outperforms previous models in cumulative returns and Sharp ratio.
Paper proposes a new method to handle missing data in medical records using sequential variational autoencoders.
problem Missing data in medical records due to sensor off-times and uneven data collection.
method Sequential variational autoencoders (VAEs) with a new methodology called Shi-VAE.
result Shi-VAE achieves the best performance in terms of both metrics compared to state-of-the-art methods.
Study examines how disturbances affect financial returns in Austrian forests.
problem Financial impact of disturbances on timberland returns in Austria.
method Applied probability theory to analyze two management regimes: even-aged and semi-stationary.
result Severe disturbances can lead to a shift from continuous-cover to even-aged forestry, affecting financial sensitivity.
CNNs can develop blind spots due to uneven padding in feature maps.
problem Spatial bias in convolutional networks leads to blind spots in certain tasks.
method Identified and analyzed the role of padding in convolutional networks, proposing solutions to mitigate bias.
result Mitigating spatial bias improves model accuracy, especially in tasks like small object detection.
OMTL uses ontology to learn from imbalanced EHR data.
problem Imbalanced and small usable data for phenotypes in EHRs.
method Ontology-driven multi-task learning framework.
result Improved learning performance on phenotypes.
Develops methods for valid and validated confidence sets in multiclass and multilabel prediction.
problem Challenges of typical conformal prediction methods in multiclass and multilabel problems, especially uneven coverage.
method Leverages quantile regression to build methods that always guarantee correct coverage and asymptotically optimal conditional coverage, addressing label interactions with tree-structured classifiers.
result Empirical evaluation suggests more robust coverage of confidence sets.
The aim of this work is to establish the personal income distribution from the elementary constituents of a free market; products of a representative good and agents forming the economic network. The economy is treated as a self-organized system. Based on the idea that the dynamics of an economy is governed by slow mod…
Improved diffusion models for image synthesis with better training dynamics.
problem Uneven and ineffective training in diffusion models.
method Redesigned network layers to preserve activation, weight, and update magnitudes.
result Significantly better networks at equal computational complexity, improving FID to 1.81.
To train good supervised and semi-supervised object classifiers, it is critical that we not waste the time of the human experts who are providing the training labels. Existing active learning strategies can have uneven performance, being efficient on some datasets but wasteful on others, or inconsistent just between ru…
New method generates geolocated synthetic populations from real data.
problem Generating synthetic populations with explicit geographic coordinates.
method Mapping coordinates into a latent space using Normalizing Flows (NF), then combining with other features in a Variational Autoencoder (VAE).
result NF+VAE architecture outperforms existing methods in generating geolocated synthetic populations.
Emergent misalignment is influenced by training dynamics, model priors, and data.
problem Emergent misalignment in models
method Exploring training dynamics, model priors, and data
result Activation deltas before and after narrow fine-tuning correlate with their similarities when measured with the last prompt-token activations.
LOOCV is often useful for analyzing small, structured experimental designs.
problem The effectiveness of cross-validation in analyzing designed experiments.
method Empirical study comparing LOOCV and other model selection methods.
result LOOCV is often useful in the analysis of small, structured experimental designs.
The paper addresses bias in visual recognition models by reweighting observations.
problem Bias in deep neural networks trained on biased image databases.
method Reweighting observations based on known biasing mechanisms to form a nearly debiased estimator.
result The approach can remedy representativeness issues in visual recognition systems.
This work improves communication efficiency in federated learning over wireless networks by optimizing energy consumption.
problem Optimizing energy consumption in federated learning over wireless networks.
method Adopting SignSGD for gradient sign exchange, considering channel capacity with outage, and proposing a stochastic sign-based algorithm for uneven data distribution.
result Proposed methods achieve a balance between learning performance and energy consumption.
The study analyzes how wartime controls influenced zaibatsu stock prices in Japan.
problem How wartime economic controls affected zaibatsu stock prices in Japan.
method Developed a four-portfolio asset-pricing model and used a CAPM-AR(p)-SV event-study framework.
result Wartime economic controls influenced stock prices through financing wedges and zaibatsu affiliation.
Recent research on Bitcoin Transaction Networks reveals a growing, sparse, and core-periphery structure.
problem Understanding the evolution of Bitcoin's network structure and user behavior.
method Review of recent results on Bitcoin Transaction Networks, including Address Network, User Network, and Lightning Network.
result Bitcoin Transaction Networks exhibit a core-periphery structure, indicating increasing centralization.
Fake engagement is one of the significant problems in Online Social Networks (OSNs) which is used to increase the popularity of an account in an inorganic manner. The detection of fake engagement is crucial because it leads to loss of money for businesses, wrong audience targeting in advertising, wrong product predicti…
DeepMaxent uses neural networks to improve species distribution models.
problem Sampling biases and lack of absence data in presence-only observations.
method DeepMaxent employs neural networks to learn shared features among species using the maximum entropy principle.
result DeepMaxent outperforms traditional methods in predicting species distributions, especially in unevenly sampled regions.
Success stories of applied machine learning can be traced back to the datasets and environments that were put forward as challenges for the community. The challenge that the community sets as a benchmark is usually the challenge that the community eventually solves. The ultimate challenge of reinforcement learning rese…
In image classification, visual separability between different object categories is highly uneven, and some categories are more difficult to distinguish than others. Such difficult categories demand more dedicated classifiers. However, existing deep convolutional neural networks (CNN) are trained as flat N-way classifi…
Proposes dynamic model type recommendation for OLP technique.
problem Limited local competence of base-classifiers in uneven data distributions.
method Builds a multi-label meta-classifier to recommend model types based on local data complexity.
result Statistically similar performance to original OLP with fixed base-classifier model.