There are few papers about the consumption pattern of the Portuguese wine, using econometrics techniques. This work, pretend to analyze the consumers behavior of the wine produced in Portugal, determining the demand equation with panel data methods. There were used statistical data available in the Alentejo Regional Wi…
DL models can outperform regionalized models in hydrology by pooling diverse data.
problem Traditional wisdom in hydrology suggests regionalization improves model performance, but DL models can unify data for better performance.
method Used DL models on pooled data from different regions, showing improved performance compared to regionalized models.
result DL models can improve performance by pooling diverse data, highlighting the 'data synergy' effect.
New method combines regional HIV prevention trial data without sharing individual patient info.
problem Regional differences in HIV prevention efficacy, privacy concerns, and data sharing limitations.
method Federated learning approach that combines site-specific estimators via L1-regularization.
result Improved precision in estimating region-specific survival curves.
The paper develops optimal confidence regions for categorical data.
problem Constructing tight confidence regions for categorical data.
method Develops new theory for minimum average volume confidence regions.
result Shows optimality of the regions for categorical data and its implications for machine learning.
The paper contains a short review of techniques examining regional wealth inequalities based on recently published research work but is also presenting unpublished features. The data pertains to Italy (IT), over the period 2007-2011: the number of cities in regions, the number of inhabitants in cities and in regions, a…
Adaptive region-based active learning seeks labels for complex data.
problem Efficiently label complex datasets with minimal human effort.
method Adaptive region partitioning and active learning for distinct predictors.
result Substantial empirical benefits over existing methods.
Employing profits data of Japanese companies in 2002 and 2003, we identify the non-Gibrat's law which holds in the middle profits region. From the law of detailed balance in all regions, Gibrat's law in the high region and the non-Gibrat's law in the middle region, we kinematically derive the profits distribution funct…
Geotagged data can be used to describe regions in the world and discover local themes. However, not all data produced within a region is necessarily specifically descriptive of that area. To surface the content that is characteristic for a region, we present the geographical hierarchy model (GHM), a probabilistic model…
Spatially constrained clustering divides landscapes into homogeneous regions with spatial contiguity.
problem Dividing landscapes into homogeneous patches with spatial contiguity and hierarchy.
method Developed a spatially constrained spectral clustering framework using a flexible kernel and recursive bisection.
result The proposed framework outperforms baseline methods in balancing region contiguity and homogeneity.
Stochastic partition models tailor a product space into a number of rectangular regions such that the data within each region exhibit certain types of homogeneity. Due to constraints of partition strategy, existing models may cause unnecessary dissections in sparse regions when fitting data in dense regions. To allevia…
Estimates the upper bound of linear regions in spheres centered at specific data points in ReLU neural networks.
problem Bounding the number of linear regions in specific areas of neural networks using ReLU activations.
method Developed a method to estimate the upper bound of linear regions in any sphere within the input space of a ReLU neural network.
result The boundaries of linear regions move away from training data points during training, and spheres centered at these points contain more regions than arbitrary points.
Scalable method for regionalizing and extracting temporal patterns from time series data.
problem Static spatial snapshots and ad hoc regularization limit effective spatial analysis and resource management.
method Minimum description length principle for fully nonparametric spatial partitioning and time series archetypes.
result Accurately recovers planted regional structure and drivers in synthetic and empirical data.
A new method selects regions of interest in GC-MS data without prior target selection.
problem Challenges in GC-MS data analysis due to fragmentation and shared fragment ions.
method Uses a pseudo F-ratio moving window (ψFRMV) to automatically select regions of interest. result Algorithm can accurately identify signal regions in GC-MS data.
sBayFDNN bridges deep learning and functional data analysis for complex, structured data.
problem Challenges in functional data analysis, especially for complex, continuously structured data.
method Sparse Bayesian functional deep neural network (sBayFDNN) that learns adaptive functional embeddings and interpretable region selection.
result First theoretical guarantees for a Bayesian deep functional model, ensuring reliability and statistical rigor.
A new method for efficient diffusion geometry computation from data regions.
problem Heavy computational load in diffusion maps for modern data analysis.
method Compressed diffusion process between data regions using an adapted MGC kernel.
result Efficient pointwise diffusion map embedding from data regions.
Paper defines ε-Safe Decision Regions for exponential family distributions and approximates them for unbalanced data.
problem Need probabilistic guarantees for reliable predictions in machine learning.
method Formalizes ε-Safe Decision Regions, proves their form for exponential family distributions, and develops Multi Cost SVM for unbalanced data.
result Formal definition and analytical determination of ε-Safe Decision Regions for exponential family distributions.
JANET improves time series prediction with adaptive uncertainty regions.
problem Time series data's lack of exchangeability and multi-step prediction challenges.
method Proposes JANET, a framework for joint adaptive prediction regions with controlled error rates.
result Demonstrates superior performance in multi-step prediction tasks across diverse datasets.
DiwE uses regional distribution changes to create diverse ensemble classifiers for concept drift.
problem Handling concept drift in evolving data streams.
method DiwE measures diversity based on regional distribution disagreement and uses it to weight instances and select classifiers.
result DiwE outperforms other algorithms on various synthetic and real-world data stream benchmarks.
Meta-learning improves few-shot land cover classification across diverse regions.
problem Capturing diversity in land cover classification across different geographic regions.
method Model-agnostic meta-learning (MAML) algorithm applied to classification and segmentation tasks.
result Few-shot model adaptation outperforms traditional methods in diverse land cover classification tasks.
Paper develops flood forecasting system for data scarce regions.
problem Flood forecasting in developing countries with data scarcity.
method Operational system for flood extent forecasts in India.
result Scalable and cost-efficient flood forecasting system.
Develops a method to segment high-dimensional data into regions with different local intrinsic dimensions.
problem Data often has varying intrinsic dimensions within the same dataset, challenging traditional analysis.
method Discovers regions with different local intrinsic dimensions and segments the data accordingly.
result Many real-world data sets contain regions with widely heterogeneous dimensions, which can be segmented using local intrinsic dimension.
Rectangular Bounding Process (RBP) improves partitioning efficiency in multi-dimensional spaces.
problem Creating many unnecessary divisions in sparse regions when describing dense regions.
method Introduces Rectangular Bounding Process (RBP) to efficiently partition multi-dimensional spaces using a bounding strategy.
result The RBP is self-consistent and can be extended to infinite space, offering rich yet parsimonious expressiveness.
New methods improve prediction regions for high-dimensional data.
problem Creating effective prediction regions for high-dimensional data.
method CD-split and HPD-split methods that combine split method and data-driven partition.
result CD-split and HPD-split converge to oracle highest predictive density set and satisfy local and asymptotic conditional validity.
In this paper an approach based on expectation maximization (EM) clustering to find the climate regions and a support vector machine to build a predictive model for each of these regions is proposed. To minimize the biases in the estimations a ten cross fold validation is adopted both for obtaining clusters and buildin…
The growing conflicts in and about oil exporting regions and speculations about volatile oil prices during the last decade have renewed the public interest in predictions for the near future oil production and consumption. Unfortunately, studies from only 10 years ago, which tried to forecast the oil production during …
Employing data on the assessed value of land in 1974--2007 Japan, we exhibit a quasistatically varying log-normal distribution in the middle scale region. In the derivation, a Non-Gibrat's law under the detailed quasi-balance is adopted together with two approximations. The resultant distribution is power-law with the …
Study uses open data to improve traffic emissions estimation.
problem Estimating accurate link level traffic emissions.
method Data-driven framework integrating MOVES, GPS, OSM, and satellite imagery.
result Neural network reduces RMSE by over 50% for key pollutants.
We present an approach for polarimetric Synthetic Aperture Radar (SAR) image region boundary detection based on the use of B-Spline active contours and a new model for polarimetric SAR data: the GHP distribution. In order to detect the boundary of a region, initial B-Spline curves are specified, either automatically or…
Proposes a method to accelerate safe sequential learning using offline data.
problem Limited exploration due to disconnected safe regions and slow task learning.
method Safe transfer sequential learning using Gaussian processes and offline data.
result Enhances global exploration across multiple disjoint safe regions with lower data consumption.
Study uses social network data to analyze regional inflation trends.
problem Analyzing inflation trends using social media data.
method BERT neural networks for identifying pro-inflationary and disinflationary keywords.
result Models can visualize and classify inflationary keywords in different contexts.
Method constructs confidence regions for linear models with arbitrary predictors.
problem Constructing confidence regions for linear models with non-linear predictors.
method Mixed Integer Linear Programming for constraints.
result Empty confidence regions for hypothesis testing.
The paper models crime risk using Foursquare check-ins and mobility data.
problem Understanding and predicting crime risk in urban areas.
method Directed graph of aggregated movement data, region risk factor derivation, DIFFER features.
result Reliable correlations between DIFFER features and crime count observed.
Automatic detection of anomalies in space- and time-varying measurements is an important tool in several fields, e.g., fraud detection, climate analysis, or healthcare monitoring. We present an algorithm for detecting anomalous regions in multivariate spatio-temporal time-series, which allows for spotting the interesti…
One-hot CNN (convolutional neural network) has been shown to be effective for text categorization (Johnson & Zhang, 2015). We view it as a special case of a general framework which jointly trains a linear model with a non-linear feature generator consisting of `text region embedding + pooling'. Under this framework, we…
This work aims to test the Verdoorn Law, with the alternative specifications of (1)Kaldor (1966), for five regions (NUTS II) Portuguese from 1986 to 1994 and for the 28 NUTS III Portuguese in the period 1995 to 1999. Will, therefore, to analyze the existence of increasing returns to scale that characterize the phenomen…
Article addresses practical challenges in conformal prediction.
problem Challenges in determining, computing, and controlling conformal prediction regions.
method Proposes a quadratic-polynomial non-conformity measure.
result Allows circumventing three challenges in full conformal prediction framework.
The paper introduces CoCoCat bonds for multi-region natural catastrophes, accounting for complex dependencies.
problem Valuation of multi-region contingent convertible bonds under complex dependencies.
method Developed a model accounting for inter-regional dependencies using change-of-measure techniques.
result Significant impact of inter-regional dependencies on CoCoCat bond pricing.
Reweighting improves risk bounds in certain data regions.
problem Improving risk bounds in classification and heteroscedastic regression.
method Weighted empirical risk minimization with a data-dependent weight function.
result A weighted ERM estimator can achieve superior performance in specific sub-regions.
This paper explores a real-world fundamental theme under a data science perspective. It specifically discusses whether fraud or manipulation can be observed in and from municipality income tax size distributions, through their aggregation from citizen fiscal reports. The study case pertains to official data obtained fr…
A new method for time-series data provides guaranteed coverage and adapts to non-exchangeable data.
problem Guaranteed coverage for time-series data prediction intervals.
method Sequential Conformalized Density Regions (SCDR) using quantile random forest.
result SCDR achieves guaranteed asymptotic coverage and outperforms existing methods in simulations.
Method identifies regions of maximum dissimilarity in stochastic processes.
problem Comparing local characteristics of two random processes to find periods of maximum dissimilarity.
method Bayesian inference with integrated nested Laplace approximation for stochastic processes.
result Identifies regions of maximum dissimilarity with a certain volume.
Heavy-tailed distributions are frequently used to enhance the robustness of regression and classification methods to outliers in output space. Often, however, we are confronted with "outliers" in input space, which are isolated observations in sparsely populated regions. We show that heavy-tailed stochastic processes (…
New method uses conformalization to create classification regions from ambiguous labels.
problem Creating provable guarantees in classification with uncertain labels.
method Conformal methods applied to credal regions for classification problems.
result New method provides smaller and more disentangled prediction sets.
Proposes a method to estimate acceptance regions for many classes, including new ones.
problem Lack of methods to handle new classes in set-valued classification.
method Generalized Prediction Set (GPS) approach to estimate acceptance regions.
result Achieves a good balance between accuracy, efficiency, and anomaly detection.
Proposes CPO framework for robust decision-making with explainable uncertainty regions.
problem Overly conservative uncertainty regions in data-driven optimization lead to suboptimal decisions.
method Conformal-Predict-Then-Optimize (CPO) framework using conditional generative models and visual summaries.
result Demonstrates improved robustness and explainability in decision-making.
Novel BSG method for efficient stochastic optimization.
problem Efficient optimization of non-convex surfaces in stochastic settings.
method Binary search combined with first order gradient optimization.
result BSG produces more promising results and better generalization than other methods.
Transfer learning improves highway traffic forecasting using graph neural networks.
problem Lack of historical data for traffic forecasting on large highway networks.
method Developed a transfer learning approach for DCRNN, a graph neural network for highway forecasting.
result TL-DCRNN can forecast traffic on unseen regions of the highway network with high accuracy.
Stochastic variational inference allows for fast posterior inference in complex Bayesian models. However, the algorithm is prone to local optima which can make the quality of the posterior approximation sensitive to the choice of hyperparameters and initialization. We address this problem by replacing the natural gradi…