Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

3026049061,208 · Jun 202019922001200920172026
48 results for internet search data

The study finds that Chinese internet users have different search behaviors and attention patterns.

problem Heterogeneity in search behavior and attention among Chinese internet users.
method Data extraction technology to analyze Baidu Index keyword search volume data.
result Chinese internet users exhibit different search behaviors and attention patterns.

Since time immemorial, people have been looking for ways to organize scientific knowledge into some systems to facilitate search and discovery of new ideas. The problem was partially solved in the pre-Internet era using library classifications, but nowadays it is nearly impossible to classify all scientific and popular…

2018-11-15abs ↗pdf ↗

Machine learning predicts COVID-19 activity in China.

problem Real-time forecasting of COVID-19 activity in Chinese provinces.
method Combines mechanistic disease models with digital traces (internet searches, news alerts). Uses clustering and data augmentation techniques.
result Stable and accurate forecasts 2 days ahead of current time, outperforming baseline models in 27 out of 32 provinces.

The study aims to explore the strength of causal relationship between stock price search interest and real stock market outcomes on worldwide equity market indices. Such a phenomenon could also be mediated by investor behavior and extent of news coverage. The stock-specific internet search trends data and corresponding…

2018-04-05abs ↗pdf ↗

Study finds 'happiness' search data predicts stock returns, suggesting utility needs impact firm performance.

problem Investing in firms that meet societal utility needs.
method Used Google Trends data on 'happiness' search volume to predict stock returns.
result Happiness search exposure (HSE) explains future stock returns, particularly for big and value firms.

Efficiently selects nearest neighbors for labeling to speed up active learning.

problem Intractable active learning and search for large-scale unlabeled data.
method Restricts candidate pool to nearest neighbors of labeled set.
result Achieved similar performance to global approach but reduced computational cost by up to 3 orders of magnitude.

Large speech dataset for commercial use with 9.98% word error rate.

problem Creating a diverse speech recognition dataset for commercial purposes.
method Internet search for licensed audio data with transcriptions, training model on the dataset.
result Model trained on dataset achieves 9.98% word error rate on Librispeech's test-clean test set.

This research examines relationship between staging of Venture Capital (VC) investments and social feedback visible in publicly available data on the Web. We address the question of Venture Capital investment sensitivity to performance and prospects of new venture, given as likelihood of obtaining future financing, ava…

2012-12-30abs ↗pdf ↗

Study investor attention using search volume data before and after mobile device popularity.

problem Accurately measure investor attention in a fast-paced market.
method Compare investor attention using search volume data before and after mobile device popularization.
result Investor attention measured using search volume data is more accurate and faster after mobile device popularization.

Deep learning approach for efficient IoT task scheduling in MEC networks.

problem Minimizing task latency in IoT users with large-scale MEC systems.
method Stacked auto-encoder for data compression, adaptive simulated annealing, experience replay.
result Near-optimal performance with significantly reduced computational time.

Study predicts internet-based treatment effects for GPPPD based on dyadic coping.

problem Identifying which patients will benefit most from internet-based GPPPD treatment.
method Developed a multivariable decision tree model using recursive partitioning.
result Predicts large effects for high dyadic coping patients, small effects for low dyadic coping patients.

We analyze two communication-efficient algorithms for distributed statistical optimization on large-scale data sets. The first algorithm is a standard averaging method that distributes the NN data samples evenly to $\nummac$ machines, performs separate minimization on each subset, and then averages the estimates. We p…

2012-09-19abs ↗pdf ↗

Feedback loops amplify dataset biases, affecting future model performance.

problem Feedback loops amplify biases in datasets, risking future model reliability.
method Formalized system where model interactions are recorded and reused, analyzed for bias amplification.
result Models that behave like samples from the training distribution are more stable and calibrated.

In this paper we present a review of the existing typologies of Internet service users. We zoom in on social networking services including blogs and crowdsourcing websites. Based on the results of the analysis of the considered typologies obtained by means of FCA we developed a new user typology of a certain class of I…

2013-11-30abs ↗pdf ↗

The Efficient Market Hypothesis (EMH) is widely accepted to hold true under certain assumptions. One of its implications is that the prediction of stock prices at least in the short run cannot outperform the random walk model. Yet, recently many studies stressing the psychological and social dimension of financial beha…

2013-10-20abs ↗pdf ↗

FinGPT democratizes financial data for LLMs, enabling innovation.

problem Limited financial text datasets and disparities between general and financial text data.
method Automates collection and curation of real-time financial data from diverse Internet sources, fine-tuning with RLSP and LoRA.
result Democratizes access to financial data for LLMs, enabling innovation.

Conventional surveillance systems for monitoring infectious diseases, such as influenza, face challenges due to shortage of skilled healthcare professionals, remoteness of communities and absence of communication infrastructures. Internet-based approaches for surveillance are appealing logistically as well as economica…

2018-11-27abs ↗pdf ↗

Paper develops security model and pricing for stable digital currency in quantum blockchain network.

problem Securing and pricing stable digital currency in a quantum blockchain network.
method Developed a block-based quantum channel networking technology and a FinTech platform model with dynamic pricing.
result Established a generalized IoB security model using quantum channel networking and QKD.

Deep learning model detects and corrects outliers in crowd-sourced weather data.

problem Data quality issues in crowd-sourced weather data.
method Bayesian deep learning approach with Gaussian-uniform mixture density network.
result Automated outlier detection in spatio-temporal environmental modeling.

Network detection is an important capability in many areas of applied research in which data can be represented as a graph of entities and relationships. Oftentimes the object of interest is a relatively small subgraph in an enormous, potentially uninteresting background. This aspect characterizes network detection as …

2013-03-22abs ↗pdf ↗

We introduce a mathematical criterion defining the bubbles or the crashes in financial market price fluctuations by considering exponential fitting of the given data. By applying this criterion we can automatically extract the periods in which bubbles and crashes are identified. From stock market data of so-called the …

2006-08-01abs ↗pdf ↗

Contextual bandits are widely used in Internet services from news recommendation to advertising, and to Web search. Generalized linear models (logistical regression in particular) have demonstrated stronger performance than linear models in many applications where rewards are binary. However, most theoretical analyses …

2017-02-28abs ↗pdf ↗