Accurate real-time tracking of influenza outbreaks helps public health officials make timely and meaningful decisions that could save lives. We propose an influenza tracking model, ARGO (AutoRegression with GOogle search data), that uses publicly available online search data. In addition to having a rigorous statistica…
Google Trends can lead to misleading forecasts if not used carefully.
problem Misleading forecasts due to the variability of Google Trends data.
method Analyzing the variability of Google Trends data and proposing solutions.
result Google Trends can be a problem if not used with caution.
Google Trends data improves economic forecasts of private consumption.
problem Improving economic forecasts of private consumption.
method Machine learning techniques applied to categorized Google search data.
result Google data can identify patterns to generate a leading indicator in real time.
Ethereum trends analyzed through blockchain transactions and Google searches.
problem Identifying market manipulation in crypto prices.
method Big data analysis of Ethereum transactions, smart contracts, and search volumes.
result Big players manipulate crypto markets after price drops.
Study finds 'happiness' search data predicts stock returns, suggesting utility needs impact firm performance.
problem Investing in firms that meet societal utility needs.
method Used Google Trends data on 'happiness' search volume to predict stock returns.
result Happiness search exposure (HSE) explains future stock returns, particularly for big and value firms.
Study quantifies how COVID-19 spread affects US stock markets.
problem Impact of COVID-19 on US stock market during pandemic.
method Developed a novel temporal complex network approach using econometric and ML models.
result Local spread of COVID-19 and Google searches impact abnormal stock prices.
We study power-law correlations properties of the Google search queries for Dow Jones Industrial Average (DJIA) component stocks. Examining the daily data of the searched terms with a combination of the rescaled range and rescaled variance tests together with the detrended fluctuation analysis, we show that the searche…
Efficient search methods can outperform random search on challenging tasks.
problem Comparing the performance of efficient and random search methods in neural architecture search.
method Comparison of weight sharing and random search methods on progressively larger search spaces for image classification and detection.
result Efficient search methods can provide substantial gains over random search on large, realistic tasks.
Portfolio diversification and active risk management are essential parts of financial analysis which became even more crucial (and questioned) during and after the years of the Global Financial Crisis. We propose a novel approach to portfolio diversification using the information of searched items on Google Trends. The…
Study links public concern in Italy to financial markets worldwide.
problem Understanding public concern's impact on financial markets during pandemics.
method Used Google Trends data from YouTube, News, and Search to measure public concern and correlate it with stock index returns.
result Public concern in Italy drives concerns in other countries and explains stock index returns of multiple nations.
Financial market prediction on the basis of online sentiment tracking has drawn a lot of attention recently. However, most results in this emerging domain rely on a unique, particular combination of data sets and sentiment tracking tools. This makes it difficult to disambiguate measurement and instrument effects from f…
Open-source Vizier optimizes complex systems for Google and beyond.
problem Optimizing large-scale systems with multiple objectives and constraints.
method Distributed, fault-tolerant, flexible API for blackbox optimization.
result OSS Vizier supports a wide range of optimization problems and is available as open-source.
Study evaluates clustering methods for Google Trends data.
problem Clustering high-dimensional, noisy time series data.
method Symbolic Aggregate Approximation (SAX), Enhanced SAX (eSAX), and Topological Data Analysis (TDA).
result TDA provides more balanced and meaningful groupings than SAX and eSAX.
Study uses high-frequency data to predict ruble depreciation during crisis.
problem Predicting ruble depreciation during the Russian invasion of Ukraine.
method Uses intraday high-frequency data (google searches and implied volatility) to model exchange rate fluctuations.
result Implied volatility is more effective than attention in predicting ruble depreciation.
Paper optimizes KWS models using NAS and quantization for limited resources.
problem Developing efficient keyword spotting models in resource-constrained environments.
method Neural Architecture Search (NAS) for model structure optimization and quantization of weights and activations.
result Achieved high accuracy (95.55%) with minimal parameters and operations using NAS and quantization.
Study uses web search data to analyze tech startups growth.
problem Analyzing growth dynamics of tech startups.
method Utilized Google Trends data for 241 US-based tech startups.
result Web search traffic correlates positively with tech startup growth.
Robotics improves by using image search to solve new tasks.
problem Generalization in robotics.
method Combining visual and textual information to demarcate intended word meaning.
result Our approach leads to improved results compared to Google searches, treating the problem of polysemes.
Study uses neural networks to value Bitcoin options considering price jumps and sentiment.
problem Valuing Bitcoin options under price jumps and market sentiment.
method Bivariate jump-diffusion model, incorporating Google search sentiment, and artificial neural networks.
result Derives a closed formula for Bitcoin option pricing and validates using high-volatile stocks.
Among other macroeconomic indicators, the monthly release of U.S. unemployment rate figures in the Employment Situation report by the U.S. Bureau of Labour Statistics gets a lot of media attention and strongly affects the stock markets. I investigate whether a profitable investment strategy can be constructed by predic…
Studies have shown that the people depicted in image search results tend to be of majority groups with respect to socially salient attributes. This skew goes beyond that which already exists in the world - e.g., Kay et al. showed that although 28% of CEOs in US are women, only 10% of the top 100 results for CEO in Goog…
Search-based methods for hard combinatorial optimization are often guided by heuristics. Tuning heuristics in various conditions and situations is often time-consuming. In this paper, we propose NeuRewriter that learns a policy to pick heuristics and rewrite the local components of the current solution to iteratively i…
A new adversarial attack method using structured search and contextual bandits.
problem Black-box adversarial attacks on deep learning models.
method Structured search space and Bayesian optimization for contextual bandits.
result Achieves state-of-the-art success rates and query efficiencies.
Fully automating machine learning pipelines is one of the key challenges of current artificial intelligence research, since practical machine learning often requires costly and time-consuming human-powered processes such as model design, algorithm development, and hyperparameter tuning. In this paper, we verify that au…
Study examines how pandemic anxiety affects financial market trust.
problem Anxiety during pandemic and trust in financial markets.
method Used Google search volume and stock market data to create mood indicators.
result Different clusters of countries and markets in terms of pessimism and optimism emerged.
Machine learning speeds up search procedures for sorted tables.
problem Improving the speed of sorted table search procedures.
method Systematic experimental comparison of efficient implementations with learned counterparts.
result Learned data structures can significantly speed up search procedures.
Bitcoin's attention is linked to Google Trends data, not general uncertainty.
problem Bitcoin's correlation with Google Trends data was previously misunderstood.
method Analyzed bidirectional relationships between Bitcoin returns and Google Trends attention over six days.
result Information flows from Bitcoin volatility to Google Trends attention, not the other way.
Study of the forecasting models using large scale microblog discussions and the search behavior data can provide a good insight for better understanding the market movements. In this work we collected a dataset of 2 million tweets and search volume index (SVI from Google) for a period of June 2010 to September 2011. We…
DCN-V2 improves deep & cross network for web-scale learning to rank systems.
problem Efficiently learning feature interactions in large-scale recommender systems.
method Proposes DCN-V2, an improved framework for deep & cross network learning.
result DCN-V2 outperforms state-of-the-art algorithms on benchmark datasets.
Federated learning is a distributed form of machine learning where both the training data and model training are decentralized. In this paper, we use federated learning in a commercial, global-scale setting to train, evaluate and deploy a model to improve virtual keyboard search suggestion quality without direct access…
Study investor attention using search volume data before and after mobile device popularity.
problem Accurately measure investor attention in a fast-paced market.
method Compare investor attention using search volume data before and after mobile device popularization.
result Investor attention measured using search volume data is more accurate and faster after mobile device popularization.
Analyzing a comprehensive news dataset, we document that joint news coverage triggers attention contagion, causing temporarily inflated valuations for affected stocks. Tracing SEC EDGAR visits from unique IPs, we provide direct evidence of attention spillovers between stocks. Stocks with greater joint news coverage exh…
Improves search performance by transferring knowledge from recommender system.
problem Cold start and feedback loop problems in search retrieval.
method Zero-Shot Heterogeneous Transfer Learning framework.
result Significant improvements in relevance and user interactions over production system.
Empirical law predicts accuracy of Google Translate's translation chains.
problem Predicting accuracy in machine translation with multiple hops.
method Empirical testing of Google Translate's sequential translation.
result Accuracy decreases with the number of translating hops, following a power law.
The increasing inclusion of Deep Learning (DL) models in safety-critical systems such as autonomous vehicles have led to the development of multiple model-based DL testing techniques. One common denominator of these testing techniques is the automated generation of test cases, e.g., new inputs transformed from the orig…
MFIN networks improve crypto trading with multiple features.
problem Selecting and processing multiple features for effective trading.
method End-to-end framework using Multi-Factor Inception Networks (MFINs).
result MFINs learn uncorrelated, higher-Sharpe strategies not captured by traditional factors.
We present the first sublinear memory sketch that can be queried to find the nearest neighbors in a dataset. Our online sketching algorithm compresses an N element dataset to a sketch of size O(Nblog3N) in O(N(b+1)log3N) time, where b<1. This sketch can correctly report the nearest neighbors of any …
Using non-linear machine learning methods and a proper backtest procedure, we critically examine the claim that Google Trends can predict future price returns. We first review the many potential biases that may influence backtests with this kind of data positively, the choice of keywords being by far the greatest culpr…
The paper combines Bitcoin price models with expert corrections for better predictions.
problem Improving Bitcoin price predictions using statistical and expert insights.
method Linear regression models combined with expert corrections, utilizing Bayesian approach for fat-tailed distributions.
result Better price prediction results compared to using either model or expert opinion alone.
Review of childhood data for predicting overweight and obesity in later life.
problem Predicting overweight and obesity in later life using childhood data.
method Bibliographic searches and iterative searching of references.
result High performance prediction models often use short time periods or late childhood data.
We present an approach to automate the process of discovering optimization methods, with a focus on deep learning architectures. We train a Recurrent Neural Network controller to generate a string in a domain specific language that describes a mathematical update equation based on a list of primitive functions, such as…
The paper uses data science to predict stock trends of Amazon, Apple, Google, and Microsoft.
problem Short-term market movement prediction for major tech stocks.
method Combination of technical analysis and machine/deep learning for trend classification.
result Generated labels for data set: +1 (buy), 0 (hold), -1 (sell).
Optimized neural networks for Edge TPU achieve high accuracy in real-time image classification.
problem Designing neural networks for hardware accelerators to achieve optimal performance.
method Hardware-aware neural architecture search and model customization for Edge TPU.
result Improved accuracy-latency tradeoff on Pixel 4's Edge TPU compared to existing models.
LSTM outperforms traditional models in forecasting international migration.
problem Precise forecasting of international migration for policymaking.
method Replaced a gravity linear model with an LSTM approach using Google Trends data.
result LSTM approach combined with Google Trends data outperforms existing models.
Study uses Google matrix analysis to show how COVID-19 changed international trade flows.
problem Impact of COVID-19 on international trade patterns.
method Google matrix analysis of World Trade Network (WTN), including PageRank, CheiRank, and reduced Google matrix.
result Significant changes in international trade flows due to the pandemic, affecting export and import balances.
ChatGPT launch boosted AI-related crypto assets by 10.7% to 15.6%.
problem Investor perception of AI assets after ChatGPT launch.
method Synthetic difference-in-difference methodology.
result AI-related crypto assets experienced significant returns after ChatGPT launch.
MACH reduces memory usage for extreme classification by hashing.
problem Expensive training of deep models with large softmax layers.
method Merged-Average Classifiers via Hashing (MACH) using count-min sketch.
result Significant memory reduction and training speedup.
We present Listen, Attend and Spell (LAS), a neural network that learns to transcribe speech utterances to characters. Unlike traditional DNN-HMM models, this model learns all the components of a speech recognizer jointly. Our system has two components: a listener and a speller. The listener is a pyramidal recurrent ne…
New method classifies nonlinear time series using deep CNNs and bispectra.
problem Classifying nonlinear time series data effectively.
method Combines HOSA with deep CNNs.
result Effective classification of nonlinear time series data.