What books will be your bestseller? A machine learning approach with Amazon Kindle

Research output: Contribution to journalArticlepeer-review

Abstract

Purpose: With the rapid increase in internet use, most people tend to purchase books through online stores. Several such stores also provide book recommendations for buyer convenience, and both collaborative and content-based filtering approaches have been widely used for building these recommendation systems. However, both approaches have significant limitations, including cold start and data sparsity. To overcome these limitations, this study aims to investigate whether user satisfaction can be predicted based on easily accessible book descriptions. Design/methodology/approach: The authors collected a large-scale Kindle Books data set containing book descriptions and ratings, and calculated whether a specific book will receive a high rating. For this purpose, several feature representation methods (bag-of-words, term frequency–inverse document frequency [TF-IDF] and Word2vec) and machine learning classifiers (logistic regression, random forest, naive Bayes and support vector machine) were used. Findings: The used classifiers show substantial accuracy in predicting reader satisfaction. Among them, the random forest classifier combined with the TF-IDF feature representation method exhibited the highest accuracy at 96.09%. Originality/value: This study revealed that user satisfaction can be predicted based on book descriptions and shed light on the limitations of existing recommendation systems. Further, both practical and theoretical implications have been discussed.

Original languageEnglish
Pages (from-to)137-151
Number of pages15
JournalElectronic Library
Volume39
Issue number1
DOIs
StatePublished - 2020

Keywords

  • Book descriptions
  • Content-based filtering
  • Machine learning
  • Natural language processing
  • Recommendation systems

Fingerprint

Dive into the research topics of 'What books will be your bestseller? A machine learning approach with Amazon Kindle'. Together they form a unique fingerprint.

Cite this