socialreaper is a Python 3.6+ library that scrapes Facebook, Twitter, Reddit, Youtube, Pinterest, and Tumblr.
Not a programmer? Try the GUI
pip3 install socialreaper
For version 0.3.0 only
pip3 install socialreaper==0.3.0
Get the comments from McDonalds' 1000 most recent posts
fromsocialreaperimportFacebookfbk=Facebook("api_key")
comments=fbk.page_posts_comments("mcdonalds", post_count=1000, comment_count=100000)
forcommentincomments:
print(comment['message'])Save the 500 most recent tweets from the user @realDonaldTrump to a csv file
fromsocialreaperimportTwitterfromsocialreaper.toolsimportto_csvtwt=Twitter(app_key="xxx", app_secret="xxx", oauth_token="xxx", oauth_token_secret="xxx")
tweets=twt.user("realDonaldTrump", count=500, exclude_replies=True, include_retweets=False)
to_csv(list(tweets), filename='trump.csv')Get the top 10 comments from the top 50 threads of all time on reddit
fromsocialreaperimportRedditfromsocialreaper.toolsimportflattenrdt=Reddit("xxx", "xxx")
comments=rdt.subreddit_thread_comments("all", thread_count=50, comment_count=500, thread_order="top", comment_order="top", search_time_period="all")
# Convert nested dictionary into flat dictionarycomments= [flatten(comment) forcommentincomments]
# Sort by comment scorecomments=sorted(comments, key=lambdak: k['data.score'], reverse=True)
# Print the top 10forcommentincomments[:9]:
print("###\nUser: {}\nScore: {}\nComment: {}\n".format(comment['data.author'], comment['data.score'], comment['data.body']))Get the comments containing the strings prize, giveaway from
youtube channel mkbhd's videos
fromsocialreaperimportYoutubeytb=Youtube("api_key")
channel_id=ytb.api.guess_channel_id("mkbhd")[0]['id']
comments=ytb.channel_video_comments(channel_id, video_count=500, comment_count=100000, comment_text=["prize", "giveaway"], comment_format="plainText")
forcommentincomments:
print(comment)You can export a list of dictionaries using socialreaper's CSV class
fromsocialreaperimportFacebookfromsocialreaper.toolsimportCSVfbk=Facebook("api_key")
posts=list(fbk.page_posts("mcdonalds"))
CSV(posts, file_name='mcdonalds.csv')