29. Web scraping

[status: mostly-complete-needs-polishing-and-proofreading]

29.1. Motivation, prerequisites, plan

The web is full of information and we often browse it visually with a browser. But when we collect a scientific data set from the web we do not want to have a “human in the loop”, rather we want an automatic program to collect that data so that our results can be reproducible and our procedure can be fast and automatic.

Although my focus here is mainly on scientific applications, web scraping can also be used to mirror a web site.

Prerequisites

  • The 10-hour “serious programming” course.

  • The “Data files and first plots” mini-course in Section 2

  • You should install the program wget:

    $ sudo apt install wget
    

Plan

Our plan is to find some interesting data sets on the web.

In our first approach in Section 29.3 we will download them to our disk using the command line program wget and plot them with gnuplot. Then in Section 29.4 we will show how you can retrieve data in your python program.

Finally in Section 29.5 we will scratch the surface of all the amazing scientific data sets that can be found on the web.

We will try to look at both time history and image data. Time histories are data sets where we look at an interesting quantity as it changes in time.

Examples of time histories include temperature as a function of time (in fact, all sorts of weather and climate data) and stock market prices as a function of time.

Examples of image data include telescope images of the sky and satellite imagery of the earth and of the sun.

29.2. What does a web page look like underneath? (HTML)

To introduce students to the staples of a web page, remember:

  • Not everyone knows what HTML is.

  • Few people have seen HTML.

So we introduce HTML (hypertext markup language) by example first, and then point out what “hypertext” and “markup” mean.

So I type up a quick html page, and the students watch on the projector and type their own. The page I put up is a simple hello page at first, then I add a link.

<html>
    <head>
        <title>A simple web page</title>
    </head>

    <body>
        <h1>Mark's web page</h1>
        <p>This is Mark's web page</p>
        <p>Now a paragraph with some <i>text in italics</i>
           and some <b>text in boldface</b>
        </p>
    </body>
</html>

Save this to a file called, for example, myinfo.html in your home directory and then view it by pointing a web browser to file:///home/MYLOGINNAME/myinfo.html (yes, there are three slashes in the file URL file:///...).

That simple web page lets me explain what I mean by markup: bits of text like <p> and <i> and <head> are not text in the document: they specify how the document should be rendered (for example <b> and <i> specify how the text should look, <p> breaks the text into paragraphs). Some of the tags don’t affect the text at all, but tell us how the document should be understood (for example the metadata tags <html> and <title>).

Then let’s add a hyperlink: a link to the student’s school. My html page now looks like:

Listing 29.2.1 A simple web page with an anchor (hyperlink) element in it.
<html>
    <head>
        <title>A simple web page</title>
    </head>

    <body>
        <h1>Mark's web page</h1>
        <p>This is Mark's web page</p>
        <p>Now a paragraph with some <i>text in italics</i>
           and some <b>text in boldface</b>
        </p>
        <p>Mark went to high school at
           <a href="http://liceoparini.gov.it/">Liceo Parini</a>
        </p>
    </body>
</html>

Then save and reload the page in your browser.

Here I’ve introduced the hyperlink. In HTML this is made up of an element called <a> (anchor) which has an attribute called href which has the URL of the hyperlink.

So as we write programs that pick apart a web page we now know what web pages look like. If we want to find the links in a web page we can use the Python string find() method to look for <a and then for </a> and to use the text in between the two.

29.3. Command line scraping with wget

In Section 7.1 we had our first glimpse of the command wget, a wonderful program which grabs a page from the web and puts the result into a file on your disk. This type of program is sometimes called a “web crawler” or “offline browser”.

wget can even follow links up to a certain depth and reproduce the web hierarchy on a local disk.

In areas with poor network connectivity people can use wget when there is a brief moment of good newtorking: they download all they need in a hurry, then point their browser to the data on their local disk.

29.3.1. First download with wget

Let us make a directory in which to work and start getting data.

$ mkdir scraping
$ cd scraping
$ wget https://raw.githubusercontent.com/fivethirtyeight/data/master/alcohol-consumption/drinks.csv

We now have a file called drinks.csv - how do we explore it?

I would first use simple file tools:

less drinks.csv

shows lines like this:

country,beer_servings,spirit_servings,wine_servings,total_litres_of_pure_alcohol
Afghanistan,0,0,0,0.0
Albania,89,132,54,4.9
Algeria,25,0,14,0.7
Andorra,245,138,312,12.4
Angola,217,57,45,5.9
## ...

If you like to see data in a spreadsheet you could try to use libreoffice or gnumeric:

libreoffice drinks.csv

29.3.2. Simple analysis of the drinks.csv file

Sometimes you can learn quite a bit about what’s in a file with simple shell tools, without using a plotting program or writing a data analysis program. I will show you a some things you can do with one line shell commands.

Looking at drinks.csv we see that the fourth column is the number of wine servings per capita drunk in that country. Let us use the command sort to order the file by wine consumption.

A quick look at the sort documentation with man sort shows us that the -t option can be used to use a comma instead of white space to separate fields. We also find out that the -k option can be used to specify a key and -g to sort numerically (including floating point). Put these together to try running:

sort -t , -k 4 -g drinks.csv

this will show you all those countries in order of increasing wine consumption, rather than in alphabetical order. To see just the last few 15 lines you can run:

sort -t , -k 4 -g drinks.csv | tail -15

This is a great opportunity to laugh at the confirmation of some stereotypes and the negation of others.

If you look at the last few lines you see that the French consume the most wine per capita, followed by the Portuguese.

If you sort by the 5th column you will see the overall use of alcohol and the 3rd column will show you the use of spirits (hard liquor) while the 2nd column shows consumption of beer.

29.3.3. Looking at birth data

$ wget https://raw.githubusercontent.com/fivethirtyeight/data/master/births/US_births_2000-2014_SSA.csv
$ tr '\r' '\n' < US_births_2000-2014_SSA.csv > births_2000-2014_SSA-newline.csv
$ gnuplot
 # and at the gnuplot> prompt:
set datafile separator ","
plot 'births_2000-2014_SSA-newline.csv' using 5 with lines

29.4. Scraping from a Python program

29.4.1. Brief interlude on string manipulation

$ python3
>>> s = 'now is the time for all good folk to come to the aid of the party'
>>> s.split()
['now', 'is', 'the', 'time', 'for', 'all', 'good', 'folk', 'to', 'come', 'to', 'the', 'aid', 'of', 'the', 'party']
# now we've seen what that looks like, save it into a variable
>>> words = s.split()
>>> words
['now', 'is', 'the', 'time', 'for', 'all', 'good', 'folk', 'to', 'come', 'to', 'the', 'aid', 'of', 'the', 'party']
>>>
# now try to split where the separator is a comma
>>> csv_str = 'name,age,height'
>>> words = csv_str.split()
>>> csv_str = 'name,age,height'
>>> words = csv_str.split()
>>> words
['name,age,height']
# didn't work; try telling split() to use a comma
>>> words = csv_str.split(',')
>>> words
['name', 'age', 'height']

29.4.2. The birth data from Python

Listing 29.4.1 get-birth-data.py - A program which downloads birth data.
#! /usr/bin/env python3

import urllib.request

day_map = {1: 'mon', 2: 'tue', 3: 'wed', 4: 'thu', 5: 'fri', 
           6: 'sat', 7: 'sun'}

def main():
    f = urllib.request.urlopen('https://raw.githubusercontent.com/fivethirtyeight/data/master/births/US_births_2000-2014_SSA.csv')
    ## this file has carriage returns instead of newlines, so
    ## f.readlines() won't work in all cases.  I read the whole
    ## file in, and then split it into lines
    entire_file = f.read()
    f.close()
    lines = entire_file.split()
    print('lines:', lines[:3])
    dataset = []
    for line in lines[1:]:
        # print('line:', line, str(line))
        line = line.decode('utf-8')
        words = line.split(',')
        # print(words)
        values = [int(w) for w in words]
        dataset.append(values)
    day_of_week_hist = process_dataset(dataset)
    print_histogram(day_of_week_hist)

def process_dataset(dataset):
    ## NOTE: the fields are:
    ## year,month,date_of_month,day_of_week,births
    print('dataset has %d lines' % len(dataset))
    ## now we form a histogram of births according to the day of the
    ## week
    day_of_week_hist = {}
    for i in range(1, 8):
        day_of_week_hist[i] = 0
    for row in dataset:
        day_of_week = row[3]
        month = row[1]
        n_births = row[4]
        day_of_week_hist[day_of_week] += n_births
    return day_of_week_hist

def print_histogram(hist):
    print(hist)
    keys = list(hist.keys())
    keys.sort()
    print('keys:', keys)
    for day in keys:
        print(day, day_map[day], hist[day])

main()

29.5. Finding neat scientific data sets

https://www.dataquest.io/blog/free-datasets-for-projects/ (they mention fivethirtyeight)

https://github.com/fivethirtyeight/data

29.5.1. Time histories

Temperature

Births

wget https://raw.githubusercontent.com/fivethirtyeight/data/master/births/US_births_2000-2014_SSA.csv

29.5.2. Images

NASA nebulae

Goes images of the sun

29.6. Beautiful Soup

Beautiful Soup is a powerful python package that allows you to scrape web pages in a structured manner. Unlike the code we have seen so far, which does brute-force parsing of html text chunks in Python, beautiful soup is aware of the “document object model” (DOM).

Start by installing the python package. You can probably install with pip, or on debian-based distributions you can run:

sudo apt install python3-bs4

Now enter the program billboard_hot_100_scraper_2023.py in Listing 29.6.1:

Listing 29.6.1 billboard_hot_100_scraper_2023.py - the Billboard Hot 100 list using Beautiful Soup form the web site https://www.billboard.com/charts/hot-100
#! /usr/bin/env python3

"""This program was inspired by Jaimes Subroto who had written a
program that worked with the 2018 billboard html format.  Billboard
has changed its html format quite completely in 2023, so this is a
re-implementation that handles the new format.
"""

import urllib.request
from bs4 import BeautifulSoup as soup

def main():
    url = 'https://www.billboard.com/charts/hot-100'
    # url = 'https://web.archive.org/web/20180415100832/https://www.billboard.com/charts/hot-100/'

    # boiler plate stuff to load in an html page from its URL
    url_client = urllib.request.urlopen(url)
    page_html = url_client.read()
    url_client.close()

    # let us save it to a local html file, using utf-8 decoding so
    # that we turn the byte stream into simple ascii text
    open('page_saved.html', 'w').write(page_html.decode('utf-8'))

    # boiler plate use of beautiful soup: use the html parser on the file
    page_soup = soup(page_html, "html.parser")

    # now for the part where you need to know the structure of the
    # html file.  by inspection I found that in 2023 they use <ul>
    # list elements with the attribute "o-chart-restults-list-row", so
    # this is how you find those elements in beautiful soup:
    list_elements = page_soup.select('ul[class*=o-chart-results-list-row]') # *= means contains
    # now that we have our list are ready to read things in, we also prepare 
    outfname = 'billboard_hot_100.csv'
    with open(outfname, 'w') as fp:
        headers = 'Song, Artist, Last Week, Peak Position, Weeks on Chart\n'
        fp.write(headers)
        # Loops through each list element
        for element in list_elements:
            handle_single_row(element, fp)
    print(f'\nBillboard hot 100 table saved to {outfname}')

def handle_single_row(element, fp):
    all_list_items = element.find_all('li')
    title_and_artist = all_list_items[4]
    # try to separate out the title and artist.  title should be an
    # <h3> element, artist is a <span> element
    title = title_and_artist.find('h3').text.strip()
    artist = title_and_artist.find('span').text.strip()
    # now the rest of the columns
    last_week = all_list_items[7].text.strip()
    peak_pos = all_list_items[8].text.strip()
    weeks_on_chart = all_list_items[9].text.strip()
    # we have enough to write an entry in the csv file
    csv_line = f'"{title}", "{artist}", {last_week}, {peak_pos}, {weeks_on_chart}'
    print(csv_line)
    fp.write(csv_line + '\n')


if __name__ == '__main__':
    main()

If you run:

$ chmod +x billboard_hot_100_scraper_2023.py
$ ./billboard_hot_100_scraper_2023.py

The results can be seen in the CSV file billboard_hot_100.csv:

Table 29.6.1 Billboard Hot 100

Song

Artist

Last Week

Peak Position

Weeks on Chart

Choosin’ Texas

Ella Langley

1

1

24

Boston

Stella Lefty

2

2

25

Been By Now

Morgan Wallen

3

2

9

Hate That I Made You Love Me

Ariana Grande

4

1

1

Dracula

Tame Impala & JENNIE

5

5

52

BbY WOW

Karol G With Judeline & rusowsky

7

6

7

So Easy (To Fall In Love)

Olivia Dean

8

5

52

Man I Need

Olivia Dean

9

2

57

Last Thing You Need

Morgan Wallen

9

1

Be Her

Ella Langley

10

2

32

Nicole Kidman

ADELA

22

11

3

I Knew It, I Knew You

Taylor Swift

6

1

2

Stupid Song

Olivia Rodrigo

12

3

15

I Just Might

Bruno Mars

13

1

3

Risk It All

Bruno Mars

11

4

30

Midnight Sun

Zara Larsson

14

13

36

I Can’t Love You Anymore

Ella Langley & Morgan Wallen

15

4

22

Babydoll

Dominic Fike

19

16

33

Janice STFU

Drake

18

1

2

Earrings

Malcolm Todd

20

19

30

Drop Dead

Olivia Rodrigo

16

1

1

Be By You

Luke Combs

17

12

32

Dead Fresh

Lil Baby

21

18

10

Loving Life Again

Ella Langley

24

21

27

Jaded

Koe Wetzel & Ella Langley

25

19

7

Loser

Tame Impala

23

18

9

Ain’t In LA

ADELA

30

27

8

The Cure

Olivia Rodrigo

27

5

18

What You Need

Tems

29

29

27

Animal

KATSEYE

28

24

9

Think As You Drunk

Riley Green

58

31

13

Oh Yeah?

Steve Lacy

31

19

8

Mr. Know It All

Teddy Swims

38

33

23

Orbiter

Noah Kahan

33

33

22

Cinderella

Mac Miller Featuring Ty Dolla $ign

35

25

20

Take Me Back (Leave Me There)

Cody Johnson

36

33

6

Rethink Some Things

Luke Combs

39

37

26

Dai Dai (FIFA World Cup Official Song 2026)

Shakira X Burna Boy

32

17

15

My Body Isn’t Ready

sombr

55

30

13

AH HA

Cardi B

43

25

8

Self Aware

Temper City

44

35

24

Lush Life

Zara Larsson

35

27

Spend Dat

Yung Miami

34

17

18

So Good

Jhene Aiko Featuring Kendrick Lamar

26

26

2

September

Earth, Wind & Fire

8

18

Shabang

Drake

40

4

19

Morning Dew (Donk)

Beyonce

37

26

12

Bloodstream

Alyssa Grace

45

45

12

Petal

Ariana Grande

46

4

8

Freakin’ Out

Dexter And The Moonrocks

47

33

27

Carry On

Kenny Chesney

56

51

14

Bass Persuades

Miley

63

35

3

WTF GOIN

Belly Gang Kushington & 21 Savage

53

53

5

Phone, Keys, Wallet

Lainey Wilson & John Mayer

48

48

16

Miss My Dawg

Yeat & Drake

55

1

Rein Me In

Sam Fender & Olivia Dean

51

45

27

Something To Lose

Stella Lefty With Vincent Mason

50

40

17

Great Expectation

Sienna Spiro

57

52

12

Hit The Wall

Gracie Abrams

52

26

19

My Way

Riley Green

60

10

Kid Myself

John Morgan

62

61

6

Sue Me

Audrey Hobert

64

57

9

Cowgirl

Shaboozey

49

42

12

Ghetto Love Story

BabyChiefDoit

67

44

12

That’s Just Me

Riley Green

65

1

Empty Words

Corey Kent

65

65

15

Mexico Honey

Kacey Musgraves

61

39

17

onsra

comehelpglo

92

68

2

South Of Sanity

Zach Top

71

69

11

McArthur

HARDY, Eric Church, Morgan Wallen & Tim McGraw

70

31

34

Noble

F3miii

60

51

20

All My Exes

Lauren Alaina Featuring Chase Matthew

76

72

5

The Feeling

Steve Lacy

78

67

10

Chevy Silverado

Bailey Zimmerman

66

41

14

Another Drink

Marshmello & Kelsea Ballerini

68

60

8

Country And She Knows It

Luke Bryan

85

76

3

Stop The Wedding!

Ashe

84

77

3

String By

Mack Geiger

75

71

6

Painted You Pretty

Hudson Westbrook

83

79

5

Get To Drinkin’

Zach John King

87

80

3

Maggots For Brains

Olivia Rodrigo

77

12

15

Joseph

Falling in Reverse, Corey Taylor & Serj Tankian

82

1

Finders Keepers

Luke Bryan & Luke Combs

83

1

RHYNO

Travis Scott

84

1

Honeybee

Olivia Rodrigo

81

9

15

Ahi

Karol G With Drake

88

44

7

Hands Up

Jelly Roll

86

82

6

God I’m Just Grateful

Elevation Worship & Chandler Moore

90

88

3

Say So

Dan + Shay

98

89

3

He Belongs

Jhene Aiko

41

41

2

Still

Karol G With Bruno Mars

79

41

7

Babydoll

Jamie Miller

92

1

Duvet

boa

93

1

Nothing

Steve Lacy

94

1

Dope Girl

Rod Wave

70

3

Ghost

Jhene Aiko

42

42

2

Bet On That

Blake Whiten

94

84

5

Expectations

Olivia Rodrigo

80

20

15

Let’s Get Married

Miley Cyrus

99

1

Kingdom Of Fear

Cameron Whitcomb

94

7