Think Stats · Chapter 1 · · 4 min read

Statistical Thinking for Programmers

1.1 Do first babies arrive late?

Anecdotal Evidence is based on data that is unpublished and usually personal. For example,

“My two friends that have given birth recently to their first babies, BOTH went almost 2 weeks overdue before going into labour or being induced.”
Anecdotal Evidence usually fail because of Small number of observations, Selection bias (People who join a discussion of this question might be interested because their first babies were late.), Confirmation bias (People who believe the claim might be more likely to contribute examples that confirm it) and Inaccuracy.

1.2 A Statistical Approach

Limitations of Anecdotal Evidence can be addressed by using the tools of statistics, which include Data Collection, Descriptive Statistics, Exploratory Data Analysis, Hypothesis Testing and Estimation.

1.3 The National Survey of Family Growth (NSFG)

NSFG is a cross-sectional study (it captures a snapshot of a group at a point in time). The alternative is a longitudinal study which observes a group repeatedly over a period of time. The people who participate in a survey are called respondents. Cross-sectional studies are meant to be representative, which means that every member of the target population has an equal chance of participating. NSFG is deliberately oversampled as certain groups are sampled at higher rates compared to their representation in US population. Drawback of oversampling is that it is hard to arrive at a conclusion based on statistics from the survey.

Exercise 1.2 Download data from the NSFG:

import pandas as pd
# Reference to extract the columns: http://greenteapress.com/thinkstats/survey.py
pregnancies = pd.read_fwf("2002FemPreg.dat",
                         names=["caseid", "nbrnaliv", "babysex", "birthwgt_lb",
                               "birthwgt_oz", "prglength", "outcome", "birthord",
                               "agepreg", "finalwgt"],
                         colspecs=[(0, 12), (21, 22), (55, 56), (57, 58), (58, 60),
                                (274, 276), (276, 277), (278, 279), (283, 285), (422, 439)])
pregnancies.head()

caseidnbrnalivbabysexbirthwgt_lbbirthwgt_ozprglengthoutcomebirthordagepregfinalwgt
011.01.08.013.03911.033.06448.271112
111.02.07.014.03912.039.06448.271112
223.01.09.02.03911.014.012999.542264
321.02.07.00.03912.017.012999.542264
421.02.06.03.03913.018.012999.542264

The description for the fields are as follows:

caseidprglengthoutcomebirthordfinalwgt
Integer ID of RespondentInteger Duration of pregnancy in weeks1 indicates a live birthcode for first child: 1Number of people in US population this respondant represents

Exercise 1.3 Explore the data in the Pregnancies table. Count the number of live births and compute the average pregnancy length (in weeks) for first babies and others for the live births.

pregnancies.describe()

caseidnbrnalivbabysexbirthwgt_lbbirthwgt_ozprglengthoutcomebirthordagepregfinalwgt
count13593.0000009148.0000009144.0000009144.0000009087.00000013593.00000013593.0000009148.00000013241.00000013593.000000
mean6216.5265951.0259071.4945326.6537627.40387429.5312291.7639961.82455224.2309498196.422280
std3645.4173410.2528640.5152951.5888098.09745413.8025231.3159301.0370535.8243029325.918114
min1.0000001.0000001.0000000.0000000.0000000.0000001.0000000.00000010.000000118.656790
25%3022.0000001.0000001.0000006.0000003.00000013.0000001.0000001.00000020.0000003841.375308
50%6161.0000001.0000001.0000007.0000007.00000039.0000001.0000002.00000023.0000006256.592133
75%9423.0000001.0000002.0000008.00000011.00000039.0000002.0000002.00000028.0000009432.360931
max12571.0000009.0000009.0000009.00000099.00000050.0000006.0000009.00000044.000000261879.953864
live_births = pregnancies[pregnancies['outcome'] == 1]
print("Number of live births is: " + str(live_births.shape[0]))
mean_first = live_births[live_births['birthord'] == 1]['prglength'].mean()
mean_other = live_births[live_births['birthord'] != 1]['prglength'].mean()
print("Mean Pregnancy length for live births of first babies is: " + str(mean_first))
print("Mean Pregnancy length for live births of other babies is: " + str(mean_other))
print("Difference in Mean Pregnancy length for first and other babies is : " + str(mean_first - mean_other))
Number of live births is: 9148
Mean Pregnancy length for live births of first babies is: 38.6009517335
Mean Pregnancy length for live births of other babies is: 38.5229144667
Difference in Mean Pregnancy length for first and other babies is : 0.0780372667775

1.5 Significance

From the above analysis, it is evident that the difference in mean pregnancy lengths of first and other babies is 13.11 hours. A difference like this is called an apparent effect which means that there must be something going on but we are not sure yet. If the difference occurred by chance, we can conlcude that thet effect was not statistically significant. An apparent effect that is caused by bias, measurement error, or some other kind of error is called artifact.