<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Think Stats on Amit Rajan</title><link>https://amitrajan012.github.io/topics/think-stats/</link><description>Recent content in Think Stats on Amit Rajan</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 10 Sep 2018 11:13:31 +0100</lastBuildDate><atom:link href="https://amitrajan012.github.io/topics/think-stats/index.xml" rel="self" type="application/rss+xml"/><item><title>Correlation</title><link>https://amitrajan012.github.io/post/chapter-9-correlation/</link><pubDate>Mon, 10 Sep 2018 11:13:31 +0100</pubDate><guid>https://amitrajan012.github.io/post/chapter-9-correlation/</guid><description>&lt;h3 id="91-standard-scores"&gt;9.1 Standard scores&lt;/h3&gt;&#10;&lt;p&gt;The main challenge in measuring correlation is that the variables we want to compare might not be expressed in the same units. There are two common solutions to this problem:&lt;/p&gt;&#10;&lt;ul&gt;&#10;&lt;li&gt;Transform all values to &lt;b&gt;standard scores&lt;/b&gt;. This leads to &lt;b&gt;Pearson coefficient of correlation&lt;/b&gt;.&lt;/li&gt;&#10;&lt;li&gt;Transform all values to their &lt;b&gt;percentile ranks&lt;/b&gt;. This leads to &lt;b&gt;Spearman coefficient&lt;/b&gt;.&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;&lt;b&gt;Normalizing&lt;/b&gt; the score means subtracting mean from every value and dividing it by standard deviation.&lt;/p&gt;</description></item><item><title>Estimation</title><link>https://amitrajan012.github.io/post/chapter-8-estimation/</link><pubDate>Tue, 04 Sep 2018 14:07:28 +0100</pubDate><guid>https://amitrajan012.github.io/post/chapter-8-estimation/</guid><description>&lt;h3 id="81-the-estimation-game"&gt;8.1 The estimation game&lt;/h3&gt;&#10;&lt;p&gt;In a rudimentary way, &lt;b&gt;sample mean&lt;/b&gt; can be used to estimate a distribution. The process is called &lt;b&gt;estimation&lt;/b&gt; and the statistic we used (sample mean) is called an &lt;b&gt;estimator&lt;/b&gt;. If there are no &lt;b&gt;outliers&lt;/b&gt;, sample mean minimizes the &lt;b&gt;mean squared error (MSE)&lt;/b&gt;. A &lt;b&gt;maximum likelihood estimator (MLE)&lt;/b&gt; is an estimator that has the highest chance of being right (value with highest probability).&lt;/p&gt;&#10;&lt;/br&gt;&#10;&lt;p&gt;&lt;b&gt;Exercise 8.1:&lt;/b&gt; Write a function that draws 6 values froma normal distribution with \(\mu\) = 0 and \(\sigma\) = 1. Use the sample mean to estimate m and compute the error \(\bar{x} - \mu\). Run the function 1000 times and compute MSE.&lt;/p&gt;</description></item><item><title>Hypothesis Testing</title><link>https://amitrajan012.github.io/post/chapter-7-hypothesis-testing/</link><pubDate>Thu, 30 Aug 2018 06:17:39 +0100</pubDate><guid>https://amitrajan012.github.io/post/chapter-7-hypothesis-testing/</guid><description>&lt;/br&gt;&#10;The process of &lt;b&gt;Hypothesis Testing&lt;/b&gt; can be summarized by following three steps:&#10;&lt;ul&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;b&gt;Null Hypothesis:&lt;/b&gt; The null hypothesis is a mode of the system based on the assumption that apparent effect was actually due to chance.&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;b&gt;p-value:&lt;/b&gt; The p-value is the probability of the apparent effect under the null hypothesis.&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;li&gt;&#10;&lt;p&gt;&lt;b&gt;Interpretation:&lt;/b&gt; Based on the p-value, we can conclude that whether the effect is statistically significant or not.&lt;/p&gt;&#10;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;p&gt;Hence to test a hypothesis, we assume that it is not real (called as &lt;b&gt;null hypothesis&lt;/b&gt;). Based on the assumption (i.e. null hypothesis), we compute the probability of the apparent effect (this is called &lt;b&gt;p-value&lt;/b&gt;). If p-value is low enough, we can conclude that the null hypothesis is unlikely to be true.&lt;/p&gt;</description></item><item><title>Operations on Distributions</title><link>https://amitrajan012.github.io/post/chapter-6-operations-on-distributions/</link><pubDate>Sun, 26 Aug 2018 16:27:09 +0100</pubDate><guid>https://amitrajan012.github.io/post/chapter-6-operations-on-distributions/</guid><description>&lt;h3 id="61-skewness"&gt;6.1 Skewness&lt;/h3&gt;&#10;&lt;p&gt;&lt;b&gt;Skewness&lt;/b&gt; is a statistic that measures the assymetry of a distribution. Sample skewness is defined as:&lt;/p&gt;&#10;\[Skewness = \frac{Mean\ Cubed\ Deviation}{Mean\ Squared\ Deviation^{\frac{3}{2}}}\]&lt;p&gt;Negative skewness means that the distribution skews left and positive skewness means that the distribution is skewed to right. &lt;b&gt;Outliers&lt;/b&gt; have a disproportionate effect on skewness.&lt;/p&gt;&#10;&lt;p&gt;Another way to evaluate the asymmetry of a distribution is to compare mean and median. For distributions that skews left, mean is less than the median. This statistic is denoted by &lt;b&gt;Pearson&amp;rsquo;s Median Skewness Coefficient&lt;/b&gt; which is given as:&lt;/p&gt;</description></item><item><title>Probability</title><link>https://amitrajan012.github.io/post/chapter-5-probability/</link><pubDate>Wed, 22 Aug 2018 09:12:36 +0100</pubDate><guid>https://amitrajan012.github.io/post/chapter-5-probability/</guid><description>&lt;p&gt;The belief in which probability is expressed in terms of frequencies is called as &lt;b&gt;frequentism&lt;/b&gt;. Frequentists believe that if there is no set of identical trials, there is no probability.&lt;/p&gt;&#10;&lt;p&gt;An alternative is &lt;b&gt;Bayesianism&lt;/b&gt;, which defines probability as a degree of belief that an event will occur. By this definition, the notion of probability can be applied in almost any cicumstance. One drawback of Bayesian probability is that it depends on person&amp;rsquo;s state of knowledge. People with different information might have different degree of belief about the same event. This is the reason that many people think that Bayesian probability is more subjective than frequentists.&lt;/p&gt;</description></item><item><title>Continuous Distributions</title><link>https://amitrajan012.github.io/post/chapter-4-continuous-distributions/</link><pubDate>Fri, 17 Aug 2018 19:02:16 +0100</pubDate><guid>https://amitrajan012.github.io/post/chapter-4-continuous-distributions/</guid><description>&lt;h3 id="41-the-exponential-distribution"&gt;4.1 The exponential distribution&lt;/h3&gt;&#10;&lt;p&gt;CDF of the exponential distribution is defined as (where lambda determines the shape of the distribution):&#10;&lt;/p&gt;&#10;\[CDF(x) = 1-e^{-\lambda x}\]&lt;p&gt;&#10;Mean and Median of exponential distribution is computed as follows:&#10;&lt;/p&gt;&#10;\[Mean = \frac{1}{\lambda}\]&lt;p&gt;&#10;&lt;/p&gt;&#10;\[Median = \frac{ln(2)}{\lambda}\]&lt;p&gt;&#10;In the real world, exponential distribution is encountered when we look at the series of events and measure the time between them which is called as &lt;b&gt;interval times&lt;/b&gt;. The exponential distribution for different values of lambda is shown below. It can be seen that lambda decides the shape of the distribution.&lt;/p&gt;</description></item><item><title>Cumulative Distribution Functions</title><link>https://amitrajan012.github.io/post/chapter-3-cumulative-distribution-functions/</link><pubDate>Sun, 12 Aug 2018 09:22:46 +0100</pubDate><guid>https://amitrajan012.github.io/post/chapter-3-cumulative-distribution-functions/</guid><description>&lt;h3 id="31-the-class-size-paradox"&gt;3.1 The class size paradox&lt;/h3&gt;&#10;&lt;p&gt;For a probability distribution, the mean calculated from its PMF is lower than the one calculated by taking a sample from it. This happens because the larger classes tend to get oversampled.&#10;&lt;br&gt;&lt;br&gt;&lt;/p&gt;&#10;&lt;p&gt;&lt;b&gt;Exercise 3.1&lt;/b&gt; Build the PMF of the college class-size data and compute the mean as perceived by the Dean. Find the distribution of class sizes as perceived by students and compute&#10;its mean.&lt;/p&gt;</description></item><item><title>Descriptive Statistics</title><link>https://amitrajan012.github.io/post/chapter-2-descriptive-statistics/</link><pubDate>Wed, 08 Aug 2018 04:17:12 +0100</pubDate><guid>https://amitrajan012.github.io/post/chapter-2-descriptive-statistics/</guid><description>&lt;h3 id="21-means-and-averages"&gt;2.1 Means and averages&lt;/h3&gt;&#10;&lt;p&gt;&lt;b&gt;Mean&lt;/b&gt; of a sample is a summary statistics that can be computed as (where n is the total number of samples):&#10;&lt;/p&gt;&#10;\[\mu = \frac{1}{n} \sum_i{x_i}\]&lt;p&gt;&#10;An &lt;b&gt;average&lt;/b&gt; is one of many summary statistics that can be used to describe the typical value or the &lt;b&gt;central tendency&lt;/b&gt; of a sample.&lt;/p&gt;&#10;&lt;h3 id="22-variance"&gt;2.2 Variance&lt;/h3&gt;&#10;&lt;p&gt;As mean describes the central tendency of a sample, &lt;b&gt;Variance&lt;/b&gt; is intended to describe the &lt;b&gt;spread&lt;/b&gt;. The variance is defined as:&#10;&lt;/p&gt;</description></item><item><title>Statistical Thinking for Programmers</title><link>https://amitrajan012.github.io/post/chapter-1-statistical-thinking-for-programmers/</link><pubDate>Sun, 05 Aug 2018 10:02:08 +0100</pubDate><guid>https://amitrajan012.github.io/post/chapter-1-statistical-thinking-for-programmers/</guid><description>&lt;h3 id="11-do-first-babies-arrive-late"&gt;1.1 Do first babies arrive late?&lt;/h3&gt;&#10;&lt;p&gt;&lt;b&gt;Anecdotal Evidence &lt;/b&gt; is based on data that is unpublished and usually personal. For example,&#10;&lt;i&gt;&lt;center&gt;&amp;ldquo;My two friends that have given birth recently to their first babies,&#10;BOTH went almost 2 weeks overdue before going into&#10;labour or being induced.” &lt;/center&gt;&lt;/i&gt;&#10;Anecdotal Evidence usually fail because of &lt;b&gt;Small number of observations&lt;/b&gt;, &lt;b&gt;Selection bias&lt;/b&gt; (People who join a discussion of this question might be interested because their first babies were late.), &lt;b&gt;Confirmation bias&lt;/b&gt; (People who believe the claim might be more likely to contribute examples that confirm it) and &lt;b&gt;Inaccuracy&lt;/b&gt;.&lt;/p&gt;</description></item></channel></rss>