Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Sunday, October 26, 2008

The normal distribution

This beautiful image comes courtesy of W. J. Youden.

I've been reading Edward Tufte's superb The Visual Display of Quantitative Information - perhaps the most perfect book I've ever come across. Almost every page is a revelation, so expect me to be posting more of the wonderful graphs Tufte has collected in the future.

Monday, October 20, 2008

The importance of being clear

A formatting fubar involving an Excel spreadsheet has left Barclays Capital with contracts involving collapsed investment bank Lehman Brothers than it never meant to acquire.

Working to a tight deadline, a junior law associate at Cleary Gottlieb Steen & Hamilton LLP converted an Excel file into a PDF format document. The doc was to be posted on a bankruptcy court's website before a midnight purchase offer deadline on 18 September, just four hours after Barclays sent the spreadsheet to the lawyers. The Excel file contained 1,000 rows of data and 24,000 cells.

Some of these details on various trading contracts were marked as hidden because they were not intended to form part of Barclays' proposed deal. However, this "hidden" distinction was ignored during the reformatting process so that Barclays ended up offering to take on an additional 179 contracts as part of its bankruptcy buyout deal, Finextra reports.


The Register has the full story. As Merv has always warned, 'horrible things happen when you hide cells in excel'.

I see this as a manifestation of a wider lack of education on the importance of communicating information efficiently. The Spartans, Tufte, Strunk and White, the Economist, Picasso and numerous econometricians have done a lot to improve things, but management-speak, TV advertising and other such phenomena show we still have a long way to go.

Monday, September 29, 2008

Stata lessons & other resources

A friend asked for a quick list, so here goes:

UCLA's excellent resources to help you learn and use Stata

Another great collection of Stata Resources by Park Hun Myoung, as well as a stata command cheat-card

London School of Economics Stata Resources

Syracuse University's Stata tutorial

Program in Statistics and Methodology by the pol sci's at Ohio State University

Duke Stata tutorials

Tuesday, September 9, 2008

A lesson in statistics

Andrew Gelman:

[...] when he talks with people about statistical procedures, engineers focus on the algorithm being applied to the data, whereas statisticians are always thinking about the psychology of the person doing the analysis.

This is a big topic for another day - the day I start posting on my main area of expertise, econometrics - but these are words worth pondering.

Tuesday, September 2, 2008

Fun with statistics, "which Palin is the mother?" edition

Well, Sarah, I'm calling you a liar. And not even a good one. Trig Paxson Van Palin is not your son. He is your grandson.

The Daily Kos has an 'interesting' story claiming that Sarah Palin's youngest son, who has been diagnosed with Down's syndrome, is really her daughter's son. There is much evidence presented, mainly consisting of pictures where certain bellies appear to be too large while others too small. But statistics deliver the killer argument:

The final point of interest is that Trig Palin has been diagnosed with Down's syndrome (aka trisomy 21). This is an interesting point, as chances of having offspring with Down's Syndrome increases from under 1% to 3% after a mother reaches the age of 40. However, 80% of the cases of Down's Syndrome are in mother's (sic) under the age of 35, through sheer quantities of births in this age group.

And of course, 99% of deaths are not due to suicide, so killing yourself is safe!

Thanks to Scatterplot for the pointer.

Tuesday, April 1, 2008

I hate seasonally adjusted series

To all data providers, wherever they may be:

Please, please stop seasonally adjusting the series you give me. I can run a regression with seasonal dummies myself, thank you very much. So can everyone else. The difficult thing is to get from the seasonally adjusted series to the original one, and I don't see why you should deprive me of the privilege.

Yes, I know I shouldn't want to in most cases, but let me be the judge of that, OK? What's your problem anyway, what is it to you? Please, just do me this favour.

Thursday, March 13, 2008

What are the odds?

You are on holiday in some strange land, and you bump into Sue, an old friend. What are the odds?

Well, the probability is 1. 100%. It's certain. It bloody happened.

OK, you say, fair point - but that's not what you meant. What you meant is what is the probability you would bump into Sue, assuming you hadn't just bumped into him. I reply that that's a silly assumption to make as you just did bump into him, but you insist.

Well then, the probability is whatever you want it to be - pick a number, and I'll explain to you why it's plausible. You look perplexed and ask what I mean.

I explain: first of all, you need to specify the probability of what you are interested in.
-Do you want to know the probability you'd bump into Sue at the time and place you did
-Do you want the probability you would bump into an old friend while on holiday in general
-Or do you want the probability that something 'remarkable' enough would happen to you at some point in your life that would make you start asking silly questions about 'what is the probability of that happening'?

You say the first, obviously, and that I should stop being clever. I wasn't done, I say, and proceed to ask what is the information set I should base my probability estimate at - quantum mechanics aside, randomness is in the eye of the beholder after all, and if I was all-knowing God the probability of whatever it is that happened would be 1 even before it happened.

You throw your pina colada on my head and vow never to speak to me again.

A few days later, you are kind of missing me but don't feel like talking to me yet, so you visit bluematter. as a first step in rebuilding the relationship. And the first thing you see is this delightful little story, via Andrew Gelman:


In the city of Syracuse, the strangest thing happened in Tuesday's Democratic presidential primary.

Sen. Hillary Clinton and Sen. Barack Obama received the exact same number of votes, according to unofficial Board of Election results.

Clinton: 6,001.

Obama: 6,001.

The odds of Clinton and Obama tying were less than one in 1 million, said Syracuse University mathematics Professor Hyune-Ju Kim.

Elaborating on Thursday, she [Professor Hyune-Ju Kim] noted: "The "almost impossible" odd is obtained when we assume the Syracuse voter distribution follows the New York state distribution. Since it is almost impossible to observe what we have observed, statistically we can conclude that Syracuse voter distribution is significantly different from the New York state distribution."

There would be less than one in 1 million chance of a tie occurring between Clinton and Obama in voting by a randomly selected group of 12,346 New York Democratic voters, she said.


To which Andrew replies:

Not to pick on some harried mathematics professor who'd probably rather be out proving theorems, but . . . of course Syracuse voters are not a randomly selected group of New Yorkers. You don't need a statistical test to see that. Regarding the probability of an exact tie: I don't think that's so low: a quick calculation might say that either Clinton or Obama could've received between, say, 5000 and 7000 votes, giving something like a 1/2000 chance of an exact tie. That's gotta be the right order of magnitude.

If there was one thing you were ever certain about, it is that you don't want to read what I have to say on this. A baseball bat happens to lie next to you (what are the odds!?). You grab it with both your shaky hands and smash the computer monitor to pieces.

Wednesday, March 12, 2008

Roses are red, violets are blue, and correlation is not causation

Merv emails me this article:

Football clubs with red team strips are more successful than those with other colours, according to a study released Wednesday.

The fact that English clubs Manchester United, Liverpool and Arsenal regularly top league tables is not a coincidence, say the experts from Durham University and the University of Plymouth.

Red shirts give the team an advantage due to deep-rooted biological response to the colour. "In nature, red is often associated with male aggression and display," they said, giving the example of the red-breasted robin.

"It is a testosterone-driven signal of male quality, and its striking effect has even been harnessed by soldiers in the past," added the researchers, after analyzing data on English football league results since World War II.


The red-breasted robin? That's the most fearsome red beast they can think of?

I can't access the paper, but just reading the abstract is enough to convince me it's bonkers. The authors also have a 2005 paper - in Nature no less - entitled 'Red enhances human performance in contests'. If any reader has a access to Nature, would you be kind enough to email me the article so I can -ahem- review it?

Wednesday, February 20, 2008

Planning an informed jump off a bridge

Zubin Jelveh has the inside scoop on where to go to avoid the crowds:

Monday, February 11, 2008

Econometric causality

James Heckman has an excellent paper on the subject (free access).

Thursday, December 6, 2007

Stop abusing statistical significance

I just made my first edit on Wikipedia, on the article on 'statistical power'. Here's the old text, with the deleted parts in bold:

There are times when the recommendations of power analysis regarding sample size will be inadequate. Power analysis is appropriate when the concern is with the correct acceptance or rejection of a null hypothesis. In many contexts, the issue is less about determining if there is or is not a difference but rather with getting a more refined estimate of the population effect size. For example, if we were expecting a population correlation between intelligence and job performance of around .50, a sample size of 20 will give us approximately 80% power (alpha = .05, two-tail). However, in doing this study we are probably more interested in knowing whether the correlation is .30 or .60 or .50. In this context we would need a much larger sample size in order to reduce the confidence interval of our estimate to a range that is acceptable for our purposes. These and other considerations often result in the true but somewhat simplistic recommendation that when it comes to sample size, "More is better!"

However, huge sample sizes can lead to statistical tests becoming so powerful that the null hypothesis is always rejected for real data. This is a problem in studies of differential item functioning.


Leaving the cost of collecting data aside, larger (appropriately collected) samples are ALWAYS BETTER. At the end of the day, if your sample is *too* large (for example if your statistical software restricts the amount of information you can load on it and you don't need the extra information anyways) you can always obtain a smaller random sample from your larger random sample. So, the 'more is better' recommendation is simple, but not simplistic.

The last paragraph reveals a fundamental misconception about statistical significance that refuses to go away. If the effect of an independent variable on the dependent variable is zero, using a very large sample will result to an estimated effect that is 0 to many decimal places; as the sample size increases further, the effect will approach *exactly* zero even more. NEVER USE STATISTICAL SIGNIFICANCE AS A PROXY FOR PRACTICAL SIGNIFICANCE. I have no clue whether large sample sizes have been seen as a problem in the past in studies of differential item functioning, but if that is the case then the researchers are idiots.

Here is another post on problematic applications of statistical significance.

Sunday, December 2, 2007

Race and brains, once more unto the breach

If the previous post did not persuade you to stop wasting your time thinking about it, McMegan links to Jim Manzi who addresses the question. He concludes his well researched piece with this:

Do genetic differences accounts for any material portion of the difference in IQ scores by self-identified racial groups in the US? The only honest answer is that we don’t know. This, not political correctness is why the American Psychological Association’s formal consensus point of view on this question is stated without qualification: “At present, this question has no scientific answer.”

All right and proper, but that's not the right question to ask. What you really want to know is this: are 'black genes' leading to materially less intelligence than 'white genes'? And the answer is simple: IQ tests can't tell you that.

My understanding is that IQ scores say nothing about 'absolute' intelligence, they only provide a ranking. Or to use terminology more familiar to some of my readers, IQ scores only have an ordinal, not a cardinal meaning. It is very well likely that someone scoring 110 is only trivially more intelligent than someone scoring 90; what the difference between 110 and 90 actually means in terms of 'amount of intelligence' is anyone's guess.

OK, I hear you say, but don't we use quasi-cardinal interpretations for IQ scores? (e.g. isn't 'normal intelligence' supposed to lie between 90 and 109?) Quoting Manzi again:

There are statistically significant differences in IQ test performance between self-identified racial and ethnic groups in the US, and these differences have been sustained over long periods of time. The specific difference that is most widely discussed is the fact that in the US Non-Hispanic whites score, on average, about 15 points (~1 STDEV) higher than African-Americans. (Leaving aside the complication that it matters exactly how we define “long periods of time”, since, for example, there is circumstantial evidence that the black-white IQ gap may have been reduced substantially over the past several decades.)

So, 15 points is the maximum possible difference between the races. We also know with certainty that environment plays a role in determining intelligence, so the difference that can be attributed to genetics is a maximum of 10 points or so, with the actual difference (if it exists) likely to be much smaller. That's nothing: let me remind you that the average (and I think also median) person scores 100 by design, that 'average' or 'normal' intelligence is a 20 point band, and that in any case IQ is a flawed measure of 'intelligence' as used in everyday language (and a 'better than random' - but not by much - predictor of 'success' in life).

To bring the human element into this, my own results from several IQ tests are uniformly distributed across a 35 points range, and while I'm pretty good at arriving to answers to almost every IQ test type question thrown at me - questions that the average person won't answer at all - I take more time than the average person to do so (does that make me more or less intelligent?). And since the human element sells, I got more for you: here is a long list of highly successful and intelligent people who were very likely autistic, and here's the corresponding long list of people with dyslexia. What's the point? Intelligence is not a uni-dimensional variable.

And just to make sure, I should also mention the Flynn effect, quoting from this (excellent) paper (free access):

Since 1932 and probably prior to that, test scores have been increasing at a rate of 3 to 6 IQ points per decade, depending on the IQ test used. The preponderance of evidence indicates that scores are continuing to rise at a constant rate (Flynn, 2006b).

And since you apparently have to be a genius to read Bluematter., I won't even draw out the implications of the following observation on the likely (non)persistence of any currently observed differences amongst races (ala Manzi's circumstantial evidence):

There is at least one exception, however. The Scandinavian countries currently are showing little or no rise in their test scores (Flynn, 2006a). As large IQ increases were seen in Norway prior to 1968, Flynn suggests that Scandinavia might have experienced early increases that have since abated. This raises the possibility that IQ increases in other industrialized nations will also end.

So in IQ we have a metric with no cardinal interpretation, with a weak correlation to 'general intelligence' and an even weaker one to 'success in life', with the observed differences between races being pretty small and most likely diminishing even before controlling for environmental characteristics.

What's the issue again?

Saturday, December 1, 2007

Assume blacks have lower IQ than whites and DNA is to blame

Now name one thing you would do differently - as a politician, as a citizen or as a human being. I'll be damned if you can come up with a single example.

I really, really can't understand what this debate is all about (other than in an immature 'I did not evolve from the apes' or 'I am really frustrated the earth is not at the centre of the solar system' kind of a way). Hell, even this debate is more relevant.

Now can we please, as a culture, move on?

Disaggregating annual variables

A reader emails me:

When doing econometrics on quarterly time series data, if there are some key variables that are available only annually, is there merit in interpolating the annual data to create a quarterly series or should the variables be discarded? What is the general advice on interpolation?


As a general rule, you should not discard the annual data. As with any econometric problem, the question is: do the additional data contain potentially useful information? If the answer is yes, then the next step is to find the best way to disaggregate the annual observations into quarterly ones.

There are many ways to do this, and the most appropriate one will depend on the problem at hand. Also, keep in mind that determining the appropriate standard errors for your included variables, especially the disaggregated ones, can be a bit tricky in this setting.

Here are some relevant papers (only the first free access)

Also keep in mind that, depending on the problem you are facing, it may even make sense to aggregate variables – e.g. making annual variables out of quarterly ones. Yes, you shed information in that case, but an even more critical question to ask is whether your assumptions are satisfied (usually E(u|X)=0). In many cases, you would have reason to expect the error to be correlated with your dependent variables in a ‘quarterly’ model but not in an ‘annual’ model, in which case it would be most probably preferable to use the latter.

Tuesday, November 27, 2007

What's in a name?

Andrew Gelman quotes this paper (free access) by Leif Nelson and Joseph Simmons:


In five studies, we found that people like their names enough to unconsciously pursue consciously avoided outcomes that resemble their names. Baseball players avoid strikeouts, but players whose names begin with the strikeout-signifying letter K strike out more than others (Study 1). All students want As, but students whose names begin with letters associated with poorer performance (C and D) achieve lower grade point averages (GPAs) than do students whose names begin with A and B (Study 2), especially if they like their initials (Study 3). Because lower GPAs lead to lesser graduate schools, students whose names begin with the letters C and D attend lower-ranked law schools than students whose names begin with A and B (Study 4). Finally, in an experimental study, we manipulated congruence between participants’ initials and the labels of prizes and found that participants solve fewer anagrams when a consolation prize shares their first initial than when it does not (Study 5). These findings provide striking evidence that unconsciously desiring negative name-resembling performance outcomes can insidiously undermine the more conscious pursuit of positive outcomes.

The explanation? (Keep in mind this is a paper published in Psychological Science)

People like their names and initials (Nuttin, 1987). In fact, this name-letter effect (NLE) is influential enough to encourage the pursuit of name-resembling life outcomes and partners. [...]

Do people consciously or unconsciously pursue name-resembling outcomes? Do a few people named Jack deliberately move to Jacksonville for its Jack-resembling appeal, or are they driven by an unconscious desire? Researchers have certainly argued that the latter is true. The NLE is described as an indicator of implicit egotism (e.g., Koole, Dijksterhuis, & van Knippenberg, 2001; Jones et al., 2004; Pelham, Carvallo, & Jones, 2005; Pelham et al., 2002; Sherman & Kim, 2005), as own-name liking is thought to indicate unconscious self-liking.

I'm not convinced. To refer back to one of the quoted studies, how about students whose surnames start with 'F'? Shouldn't they be performing much worse than the C's and D's?

Two potential explanations here:

1. Omitted variable bias. For example, say that names that start with C or D are way less frequent in the population of Asian students compared to Anglo-Saxon surnames. Further, assume that Asians are discriminated against when it comes to college admission, perhaps due to uncertainty about the quality of the schools they attend. That way, the average Asian in college will be a better student than the average Anglo-Saxon, and he will also be less likely to have a name that starts with C or D.

2. There are an infinite number of hypotheses, and a finite but very large number of original datasets. In other words, datamining - or if we want to be somewhat less harsh on the researcher, pure luck.

And talking of names, here's Levitt and Dubner approaching the issue from a completely different angle.

Sunday, October 28, 2007

Does being beautiful mean being average?

Yes, according to this research (via Andrew Gelman):

The debate over the definition of beauty has been waged by both scientists and philosophers for centuries. We tested the idea that a facial configuration close to the population mean is fundamental to attractiveness.

First, we digitized images of faces of male and female college students (i.e., transformed the facial images into little dots of lightness and darkness called "pixels"). Each face is represented by a matrix of pixel values that can be mathematically averaged with the matrices of other faces. Once digitized and averaged together, we can turn the averaged pixel values back into images and have the composite faces rated for attractiveness.

College students rated the male and female composite faces as significantly higher in attractiveness than the individual faces used to create them, if the composites had at least 16 different faces in them. Thus, averaged faces are attractive. Note that when we use the word, "average," we mean the arithmetical mean, and not an average-looking person. If, for example, you take a female composite (averaged) face made of 32 different faces and overlay it on the face of an extremely attractive female model, the two images line up almost perfectly indicating that the model's facial configuration is very similar to the composites' facial configuration.

[...] we view averageness as fundamental and necessary to facial attractiveness. Averageness is not the only component of attractiveness, but without it, no face will be attractive.


Here are some selected publications on the matter.

My two cents: what makes you ugly are extreme characteristics (e.g. big nose or ears); averaging simply takes care of these 'large errors'. The same principle is behind the frequently superior performance of composite forecasts (e.g. of economic variables), where the arithmetic mean of a number of forecasts is often more accurate than any of the individual components.

Postscript: Note that averaging doesn't quite work with hair.

Saturday, October 20, 2007

Man, good effort but then you mess it up

Hai hai (as my Urdu-speaking partner would say), Derek Lowe starts well but messes up his conclusion (via Megan McArdle):

The news of a possible diagnostic test for Alzheimer’s disease is very interesting [...]

But let’s run some numbers. The test was 91% accurate when run on stored blood samples of people who were later checked for development of Alzheimer’s, which compared to the existing techniques is pretty good. Is it good enough for a diagnostic test, though? We’ll concentrate on the younger elderly, who would be most in the market for this test.The NIH estimates that about 5% of people from 65 to 74 have AD. According to the Census Bureau (pdf), we had 17.3 million people between those ages in 2000, and that’s expected to grow to almost 38 million in 2030. Let’s call it 20 million as a nice round number.

What if all 20 million had been tested with this new method? We’ll break that down into the two groups – the 1 million who are really going to get the disease and the 19 million who aren’t. When that latter group gets their results back, 17,290,000 people are going to be told, correctly, that they don’t seem to be on track to get Alzheimer’s. Unfortunately, because of that 91% accuracy rate, 1,710,000 people are going to be told, incorrectly, that they are. You can guess what this will do for their peace of mind. Note, also, that almost twice as many people have just been wrongly told that they’re getting Alzheimer’s than the total number of people who really will.

Meanwhile, the million people who really are in trouble are opening their envelopes, and 910,000 of them are getting the bad news. But 90,000 of them are being told, incorrectly, that they’re in good shape, and are in for a cruel time of it in the coming years.

The people who got the hard news are likely to want to know if that’s real or not, and many of them will take the test again just to be sure. But that’s not going to help; in fact, it’ll confuse things even more. If that whole cohort of 1.7 million people who were wrongly diagnosed as being at risk get re-tested, about 1.556 million of them will get a clean test this time. Now they have a dilemma – they’ve got one up and one down, and which one do you believe? Meanwhile, nearly 154,000 of them will get a second wrong diagnosis, and will be more sure than ever that they’re on the list for Alzheimer’s.

Meanwhile, if that list of 910,000 people who were correctly diagnosed as being at risk get re-tested, 828 thousand of them will hear the bad news again and will (correctly) assume that they’re in trouble. But we’ve just added to the mixed-diagnosis crowd, because almost 82,000 people will be incorrectly given a clean result and won’t know what to believe.

I’ll assume that the people who got the clean test the first time will not be motivated to check again. So after two rounds of testing, we have 17.3 million people who’ve been correctly given a clean ticket, and 828,000 who’ve been correctly been given the red flag. But we also have 154,000 people who aren’t going to get the disease but have been told twice that they will, 90,000 people who are going to get it but have been told that they aren’t, and over 1.6 million people who have been through a blender and don’t know anything more than when they started.

Sad but true: 91% is just not good enough for a diagnostic test.

Yes, doctors need to be able to calculate the probability a patient has a given disease taking into account not only the accuracy of the test but also other available information (e.g., for random testing, prevalence of the disease amongst an age-group); and they need to communicate this information clearly to the patient. This misunderstanding is a real problem, and something that doctors and everyone else need to be educated about.

But to go from that to '91% is just not good enough' is a huge leap.

As long as there isn't a 100% accurate test, we can never be certain whether the disease is present or not; but the test does give a lot of relevant information and we can lower the probability of a false alarm as much as we like by administering the test again and again.

If a disease affects 1 in 20 people and the test is 90% accurate, a 'positive' result means you have a mere 32% probability you are actually ill. If you administer the test a second time and you get a second positive, this probability jumps to 81%, and this keeps rising with the number of positive results. For a negative test result, the news are even better: the first negative result translates to a 99.5% you are healthy, the second negative to a .999% that you are.

(18% of the people will get one positive and one negative, which simply means there is a 95% probability they are healthy - i.e. the same as before taking any tests. Instead of 'not knowing what to believe', as Lowe speculates, their doctors should just explain to them that they need more testing if they want to increase the accuracy of the standard, pre-test prediction (healthy) above 95%)

Pay attention now, here comes the correct conclusion: If you don't have any symptoms, a positive test result for most diseases doesn't mean much - in most cases, you are still more likely to be healthy than not.

Next time you take a test, ask your doctor to calculate the probability you are actually ill or healthy; and if you want more certainty, take the test again, and again, until you are content with the degree of certainty on offer. And thank all those nice researchers for them 90% accurate tests - at least if they are not painful.

Statistical significance

...is not a measure of confidence in the point estimate; it says nothing about accuracy.

Statistical significance simply means that the true value of the statistic being estimated has a higher than 5% or 1% probability (the levels conventionally chosen) to be away from a range around zero (the range being determined by sample size and degrees of freedom of the estimator in question). Any statistic which is large enough will be found to be statistically significant even in small samples; this doesn't mean, however, that the accuracy of the point estimate can't be very poor.

Sunday, October 7, 2007

Multicollinearity

Advance warning: This is a tedious post, and it is extremely unlikely you will find it either interesting or informative.

Santosh Anagol is an economics PhD student at Yale and he blogs at Brown Man's Burden. Going through his stuff, I came across a short paper he wrote back in 2004 about the implications of multicollinearity (I won't link to Wikipedia on this, as the article on multicollinearity is lacking and potentially misleading. For more information, read a standard econometrics textbook.)

What he does is simple enough:


with the error normally distributed and uncorrelated with the x's, etc. He then proceeds to run the regression three times, with σ12 (the covariance of x1 with x2) going from zero to .99.

At correlations below .999 our statistical model nails the point estimates and has large t-values. So we don’t need to worry about correlated regressors unless the correlation is EXTREMELY high.
Talking about variables with a correlation of .99 is not very relevant for practical purposes (For many popular datasets, I doubt the correlation between the recorded values and their true values is even as high as .95). In any case, the sample size chosen (1000) is large, and it is not surprising that the OLS estimators yield estimates close to the true value even in the presence of .95 correlation (it is not surprising to an experienced econometrician; see the conclusion to the post). What is more interesting to observe is how the confidence interval around these point estimates changes as σ12 is chosen to be higher. With x's barely correlated, x1 is roughly .13 points wide, with σ12=.5 it goes to .15 units and at σ12=.95 it reaches almost .4 units.

Continuing with the results:
With regressors that have correlations around .99, we get some bad results. In this case the point estimates are off, and one of them is significant. This would obviously be the wrong conclusion about the DGP.



'Statistical significance' is often misunderstood to be a measure of confidence in the point estimate, but it is nothing of the sort. Finding an estimate to be 'statistically significant' simply means that the (95% in this case) confidence interval does not include zero - in other words, there's a low chance that the true value of the statistic in the population is zero, and thus the variable of interest is likely to have an effect on y.

So, the conclusions we would draw about the DGP from the above results are actually the right ones: x1 is not likely to be equal to zero (and it isn't; it equals 2 by construction), and there's a 95% chance that x2 lies between -3.12 and 2.9897 (which it does; by construction, x2=1). The only reason β1 is found to be statistically significant and β2 not is the fact that x1 was picked to equal 1 and x2 was picked to equal 2, so we need to feed our estimators with more information in order to establish that x1 has an effect on y than is the case with establishing the same thing for x2.

The point estimates are indeed off, but this is purely due to the particular random sample - and the large confidence intervals alert us as to the possibility this is the case. Run the same model with a larger sample size (or pick a large number of other random samples and draw the probability distribution of your estimators), and the OLS estimates will be spot on.

And a final observation:

If two variables are highly correlated, will it screw up coefficients on other, exogenous variables? I ran the model with another regressor x3 that was uncorrelated with x1 and x2 , and with a coefficient of 3 in the data generating process. The degree of correlation between x1 and x2 DOES NOT CHANGE point estimates and t-stats of our coefficient on x3.


...which is to be expected from theory. Any explanatory variable that is not correlated with the x's of interest does not need to enter the model at all - it can safely reside in the error term without any bias being introduced as a result. (the Gauss Markov assumptions only call for the error to equal zero given x). The coefficient on x3 would be the same even if x1 and x2 were not included in the regression, and the coefficients on x1 and x2 are not affected by the inclusion of x3 in the specification.

Before leaving this post, I should make clear that I am not critical of Santosh's note; in fact, I think it's great and his effort is to be applauded. From the introduction to the paper:

I’ve been confused for a while about the effects of having x variables that are correlated. This is pretty embarrassing, given this is undergrad metrics stuff. But I’ve also seen enough grad students and professors throw around ”multicollinearity” without really understanding its implications that its worth straightening out.

This is not 'undergrad metrics stuff' at all. It is true that economics undergrads learn about the qualitative effect of 'multicollinearity', but developing an understanding of its significance in practice only comes after substantial exposure to the literature and hands-on experience (as with so many things econometrics). Santosh's attitude is the right one, and playing around with simulated data is a great, low cost way to digest the theory and really understand econometrics - one that tutors should be encouraging far more than is currently the case.

Saturday, October 6, 2007

World Freedom Atlas

The World Freedom Atlas is a “geovisualization tool” for world statistics. The amount of information is impressive - Peter Klein went into the trouble of linking to some of the sources:

It includes the most important variables used by economists including income and purchasing power from the Penn World Table, legal origin from LLSV, economic freedom from the Fraser Institute and the Heritage Foundation, policy constraints from Witold Henisz, the World Bank’s governance indicators, and a host of other variables from Acemoglu, Johnson and Robinson; Barro and Lee; Easterly and Levine; Persson and Tabellini; and several others.