Showing posts with label econometrics. Show all posts
Showing posts with label econometrics. Show all posts

Sunday, October 26, 2008

The normal distribution

This beautiful image comes courtesy of W. J. Youden.

I've been reading Edward Tufte's superb The Visual Display of Quantitative Information - perhaps the most perfect book I've ever come across. Almost every page is a revelation, so expect me to be posting more of the wonderful graphs Tufte has collected in the future.

Monday, October 20, 2008

The importance of being clear

A formatting fubar involving an Excel spreadsheet has left Barclays Capital with contracts involving collapsed investment bank Lehman Brothers than it never meant to acquire.

Working to a tight deadline, a junior law associate at Cleary Gottlieb Steen & Hamilton LLP converted an Excel file into a PDF format document. The doc was to be posted on a bankruptcy court's website before a midnight purchase offer deadline on 18 September, just four hours after Barclays sent the spreadsheet to the lawyers. The Excel file contained 1,000 rows of data and 24,000 cells.

Some of these details on various trading contracts were marked as hidden because they were not intended to form part of Barclays' proposed deal. However, this "hidden" distinction was ignored during the reformatting process so that Barclays ended up offering to take on an additional 179 contracts as part of its bankruptcy buyout deal, Finextra reports.


The Register has the full story. As Merv has always warned, 'horrible things happen when you hide cells in excel'.

I see this as a manifestation of a wider lack of education on the importance of communicating information efficiently. The Spartans, Tufte, Strunk and White, the Economist, Picasso and numerous econometricians have done a lot to improve things, but management-speak, TV advertising and other such phenomena show we still have a long way to go.

Monday, September 29, 2008

Stata lessons & other resources

A friend asked for a quick list, so here goes:

UCLA's excellent resources to help you learn and use Stata

Another great collection of Stata Resources by Park Hun Myoung, as well as a stata command cheat-card

London School of Economics Stata Resources

Syracuse University's Stata tutorial

Program in Statistics and Methodology by the pol sci's at Ohio State University

Duke Stata tutorials

Tuesday, April 1, 2008

I hate seasonally adjusted series

To all data providers, wherever they may be:

Please, please stop seasonally adjusting the series you give me. I can run a regression with seasonal dummies myself, thank you very much. So can everyone else. The difficult thing is to get from the seasonally adjusted series to the original one, and I don't see why you should deprive me of the privilege.

Yes, I know I shouldn't want to in most cases, but let me be the judge of that, OK? What's your problem anyway, what is it to you? Please, just do me this favour.

Monday, February 11, 2008

Econometric causality

James Heckman has an excellent paper on the subject (free access).

Thursday, December 6, 2007

Stop abusing statistical significance

I just made my first edit on Wikipedia, on the article on 'statistical power'. Here's the old text, with the deleted parts in bold:

There are times when the recommendations of power analysis regarding sample size will be inadequate. Power analysis is appropriate when the concern is with the correct acceptance or rejection of a null hypothesis. In many contexts, the issue is less about determining if there is or is not a difference but rather with getting a more refined estimate of the population effect size. For example, if we were expecting a population correlation between intelligence and job performance of around .50, a sample size of 20 will give us approximately 80% power (alpha = .05, two-tail). However, in doing this study we are probably more interested in knowing whether the correlation is .30 or .60 or .50. In this context we would need a much larger sample size in order to reduce the confidence interval of our estimate to a range that is acceptable for our purposes. These and other considerations often result in the true but somewhat simplistic recommendation that when it comes to sample size, "More is better!"

However, huge sample sizes can lead to statistical tests becoming so powerful that the null hypothesis is always rejected for real data. This is a problem in studies of differential item functioning.


Leaving the cost of collecting data aside, larger (appropriately collected) samples are ALWAYS BETTER. At the end of the day, if your sample is *too* large (for example if your statistical software restricts the amount of information you can load on it and you don't need the extra information anyways) you can always obtain a smaller random sample from your larger random sample. So, the 'more is better' recommendation is simple, but not simplistic.

The last paragraph reveals a fundamental misconception about statistical significance that refuses to go away. If the effect of an independent variable on the dependent variable is zero, using a very large sample will result to an estimated effect that is 0 to many decimal places; as the sample size increases further, the effect will approach *exactly* zero even more. NEVER USE STATISTICAL SIGNIFICANCE AS A PROXY FOR PRACTICAL SIGNIFICANCE. I have no clue whether large sample sizes have been seen as a problem in the past in studies of differential item functioning, but if that is the case then the researchers are idiots.

Here is another post on problematic applications of statistical significance.

Saturday, December 1, 2007

Disaggregating annual variables

A reader emails me:

When doing econometrics on quarterly time series data, if there are some key variables that are available only annually, is there merit in interpolating the annual data to create a quarterly series or should the variables be discarded? What is the general advice on interpolation?


As a general rule, you should not discard the annual data. As with any econometric problem, the question is: do the additional data contain potentially useful information? If the answer is yes, then the next step is to find the best way to disaggregate the annual observations into quarterly ones.

There are many ways to do this, and the most appropriate one will depend on the problem at hand. Also, keep in mind that determining the appropriate standard errors for your included variables, especially the disaggregated ones, can be a bit tricky in this setting.

Here are some relevant papers (only the first free access)

Also keep in mind that, depending on the problem you are facing, it may even make sense to aggregate variables – e.g. making annual variables out of quarterly ones. Yes, you shed information in that case, but an even more critical question to ask is whether your assumptions are satisfied (usually E(u|X)=0). In many cases, you would have reason to expect the error to be correlated with your dependent variables in a ‘quarterly’ model but not in an ‘annual’ model, in which case it would be most probably preferable to use the latter.

Tuesday, November 27, 2007

What's in a name?

Andrew Gelman quotes this paper (free access) by Leif Nelson and Joseph Simmons:


In five studies, we found that people like their names enough to unconsciously pursue consciously avoided outcomes that resemble their names. Baseball players avoid strikeouts, but players whose names begin with the strikeout-signifying letter K strike out more than others (Study 1). All students want As, but students whose names begin with letters associated with poorer performance (C and D) achieve lower grade point averages (GPAs) than do students whose names begin with A and B (Study 2), especially if they like their initials (Study 3). Because lower GPAs lead to lesser graduate schools, students whose names begin with the letters C and D attend lower-ranked law schools than students whose names begin with A and B (Study 4). Finally, in an experimental study, we manipulated congruence between participants’ initials and the labels of prizes and found that participants solve fewer anagrams when a consolation prize shares their first initial than when it does not (Study 5). These findings provide striking evidence that unconsciously desiring negative name-resembling performance outcomes can insidiously undermine the more conscious pursuit of positive outcomes.

The explanation? (Keep in mind this is a paper published in Psychological Science)

People like their names and initials (Nuttin, 1987). In fact, this name-letter effect (NLE) is influential enough to encourage the pursuit of name-resembling life outcomes and partners. [...]

Do people consciously or unconsciously pursue name-resembling outcomes? Do a few people named Jack deliberately move to Jacksonville for its Jack-resembling appeal, or are they driven by an unconscious desire? Researchers have certainly argued that the latter is true. The NLE is described as an indicator of implicit egotism (e.g., Koole, Dijksterhuis, & van Knippenberg, 2001; Jones et al., 2004; Pelham, Carvallo, & Jones, 2005; Pelham et al., 2002; Sherman & Kim, 2005), as own-name liking is thought to indicate unconscious self-liking.

I'm not convinced. To refer back to one of the quoted studies, how about students whose surnames start with 'F'? Shouldn't they be performing much worse than the C's and D's?

Two potential explanations here:

1. Omitted variable bias. For example, say that names that start with C or D are way less frequent in the population of Asian students compared to Anglo-Saxon surnames. Further, assume that Asians are discriminated against when it comes to college admission, perhaps due to uncertainty about the quality of the schools they attend. That way, the average Asian in college will be a better student than the average Anglo-Saxon, and he will also be less likely to have a name that starts with C or D.

2. There are an infinite number of hypotheses, and a finite but very large number of original datasets. In other words, datamining - or if we want to be somewhat less harsh on the researcher, pure luck.

And talking of names, here's Levitt and Dubner approaching the issue from a completely different angle.

Sunday, October 28, 2007

Does being beautiful mean being average?

Yes, according to this research (via Andrew Gelman):

The debate over the definition of beauty has been waged by both scientists and philosophers for centuries. We tested the idea that a facial configuration close to the population mean is fundamental to attractiveness.

First, we digitized images of faces of male and female college students (i.e., transformed the facial images into little dots of lightness and darkness called "pixels"). Each face is represented by a matrix of pixel values that can be mathematically averaged with the matrices of other faces. Once digitized and averaged together, we can turn the averaged pixel values back into images and have the composite faces rated for attractiveness.

College students rated the male and female composite faces as significantly higher in attractiveness than the individual faces used to create them, if the composites had at least 16 different faces in them. Thus, averaged faces are attractive. Note that when we use the word, "average," we mean the arithmetical mean, and not an average-looking person. If, for example, you take a female composite (averaged) face made of 32 different faces and overlay it on the face of an extremely attractive female model, the two images line up almost perfectly indicating that the model's facial configuration is very similar to the composites' facial configuration.

[...] we view averageness as fundamental and necessary to facial attractiveness. Averageness is not the only component of attractiveness, but without it, no face will be attractive.


Here are some selected publications on the matter.

My two cents: what makes you ugly are extreme characteristics (e.g. big nose or ears); averaging simply takes care of these 'large errors'. The same principle is behind the frequently superior performance of composite forecasts (e.g. of economic variables), where the arithmetic mean of a number of forecasts is often more accurate than any of the individual components.

Postscript: Note that averaging doesn't quite work with hair.

Saturday, October 20, 2007

Statistical significance

...is not a measure of confidence in the point estimate; it says nothing about accuracy.

Statistical significance simply means that the true value of the statistic being estimated has a higher than 5% or 1% probability (the levels conventionally chosen) to be away from a range around zero (the range being determined by sample size and degrees of freedom of the estimator in question). Any statistic which is large enough will be found to be statistically significant even in small samples; this doesn't mean, however, that the accuracy of the point estimate can't be very poor.

Sunday, October 7, 2007

Multicollinearity

Advance warning: This is a tedious post, and it is extremely unlikely you will find it either interesting or informative.

Santosh Anagol is an economics PhD student at Yale and he blogs at Brown Man's Burden. Going through his stuff, I came across a short paper he wrote back in 2004 about the implications of multicollinearity (I won't link to Wikipedia on this, as the article on multicollinearity is lacking and potentially misleading. For more information, read a standard econometrics textbook.)

What he does is simple enough:


with the error normally distributed and uncorrelated with the x's, etc. He then proceeds to run the regression three times, with σ12 (the covariance of x1 with x2) going from zero to .99.

At correlations below .999 our statistical model nails the point estimates and has large t-values. So we don’t need to worry about correlated regressors unless the correlation is EXTREMELY high.
Talking about variables with a correlation of .99 is not very relevant for practical purposes (For many popular datasets, I doubt the correlation between the recorded values and their true values is even as high as .95). In any case, the sample size chosen (1000) is large, and it is not surprising that the OLS estimators yield estimates close to the true value even in the presence of .95 correlation (it is not surprising to an experienced econometrician; see the conclusion to the post). What is more interesting to observe is how the confidence interval around these point estimates changes as σ12 is chosen to be higher. With x's barely correlated, x1 is roughly .13 points wide, with σ12=.5 it goes to .15 units and at σ12=.95 it reaches almost .4 units.

Continuing with the results:
With regressors that have correlations around .99, we get some bad results. In this case the point estimates are off, and one of them is significant. This would obviously be the wrong conclusion about the DGP.



'Statistical significance' is often misunderstood to be a measure of confidence in the point estimate, but it is nothing of the sort. Finding an estimate to be 'statistically significant' simply means that the (95% in this case) confidence interval does not include zero - in other words, there's a low chance that the true value of the statistic in the population is zero, and thus the variable of interest is likely to have an effect on y.

So, the conclusions we would draw about the DGP from the above results are actually the right ones: x1 is not likely to be equal to zero (and it isn't; it equals 2 by construction), and there's a 95% chance that x2 lies between -3.12 and 2.9897 (which it does; by construction, x2=1). The only reason β1 is found to be statistically significant and β2 not is the fact that x1 was picked to equal 1 and x2 was picked to equal 2, so we need to feed our estimators with more information in order to establish that x1 has an effect on y than is the case with establishing the same thing for x2.

The point estimates are indeed off, but this is purely due to the particular random sample - and the large confidence intervals alert us as to the possibility this is the case. Run the same model with a larger sample size (or pick a large number of other random samples and draw the probability distribution of your estimators), and the OLS estimates will be spot on.

And a final observation:

If two variables are highly correlated, will it screw up coefficients on other, exogenous variables? I ran the model with another regressor x3 that was uncorrelated with x1 and x2 , and with a coefficient of 3 in the data generating process. The degree of correlation between x1 and x2 DOES NOT CHANGE point estimates and t-stats of our coefficient on x3.


...which is to be expected from theory. Any explanatory variable that is not correlated with the x's of interest does not need to enter the model at all - it can safely reside in the error term without any bias being introduced as a result. (the Gauss Markov assumptions only call for the error to equal zero given x). The coefficient on x3 would be the same even if x1 and x2 were not included in the regression, and the coefficients on x1 and x2 are not affected by the inclusion of x3 in the specification.

Before leaving this post, I should make clear that I am not critical of Santosh's note; in fact, I think it's great and his effort is to be applauded. From the introduction to the paper:

I’ve been confused for a while about the effects of having x variables that are correlated. This is pretty embarrassing, given this is undergrad metrics stuff. But I’ve also seen enough grad students and professors throw around ”multicollinearity” without really understanding its implications that its worth straightening out.

This is not 'undergrad metrics stuff' at all. It is true that economics undergrads learn about the qualitative effect of 'multicollinearity', but developing an understanding of its significance in practice only comes after substantial exposure to the literature and hands-on experience (as with so many things econometrics). Santosh's attitude is the right one, and playing around with simulated data is a great, low cost way to digest the theory and really understand econometrics - one that tutors should be encouraging far more than is currently the case.

Sunday, September 16, 2007

Most research findings are false

It can be proven that most claimed research findings are false.

So says John Ioannidis in a short and highly readable article in PLoS Medicine (free access). He postulates that the probability of research findings being wrong is higher when:

1. The sample sizes are small
2. The estimated effects are small
3. There is flexibility in research design, definitions, outcomes etc
4. There are financial or other interests and prejudices
5. The scientific field is 'hot'. i.e. many scientific teams work against the clock to beat the competition.

He is right, but there is less to his thesis than meets the eye.

1. Not all scientific findings are born equal. Just because a few new papers based on a handful of observations have been published claiming X has an effect on Y does not mean that their findings are accepted as the God-given truth. Some papers are more convincing than others, and readers will apply a probability that any research finding is wrong. The five factors above do increase the probability that a published finding is wrong, but they also increase the probability the results will be taken with a pinch of salt.

2. Small effects are rarely important in a 'practical' sense - as economists would say, small effects are unlikely to have important 'policy implications'. The probability that a finding is 'false' diminishes when the estimated effect is large (as per Ioannidis's second corrolary). Even if you argue that scientific results are often consumed by readers that are unable to critically assess them, it is unlikely that findings which are 'wrong' will have damaging consequences.

To cut a long story short, most research findings are indeed wrong. But is the number of false findings divided by the total number of findings saying much? Simply weigh these findings by their credibility and importance, and the world of scientific discovery looks rosy again.

Here's a WSJ article on the matter. HT to Ben Muse.

Wednesday, August 29, 2007

Why bother posting on a bank holiday? Part 2

I now have the answer to this question, courtesy of the comments section and Tim Worstall, which is:

1) To reward loyal readers. A fair point and a regrettable part of the moral hazard that datacharmer accepted with his holiday cover, i.e. that we have far less incentive to cater for his loyal readers.

2) To create content for people that visit more infrequently but read older posts too. I'll begrudgingly accept this one, although I could have created the same amount of content for occasional visitors by posting 4 times on Tuesday and skipped posting for the 3-day weekend.

3) Because you might get a lot more visitors than you were expecting, possibly due to a link from another site.

On the subject of (3), I've revised my forecasting model to:
Visits = 119.8 - 33.2 * Weekend + 205.4 * TW, where Weekend is a dummy variable that takes the value of 1 for a weekend day and zero otherwise and TW is a dummy that takes the value of 1 if Tim Worstall links to this blog and zero otherwise, with all coefficients statistically significantly different from zero at the 95% level (and for the data miners amongst you, an R-squared of 44%).

Monday, August 27, 2007

Why bother posting on a bank holiday?

Site visits for bluematter show a very pronounced weekly cycle, with most visits at the start of the week (normally peaking on Mondays, including on this graph the 6th, 13th and 20th) followed by a decline down to least visits on Sundays:


Using the last 30 days data before today, Visits = 119.8 - 33.2 * Weekend, where Weekend is a dummy variable that is 1 for weekend days but zero otherwise. The coefficient on Weekend is statistically significantly different from zero at the 95% level.

Given that most of our readers are probably English and it's a bank holiday, I can expect materially less than 86.6 people to read this today.

Saturday, August 25, 2007

The 10 Commandments of Applied Econometrics

A few posts ago, datacharmer gave a plug for Peter Kennedy's outstanding econometrics textbook. To follow up that plug, this is a paraphrased version of Kennedy's 10 Commandments of Applied Econometrics, a recipe for good research:

1. Thou shalt use common sense and economic theory
2. Thou shalt ask the right question
3. Thou shalt know the context
4. Thou shalt inspect the data
5. Thou shalt not worship complexity
6. Thou shalt look long and hard at thy results
7. Thou shalt beware the costs of data mining
8. Thou shalt be willing to compromise
9. Thou shalt not confuse statistical significance with substance
10. Thou shalt confess in the presence of sensitivity

Sunday, August 19, 2007

Are data-miners made or born?

Green (1990) gave 199 students the same data but with different errors, and asked them to find an appropriate specification. All students had been taught that models should be specified in a theoretically sensible fashion, but some were also taught about how to use F tests and goodness-of-fit for this purpose. These latter students were quick to abandon common sense, perhaps because using clearly defined rules and procedures is so attractive when faced with finding a specification.

As with driving and flying fighter jets, the risk of 'accidents' when doing econometrics is at its highest amongst those with some, but not much, experience. The novice pilot, or econometrician for that matter, will take extra care to make sure everything is as it should be. Being concious of the possibility of disaster, she will not attempt flashy tricks, whether that's flying at mach-2 or drawing inferences from not-very-well-understood statistics.

As the student becomes more comfortable and builds up some confidence in her abilities, she will tend to underestimate the amount of care that needs to be applied when performing a given manoeuvre. Flight instructors are well aware of this tendency, but it is still the case that most accidents involve pilots that have had between 500-1000 hours of flying experience. A very similar risk profile holds for budding econometricians; we should be thankful that the aftermath of a 'crash' for the analyst is far less painful than for the pilot.

The quoted excerpt is from Peter Kennedy's excellent econometrics textbook, now in its fifth edition. As far as I'm concerned, the book is unique in its approach; it focuses on the intuition behind the various methods and techniques used in econometrics using simple English, with the underlying equations relegated to technical annexes.

I have long thought that the way undergraduates are being taught econometrics is far from ideal: what good is it learning the proof of why OLS is BLUE in your second week in an introductory econometrics course? Kennedy's approach is promising, and purchasing the book is a good idea for students that have had little experience with econometrics and are not yet proficient in 'translating' the maths into something that makes intuitive sense.

The issue of the Political Methodologist where Green's article comes from is here (free access), but I have to warn you that the quality of the scanning is poor.