I got 37,300. They claim this is not quite 95th percentile, which I am a tad skeptical accurately represents my vocabulary-size percentile relative to the general population. Perhaps this survey is being forwarded around unusually literate people at the top end, or more than 5% of responders are cheating. Where are the fake words to catch cheaters? I Googled a lot of what I didn't recognize, and everything I checked was real.
I just counted, and you've used at least 12685 dictionary words in HP:MOR [1]. I was expecting it to be more to be honest, but it didn't seem right not to post just because it didn't match my expectations.
It certainly seems unlikely that someone who can produce an excellent 500,000 word work of fiction (aside: thanks very much by the way,) in addition to reams of technical writing, has a vocabulary not in the 95th percentile of the population. OTOH, HP:MOR has fewer words in it than I expected, and even the upper bound of 14795 seems low. Maybe the working is wrong; it's shown below.
There's a huge difference between writing something targeted at a selected audience, in a given time period, limited range of topics the characters will discuss, etc. and listing out all of the vocabulary words you know personally across all domains of knowledge. Even though HP:MoR Harry has a broad vocabulary, for example, and may use some words not all readers will know, there are still tons of topics (with their own specialized vocabulary...) that will never come up in the written storyline -- even if Harry would know them well.
Harry also presumably does some practical limiting of his vocabulary in conversation, because only shared vocabulary is useful if you're trying to actually communicate and don't want to stop to give definitions all the time.
It might be more interesting to compare the unique word count of HP:MoR against some of the "real" Harry Potter books, if you can get your hands on the text.
Thanks to the internet†, I can reveal a surprise: with the same methodology, the dictionary words in the seven real Harry Potter books concatenated is 19,245 and the total unique words is 21,441.
Total word count is 1,122,131 which is longer than HP:MoR by a factor of three. Plotting mean unique word count for the whole, halves and quarters of MoR gives a fit of uniques=168*length^0.3357, which makes sense given Zipf's law. That formula predicts about 18,050 words for a work of the same length as the original HP.
(Edit to add obvious test in the other direction.) The first 386,829 words of the original HP contain 12,255 unique words. The last 386,829 words contain 13,635 uniques. So, its comparable but perhaps slightly more varied (MoR had 12,685).
In light of those figures, is it possible Eliezer's vocabulary is less good than he thinks (Dunning-Kruger)? Especially as the Harry Potter book were written for children and presumably edited as such.
On the other hand, the fact that Eliezer seems to have used fewer words in his writing than you'd expect if his vocab was excellent doesn't mean that his known vocab is poor — he might just not use all the words he knows in writing.
Additionally, given the success of J K Rowling as an author, you might expect her vocabulary to be excellent, so it is conceivable that he's good and she's better.
† I have all the Harry Potter books on a shelf at home. Is torrenting the pdfs at work so I can word count them infringing copyright? I could have done it manually, it just would have taken longer.
I thought Eliezer's Lesswrong sequences might give different results. Applying your tests to those (from http://jb55.com/lesswrong/), I get 257,646 total words, 11,666 unique dictionary words, and 12,721 unique words (I'm surprised there aren't more unique words, given that the quantum physics sequence is in there).
168*257,646^.3357 = 11,010, so the sequences seem to be at about HP level.
Excellent work, by the way; thanks for the analysis.
Site creator here -- you're right, survey participants are incredibly literate. I suppose that's Internet users in general, disregarding YouTube commenters :), or else the particular people who have spread the test, or are interested in taking it. Average verbal SAT score on the site is 700 (out of 800), far above the population's average of around 500.
Right; most people with "normal" or worse vocabularies are emphatically not going to see a "test your vocabulary" link and say "Hey, instead of doing something fun, let's see how poor my vocabulary really is! Then I'll tell all my friends!"
And when they see some egghead friend on Facebook has posted their vocabulary score and is challenging them to respond... they'll roll their eyes, and move on to their Farmville updates.
Don't get me wrong -- I love these things, and it came back with 37K for me -- but there's no way I'm posting that score, or even the link, to Facebook. I know how to maintain friendships, and saying "look how smart I am; I'm probably smarter than you" does not figure into it.
I can't imagine that the percentiles are reflective of the general population...I got 27,000 which it claims is the median score, and from practical experience my vocabulary is quite a bit larger than most people's. If this is measuring, say, the HN crowd, then perhaps it's accurate.
I found the same thing. I suspect a combination of the early respondents being quite a bit above average and the possibility that some people are checking words they don't actually know, or simple believe they know a correct definition for.
I think it would be a good idea to weight the survey with some test questions that ask if you know a definition to some of the less common words and then ask you to pick a correct definition from a list of 5 with 4 incorrect answers. At least this way they can approximate how much someone may exaggerate their knowledge.
However as someone who answered as honestly as I could (without spending the time to verify my definition of each word) it is cool to know what my personal vocabulary is.
I only got ~22k and I managed to score in the top 15% on the GRE verbal portion not too long ago. So either people who take the GRE are on average below median or the results aren't quite accurate.
Likewise. I answered honestly and got a little over 22k. I've always tested extremely well on verbal portions of standardized tests and feel that my vocabulary is well above normal but this would put me at about the 20th percentile.
Let's think about it logically. If you were to actually test the users instead of asking them 'which of these do you know?', the ONLY option is multiple choice since analysing text-field input from the user to determine if their definition is correct is at best extremely difficult, and more than likely impossible ... right ?
I agree that the quality of the result is highly dependent on the honesty
of the person tested, but with multiple choice, wouldn't you be able to
deduct the meaning of the word from choices you are given?
E.g.: One of the words I encountered in the test was 'terpsichorean'. While I knew that Terpsichore is one of the Muses, I did not know which one and left the box unticked. Had there been multiple choices, I might have guessed the correct solution.
I ticked 'terpsichorean', because it made me think "dreamy travelogue writing of a scenic beach with either terpsichorean sky or sea, it means X looks a pale shade of blue-green".
Google tells me it means dancing so I was way way off (maybe conflating turquoise and cerulean?).
But I have no way of knowing how many words that I feel comfortable defining are actually nowhere near correct, so to be any kind of accurate, they need to do some verification of correctness. All 'honesty' means is 'don't deliberately cheat' not 'don't be dumb'.
I had "discomfit" as one of my words that I wasn't clear of the definition of, I was pretty close when I looked it up but couldn’t have guaranteed it. It's probably easily confused with discomfort ... which made me think that this needs to be a little more tested. Commonly misread words could easily inflate scores.
However, I think a multiple choice test could also inflate scores unless the definitions were very cunningly constructed.
I scored 75-80th percentile (32,800) which surprised me. It seems quite a lot of words, for one. For another I consider my vocab' to be very good and I don't think I'm being bigheaded in that. Ergo I expected to be ranked higher.
On the second page there was an entire column of words of which I recognised only three sufficiently to provide a guaranteed accurate definition. One of that column was terpischorean, another tatterdemalion.
Whilst looking up tatterdemalion I found little use of it after the 1930s except as a proper noun (a Marvel Comics character for example). What I did find however is that Google Books is useless for finding dates. One citation from an author Sir Edward Bulwer Lytton is given a date of 1999. That's a reprint date, the author died in the 19th century.
Actually, the SAT also employs multiple choice, and quite successfully. You see, multiple choice can not only give you hints, but can also be used to lead you astray. In the end, it balances out.
True, but it's still much better than just relying on people submitting accurate results without even testing them ... at least with multiple choice everyone is being tested to approximately the same metric, and you can get a more accurate percentile
To my defense, they are both derived from 'deducere', to lead away, and 'were not distinguished in sense until the mid 17th cent' according to my system’s dictionary
Eliezer, I think you appreciate some of the ideas behind the FAQ I'll repost here with adaptation to the current situation:
VOLUNTARY RESPONSE POLLS
As I commented previously when we had a poll on the ages of HNers, the data can't be relied on to make such an inference. That's because the data are not from a random sample of the relevant population. One professor of statistics, who is a co-author of a highly regarded AP statistics textbook, has tried to popularize the phrase that "voluntary response data are worthless" to go along with the phrase "correlation does not imply causation." Other statistics teachers are gradually picking up this phrase.
-----Original Message----- From: Paul Velleman [SMTPfv2@cornell.edu] Sent: Wednesday, January 14, 1998 5:10 PM To: apstat-l@etc.bc.ca; Kim Robinson Cc: mmbalach@mtu.edu Subject: Re: qualtiative study
Sorry Kim, but it just aint so. Voluntary response data are worthless. One excellent example is the books by Shere Hite. She collected many responses from biased lists with voluntary response and drew conclusions that are roundly contradicted by all responsible studies. She claimed to be doing only qualitative work, but what she got was just plain garbage. Another famous example is the Literary Digest "poll". All you learn from voluntary response is what is said by those who choose to respond. Unless the respondents are a substantially large fraction of the population, they are very likely to be a biased -- possibly a very biased -- subset. Anecdotes tell you nothing at all about the state of the world. They can't be "used only as a description" because they describe nothing but themselves.
I think Professor Velleman promotes "Voluntary response data are worthless" as a slogan for the same reason an earlier generation of statisticians taught their students the slogan "correlation does not imply causation." That's because common human cognitive errors run strongly in one direction on each issue, so the slogan has take the cognitive error head-on. Of course, a distinct pattern in voluntary responses tells us SOMETHING (maybe about what kind of people come forward to respond), just as a correlation tells us SOMETHING (maybe about a lurking variable correlated with both things we observe), but it doesn't tell us enough to warrant a firm conclusion about facts of the world. The Literary Digest poll
is a spectacular historical example of a voluntary response poll with a HUGE sample size and high response rate that didn't give a correct picture of reality at all.
When I have brought up this issue before, some other HNers have replied that there are some statistical tools for correcting for response-bias effects, IF one can obtain a simple random sample of the population of interest and evaluate what kinds of people respond. But we can't do that here on HN, nor can we for the online vocabulary estimation.
Another reply I frequently see when I bring up this issue is that the public relies on voluntary response data all the time to make conclusions about reality. To that I refer careful readers to what Professor Velleman is quoted as saying above (the general public often believes statements that are baloney) and to what Google's director of research, Peter Norvig, says about research conducted with better data,
that even good data (and Norvig would not generally characterize voluntary response data as good data) can lead to wrong conclusions if there isn't careful thinking behind a study design. Again, human beings have strong predilections to believe certain kinds of wrong data and wrong conclusions. We are not neutral evaluators of data and conclusions, but have predispositions (cognitive illusions) that lead to making mistakes without careful training and thought. Here, the conclusion "those other guys are cheating and that dragged down my vocabulary percentile score" is an example of a conclusion resulting from human predispositions.
Another frequently seen reply is that sometimes a "convenience sample" (this is a common term among statisticians for a sample that can't be counted on to be a random sample) of a population offers just that, convenience, and should not be rejected on that basis alone. But the most thoughtful version of that frequent reply I recently saw did correctly point out that if we know from the get-go that the sample was not done statistically correctly, then even if we are confident (enough) that HN participants are young or that their vocabularies are large, we wouldn't want to extrapolate from that to conclude that the users of any technology site are young, or that respondents to online surveys as a whole have large vocabularies.
On my part, I wildly guess that most HNers are younger than I am in part because this kind of poll recurs often on HN. I similarly guess that participants in online surveys of vocabulary size are likely to have larger vocabularies than average people in the general public because most people I meet find discussions of word meanings boring. But neither guess gives me a good quantitative basis for estimating how much users here differ from the general population.
I'm questioning the way respondents are classified.
I chose 'Canada' as my region since I'm from Montreal. My first language is English and I'm fluent French. I did the first half of elementary school in French. Firstly non-Quebec anglophones tend to have better grammar and larger vocabularies than anglophone Quebecers. Secondly it doesn't take into account that English can be a 3rd language . Most immigrants to Quebec are required by law to attend French language elementary and high schools (there are exceptions). Immigrant children who's first language isn't English or French (the majority) take on two new languages, English being their 3rd after French. English tends to be the social language for many.
Montreal has a strong tech industry employing bilingual/multilingual people many of which read HN and possibly took part in the survey. My gut feeling is that English speaking Quebecers are skewing the stats. More granular control over region will be useful; show some insight to this reality.
Note: I traveled through China and south east Asia last year and found the quality of English to be much better than I expected. Considering Indochina ruled by the French I didn't find a person who could speak it. To possibly classify any country as "non-English-speaking" is kind of silly. Every country is "other-English-speaking" but then again it's a subjective classification isn't it. Doesn't China have the largest English speaking population now...
By contrast, I scored 23,300 putting me on par with the average result for a 15 year old, if this test is to be believed. I was brutally honest about not selecting words I only recognised and might be able infer meaning from in context, but couldn't articulate a clear definition for (of which there were a surprising number).
And yet I look at the vocabulary used in the comment posts of those claiming high scores, and wonder how they ever scored so highly (accepting that comments are not neccessarily reflective of ones general writing or vocabulary). I believe that as this test is so open to cheating, that using responses to it as the corpus for determining median scores renders the entire exercise completely meaningless.
None of the words were domain-specific; they were all general vocabulary. You're a polymath, but a lot of your knowledge is focused into technical areas that weren't represented on the test.