I tried to organize vocab by difficulty level for an English language-learning app once.
It shocked me how there is absolutely no "right" answer.
If you are teaching English for travel, then you're prioritizing a lot of stuff around bathrooms, transportation, menu items, etc.
If it's for understanding TV, it's a lot of words like "murder", etc. Depending on which TV shows you want to understand.
If it's for reading the newspaper, you don't ever need to know "bathroom", but you sure do need to know words like "congressman".
While if you are living somewhere, it's really important to know a lot of basic supermarket items that you wouldn't prioritize for other usages.
Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed. And the substitutes -- transcribed speech from TV, radio, podcasts, etc. -- is not the same context as the random stuff you say at home and during an average day.
adrianN 24 minutes ago [-]
There is also a whole different language used to talk to toddlers and small children that native speakers know but very few adult learners ever pick up. And slang used among teenagers and young adults that often times not even their parents understand fully.
dspillett 15 minutes ago [-]
The problem with adding slang to any organised education program is two-fold: some of it changes much more quickly than you course materials can, and much of it is quite localised.
Sam6late 59 minutes ago [-]
I tried to teach 'magicE vocab', sorting them by difficulty level for an English language plan, and got them easily arranged. For example, sham/shame and slide/slid are for hard to learn level, while ate, pale, kite, are for the easy level.
DC-3 38 minutes ago [-]
I think a lot about this. It's funny how there are certain domains of language, familiar to all native speakers, but that you are simply not likely to learn as an adult language learner, at least not without diligent and focused study.
I live and work in a foreign country, and have a modest but functional grasp of the language here. Which is to say, I know how to say 'sustainability', 'union-negotiated collective agreement', and 'offensive conduct in the workplace'.... but if you asked me the words for 'cedar', 'robin', 'pond', or 'linen' I would be struck dumb. And yet presumably every ten year old I walk past in the street would know those words as comfortably as I knew them in English as a ten year old in England.
flyingshelf 20 minutes ago [-]
It's kinda funny. My life switched to English when I was around 18 and, while my family still speaks my native language exclusively, everything I've learned since then I don't know how to say it in my native language.
Just earlier I was talking to them, trying to say "when life gives you lemons" and being unable to find an equivalent phrase. I hope they understood my lemons reference regardless.
coliveira 23 minutes ago [-]
When you're a foreigner learning a language your experience is mostly from reading. There are words, however, that are much more common at home and in spoken language. So you'll probably never learn the words that a 10 year old kid knows.
watwut 12 minutes ago [-]
It is changing. A bulk of learners now get their vocabularies from youtube, netflix and podcasts. Watching kids shows is common recommendation (I disagree with it somewhat, but that does not make it not common), so I think people will learn also those 10 years old world words.
ReactiveJelly 44 minutes ago [-]
Kinda like how there's no average-sized airplane pilot
thaumasiotes 1 hours ago [-]
> While if you are living somewhere, it's really important to know a lot of basic supermarket items that you wouldn't prioritize for other usages.
Why? Having spent a good amount of time living in Shanghai, I found it important to be able to understand menus. But there's no pressure to know the words for supermarket items; you can just go to the supermarket and look for the item.
Otherwise your point is correct; all semantic words are equally difficult and which ones you know depends on the things you like to talk about. Grammatical words are more difficult, and more important, but this is so widely understood that language-learning material already treats them as an entirely separate class of things to learn.
> Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed.
(1) You seem to want COCA, which includes a bunch of transcribed telephone calls.
(2) Word frequencies are still the wrong concept. If you want to understand a particular document, you need to understand almost all of the words that appear in that document. (You'll be able to learn some of them from their use in the document.) If you decide to learn a list of "frequent" words, you're unlikely to be able to understand more than a couple of isolated sentences in any given document.
Frieren 1 hours ago [-]
> The “Social-Communicative” level barely changed in size. But nearly a quarter of the words in the 1953 list are gone, and 39% of the 2023 words are new. Humble, loyalty, fellowship, generous, polite, and companionship gave way to community, identity, organization, ethnic, gender, and narrative. ...It offers fewer words for the people directly around you, but more for belonging at a distance.
I would blame inequality on this one. In a more unequal world tribalization is a survival strategy and language follows.
When you see everybody else as your equals then focusing on describing that individual person, instead of their group, makes more sense.
Economic inequality affects deeply how we think about others.
SoftTalker 37 minutes ago [-]
I can't quite follow your point. Are you saying there is more inequality today than in 1953?
theproblemisyou 16 minutes ago [-]
[flagged]
kuboble 1 hours ago [-]
I would blame it more on globalization.
In 1953 people were not exposed as much to different groups of people far away.
kuboble 2 hours ago [-]
I tried to build a similar list myself for German and it's not easy as just taking a lot of content and counting frequency. I also haven't found existing curated lists of most useful vocabulary.
There are some databases but e.g. they are biased towards Wikipedia and web which makes some very obscure words at the top of popularity (like some technical words which are present on each wiki page like Datenschutz or Impressum).
8 minutes ago [-]
anymouse123456 1 hours ago [-]
Useless, obnoxious, absolutely frustrating Scrolljacker? Straight to jail.
Jtarii 34 minutes ago [-]
always hated this style of web presentation.
vanderZwan 45 minutes ago [-]
The author mentions various categories grew or shrank, but always in percentages. Since the list also grew from 2300 to 2800 words that feels like might distort things a bit: in absolute count, a category that shrank by 1% lost fewer words than a category that grew by 1% would have gained, no?
Having said that, the categories that shrank all did so by a big enough percentage to also shrink in absolute number of words, so at least that isn't a problem.
dredmorbius 21 minutes ago [-]
Trying to read this article by paging through it with the spacebar is ... absolutely impossible.
I'm not going to finger-scroll or down-arrow the whole thing.
Don't break scrolling. Please.
bluGill 7 minutes ago [-]
I thought it was one of those paywalls that give you a taste. Since it wasn't enough for me to pay I was going for close when more came up. Only then did I realize what was going on. It was still annoying enough give up after a few more pages.
wseqyrku 50 minutes ago [-]
It's changed so much I can't even parse the title.
bethekidyouwant 1 hours ago [-]
Why keen?
OJFord 1 hours ago [-]
Interesting one if you look at Google Ngram Viewer – usage dropped off massively to 70s/80s, and it's picked back up since but not to 1953 or earlier levels, so even that doesn't explain it.
Must just be the combination of that increase as well as other words decreasing in usage I suppose. E.g. perhaps we're a bit less keen, but also much less passionate, so keen ends up making the cut.
ghaff 1 hours ago [-]
As a native English speaker, keen feels like a very 1950s TV word. Don't know if that's actually true but feels like that way. I expect you're more likely to hear like cool today are some more contemporary word.
OJFord 28 minutes ago [-]
I'm also a native English speaker, sounds perfectly normal & commonplace to me.
mixmastamyk 1 hours ago [-]
Brits still use the word somewhat frequently, though I haven’t bumped into many of them in the last ten years.
saltcured 32 minutes ago [-]
Right, as far as live exposure, I associate its usage with UK/AU/NZ expats in my western US, academic R&D microcosm.
ghaff 30 minutes ago [-]
Yeah, I guess I can also see it as more of a Britishism (etc.) currently.
bethekidyouwant 56 minutes ago [-]
“It offers fewer words for the people directly around you, but more for belonging at a distance” - the further I get into this article the more it just seems like sampling bias as a narrative.
OJFord 17 minutes ago [-]
It's a wordy (/AI?) way to say it, but I took it to be commenting on our lives being more 'abstract'/digital than they were, so naturally a lot of our language is less physical.
Oughtn't really affect emotive language like 'keen', though.
morninglight 1 hours ago [-]
Point your Duck at "VOA Special English"
ksec 36 minutes ago [-]
I mean the subject is interesting but I have the say the design of the page is not only janky and needlessly complicated with animations.
thomastjeffery 42 minutes ago [-]
> It’s as if the world now requires you to be more precise about everything.
This is something I've been thinking a lot about. We have trended from subjective language to objective language. Why?
Computing. Software is written with objective language. Everything is clearly unambiguously defined. Blue is no longer a category, it's #0000FF. Logic must always reduce to a binary truth value. Most of what we have to talk about is somehow relative to software. Software even structures most of what we write! We don't just talk to each other, we tweet, email, message, post, search, etc. These structures each imply a specific set of phrase structures that can make sense.
Lately, it's hard to go even a day without reading some complaint that such and such was written by "AI" (an LLM). Why is this so obvious? Well, the core advantage that LLMs provide is that they don't compute. Inside an LLM, there is no arithmetic, no logical branches, no truth values. Phrases aren't generated to define or to resolve. They are generated to continue. Sure, we can direct the story to follow the steps of logical deduction, but that isn't anything like calculation. An LLM simply isn't invested in logic, precision, correctness, etc. the way we expect modern writers to be. It's not the em-dashes or the word choice that illustrates this, it's the fundamental perspective of the system.
We are sorely missing subjectivity. Natural language never was, and never will be, computable. You can't reduce a natural story to binary truth values without choosing an arbitrary perspective that resolves its ambiguity. The more precisely abstract our language gets, the more detached from reality our stories become. The more objective our assertions about reality are, the less relevant they can be.
My answer to this is to make the arbitrary choice of perspective a first-class feature. If we can explicitly decide what meaning is relevant, we should be able to weakly solve natural language processing. It seems like a pretty simple and obvious idea, but so far is easier said than done.
pixl97 12 minutes ago [-]
Eh, this seems somewhat right, but somewhat not right either.
Language has to do with our Monkeysphere, that humans over long periods of time in the past were limited to a very small subset of people they interacted with socially and closely. A few institutions likely had a large effect on your language like the church depending on where you were. After that it was the people you interacted with to stay alive. Because almost everything was in person or person to person transfer of information a lot of socially encoded clues were involved which lessened the need for well defined words.
Books were the first stage of homogenizing language as they could be shared over long distances and to many people, but more sequentially than latter forms of communication. After that radio and TV had a huge effect, for example the 'General American' used in broadcasts that was based heavily on a midwestern accent.
As we encroached on the 70s and 80s the previous technological advances and things like high speed interstates and trucking shrank America to something you could drive across in less than a week, and you could reach anywhere by voice nearly instantly. Suddenly people in California, Texas and New York all could be in the same meetings and local colloquialisms would need explained, so people would trend to a shared vocabulary.
It's also odd to me to say an LLM isn't subjective. Each LLM has it's own behavior, it's that there are like 20 or 30 big LLMs in all, and people are using them millions to billions of time so we're getting that one LLMs language everywhere. And that's why I disagree and will say natural language is computable, but it's also lossy and probabilistic. And for the most part it's single prompt and being ran by the user for the cheapest price possible.
DenisDolya 1 hours ago [-]
Thank you, just at the right time.
calvinmorrison 1 hours ago [-]
My mum learned watching cowboy TV in the 60s and 70s. She's got some uh colloquialisms for sure
KinetiNode 2 hours ago [-]
[dead]
delichon 2 hours ago [-]
At the moment this story has 2 upvotes in a half hour and is in the 8th position on the front page. Apparently HN has a fairy godmother algorithm that randomly promotes posts.
It shocked me how there is absolutely no "right" answer.
If you are teaching English for travel, then you're prioritizing a lot of stuff around bathrooms, transportation, menu items, etc.
If it's for understanding TV, it's a lot of words like "murder", etc. Depending on which TV shows you want to understand.
If it's for reading the newspaper, you don't ever need to know "bathroom", but you sure do need to know words like "congressman".
While if you are living somewhere, it's really important to know a lot of basic supermarket items that you wouldn't prioritize for other usages.
Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed. And the substitutes -- transcribed speech from TV, radio, podcasts, etc. -- is not the same context as the random stuff you say at home and during an average day.
I live and work in a foreign country, and have a modest but functional grasp of the language here. Which is to say, I know how to say 'sustainability', 'union-negotiated collective agreement', and 'offensive conduct in the workplace'.... but if you asked me the words for 'cedar', 'robin', 'pond', or 'linen' I would be struck dumb. And yet presumably every ten year old I walk past in the street would know those words as comfortably as I knew them in English as a ten year old in England.
Just earlier I was talking to them, trying to say "when life gives you lemons" and being unable to find an equivalent phrase. I hope they understood my lemons reference regardless.
Why? Having spent a good amount of time living in Shanghai, I found it important to be able to understand menus. But there's no pressure to know the words for supermarket items; you can just go to the supermarket and look for the item.
Otherwise your point is correct; all semantic words are equally difficult and which ones you know depends on the things you like to talk about. Grammatical words are more difficult, and more important, but this is so widely understood that language-learning material already treats them as an entirely separate class of things to learn.
> Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed.
(1) You seem to want COCA, which includes a bunch of transcribed telephone calls.
(2) Word frequencies are still the wrong concept. If you want to understand a particular document, you need to understand almost all of the words that appear in that document. (You'll be able to learn some of them from their use in the document.) If you decide to learn a list of "frequent" words, you're unlikely to be able to understand more than a couple of isolated sentences in any given document.
I would blame inequality on this one. In a more unequal world tribalization is a survival strategy and language follows.
When you see everybody else as your equals then focusing on describing that individual person, instead of their group, makes more sense.
Economic inequality affects deeply how we think about others.
In 1953 people were not exposed as much to different groups of people far away.
There are some databases but e.g. they are biased towards Wikipedia and web which makes some very obscure words at the top of popularity (like some technical words which are present on each wiki page like Datenschutz or Impressum).
Having said that, the categories that shrank all did so by a big enough percentage to also shrink in absolute number of words, so at least that isn't a problem.
I'm not going to finger-scroll or down-arrow the whole thing.
Don't break scrolling. Please.
Must just be the combination of that increase as well as other words decreasing in usage I suppose. E.g. perhaps we're a bit less keen, but also much less passionate, so keen ends up making the cut.
Oughtn't really affect emotive language like 'keen', though.
This is something I've been thinking a lot about. We have trended from subjective language to objective language. Why?
Computing. Software is written with objective language. Everything is clearly unambiguously defined. Blue is no longer a category, it's #0000FF. Logic must always reduce to a binary truth value. Most of what we have to talk about is somehow relative to software. Software even structures most of what we write! We don't just talk to each other, we tweet, email, message, post, search, etc. These structures each imply a specific set of phrase structures that can make sense.
Lately, it's hard to go even a day without reading some complaint that such and such was written by "AI" (an LLM). Why is this so obvious? Well, the core advantage that LLMs provide is that they don't compute. Inside an LLM, there is no arithmetic, no logical branches, no truth values. Phrases aren't generated to define or to resolve. They are generated to continue. Sure, we can direct the story to follow the steps of logical deduction, but that isn't anything like calculation. An LLM simply isn't invested in logic, precision, correctness, etc. the way we expect modern writers to be. It's not the em-dashes or the word choice that illustrates this, it's the fundamental perspective of the system.
We are sorely missing subjectivity. Natural language never was, and never will be, computable. You can't reduce a natural story to binary truth values without choosing an arbitrary perspective that resolves its ambiguity. The more precisely abstract our language gets, the more detached from reality our stories become. The more objective our assertions about reality are, the less relevant they can be.
My answer to this is to make the arbitrary choice of perspective a first-class feature. If we can explicitly decide what meaning is relevant, we should be able to weakly solve natural language processing. It seems like a pretty simple and obvious idea, but so far is easier said than done.
Language has to do with our Monkeysphere, that humans over long periods of time in the past were limited to a very small subset of people they interacted with socially and closely. A few institutions likely had a large effect on your language like the church depending on where you were. After that it was the people you interacted with to stay alive. Because almost everything was in person or person to person transfer of information a lot of socially encoded clues were involved which lessened the need for well defined words.
Books were the first stage of homogenizing language as they could be shared over long distances and to many people, but more sequentially than latter forms of communication. After that radio and TV had a huge effect, for example the 'General American' used in broadcasts that was based heavily on a midwestern accent.
As we encroached on the 70s and 80s the previous technological advances and things like high speed interstates and trucking shrank America to something you could drive across in less than a week, and you could reach anywhere by voice nearly instantly. Suddenly people in California, Texas and New York all could be in the same meetings and local colloquialisms would need explained, so people would trend to a shared vocabulary.
It's also odd to me to say an LLM isn't subjective. Each LLM has it's own behavior, it's that there are like 20 or 30 big LLMs in all, and people are using them millions to billions of time so we're getting that one LLMs language everywhere. And that's why I disagree and will say natural language is computable, but it's also lossy and probabilistic. And for the most part it's single prompt and being ran by the user for the cheapest price possible.
https://news.ycombinator.com/pool
But not in this case. It's a slow Sunday afternoon and getting a few upvotes quickly is enough