







English is a West Germanic language of the Indo-European language family that emerged in early medieval England and has since become a global lingua franca. The language is named after the Angles, one of the Germanic peoples who migrated to Britain after the end of Roman rule. English is the most spoken language in the world, primarily due to the global influence of the British Empire and the rise of the United States. It is the most widely learned second language in the world, with more second‑language speakers than native speakers. However, English is only the third‑most spoken native language, after Mandarin Chinese and Spanish.
1,000 most common Dutch words – Understanding Dutch
The 1000 most common words in the Dutch language, with English translations
言語データベースとソフトウェア - 言語データベースとソフトウェア
How linguistics learned to stop worrying and love the language models
Language models (LMs) can produce fluent, grammatical text. Nonetheless, some maintain that language models don’t really learn language and also, even if they did, that would not be informative for the study of human learning and processing. On the other side, there have been claims that the success of LMs obviates the need for studying linguistic theory and structure. We argue that both extremes are wrong. LMs can contribute to fundamental questions about linguistic structure, language processing, and learning. They force us to rethink arguments and ways of thinking that have been foundational in linguistics. While they do not replace linguistic structure and theory, they serve as model systems and working proofs of concept for gradient, usage-based approaches to language. We offer an optimistic take on the relationship between language models and linguistics.

Leshem (Legend) Choshen 🤖🤗 on Twitter / X
We know models fail to generalize across languages.But Adam found they fail to generalize when you just call half your English pretraining by another name, like "En2"😵💫 https://t.co/DfWJRP51Fz— Leshem (Legend) Choshen 🤖🤗 (@LChoshen) April 30, 2026
Welcome | Yale Grammatical Diversity Project: English in North America
This Project explores syntactic diversity found in varieties of English spoken in North America. By documenting the subtle, but systematic, differences in the syntax of English varieties, it provides a crucial source of data for the development of theories of human linguistic knowledge.
Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
LLMs have been showing limitations when it comes to cultural coverage and competence, and in some cases show regional biases such as amplifying Western and Anglocentric viewpoints. While there have been works analysing the cultural capabilities of LLMs, there has not been specific work on highlighting LLM regional preferences when it comes to cultural-related questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ). The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan. Moveover, our results show that when prompting in languages such as English or other high-resource ones, LLMs tend to provide more diverse outputs and show less inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised fine-tuning, and not during pre-training.

Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
LLMs have been showing limitations when it comes to cultural coverage and competence, and in some cases show regional biases such as amplifying Western and Anglocentric viewpoints. While there have been works analysing the cultural capabilities of LLMs, there has not been specific work on highlighting LLM regional preferences when it comes to cultural-related questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ). The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan. Moveover, our results show that when prompting in languages such as English or other high-resource ones, LLMs tend to provide more diverse outputs and show less inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised fine-tuning, and not during pre-training.

Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via...
Despite their impressive capabilities, multilingual large language models (MLLMs) frequently exhibit inconsistent behaviour when the prompt's language changes. While such adaptation is generally...

Best Language School in Massachusetts | Rola Languages
Rola Languages is one of the best language institutes in Boston. Offering small group and private intro-advanced classes in Spanish, French, Portuguese, Italian, Mandarin, and English in-person and online.
ChatGPT Translate | Fast, Natural, 40+ Languages
ChatGPT translates across 40+ languages with accuracy, tone, and cultural nuance. Translate text, voice, or photos for everyday use, travel, school, and work — and learn grammar or phrasing as you go.

Improving Cross-Lingual Factual Recall via Consistency-Driven Reinforcement Learning
Large language models (LLMs) trained predominantly on English data encode substantial world knowledge, yet often fail to express it reliably in other languages, a phenomenon known as cross-lingual...
Large language models are cultural technologies. What might that mean?
Four different perspectives

Stephen Krashen: Language Acquisition and Comprehensible Input
Rethinking the Multilingual Reasoning Gap with Layer Swap
Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior work suggests that forcing the CoT to remain in the input language (\emph{native reasoning}) substantially degrades performance relative to allowing the model to reason in English before answering in the input language (\emph{English-pivoted reasoning}). However, most studies of this native reasoning gap rely on inference-time interventions or limited native-language training data. We revisit this comparison at a larger scale and under comparable supervision. We construct long multilingual reasoning datasets across six languages (English, French, German, Spanish, Chinese and Swahili); fine-tune specialists in both native and English-pivoted regimes on top of \texttt{Qwen/Qwen3-8B-Base}, and evaluate across mathematics, science, general knowledge, and code. In this setting, the average native reasoning gap shrinks to 1.9--3.5\% across the five non-English languages, considerably smaller than previously reported. Weight-space analysis of the native specialists reveals aligned fine-tuning updates in the middle layers and divergence in the outer layers. This points to a largely language-agnostic reasoning core surrounded by language-specific layers. Exploiting this structure, we introduce a Layer Swap: transferring the English specialist's stronger reasoning mid-layers into each native specialist, closing most of the native reasoning gap across the five non-English languages while preserving CoT in the target language. We release all models and datasets.

Mango Languages | Vancouver Public Library
Over 70 self-paced language learning courses and dialects. Learners are introduced to vocabulary, cultural insights, grammar and more. The courses are delivered by native speakers. They include critical-thinking and memory-building exercises. Sign up for a Mango profile so you can track your progress & continue where you left off.Or use “Mango as a guest” to access the resource. New to Mango? View a video tutorial.
When I worked in software dev, I was taught to check three non-English languages when testing text strings: German to test the largest possible version of a string, Japanese or Chinese to test the shortest possible version, and Thai to check the TALLEST possible version.