







When I worked in software dev, I was taught to check three non-English languages when testing text strings: German to test the largest possible version of a string, Japanese or Chinese to test the shortest possible version, and Thai to check the TALLEST possible version.
Mar 30, 2026 at 5:01 AM

ChatGPT Translate | Fast, Natural, 40+ Languages
ChatGPT translates across 40+ languages with accuracy, tone, and cultural nuance. Translate text, voice, or photos for everyday use, travel, school, and work — and learn grammar or phrasing as you go.

Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via...
Despite their impressive capabilities, multilingual large language models (MLLMs) frequently exhibit inconsistent behaviour when the prompt's language changes. While such adaptation is generally...

How I Learned Spanish in 3 Months | Loïs Talagrand
Complete breakdown of my 90-day Spanish challenge: constraints, tools, timeline, and final 30-minute conversation result.
Please
Please is a cross-language build system with an emphasis on high performance, portability, extensibility and correctness.

Welcome | Yale Grammatical Diversity Project: English in North America
This Project explores syntactic diversity found in varieties of English spoken in North America. By documenting the subtle, but systematic, differences in the syntax of English varieties, it provides a crucial source of data for the development of theories of human linguistic knowledge.
Cross-Lingual Exploration for Parametric Knowledge
Parametric knowledge in Large Language Models is not equally accessible across languages. As a result, standard inference techniques often struggle to surface localized facts, leading to failures in...
chad/whichlang
What programming language do LLMs default to when you don't tell them? A small benchmark.
Rethinking the Multilingual Reasoning Gap with Layer Swap
Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior work suggests that forcing the CoT to remain in the input language (\emph{native reasoning}) substantially degrades performance relative to allowing the model to reason in English before answering in the input language (\emph{English-pivoted reasoning}). However, most studies of this native reasoning gap rely on inference-time interventions or limited native-language training data. We revisit this comparison at a larger scale and under comparable supervision. We construct long multilingual reasoning datasets across six languages (English, French, German, Spanish, Chinese and Swahili); fine-tune specialists in both native and English-pivoted regimes on top of \texttt{Qwen/Qwen3-8B-Base}, and evaluate across mathematics, science, general knowledge, and code. In this setting, the average native reasoning gap shrinks to 1.9--3.5\% across the five non-English languages, considerably smaller than previously reported. Weight-space analysis of the native specialists reveals aligned fine-tuning updates in the middle layers and divergence in the outer layers. This points to a largely language-agnostic reasoning core surrounded by language-specific layers. Exploiting this structure, we introduce a Layer Swap: transferring the English specialist's stronger reasoning mid-layers into each native specialist, closing most of the native reasoning gap across the five non-English languages while preserving CoT in the target language. We release all models and datasets.

English language
English is a West Germanic language in the Indo-European language family. It emerged in early medieval England and has since become a global lingua franca. The namesake of the language is the Angles, one of the Germanic peoples who migrated to Britain after the end of Roman rule. English is the most spoken language in the world, primarily due to the global influences of the former British Empire and the United States. It is the most widely learned second language in the world, with more second-language speakers than native speakers. However, English is only the third-most spoken native language, after Mandarin Chinese and Spanish.
Are you Slop??
Are you Slop? Analyze your 'slop score' - a measure of how generic or unique your text appears to language models.
How Large Language Models Actually Work
Targeted Multilingual Adaptation for Low-resource Language Families
The "massively-multilingual" training of multilingual models is known to limit their utility in any one language, and they perform particularly poorly on low-resource languages. However, there is evid
The voice safety classifier supports 30 languages and checks for profanity, asking for PII, harassment, sexual content, dating and romantic, discriminatory speech, illegal and regulated content, and disruptive audio. github.com/roostorg/model-community/tree… You can test it out here! huggingface.co/spaces/Roblox/voice-safety-cl…
model-community/roblox-voice-safety-classifier at main · roostorg/model-community
github.com