







Complete breakdown of my 90-day Spanish challenge: constraints, tools, timeline, and final 30-minute conversation result.
Online Spanish Classes San Diego | Culture & Language Center
Learn Spanish and gain confidence through real conversations and tailored online Spanish classes with friendly, certified teachers from Latin America. Visit us!

Dreaming Spanish – Learn with Comprehensible Input
Master Spanish with comprehensible input. Dreaming Spanish gives you real, engaging content at every level to build lasting fluency naturally.

Introduction
A multigrade textbook of Spanish through Caribbean music: cumbia, reggaeton, salsa.

OH NOAH! | Learn Spanish with OH NOAH! | PBS KIDS
Stanford CS336 | Language Modeling from Scratch (Spring 2025 Archive)
Archived course website for Stanford CS336: Language Modeling from Scratch (Spring 2025), including schedule, assignments, logistics, and materials.

REGOSH – Rede Latino Americana de Tecnologias Livres
Join the conversation at the reGOSH forum: introduce yourself, tell us about your projects, ask questions, collaborate with others
Stephen Krashen: Language Acquisition and Comprehensible Input
Mango Languages | Vancouver Public Library
Over 70 self-paced language learning courses and dialects. Learners are introduced to vocabulary, cultural insights, grammar and more. The courses are delivered by native speakers. They include critical-thinking and memory-building exercises. Sign up for a Mango profile so you can track your progress & continue where you left off.Or use “Mango as a guest” to access the resource. New to Mango? View a video tutorial.
Extracting Training Data from Large Language Models
Nicholas Carlini, Google; Florian Tramèr, Stanford University; Eric Wallace, UC Berkeley; Matthew Jagielski, Northeastern University; Ariel Herbert-Voss, OpenAI and Harvard University; Katherine Lee and Adam Roberts, Google; Tom Brown, OpenAI; Dawn Song, UC Berkeley; Úlfar Erlingsson, Apple; Alina Oprea, Northeastern University; Colin Raffel, Google

Welcome | Yale Grammatical Diversity Project: English in North America
This Project explores syntactic diversity found in varieties of English spoken in North America. By documenting the subtle, but systematic, differences in the syntax of English varieties, it provides a crucial source of data for the development of theories of human linguistic knowledge.
MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a single framework. As these models are rapidly evolving toward general-purpose instruction following across diverse and complex tasks, a key frontier is evaluating their crosslingual and multimodal capabilities over both short- and long-form inputs. However, existing benchmarks fall short in evaluating these dimensions jointly: they are often limited to English, mostly focus on a single modality at a time, rely on short-form inputs, or lack human annotations--hindering comprehensive assessment of model performance across languages, modalities, and task complexity. To address these gaps, we introduce MCIF (Multimodal Crosslingual Instruction Following), the first crosslingual human-annotated benchmark based on scientific talks on NLP and beyond. MCIF evaluates instruction following in crosslingual, multimodal settings over different input lengths and spans four macro-tasks: recognition, translation, question answering, and summarization. It covers three core modalities (speech, vision, and text) and four diverse languages (English, German, Italian, and Chinese), fully aligned across all dimensions. This parallel design enables a systematic evaluation of MLLMs' abilities to interpret instructions across languages and effectively integrate multimodal contextual information. Our benchmarking and analysis of 23 models highlight universal challenges across modalities and tasks, indicating substantial room for improvement in future MLLMs development. MCIF is released under CC-BY 4.0 license to promote open research.

Improving Cross-Lingual Factual Recall via Consistency-Driven Reinforcement Learning
Large language models (LLMs) trained predominantly on English data encode substantial world knowledge, yet often fail to express it reliably in other languages, a phenomenon known as cross-lingual...
How Large Language Models Actually Work
Rethinking the Multilingual Reasoning Gap with Layer Swap
Recent reasoning Large Language Models produce a chain-of-thought (CoT) predominantly in English, even when prompted in non-English languages. Prior work suggests that forcing the CoT to remain in the input language (\emph{native reasoning}) substantially degrades performance relative to allowing the model to reason in English before answering in the input language (\emph{English-pivoted reasoning}). However, most studies of this native reasoning gap rely on inference-time interventions or limited native-language training data. We revisit this comparison at a larger scale and under comparable supervision. We construct long multilingual reasoning datasets across six languages (English, French, German, Spanish, Chinese and Swahili); fine-tune specialists in both native and English-pivoted regimes on top of \texttt{Qwen/Qwen3-8B-Base}, and evaluate across mathematics, science, general knowledge, and code. In this setting, the average native reasoning gap shrinks to 1.9--3.5\% across the five non-English languages, considerably smaller than previously reported. Weight-space analysis of the native specialists reveals aligned fine-tuning updates in the middle layers and divergence in the outer layers. This points to a largely language-agnostic reasoning core surrounded by language-specific layers. Exploiting this structure, we introduce a Layer Swap: transferring the English specialist's stronger reasoning mid-layers into each native specialist, closing most of the native reasoning gap across the five non-English languages while preserving CoT in the target language. We release all models and datasets.

Where is <i>Mi Gente</i> ? Codeswitching (Afro)Latinidad in the music classroom
Musical codeswitching (CS) entails mixing musical ideas and genres. The term CS originated in linguistics, based on language alternations attested in bilinguals. Switches in sociocultural behaviors also now receive scholarly attention as CS. The current multiple case study explores CS domains (music, language, and behavior) in the context of music education programs, where CS remains under-researched. This study also fills a gap by examining historically underrepresented individuals’ (HURIs) participation in music education. Here, a CS-based account provides a deeper understanding of the complex sociocultural capital, linguistic resources and lived experiences that HURIs navigate. As part of an interpretive qualitative study design, semi-structured interviews were carried out with a HURI subpopulation (bilingual, [Afro]Latina/o/x faculty, and students) in music education. Findings show participants perceive CS to be mandatory for accessing dominant U.S. music school culture. Additional findings reveal HURIs must master CS in musical, linguistic, and behavioral domains to avoid negative outcomes, yet sustained multi-CS scenarios may have psychological and even physical costs. Insights from CS are thus critical for pinpointing institutional barriers to greater HURI involvement in music education.
