







The following 200 pages are in this category, out of approximately 5,668 total. This list may not reflect recent changes.
OKA - Wiki pages tracker (oka.wiki/tracker)
Category:WikiProject lists of reliable sources
The following 61 pages are in this category, out of 61 total. This list may not reflect recent changes.
OKA – Disseminating free content on Wikipedia and open platforms through targeted funding
We are a non-profit organization dedicated to improving Wikipedia and other open platforms. We do so by providing monthly stipends to full-time contributors and translators. We leverage AI (Large Language Models) to automate most of the work.
AI Translations Are Adding ‘Hallucinations’ to Wikipedia Articles
AI translated articles swapped sources or added unsourced sentences with no explanation, while others added paragraphs sourced from completely unrelated material.
APL Wiki
498 articles about APL that anyone can edit. See the navigational overview of content.
Evaluating Multilingual Metadata Quality in Crossref
Introduction: Scholarly research spans multiple languages, making multilingual metadata crucial for organizing and accessing knowledge across linguistic boundaries. These multilingual metadata already exist and are propagated throughout scholarly publishing infrastructure, but the extent to which they are correctly recorded, or how they affect metadata quality more broadly is little understood. Methods: Our study quantifies the prevalence of multilingual records across a sample of publisher metadata and offers an understanding of their completeness, quality, and alignment with metadata standards. Utilizing the Crossref API to generate a random sample of 519,665 journal article records, we categorize each record into four distinct language types: English monolingual, non-English monolingual, multilingual, and uncategorized. We then investigate the prevalence of programmatically-detectable errors and the prevalence of multilingual records within the sample to determine whether multilingualism influences the quality of article metadata. Results: We find that English-only records are still in the vast majority among metadata found in Crossref, but that, while non-English and multilingual records present unique challenges, they are not a source of significant metadata quality issues and, in few instances, are more complete or correct than English monolingual records. Discussion & Conclusion: Our findings contribute to discussions surrounding multilingualism in scholarly communication, serving as a resource for researchers, publishers, and information professionals seeking to enhance the global dissemination of knowledge and foster inclusivity in the academic landscape.

Wikidata
the free knowledge base with 120,900,955 data entities that anyone can edit.
Wikipedia:Writing articles with large language models
Text generated by large language models (LLMs)[a] often violates several of Wikipedia's core content policies. For this reason, the use of LLMs to generate or rewrite article content is prohibited,[b] save for these two exceptions:
How To Argue With An AI Booster
Editor's Note: For those of you reading via email, I recommend opening this in a browser so you can use the Table of Contents. This is my longest newsletter - a 16,000-word-long opus - and if you like it, please subscribe to my premium newsletter. Thanks for reading! In

Wikipedia Bans AI-Generated Content
“In recent months, more and more administrative reports centered on LLM-related issues, and editors were being overwhelmed.”
Introducing Oktana
The introductory text for our newly formed collective, Oktana. Read about why we came together, our critique and vision for work and technology and our desires for the present and future of our initiative.

Wiki-40B: Multilingual Language Model Dataset
We propose a new multilingual language model benchmark that is composed of 40+ languages spanning several scripts and linguistic families. With around 40 billion characters, we hope this new resource will accelerate the research of multilingual modeling. We train monolingual causal language models using a state-of-the-art model (Transformer-XL) establishing baselines for many languages. We also introduce the task of multilingual causal language modeling where we train our model on the combined text of 40+ languages from Wikipedia with different vocabulary sizes and evaluate on the languages individually. We released the cleaned-up text of 40+ Wikipedia language editions, the corresponding trained monolingual language models, and several multilingual language models with different fixed vocabulary sizes.
Google's OKF - The New Way to Structure Your Knowledge for Agents