







Evaluates typed decisions (choice, score, noul) over 100+ languages in a single forward pass with calibrated probabilities. Outperforms TypeSafe Jev.
GitHub - AbdelStark/jev-benchmarks at 0d610cc53e79bcbec691312b0c4adb4a0e371642
Probability-aware evaluation for typed decision models: calibration, selective risk, latency, and reproducible benchmarks. - AbdelStark/jev-benchmarks

Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via...
Despite their impressive capabilities, multilingual large language models (MLLMs) frequently exhibit inconsistent behaviour when the prompt's language changes. While such adaptation is generally...

chad/whichlang
What programming language do LLMs default to when you don't tell them? A small benchmark.
Targeted Multilingual Adaptation for Low-resource Language Families
The "massively-multilingual" training of multilingual models is known to limit their utility in any one language, and they perform particularly poorly on low-resource languages. However, there is evid
System One: fast judgments and deliberate checks
Explore System One thinking, its limits, and how fast judgments can be used in software. Interactive examples, with Jev as a case study.
Jev introduces a new shape of LLM—System One, aka Decision Models
Last week TypeSafe AI unveiled Jev, their first example of a new category of model that they are calling “System One models” (I’m with Maggie Appleton, I think “decision models” …
How Large Language Models Actually Work
Computation and its Connotations
A Review of Language Machines by Leif Weatherby

ChatGPT Translate | Fast, Natural, 40+ Languages
ChatGPT translates across 40+ languages with accuracy, tone, and cultural nuance. Translate text, voice, or photos for everyday use, travel, school, and work — and learn grammar or phrasing as you go.

Please
Please is a cross-language build system with an emphasis on high performance, portability, extensibility and correctness.

On-Device LLM Throughput Calculator - a Hugging Face Space by FL33TW00D-HF
This tool estimates and visualizes the throughput of Large Language Models on devices with memory bandwidth constraints. Users input device and model configurations, and the tool generates a plot s...
inanna-malick/jev-dsl
Agent-first Haskell DSL for TypeSafe's Jev judgment model: typed packets, inferred types, answers under the same labels

Best LLM for Coding 2026 | AI Coding Model Rankings & Benchmarks
Which AI model writes the best code? We rank every major LLM — open and closed source — across SWE-bench, HumanEval, LiveCodeBench, and Terminal-Bench coding benchmarks. Compare the best LLMs for coding, software engineering, and programming.

Are you Slop??
Are you Slop? Analyze your 'slop score' - a measure of how generic or unique your text appears to language models.