







What matters about DeepSeek’s V4 Flash 0731 is what did not change. Same architecture and parameter scale. DeepSeek says only post-training changed. The app, web model and V4 Pro API did not. Yet agent benchmarks moved sharply. This is not a scale story.
Jul 31, 2026 at 8:06 PM
DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, 10 points above previous DeepSeek V4 Flash
DeepSeek releases DeepSeek V4 Flash 0731

alphaXiv on Twitter / X
"Small-Scale Experiments: Are We There Yet?"Scaling laws show up even in tiny 4M parameter models, but only when their hyperparameters are tuned well enough.This Meta paper found that small models are much more hyperparameter-sensitive. As scale increases, the hyperparameter… pic.twitter.com/vQfFS1TpMA— alphaXiv (@askalphaxiv) August 15, 2026

228. DeepSeek Has Been Inevitable and Here's Why (History Tells Us)
DeepSeek was certain to happen. The only unknown was who was going to do it. The choices were a startup or someone outside the current center of US AI leadership and innovation.

Deep Learning Scaling is Predictable, Empirically
Deep learning (DL) creates impactful advances following a virtuous recipe: model architecture search, creating large training data sets, and scaling computation. It is widely believed that growing training sets and models should improve accuracy and result in better products. As DL application domains grow, we would like a deeper understanding of the relationships between training set size, computational scale, and model accuracy improvements to advance the state-of-the-art. This paper presents a large scale empirical characterization of generalization error and model size growth as training sets grow. We introduce a methodology for this measurement and test four machine learning domains: machine translation, language modeling, image processing, and speech recognition. Our empirical results show power-law generalization error scaling across a breadth of factors, resulting in power-law exponents---the "steepness" of the learning curve---yet to be explained by theoretical work. Further, model improvements only shift the error but do not appear to affect the power-law exponent. We also show that model size scales sublinearly with data size. These scaling relationships have significant implications on deep learning research, practice, and systems. They can assist model debugging, setting accuracy targets, and decisions about data set growth. They can also guide computing system design and underscore the importance of continued computational scaling.

DeepSeek FAQ
DeepSeek has completely upended people’s expectations for AI and competition with China. What is it, and why does it matter?
DeepSeek on Twitter / X
🧩 DeepSeek Harness v0.1 is now available in Developer Preview!🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license.🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one…— DeepSeek (@deepseek_ai) August 13, 2026
The Shape of Memory Benchmarks
Why the familiar memory benchmarks are outdated, how the agent-native work looks today and why design your own.

Neural scaling law
In machine learning, a neural scaling law is an empirical scaling law that describes how neural network performance changes as key factors are scaled up or down. These factors typically include the number of parameters, training dataset size, and training cost. Some models also exhibit performance gains by scaling inference through increased test-time compute (TTC), extending neural scaling laws beyond training to the deployment phase.

The Big LLM Architecture Comparison
From DeepSeek-V3 to Kimi K2: A Look At Modern LLM Architecture Design

LLMs Can Now Write GPU Kernels That Beat torch.compile - Break AI Scaling Limits in 7 Days
We're now seeing multi-agent systems that take your PyTorch code and produce CUDA or Triton kernels with 2x to 14x speedups over torch.compile(mode='max-autotune-no-cudagraphs'). Not on toy benchmarks. On real models like Llama-3.1-8B, Whisper, and Stable Diffusion. Learn proven techniques to shift the scaling law intercept and achieve 10-50% performance gains.

GPT-5: The Reverse DeepSeek Moment
Everyone agrees that the release of GPT-5 was botched. Everyone can also agree that the direct jump from GPT-4o and o3 to GPT-5 was not of similar size to the jump from GPT-3 to GPT-4, that it was …

DeepSeek Harness developer preview: Everything is a plugin
DeepSeek Harness is now available in developer preview to developers building agent harnesses worldwide, with the source code released at the same time. Every agent capability is implemented as a plugin that can be swapped or recomposed.
LLM Leaderboard - Comparison of over 100 AI models from OpenAI, Google, DeepSeek & others
Comparison and ranking the performance of over 100 AI models (LLMs) across key metrics including intelligence, price, performance and speed (output speed - tokens per second & latency - TTFT), context window & others.

Scaling Laws Across Model Architectures: A Comparative Analysis of...
The scaling of large language models (LLMs) is a critical research area for the efficiency and effectiveness of model training and deployment. Our work investigates the transferability and...

More US firms turn to China’s DeepSeek over pricey Silicon Valley AI
DeepSeek takes top spot on ‘trending’ list as companies look for alternatives to OpenAI and Anthropic, spending tracker’s report says.

Microsoft eyes DeepSeek for enterprise AI
Microsoft will also shift to usage-based pricing for the enterprise agent.
