DeepSeek launches highly efficient V4.1-Flash model for AI agents
DeepSeek's new V4.1-Flash model reduces memory strain by keeping only 16 billion of its 552 billion parameters active per token. The update cuts KV cache needs to a quarter of its predecessor.