DeepSeek-V4.1-Flash Launch Sparks Three Key Investment Opportunities, Says CITIC SEC

Stock News09-17 08:22

A research report from CITIC SEC highlights that the release of DeepSeek-V4.1-Flash brings simultaneous enhancements in coding and agent capabilities, alongside bolstered multimodal understanding. The model innovatively employs a CED architecture, integrating CSA2 cross-layer cache sharing, index reuse, and FP4 quantization, which cuts global KV Cache usage to roughly one-quarter of that in V4-Flash. These architectural refinements unlock room for inference cost reductions, as DeepSeek has announced further API price cuts. Domestic models are poised to compete along two trajectories—high cost-effectiveness and frontier intelligence—while DeepSeek's deep adaptation benefits both domestic computing infrastructure and local models, with cost-efficient inference accelerating the deployment of FDE and enterprise-grade agents.

CITIC SEC’s core perspectives are outlined below:

DeepSeek-V4.1-Flash: New CED architecture propels model capabilities past DeepSeek-V4-Pro. On September 10, 2026, DeepSeek unveiled V4.1-Flash, built on a new causal encoder-decoder (CED) framework. It supports multimodal understanding, handles up to one million tokens of context, and has opened both API access and model weights. The model's backbone contains 552 billion parameters, with 8 billion and 16 billion activated per token during prefilling and decoding stages, respectively, thereby lowering inference overhead for long-context and agent scenarios. According to DeepSeek's official WeChat account, given that V4.1-Flash outperforms V4-Pro across all metrics—including performance, cost, speed, and total runtime—DeepSeek plans to phase out the V4 Pro model gradually. However, as V4-Pro still sustains substantial call volumes and infrastructure migration requires time, it will remain in service until at least September 15.

Significant gains in coding and agent capabilities, with multimodal understanding broadening task execution scope. Per DeepSeek's WeChat updates: 1) In coding, at maximum reasoning effort, V4.1-Flash scores 90.6 on Terminal-Bench 2.1, surpassing Opus 5 (89.1) and GPT-5.6 Sol (88.8). On the more demanding Terminal-Bench 4.0, it achieves 31.2, trailing GPT-6 Astra (57.9) and Fable 5.1 (55.8). It also posts 74.2 on DeepSWE v1.1, nearly matching GPT-6 Astra (74.1) and exceeding Fable 5.1 (67.4). 2) For agent performance, it records 54.8 on AutomationBench and 31.8 on Agent’s Last Exam, both above Opus-5 (50.3, 28.6) and GPT-5.6 Sol (45.8, 26.7), reflecting improved execution in general automation and complex tasks following enhanced visual understanding.

Architectural innovation: CED and CSA2 jointly optimize compute paths and cache reuse, sustaining extreme cost reduction in long-context inference. 1) The novel CED architecture cuts global KV Cache footprint to about one-quarter of V4-Flash’s. The model’s 40-layer Transformer network is split into 20 causal-encoder layers and 20 decoder layers, with the encoder’s final layer directly projecting global KV for the decoder, eliminating the need for layer-by-layer hidden-state generation. Most input tokens thus only require encoder computation. To preserve local detail, the decoder additionally processes the final 128 tokens of the prompt, approximating sliding-window attention (SWA) caches, nearly halving prefill compute for long inputs. Per DeepSeek’s official technical report, prefill-phase per-token activation drops from 16 billion to 8 billion parameters. For cache precision, the global primary KV Cache is reduced from FP8 to FP4, with quantization-aware training introduced post-training to curb precision loss, nearly halving storage, while SWA KV retains FP8. Combined with CSA2’s cross-layer cache sharing, global KV Cache usage falls to roughly one-quarter of V4-Flash, concurrently reducing HBM demand to one-quarter and SSD demand to one-eighth. 2) CSA2 compresses cache occupancy and floating-point operations via hierarchical KV sharing and reuse of Top-K token retrieval results. The first two layers use only SWA; the rest employ compressed sparse attention (CSA2) with Full, Reindex, and Reuse modes. In Full mode, the model computes global primary KV and index keys, selecting Top-K historical positions. Reindex shares the previous layer’s Full-mode KV but recalculates relevance to update Top-K tokens. Reuse shares both KV and Top-K results, eliminating redundant cache generation and index computation. Specifically, the encoder allocates its remaining 18 layers (beyond 2 SWA) into groups of six, configured as "1 Full + 5 Reuse" each; the decoder’s 20 layers are split into five groups—the first being "1 Full + 3 Reuse" and the last four as "1 Reindex + 3 Reuse"—so the final 20 layers share the global KV from the first Full layer. This design merges KV Cache reduction with lower indexer compute. According to DeepSeek’s technical paper, extending context from 4K to 1M tokens raises per-token decoding FLOPs by only about 25%, indicating markedly improved long-context inference efficiency.

Inference optimization passes through to API price cuts, preserving the cost-effectiveness edge of domestic open-source models. DeepSeek’s official pricing sets V4.1-Flash per-million-token prices for off-peak cached-hit input, uncached input, and output at 0.02 yuan, 1 yuan, and 4 yuan, respectively—down 60.0%, 33.3%, and 11.1% from the previous 0.05 yuan, 1.5 yuan, and 4.5 yuan. Peak-hour rates are double the off-peak levels.

Application impact: The new model continues the extreme cost-reduction direction, benefiting FDE and enterprise-grade agent deployment. With simultaneous gains in coding, agent, and multimodal capabilities, alongside lower call prices, the model is expected to ease the cost burden of multi-round tool invocations and long-context processing for enterprise agents, accelerating FDE adoption. Meanwhile, software firms with deep industry expertise, tight integration into enterprise workflows, proprietary vertical data moats, or the ability to deliver reliable results in highly regulated settings are likely to translate model capabilities into product value and commercial growth first.

Investment strategy: Three key investment themes are recommended. 1) AI infrastructure: DeepSeek’s deep adaptation to domestic computing power aligns domestic compute with local models. 2) AI applications: The model’s ongoing open-source strategy and further inference efficiency gains favor FDE and software firms with competitive barriers. 3) Model developers: On one hand, architectural innovation continues to unlock inference efficiency and cost savings. Within the past two months, DeepSeek-V4.1-Flash, GLM-5.3-Flash, and Qwen3.8-Flash have all used architecture optimization to slash long-context inference overhead, validating further cost-reduction potential. On the other hand, domestic models will compete along high cost-effectiveness and frontier intelligence paths. Flash-series models, with lower activated parameters and call costs, cover routine coding, office tasks, and general agent needs, while large-parameter flagship models are expected to push the ceiling on complex reasoning, specialized knowledge, and long-horizon tasks.

Risk factors: Underperformance in core AI technology development or application expansion; slower-than-expected compute cost reduction; severe social impact from AI misuse; data security risks; information security risks; and intensifying industry competition.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment