On September 23, during the 2026 Snapdragon Summit, Qualcomm Technologies announced a four-party collaboration with StepFun, Wulianghuo, and Longsys to jointly optimize and adapt the StepEdge-Omni 30B-MoE model for on-device deployment. The partnership leverages the 6th generation Snapdragon 8 Super Edition mobile platform, and live demonstrations showcased an on-device intelligent agent capable of autonomous service and personalization.
The 30B-MoE model runs entirely on a smartphone without cloud calls, enabling tasks like email comprehension, itinerary planning, calendar synchronization, flight and hotel recommendations, schedule sharing, and draft emails. Through coordinated optimization of the inference engine, heterogeneous scheduling across CPU/GPU/NPU, and storage solutions, memory requirements for model runtime are reduced by over 50%. Live tests achieved prefill throughput exceeding 330 tokens per second and decode throughput above 28 tokens per second.
In this effort, Longsys contributes its on-device storage expertise by enabling storage-compute synergy, loading model weights from flash memory to reduce persistent runtime memory usage. This storage support allows a 30B-class large model to operate efficiently on a handset. Longsys stated that large-scale on-device AI deployment depends on deep coordination between compute and storage, and the company will continue driving innovation in on-device AI storage solutions.
Analysts note the collaboration underscores the growing importance of storage within the on-device AI supply chain. As AI phones evolve from simple Q&A to autonomous planning and multi-step task execution agents, demand for high-speed, high-capacity, and low-power storage is set to rise. Ecosystem collaboration among storage vendors, chipmakers, model providers, and device manufacturers is expected to deepen further.
Comments