The chip for AI at the device edge is a crucial link connecting model capabilities with end products.
This year at the World Artificial Intelligence Conference (WAIC) exhibition halls, an increasing number of AI models are being integrated into physical devices. Some are housed in desktop-sized Agent hosts, others are embedded in AI PCs and mobile workstations, and some are concealed within holographic projection boxes, powering digital humans capable of sustained conversation. The dense appearance of AI devices suggests that the explosion of edge AI is imminent.
This shift indeed possesses significant technical prerequisites. The training methodologies and capabilities accumulated by large cloud models are now being transferred to models with smaller parameter scales, enabling more and more tasks to run locally on devices. As model capabilities enter the range suitable for on-device use, the primary challenge for the edge AI industry also changes. Previously, many projects were hindered by models being "not smart enough"; now, the issue is shifting towards whether the hardware can effectively accommodate these capabilities. Models can be compressed, but the space, power consumption, thermal management, and cost of devices have limits. Thus, the edge AI chip becomes the critical link connecting model capabilities with end products.
While cloud chips pursue the upper limit of computing power, edge AI chips confront the boundaries of a specific product. They must withstand the task pressure of large model inference, speech recognition, speech synthesis, visual processing, long-term memory, and multi-Agent concurrency, while also controlling temperature, noise, and overall device cost during operation for hours or even all day.
Key Considerations for Edge AI Competition
The competition in edge AI, on the surface, compares computing power parameters. However, the real differentiator is who can deliver more usable intelligence within limited power, space, and cost, and who can stably integrate this intelligence into a mass-producible and deliverable device through mature toolchains and engineering capabilities. Whoever truly enables AI devices to move beyond demos and into the physical world will gain dominance in edge AI.
The Shift from Cloud Reliance to On-Device Intelligence
Initially, the most common way for AI to enter personal devices was by adding a dialog box to existing products. Users would open an app, input a question, the device would send the request to the cloud, and then display the model-generated answer. While AI capabilities appeared on the screen, the actual computation and memory primarily occurred on cloud servers.
However, since the beginning of this year, device manufacturers have begun showcasing another form. At GTC Taipei in June, NVIDIA (NVDA) unveiled RTX Spark for Windows PCs. According to NVIDIA's vision, future personal computers will no longer just wait for users to click apps or input commands; they will also be able to run personal agents locally: reading local files, completing tasks across different software, generating images and videos, writing plugins, and continuously processing entire workflows assigned by the user.
To keep these tasks within personal devices, RTX Spark offers up to 1 Petaflop of AI performance and 128GB of unified memory, enabling local Agent operation of large language models with 120B parameters and million-token contexts. Manufacturers including ASUS, Dell, HP, Lenovo, Microsoft Surface, and MSI also plan to launch thin and light notebooks and compact desktop hosts based on this platform. NVIDIA founder and CEO Jensen Huang stated, "The PC is being redefined."
Similar changes are occurring in other device forms. In March, Lenovo introduced a concept projector, AI Workmate, at MWC 2026, touted as an always-on desktop AI assistant capable of locally processing voice control, gesture operations, etc., via on-device AI models. Equipped with a camera, screen, and projector, it can scan a paper document, understand its content, organize it into notes or presentations, and project it directly onto a wall.
Also at MWC 2026, Qualcomm pushed edge AI into even smaller devices. Its released Snapdragon Wear Elite targets smartwatches, smart glasses, and AI pendants, aiming to complete natural voice interaction, context awareness, and real-time keyword detection locally via a dedicated NPU and low-power AI unit. For these all-day wearable devices, AI must operate continuously under extremely limited battery, space, and temperature conditions.
At WAIC 2026, numerous AI device demonstrations were also seen: mobile workstations began processing larger models and professional creation tasks locally; desktop AI hosts brought model inference and Agent development into personal workspaces; companion devices and robots required continuous reception of audio, visual, and environmental information while retaining long-term states about users and tasks.
These devices appear vastly different, but they are doing the same thing: keeping parts of the perception, memory, reasoning, and execution capabilities, originally concentrated in the cloud, within the device itself.
Globally, AI devices with local computing power are also rapidly proliferating. Gartner research data shows that AI PC shipments will reach approximately 143 million units in 2026, accounting for 55% of the global PC market. A Counterpoint report released in June also predicts that smartphones with generative AI capabilities will constitute 45% of global smartphone shipments.
The reasons behind the popularity of AI devices are becoming increasingly specific. A simple Q&A can tolerate data upload to the cloud and waiting for an answer to return; however, a long-running Agent may continuously read local documents, call applications, check execution results, and re-plan. If every link relies on the cloud, latency, network stability, data privacy, and token costs accumulate along with the task chain, potentially leading to poor user experience or privacy risks.
Especially when Agents begin entering personal computers and home devices, they no longer face abstract public knowledge but rather user files, photos, meeting notes, account permissions, and long-term memory. Whether data can remain within the device and what operations an Agent can perform have become part of device product design.
On-device computing thus gains new significance. For enterprise and government clients, data compliance is often a prerequisite for deployment decisions; for home and companion products, camera footage, voice recordings, and personal memories are inherently sensitive; for Agents requiring high-frequency calls, the long-term cost of paying for cloud tokens can gradually become a significant expense.
However, this does not mean cloud models will exit. Complex, low-frequency tasks requiring stronger reasoning capabilities are still suitable for the cloud, while local devices handle high-frequency, low-latency, privacy-sensitive, and continuously running tasks. What edge AI truly promotes is a re-division of labor between the cloud and devices.
Moving Beyond the Booth: The Three Accounts of Edge AI
At the WAIC venue, when discussing with multiple AI device manufacturers, one term was repeatedly mentioned: "accounting." These teams mostly have experience in consumer electronics or digital hardware R&D and production, already accustomed to repeatedly weighing factors like chips, memory, thermal management, battery life, and overall device cost. For them, edge AI must first pass the same hurdle: whether a model can run is just the starting point; what truly determines if a product is viable is whether performance, power consumption, size, and price can be balanced simultaneously.
AI device manufacturers first need to account for capability. An AI device completing a single Q&A in an ideal environment versus operating stably in a real device are two completely different challenges. A device needs to simultaneously support language models, speech recognition, speech synthesis, visual processing, long-term memory, and multiple Agents. A person in charge of an edge AI device manufacturer mentioned encountering performance degradation issues due to multi-Agent concurrency. When a single Agent was running, model inference speed and interaction experience met expectations; but when multiple Agents simultaneously called model, audio, and memory modules, different tasks began competing for memory bandwidth, causing latency to rise.
Peak computing power can indicate what a chip can achieve under ideal conditions but often fails to fully reflect the real experience during complex tasks. Model scale, first-token latency, concurrency numbers, context length, and sustained performance are now entering device manufacturers' evaluation lists alongside TOPS.
The second account is engineering. AI PCs must find balance between performance, battery life, and weight; desktop hosts need to control heat dissipation and noise; mobile workstations face larger models and longer sustained loads; robots and companion devices also need to handle all-day standby, voice wake-up, and multimodal perception. The same chip and model placed in different devices yield completely different results.
Model deployment itself is also a lengthy process. Chips need to support different frameworks and operators; toolchains must complete model conversion, quantization, and compilation; systems need to be compatible with Windows, Linux, Android, and various domestic operating systems. Upon entering mass production, stability, upgrades, and long-term maintenance become new tasks. Many edge AI teams do not lack demo capability; what is truly scarce is the ability to turn a one-time demonstration into replicable, deliverable capability.
The third account is commercial. Edge large models typically mean larger memory, higher-spec chips, and more complex thermal design. Especially against the backdrop of rising memory prices this year, the hardware cost of some products may increase significantly. Even if a technical solution can run, if the added cost cannot translate into value perceived by users, the product still struggles to succeed.
As edge AI enters the product phase, the unit of competition has shifted from "a model" to "a complete device." Whether a model can run is just the starting point; whether the device can work long-term, be launched on schedule, and whether users are willing to buy it collectively determine whether a technological attempt has commercial viability.
As bottlenecks in the edge AI industry shift towards hardware, a key question emerges: what kind of chip can truly carry model capabilities and be integrated into AI device products within limited power, space, and cost?
The Chip Capable of Supporting AI Device Evolution
When an edge AI chip completes tape-out and release, it indicates the technical path is preliminarily established; entering different devices, adapting to different models, and accompanying clients through mass production marks the beginning of commercial implementation.
Xin Xiaoxu, Co-founder and Product Vice President of Haimo Intelligence, a large model edge AI chip manufacturer, is also paying attention to AI device products at WAIC. Over the past few years, Haimo Intelligence has also set up booths at WAIC. The change he perceives is that last year, over half of the clients approaching Haimo Intelligence were just "wanting to try." But this year, more clients come directly with model lists, asking simple and direct questions: Can these models run? What's the performance? How much does it cost?
"The proportion of clients genuinely preparing to integrate large models into edge products is rising rapidly, originally less than 50%, now probably around 80%," Xin Xiaoxu stated. While this change may not represent the entire market, it provides a cross-section from the upstream industry: edge AI clients are reducing open-ended questions and starting to discuss delivery and cost.
In the WAIC exhibition area, Liaoqu Intelligence's AI Personal Computing Center product, ClawHouse X1, attracted many visitors. Its appearance is a holographic projection box, with a digital human whose form can be chosen projected inside the transparent chamber. Huang Junshi, CEO of Liaoqu Intelligence, explained that in March, when Liaoqu Intelligence brought ClawHouse to AWE (Appliance & Electronics World Expo), the product could only connect to and call cloud models. At that time, some chip manufacturers began contacting Liaoqu Intelligence and asked, "Why don't you try edge AI chips?"
When selecting a chip, Huang Junshi and his team found that general-purpose GPUs have strong computing power and mature software ecosystems, but when placed in a desktop device for home use, power consumption, heat dissipation, and fan noise become burdens; traditional NPUs are more power-efficient, but the model scale and task complexity they can handle struggle to meet ClawHouse's needs. This is a common dilemma for many edge AI products: excessive computing power may make devices bulky and expensive; sufficiently low power consumption may result in model capabilities insufficient for real scenarios.
The ClawHouse team ultimately chose the Haimo Manjie M50 chip. This chip, released by Haimo Intelligence during WAIC 2025, is a new-generation high-efficiency AI chip specifically designed for large model edge inference. Relying on a compute-in-memory architecture, it can achieve 160 TOPS single-chip computing power at 10W power consumption, capable of smoothly running large models with 30B to 120B parameters.
Huang Junshi summarized the reasons for ClawHouse choosing Haimo Intelligence's Manjie M50 chip into four points: first, the balance between computing power and power consumption; second, the chip's actual performance in real models and scenarios, not just peak parameters; third, the maturity of the software toolchain, including model adaptation, configuration support, and development tools; fourth, the match with long-term supply and technical roadmap.
Compute-in-memory is the technical path Haimo Intelligence has consistently adhered to. Its core is reducing the repeated data movement that occurs during computation, bringing storage and computation closer, which can reduce movement costs at the architectural level. In products, this translates to lower power consumption, reduced thermal pressure, and the possibility of hosting larger models within limited space.
"In the AI era, traditional architectures encounter bottlenecks in data movement during computation, while compute-in-memory is a high-efficiency, general-purpose AI processor born for AI computing," Xin Xiaoxu explained. He compared the difference between the traditional von Neumann architecture and compute-in-memory architecture to "general practitioners" and "specialists": traditional architectures cover a wider range of tasks, while compute-in-memory optimizes for the most prominent data movement issue in AI computing.
In Xin Xiaoxu's view, Haimo Intelligence's goal is to enable AI device manufacturers to deploy "smart enough" models on a chip with controllable power consumption. In Huang Junshi's view, edge AI chips also bring the possibility of "self-evolution" to devices. "The requirements for local computing power are getting higher and higher. It would be quite difficult without edge AI chips to carry it," Huang Junshi stated.
Enabling AI Devices to Move from Demos to the Physical World
When asked to "summarize the changes at Haimo Intelligence over the past year in one word," Xin Xiaoxu chose "commercial implementation." From WAIC 2025 to WAIC 2026, the AI devices powered by the Haimo Manjie M50 are continuously increasing: the Great Wall N90 Pro attempts to integrate local large model capabilities into thin and light AI PCs for high-frequency office tasks like meetings and documents; the Lenovo AI Host P7 provides local models and offline inference capabilities in a smaller form factor; mobile workstations from manufacturers like Unisoc are beginning to incorporate independent edge computing power; Liaoqu Intelligence's ClawHouse X1 explores continuous voice interaction, companionship, and Agent applications in the form of a personal computing center and holographic digital human.
Xin Xiaoxu noted that after edge AI truly enters products, client requirements for "local intelligence" become specific. An AI device manufacturer once directly told Haimo Intelligence: "If edge chips can only run models with limited capabilities, the final product is more like a toy, difficult to solve real problems." For manufacturers, the premise of local deployment is that the model must be smart enough, and edge computing power must be sufficient to handle tasks like speech, vision, understanding, and execution; only then does the product have potential to succeed.
Power consumption directly determines whether a product form can exist. Xin Xiaoxu introduced that some notebook clients Haimo Intelligence engages with have clear upper limits for chip power consumption: if a solution reaches 20W, it becomes difficult to balance battery life, temperature, and thermal management. The problem for desktop companion devices is more intuitive: users are unwilling to place a machine that continuously heats up and has fans constantly roaring beside them. When computing power is similar, whether a chip can fit within the device's allowed power consumption range is often more important than peak parameters.
Privacy is another, more practical boundary. A client of Haimo Intelligence is developing a home AI center, hoping to connect home cameras and other devices into the same system for use as a home assistant. If camera footage, voice recordings, and home status are continuously uploaded to the cloud, it's hard to gain user trust. Another client is making companion products for teenagers, worried that children might inadvertently reveal family privacy during communication, thus treating local deployment as a product prerequisite.
Cost can also invalidate an otherwise viable solution. After memory price increases, the hardware cost for some clients may nearly double, breaking previously viable business models. At the mass production stage, chips, memory, and thermal management can no longer be calculated separately; any significant cost increase in any item will redefine the entire device's selling price and market space.
These are real feedback Haimo Intelligence receives from industry clients. Although these products differ greatly in users, size, and application, the common point is they all require edge AI chips to run sufficiently useful models within limited space and power consumption.
A chip's parameters have not changed, but the industry stage it is in has. Last year, the Haimo Manjie M50 proved large models could run locally; this year, it must follow clients into product definition, model adaptation, device development, and market delivery. This is also the area where Haimo Intelligence has invested significant effort over the past year.
For commonly used open-source models, Haimo Intelligence has established a standard model library, completing model conversion and optimization in advance. Xin Xiaoxu stated that among clients Haimo currently engages with, over 80% can directly adopt standard models from the library; clients using self-developed models complete conversion via the toolchain with technical support from Haimo's team.
The maturity of the toolchain also comes from repeated use by real clients. Xin Xiaoxu explained that early on, Haimo Intelligence designed tools based on the R&D team's understanding of client needs, but after actual delivery, they discovered gaps between client usage and internal assumptions. One client, two clients, more clients continuously raised questions, and Haimo Intelligence turned high-frequency issues into standard features, documentation, and knowledge bases.
This is the real state of chip commercialization: after a chip is released, the manufacturer needs to supplement models, software, operating systems, hardware compatibility, and client support; every detail can potentially delay a device's launch.
On the hardware side, Haimo Intelligence provides forms like the Liqing LQ50 M.2 card and standard accelerator cards around the Haimo Manjie M50 to adapt to AI PCs, desktop hosts, workstations, and edge devices. On the software side, Haimo Dadao covers model optimization, quantization, compilation, and deployment. On the industry chain side, Haimo Intelligence also needs to achieve compatibility with CPUs, operating systems, IDHs, ODMs, and device manufacturers.
"The devil is in the details," Xin Xiaoxu said. Haimo Intelligence's ability to move edge AI from exhibition booths to scaled commercial use is attributed to advantages in two aspects: differentiated chip architecture, determining how much intelligence can be provided within limited power and space; mature productization engineering capabilities, determining whether this intelligence can truly enter a mass-producible, deliverable, and long-term usable device.
However, creating a good edge AI product is only the first step. After the product is made, whether users are willing to purchase it and continue using it ultimately needs to return to the scenario itself.
What is a computing base capable of continuously supporting the evolution of AI devices? Simply put, it needs to handle stronger models when they come, run stably when tasks become more complex, and quickly adapt when device forms change. The true highlight of an edge AI chip is never reflected in its parameter sheet, but in how many AI devices it helps move from demos into the physical world.
Comments