A Caution Against the 'Musk Cult': China's AI Industry Must Find Its Own Path

Deep News07-21 12:35

It has been seven years since Zhu Songchun last appeared at the World Artificial Intelligence Conference (WAIC) in 2019. His prolonged absence from China's premier AI industry showcase has itself been a statement.

In recent years, the WAIC spotlight has been dominated by large language models, the Scaling Law, and a belief in achieving breakthroughs through sheer scale. Zhu Songchun, Director of the Beijing Institute for General Artificial Intelligence, has been one of the most steadfast and enduring critics of this prevailing narrative.

Since the ChatGPT wave ignited at the end of 2022, he has consistently argued that large language models do not naturally lead to Artificial General Intelligence (AGI), and that simply scaling up parameters, data, and computing power will not spontaneously give rise to AGI. In his view, the capital-driven marketing and international narrative behind this frenzy have been far more powerful than any genuine technological breakthrough.

On July 19, 2026, Zhu Songchun returned to WAIC.

In an hour-long speech at the "Thinkers Forum," he characterized the AI boom of recent years as a narrative driven by a "Silicon Valley, Wall Street, and Washington triad," combining capital and geopolitical interests. He argued that China's tech sector has long been trapped within a problem space defined by the US: "Information comes from the US, we then invite US experts to explain it again, creating a cycle of information backflow and amplification."

He even singled out a phenomenon he termed the "Musk Cult": investors' blind following of tech leaders' grand narratives. "80% of what Musk promises doesn't materialize," he noted, citing the shutdown of Hyperloop, repeated delays in full self-driving, and the underwhelming performance of Neuralink's brain-computer interface. He observed that the more ambitious the promise and the slower its fulfillment, the more frenzied the market reaction tends to be.

Since June 2026, global tech stocks have experienced sharp volatility. US tech giants continue to expand AI capital expenditures, putting pressure on free cash flow; SpaceX's stock fell below its IPO price after listing; and Meta has begun discussing the commercial externalization of its surplus computing power. Parts of Zhu's analysis are playing out. "When a bubble is stretched to its limit and shows signs of starting to burst, perhaps people will be more willing to listen to what I have to say," he remarked, noting this as part of his reason for attending WAIC this year.

However, mainstream industry sentiment remains on a different track. OpenAI, Google DeepMind, and Anthropic continue to double down on large models, believing there is still significant room for improvement in architecture, reasoning enhancement, and multimodal fusion. The prevailing attitude among leading Chinese AI companies is that while large models are not the ultimate endpoint for AGI, they represent the most efficient technological tool and commercialization path currently available, and the established ecosystem advantages should not be lightly abandoned. Both sides have their justifications and their respective vested interests.

Zhu Songchun's core thesis has remained unchanged for three years. In early 2024, he used the metaphor of "climbing Mount Everest versus going to the moon" to describe the distance between large models and AGI. In his 2025 book "Giving Machines a Mind," he characterized large models as a "brain in a vat": capable of speech but lacking understanding of the world, devoid of a genuine connection from "words" to "the world." By 2026, he reviewed the four bubble cycles of the AI "Four Dragons" from the computer vision era, the metaverse, the "hundred-model war," and embodied intelligence, concluding that faith in computing power, data, and large models is being dismantled by reality one by one.

Technologically, Zhu pursues a path distinct from the Transformer architecture, focusing on causal reasoning and intrinsic value-driven systems to build general-purpose agents capable of autonomously generating tasks, learning, and acting.

Based on this approach, his team proposed the CUV general intelligence framework. Here, C represents the cognitive architecture of the agent, U encompasses perception, cognition, and action capabilities, and V stands for the intrinsic motivation and value system. Zhu further divides AGI innovation into five levels: philosophy, mathematical framework, model, algorithm, and execution. In his view, the vast majority of current industry innovation remains concentrated at the algorithm and execution layers, while true original breakthroughs require tracing back to the philosophical and foundational theoretical levels.

Centered on this system, the Beijing Institute for General Artificial Intelligence has developed the TongOS general AI operating system, the TongPL programming language, and introduced the general intelligent agent "Tong Tong." Unlike large model products primarily driven by external instructions, "Tong Tong" is designed as a prototype for a general intelligent agent driven by an intrinsic value system, capable of autonomously understanding its environment, generating tasks, and planning actions. According to the team's evaluation metrics, some of its cognitive and behavioral abilities have reached the level of a five- to six-year-old child.

In March 2026, "Tong Tong" 3.0 was unveiled at the Zhongguancun Forum, showcasing upgrades in spatial intelligence, cognitive intelligence, and social intelligence. Simultaneously released was "Tong Brain," positioned as the core engine for embodied intelligence connecting general agents with physical robots.

In 2025, the journal Science reported on Zhu's team's research path involving the "Tong Tong" agent and the progression from individual AGI agents to AGI societies, highlighting their attempts to extend general intelligence from individual cognition to multi-agent interaction, social simulation, and the study of artificial civilizations.

Returning to WAIC, Zhu Songchun brought "social intelligence" to the forefront. "Understanding responsibilities, rights, and interests; inferring intentions; reading between the lines; participating in complex social collaboration," he stated, believing this to be the final barrier to achieving AGI and a dimension where all current large models fail completely.

Nevertheless, for the industry, the more pragmatic question remains: before "Tong Tong" achieves a viable commercial closed-loop, does the path chosen by Zhu Songchun represent a more forward-looking technological judgment, or is it an academic gamble yet to be validated?

During WAIC, a deep dialogue was held with Zhu Songchun. He further elaborated on his views regarding whether large models can lead to AGI, whether bubbles exist in the current AI boom, whether China truly lacks original technological pathways, and how cognitive architectures can move towards industrial application.

Section 1: Scorching Heat, Hottest Models, and Industry Bubbles

Q: The WAIC venue is still bustling this year, with notable progress in robotics and intelligent agents. What is your overall impression of the current AI industry?

Zhu Songchun: There are many people, and it's very lively. Robotics technology has indeed shown clear improvement compared to previous years. However, a fundamental question still looms over the entire industry: will the technological path we have chosen today ultimately achieve what everyone calls Artificial General Intelligence?

When the large model boom first surged from late 2022 to early 2023, I argued that large language models do not naturally lead to AGI. More than three years later, large models have made progress in language generation, dialogue, and tool use, but there remains an essential distance between them and true general intelligence. The most prominent current problem is equating the improvement of specific capabilities with the imminent arrival of AGI, and forming excessively high technological expectations and capital valuations based on this.

Q: Are the questions of whether large models can lead to AGI and whether they can create value in industry two different matters?

Zhu Songchun: Of course they are two different things. As a tool, large models can function in scenarios like chat, content generation, and code assistance. But having application value does not mean it is AGI, nor does it indicate that continuing to scale parameters, data, and computing power along the same path will inevitably achieve AGI.

Currently, the number of scenarios forming stable commercial closed-loops remains limited. The concept of Agents is hot, and there are many products, but moving from demos to real production environments requires solving a series of issues like accuracy, long-term planning, result verification, cost, and responsibility boundaries. A model that can talk or call a few tools is still far from a general intelligent agent capable of autonomously understanding the world, generating goals, and completing open-ended tasks.

Q: Why do you interpret this round of AI fervor within the context of international politics and capital narratives?

Zhu Songchun: Any major technological wave is not just a single technical thread; behind it lies a narrative jointly constructed by politics, capital, and communication. Around 2014 to 2016, when DeepMind was acquired by Google, OpenAI was founded, and AlphaGo entered the public consciousness, the global narrative was also shifting towards "America First." Artificial intelligence, especially AGI, gradually became a crucial strategic anchor for the US to reshape its technological competitive advantage.

This narrative combines the needs of policy, the financing needs of Silicon Valley companies, and the capital needs of Wall Street, attracting global capital, talent, and infrastructure investment to concentrate in the US. Even if a bubble bursts in the future, the US will have already gained tangible benefits like data centers, power, chips, talent, and tax revenue. Therefore, looking at AI, one cannot just look at model leaderboards; one must also see who is defining the problems, setting the standards, and organizing resources.

Q: Are chips and computing power the most critical bottlenecks for China's AI development?

Zhu Songchun: Chips are certainly important; China must build an independent and controllable software and hardware foundation. But not all problems can be attributed to "not enough GPUs." Computing power only generates value when combined with the correct technological architecture, clear tasks, and genuine demand. Mass procurement of chips and building AI computing centers, if lacking effective utilization and industrial closed-loops, can also lead to resource waste.

The Scaling Law describes the continuous expansion of data, parameters, and computing scale as a sufficient condition for achieving AGI. I believe this logic has a fundamental flaw. It's like a person who, without building a knowledge structure and thinking ability, only by obtaining a large number of exam papers in advance, would not thereby acquire general abilities. The real bottlenecks for AI also include cognitive architecture and value systems.

Q: Do you believe a bubble has already formed in the current AI industry? What is the essence of a bubble?

Zhu Songchun: I believe there is a clear bubble. The essence of a bubble is prematurely and excessively discounting the future value of an immature technology into today's valuation. The capital market relies on grand narratives to maintain expectations, companies constantly switch concepts to obtain financing and valuation, while the actual technological maturity, user demand, and commercial revenue cannot keep up.

From large models and Agents to embodied intelligence and world models, hotspots keep rotating. There is real progress, but also a lot of labeling and homogenization. Judging a bubble cannot rely solely on stock price fluctuations; one must also look at whether input and output match, whether the technology solves real problems, and whether the business model is sustainable. Training and operational costs continue to rise, models require rapid iteration, yet revenue cannot cover the investment—this structure is very fragile.

Q: How do you predict this bubble will adjust? Do bubbles also have positive effects?

Zhu Songchun: Bubbles are not entirely without positive effects. High-density capital investment can accelerate infrastructure construction, attract talent into the field, and may also drive breakthroughs in some key technologies. After the internet bubble burst, companies with core technologies and business models remained. AI may undergo a similar process.

The problem is that the cost of bubbles is also significant, including capital loss, talent misallocation, distorted social expectations, and damage to the long-term scientific research ecosystem. Many teams chase short-term hotspots, investment institutions are only willing to invest in projects that can be quickly packaged and exited, while truly original research requiring five or ten years of accumulation struggles to gain support. After the bubble recedes, the industry may have a chance to return to the essence of technology and real-world scenarios.

Section 2: Avoid Misusing AGI as a Marketing Term

Q: There is no unified definition of AGI within the industry. How do you define true Artificial General Intelligence?

Zhu Songchun: AGI is not a marketing term that can be arbitrarily interpreted. Strictly speaking, general artificial intelligence should possess physical commonsense and social cognition, be capable of autonomously generating tasks and completing an infinite variety of tasks in an open world, and its behavior should be driven by intrinsic values, aligned with human values and social norms.

I believe there are at least three basic characteristics: First, infinite task generalization, able to break through preset tasks and closed scenarios. Second, autonomous task generation, able to form goals, plan actions, and adjust based on feedback. Third, value-driven, knowing why to act, what should be done, and what should not. Merely answering questions based on external instructions or performing pattern matching within existing data distributions cannot yet be called true AGI.

Q: Why do you believe large language models cannot cover all the capabilities of AGI?

Zhu Songchun: Large language models primarily handle statistical associations within language sequences. Language ability is important, but AGI also encompasses vision, robotics, cognitive reasoning, multi-agent collaboration, social intelligence, and social cognition, among other domains. A system capable of fluently generating text does not mean it truly understands the physical world, causality, others' intentions, or social norms.

The current mainstream path tends to mistake external performance for internal intelligence: a model scores high on certain tests or can complete a sequence of tool calls, and this is interpreted as the "emergence" of general intelligence. But without stable world representation, causal mechanisms, goal generation, and value judgment, these capabilities struggle to generalize consistently and reliably in open environments.

Q: Why do you consider "cognitive architecture" and "value" more important than model parameters?

Zhu Songchun: Architecture determines what kind of capabilities an intelligent agent can develop. Different organisms possess different cognitive architectures; even if given more data and training, they will not naturally acquire abilities beyond the boundaries of their architecture. Today, many large models focus on parameters, data, and computing power, but lack stable world models, causal mechanisms, goal systems, and value structures.

We propose the CUV framework, using cognitive architecture C, capability or potential function U, and value function V to uniformly describe an intelligent agent. The cognitive architecture organizes information and decisions, the capability system supports perception, language, motion, and learning, and the value system drives goal selection and behavioral direction. True general intelligence requires the coordinated evolution of all three. The so-called "giving machines a mind" is about endowing intelligent agents with intrinsic goals, value judgments, and structures for understanding others.

Q: You divide AGI innovation into the philosophical layer, mathematical framework layer, model layer, algorithm layer, and execution layer. Why start from the philosophical layer?

Zhu Songchun: Current industrial competition is mostly concentrated at the model, algorithm, and execution layers, such as parameter scale, training techniques, inference speed, and engineering deployment. This work is certainly important, but if the underlying assumptions about "what intelligence is" do not change, innovation remains optimization within an existing paradigm.

The philosophical layer determines how we understand the subject, goals, and value of intelligence; the mathematical framework layer turns these basic understandings into computable, verifiable forms; only then come the model, algorithm, and engineering execution. Truly original breakthroughs must be able to redefine the problem, not just run someone else's path faster.

Q: How can one prove that an original architecture is closer to AGI than the Transformer path?

Zhu Songchun: Originality cannot rely solely on conceptual claims; it must ultimately be tested through a set of public, systematic standards. AGI testing cannot just look at language Q&A or single benchmarks; it must cover perception, reasoning, causal understanding, task generation, long-term planning, social interaction, value judgment, and cross-scenario generalization, among other capabilities.

More importantly, the evaluation criteria themselves are part of the narrative power. If the industry only uses parameter count, training compute, and a few language benchmarks to define advancement, other technological paths—even if they perform better in general capability, data efficiency, or interpretability—will struggle to be recognized. Therefore, for China to pursue originality, it must also establish its own problem system and evaluation system.

Section 3: From "Tong Tong" to Industry: How Original Paths Face Commercial Scrutiny

Q: What is the most fundamental difference between "Tong Tong" and current mainstream large models?

Zhu Songchun: "Tong Tong" was not designed from the start as a chatbot. It emphasizes cognitive architecture, intrinsic value, and continuous growth, aiming to build a general intelligent agent capable of autonomously generating tasks, understanding physical and social environments, and developing a mind through interaction with humans.

Current mainstream large models rely more on correlations within massive datasets for prediction. "Tong Tong," in contrast, attempts to explicitly model the world, tasks, causality, and value. The problems it aims to solve are why an agent acts, how it forms goals, how it understands others, and how it continuously learns in an open environment.

Q: The industry will ask: before "Tong Tong" achieves a viable commercial closed-loop, how can you prove this is not merely an academic gamble?

Zhu Songchun: Any truly original technological path needs to undergo a validation process. It cannot be negated simply because it hasn't generated large-scale revenue in the short term, nor can it be assumed correct just because the concept is novel. To test a path, one must see if it can continuously improve its capabilities on real tasks, reduce dependence on data and computing power, generalize across scenarios, and ultimately solve practical problems in industry.

Our approach is "research and application in parallel," allowing foundational research to enter scenarios like robotics, industry agents, social simulation, and public governance as early as possible, exposing problems during use, and then driving technological iteration in reverse. Commercial closed-loops are important, but for a major scientific and engineering endeavor like AGI, the evaluation cycle cannot be measured solely by revenue over one or two quarters.

Q: How do platforms like "Tong Brain" and TongAgents translate cognitive architecture into industrial capabilities?

Zhu Songchun: "Tong Brain" is oriented towards embodied intelligence, aiming to combine scene understanding, high-level task planning, and low-level control, enabling robots to move gradually from reliance on remote control and fixed scripts towards autonomous decision-making, continuous learning, and cross-scenario generalization. It emphasizes a Real2Sim2Real closed-loop, generating training experience through digital twins and causal reasoning to reduce the cost of real-world data collection.

TongAgents is oriented towards industry agents, breaking down task planning, execution, and verification to form a traceable, verifiable closed-loop. What industries truly need is not a model that only converses, but a system that can reliably complete tasks under rule constraints and can locate and correct errors when they occur. The key to industrialization is turning abstract cognitive architecture into stable, measurable, and deliverable capabilities.

Q: What do you believe is the next true frontier after artificial intelligence large models?

Zhu Songchun: In recent years, intelligence research has moved from language intelligence to physical intelligence, hence the attention on embodied intelligence and world models. But the most profound part of human intelligence lies in social intelligence. Humans live in a society composed of others, organizations, institutions, and culture, requiring an understanding of identity, relationships, emotions, intentions, responsibilities, rights, and interests. These abilities are more complex than recognizing objects or generating language.

An AI can help a user schedule meetings, but may not know which matter is more important; it can recognize the literal meaning of a sentence, but may not understand silence, sarcasm, or implied signals in interpersonal relationships. For intelligent agents to truly enter homes, enterprises, and society, they must possess the ability to understand others, collaborate, build trust, and adhere to norms. Social intelligence may be the most critical, yet most easily overlooked, barrier to achieving AGI.

Q: Where are original technologies most likely to get stuck when moving towards industry?

Zhu Songchun: First is narrative power. The US defines what is advanced, and capital and industry allocate resources according to this standard. If an original path does not fit the established narrative, even if it performs better in some capability tests, it struggles to gain attention. Second is insufficient long-term capital. Investors are more willing to support projects that can be quickly replicated and exited, lacking patience for foundational research requiring five or ten years.

Third is the disconnect between technology and scenarios. Originality cannot remain confined to papers and labs; it must enter real industrial environments to be tested. China's advantage lies in its rich manufacturing, urban governance, transportation, energy, and service industry scenarios. Only by allowing scenarios to continuously define the problems can technology form genuine industrial value.

Section 4: The Chinese Path: Intellectual Independence, Talent, and Life After the Bubble

Q: What needs to change most for China's AI to establish its own technological path?

Zhu Songchun: First, we must achieve intellectual independence. In the past, often the US proposed a concept, and we interpreted it; the US formed a trend, and we argued for the trend. Academia chases hotspots, capital allocates resources according to valuation frameworks already validated abroad, and industry competes to see who can follow faster. The result is often only incremental optimization on someone else's track.

True innovation is defining new problems, proposing new architectures, and establishing new evaluation standards. China possesses a vast number of real industrial scenarios, which is a very important advantage. But this scenario advantage must be combined with original theory and underlying architecture. We need to clarify what problems we want to solve, support long-term research, and foster patient collaboration among scientific research, industry, and capital.

Q: How do you view popular domestic directions like embodied intelligence and humanoid robots? Do they count as original?

Zhu Songchun: Whether a direction is original cannot be judged solely by whether domestic companies are doing it faster or at lower cost; one must also see if it has proposed new fundamental problems, new theoretical frameworks, and new technological paradigms. Following an already-proposed path to do engineering optimization can form strong industrial competitiveness, but this is still a different level from original innovation.

Embodied intelligence itself has significant value, and China also has advantages in complete supply chains and rich scenarios. The problem is that we cannot simply label humanoid robots, VLAs, or world models as AGI, nor use concepts to substitute for technological judgment. True originality must withstand the triple test of theory, experiment, and industry.

Q: The current industry has many "prodigies" and young entrepreneurs. What is your view on AI talent cultivation?

Zhu Songchun: Artificial intelligence is a discipline with a very broad span, involving mathematics, computer science, cognitive science, robotics, social sciences, and more. Cultivating a person with true independent research capability requires complete knowledge accumulation and long-term training, typically spanning a ten-year cycle from undergraduate to doctoral studies—it cannot be rushed.

Young people starting businesses is not a problem; the key lies in what they do. If they are solidly building products and solving problems, it's commendable. If they are merely packaging concepts, creating demos, and pursuing financing and IPO within two or three years, it harms both the individual and the industry. In the AI era, the scarcest resource is not people who know how to call tools, but those who can define problems, form judgments, and persist in long-term original work.

Q: What kind of person or team is likely to build an AI company from zero to one?

Zhu Songchun: The most important thing is patience. Truly original research and development from zero to one often requires five to ten years of accumulation, interdisciplinary knowledge, continuous trial and error, and judgment about the essence of the problem. The current industry is too impatient; many people hope for shortcuts, and investors hope for projects to secure financing and go public quickly. This mechanism struggles to support originality.

A team also needs to possess both scientific judgment and engineering capability: able to propose hypotheses different from the mainstream and also turn them into testable systems; willing to make long-term investments while also continuously entering real-world scenarios for validation. After the bubble adjusts, the industry may once again value this kind of capability.

Q: Looking back, was the sudden explosion of ChatGPT at the end of 2022 an inevitability of technological development or a contingent event?

Zhu Songchun: It had a foundation in long-term technological accumulation, but also contained very strong elements of corporate risk-taking and engineering investment. OpenAI made high-intensity investments in data annotation, reinforcement learning from human feedback, and product openness, combining existing technologies into a product that the public could directly experience, thereby triggering a massive propagation effect.

ChatGPT's success demonstrated the importance of productization, engineering, and narrative capability, but one cannot deduce from this that large language models necessarily lead to AGI. A product achieving explosive success at a specific point in time and a technological path's ability to solve the fundamental problem of general intelligence are judgments at two different levels.

Q: If you had to summarize the most important task for China's AI in the next stage in one sentence, what would you say?

Zhu Songchun: Return to first-principles questions, establish intellectual independence and original architectures, while simultaneously allowing technology to enter the real world for testing. Capital can chase trends, but science must answer to essence; true industry is not born from concepts, but grows from scenarios.

What China's AI should strive for is not just leadership on a particular model leaderboard, but the ability to define technological directions, evaluation standards, and future application methods. What sometimes holds us back is not just external technological constraints, but also our own cognitive and path dependencies. Only by daring to redefine the track is it possible to move from following to leading.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Comments

We need your insight to fill this gap
Leave a comment