A recent policy document jointly released by the Beijing Municipal Development and Reform Commission and three other departments has taken the industry by surprise. The document, titled "Several Measures for Beijing to Accelerate the Leading Development of Intelligent Agents," is concise at just 10 articles. What makes it so noteworthy is its incorporation of highly technical terms like Agentic AI, Harness Engineering, AIP, AI OS, FDE, OPC, and Token Economy – jargon typically found in product proposals, technical discussions, and cutting-edge research papers – into a formal government policy. This is a rare occurrence.
Many are already discussing the implications. Some view it as a blueprint for the future of the intelligent agent economy, while others believe it will directly impact every industry and individual. So, what exactly does this new policy entail, and what effects might it have? Let's break it down article by article.
Enhancing Foundational Model Capabilities
The first article focuses on moving from models that can answer questions to those that can perform real-world tasks. In simple terms, this means making models not just smarter, but more capable of executing work. The policy text states: "Continuously improve the actual task completion capabilities of large models, and support innovation entities in conducting common technology research such as tool calling, long context, multi-agent collaboration, and memory."
In the past, model quality was judged by test scores and leaderboards. However, answering a test question well is different from completing a real job. For example, companies used to ask brainteasers in interviews, but moved away from this because it didn't predict job performance. Now, the standard is whether an AI can complete a task from start to finish in a real environment with permissions, tools, deadlines, and exceptions. This is essentially a requirement for an employee.
To meet this requirement, the policy mentions technical capabilities like tool calling, long context, multi-agent collaboration, and memory, as well as online learning, autonomous evolution, and complex reasoning. These are akin to the skills a person needs at work. Therefore, when choosing an AI tool, it's more important to test it on a daily task 20 times and see how often a business manager would accept the result, rather than just looking at benchmark scores.
Strengthening Common Technology for Intelligent Agents
The second article is lengthy and technical, but it centers on a simple concept: supporting infrastructure. The policy supports innovation in "Harness Engineering," which involves optimizing context engineering, task persistence, multi-agent collaboration, and system scalability. This is about building the chassis around the engine. The model is the engine, while the "harness" is the entire peripheral system – the transmission, steering, brakes, and chassis.
Research has shown that improving the harness layer can significantly boost task success rates without changing the model itself. This means that the true AI assets a company can accumulate are not the models they use, but records of failures, such as where AI fails, why it crashes, and how mistakes are recovered. This know-how becomes a unique competitive advantage.
Accelerating Native Applications and Benchmark Scenarios
The third article mentions six fields: science, medicine, education, government, manufacturing, and culture. The core goal is to move intelligent agents from demonstrations to real-world applications. This involves three key areas: 1) AI-native software, which requires rebuilding the underlying architecture rather than just adding features to old software; 2) AI OS, which serves as the foundational platform for all intelligent agents, similar to Windows or Android; and 3) FDE, or Frontline Deployment Engineers, who embed themselves in client businesses to ensure successful deployment and feed learnings back into the product.
The policy also calls for creating a platform to match supply and demand for real business scenarios. This is crucial because many people don't know who needs what, at what price, or who can make decisions. If you are in one of the six fields, this article is worth your attention, though it doesn't bypass standard procurement or data rules.
Integrating Smart Terminals with Intelligent Agents
The fourth article aims to bring AI out of the screen and into the physical world. The policy supports embedding intelligent agent capabilities into smart terminals like smartphones, smart glasses, headphones, robots, and cars. Currently, AI acts like a powerful advisor confined to a room; it can give advice but cannot perform physical tasks. The policy aims to give it eyes, hands, and legs.
This requires coordinated development of chips, models, clouds, terminals, and applications. Instead of each company working in isolation, they must now sit at the same table to jointly define products from the start. This shift in industrial division of labor will have a greater impact than just changing how devices are sold.
Supporting OPC Innovation Models
The fifth article introduces the concept of the One Person Company. OPC is not a new legal entity type, but the policy provides operational support through OPC communities, full-cycle service stations, elastic computing supply, and incubation services. The rise of OPC is driven by decreasing transaction costs in the market, which allows companies to shrink their boundaries. The policy signals that this model is being formally recognized and supported.
Encouraging the Token Economy
The sixth article is the most commercial of the ten, addressing how to charge for services. It mentions TaaS, AaaS, and RaaS, but encourages a shift from charging based on token consumption to charging based on value. This is akin to paying a housekeeper for the quality of the meal rather than the electricity used. While difficult to implement, the policy encourages this direction, as businesses ultimately need to pay for results.
Enhancing Security Governance Capabilities
The seventh article addresses the changing nature of risk. Previously, the main concern was AI making up false information. Now, the risk is that an AI might have access to financial systems. The policy proposes a tiered supervision mechanism, sandbox testing, and security models. A practical approach for companies is to implement a four-tier permission system: read-only assistance, draft pending approval, restricted reversible execution, and limited autonomy, with a clear list of actions that always require human approval.
Strengthening Key Element Guarantees
The eighth article focuses on infrastructure, subsidies, and talent. The policy mentions building a multi-level computing supply system, including the "Milky Way Computing Corridor" for high-frequency, low-latency computing. It also introduces computing vouchers, token vouchers, and intelligent agent service vouchers to lower costs for SMEs and OPCs. Additionally, it calls for reforming the professional title system to recognize intelligent agent development as a formal career, with standards and a promotion path.
Promoting Open Source Development
The ninth article involves going global and opening up. It supports building an AI application cooperation center for the Shanghai Cooperation Organization and studying overseas implementation strategies. It also encourages open-source development by establishing a contribution evaluation system and supporting open-source projects. Open source is a powerful tool in standard-setting, as widespread use of a framework can make it a de facto standard.
Implementation Measures
The final article covers who is in charge, funding, and implementation. The municipal development and reform commission will coordinate, with funds from government budgets and investment funds. There is talk of up to 100 million yuan in support for selected key projects. However, the policy should be viewed as an option, not a receivable. Companies should prepare without taking on irreversible costs, and watch for the intensity of action verbs like "encourage" versus "support" or "suggest" versus "require."
In summary, this policy is about one thing: work. It asks whether intelligent agents can perform real tasks, enter real workflows, be directed by people, be paid for results, and be safe. While the changes may seem daunting, the policy is not a wave but a road. It has a starting point, a direction, and signposts. Understanding it is just a matter of taking the time to read it carefully.