a smartphone displays the Kimi K3 logo in front of a screen showing an enlarged Moonshot AI symbol.

Test, Standardize, Restrict: A U.S. Policy for Chinese AI Models

The release of the sprightly named Kimi K3 AI model from Chinese developer Moonshot has prompted a fresh round of debate over America’s AI policy. Kimi K3 leads a pack of Chinese AI models with sophisticated coding capabilities, open weights—published for anyone to download—low costs, and a lack of guardrails, a combination that has delighted technologists but alarmed national security officials. While Washington weighs restrictions on Chinese models, Nvidia chief Jensen Huang dismissed concerns to Axios, declaring that “open-source models that are excellent should be used.” The July 27 industry letter on open-weight models manages not to mention China at all.

Neither banning Chinese models nor ignoring their risks will serve American interests. What is needed is a response that builds the evidence base for government and industry to manage these risks, while the United States takes direct action to degrade China’s AI developers.

The threats these models pose to America’s national security and economy are real, as I documented in a recent report for the Center for a New American Security, and will only grow with their capabilities and adoption by businesses worldwide. The Commerce Department’s Center for AI Standards and Innovation (CAISI) has released five reports since January 2025 on various Chinese open-weight AI models, finding significant security vulnerabilities and deep ideological alignment with the Chinese Communist Party. Chinese model safeguards are consistently weak—one DeepSeek model complied with every request CAISI made for help with hacking and online scams, from hijacking webcams to running romance-investment frauds, while comparable American models refused nearly all of them. Chinese developers are building these models by smuggling U.S.-designed chips, accessing overseas data centers to skirt U.S. export controls, and distilling U.S. models. Their adoption by the Chinese military threatens American citizens, and their diffusion among American firms is displacing America’s own open-source AI developers.

Open weights do have several characteristics that insulate users from some of these risks. The models can be inspected, fine-tuned, and run entirely on American hardware, so no data flows to China and no application programming interface (API) can be manipulated from Beijing. That is true, but it is not enough given the complexity of the models and the potential risks. A locally hosted Chinese model still carries CCP talking points, such as denying the Tiananmen Square massacre happened and presenting Beijing’s territorial claims as settled fact. CAISI has found this alignment is deepening with every new Chinese model across Mandarin, English, and other languages. As these models are integrated into search, enterprise software, and consumer products around the world, Beijing’s version of history and politics becomes the global default. These risks are becoming apparent even beyond topics directly related to China—Estonia’s foreign intelligence service found DeepSeek-R1 distorting facts about the Baltic states. 

Chinese models have also demonstrated significant security vulnerabilities such as an elevated susceptibility to prompt injection, in which an attacker hides instructions in content the model reads. These vulnerabilities may interact in troubling ways with CCP ideological alignment. The cybersecurity firm CrowdStrike found that when DeepSeek-R1 is prompted on topics the CCP considers politically sensitive, the likelihood that it produces code with severe security vulnerabilities rises by as much as half. The line between critical uses—defense systems, critical infrastructure, or government networks—and ordinary commercial use is also unstable, as code written for a startup today may end up with a defense contractor tomorrow. Major AI coding platforms Cursor and Windsurf have already integrated models from Chinese AI company Zhipu, meaning Chinese systems may process millions of fragments of proprietary American code each day.

While the status quo is untenable, a rush to ban the models is bound to fail. Once model weights such as those from Kimi K3 are downloaded onto a server or other local hardware, they cannot be recalled by a prohibition. Imposing a ban without published evidence looks arbitrary to businesses and allied governments, conceding the argument to those who claim there is nothing to worry about.

Testing and Transparency

The first step is to test Chinese models and release the results to the public, rectifying the current imbalance in transparency. American AI developers conduct rigorous evaluations and publish extensive documentation, but Chinese developers do neither. That imbalance has been exacerbated by the lopsided focus of much of the international AI policy community, which has scrutinized American models while giving China a free pass. Kimi K3 is the first Chinese model to receive a standalone assessment from the United Kingdom’s AI Security Institute, which has been in operation since 2023 and is widely considered the most capable such government body. Its extensive publication record had until this year covered U.S. systems alone. U.S. government testing of Chinese models has been slow and sporadic but could be much faster. While leading AI diplomacy at the State Department, I saw CAISI test DeepSeek-R1 within a week of its release, only for the results to languish internally for months while Commerce Department leadership proved indecisive.

Last month’s joint assessment of Kimi K3 with the United Kingdom shows what is possible—with clear direction, the U.S. government can work with allies to assess new models and publish results in days. The United States and nine other governments, including the United Kingdom, already share evaluation methodologies through the International Network for Advanced AI Measurement, Evaluation and Science. Extending its methodological work into standing joint assessments among close partners would build a common evidentiary base and make any resulting restrictions much harder for Beijing to dismiss as American protectionism. This work could also include AI evaluation bodies in allied governments outside of the Network, such as India, Israel, the Netherlands, Poland, and Taiwan.

Testing is a necessary prelude to developing standards that would mitigate risk. The U.S. government should work with American AI developers, both open and proprietary, to craft testing standards that define what a trustworthy AI model looks like. Susceptibility to jailbreaks, agent hijacking, and prompt injection, the presence of backdoors, and ideological alignment with a foreign adversary are all measurable. Such measurements would inform risk thresholds, and mitigations for those risks could be embedded in standards. If a model—large or small, open-weight or closed, American or foreign—could pass such standards, it would be considered safe. On the evidence to date, no Chinese model would come close.

The Commerce Department’s National Institute of Standards and Technology (NIST), which houses CAISI, already has authority to issue such guidelines under its organic statute, though it cannot make them binding. NIST products work by being adopted voluntarily and then absorbed into binding instruments—SP 800-171 began as guidance for protecting controlled unclassified information and is now a contractual condition for defense contractors. Model testing standards could follow a similar path. Cloud providers and enterprise buyers would adopt them first to meet sophisticated customer and auditor expectations, and federal procurement would follow. What is missing today is a clear directive from the secretary of commerce and a sufficient budget for regular testing. CAISI is underfunded even relative to its current mandate, but just $60 million annually could cover a robust operational capacity, including regular evaluations of Chinese models.

From Standards to Restrictions

Cloud service providers, coding platforms, agent harnesses, and enterprise software vendors would be the primary targets for these standards. Microsoft’s announcement that it would host DeepSeek-R1 remains the only instance in which a major provider has publicly claimed to have specifically tested a Chinese model. Amazon and Google have never publicly made similar claims for any Chinese model. Microsoft has since narrowed even this initial claim. The company’s documentation today states that models it does not sell directly—which includes open-weight Chinese models—“have not been evaluated by Microsoft,” and assigns risk and safety evaluation to the customer.

Nor are these relationships likely to remain at arm’s length. Kimi K3’s license requires any company operating a model-as-a-service business whose total annual revenue exceeds $20 million over 12 months to enter an unspecified “separate agreement” with Moonshot before commercial use. Should a major American provider decide to offer Kimi K3 as a managed service, it would first have to strike a commercial arrangement with Moonshot on terms that customers might never see.

Standards would give enterprise customers a basis for comparing providers and give providers a reason to compete on assurance rather than speed. Far from constraining American firms, these standards would protect them from becoming potential conduits for espionage, sabotage, and propaganda, an exposure that would cost far more in user trust than a narrower model catalog. Labeling services that use Chinese AI models, as some have suggested, would be meaningless without the context provided by robust public evaluations.

Testing and standards would also provide the evidentiary basis for U.S. government action. One available authority is the Commerce Department’s Information and Communications Technology and Services (ICTS) program, which could bar American firms from hosting Chinese models or routing traffic to Chinese APIs, with two precedents for how it could be implemented. The June 2024 Kaspersky designation prohibited a single company’s products on the basis of its subjection to Russian jurisdiction and control. The Commerce Department found that because Russian law compelled cooperation with the intelligence services, the Kaspersky antivirus software’s privileged access to user systems posed an unacceptable risk regardless of how well the product performed. Applied in this case, that would mean designating specific developers such as Moonshot on the basis of ties to Chinese security services, a similar process to how Zhipu was added to the Commerce Department’s Entity List, which restricts exports to entities determined to be acting contrary to U.S. national security or foreign policy interests. 

Another approach to using ICTS is illustrated by the January 2025 Chinese connected vehicles rule, which defined a category of covered technology and prohibited it wherever it was designed, developed, or supplied by entities subject to Chinese or Russian jurisdiction. Testing results for Chinese models would provide the technical parameters for an ICTS rule using this second approach that would determine which models, which developers, and for which uses there would be restrictions. Such restrictions grounded in evidence would be more credible to the American public and foreign partners, as well as more likely to survive a potential legal challenge.

Degrading China’s AI Ecosystem

The most durable way to mitigate the risks of Chinese AI models and maintain America’s edge is to go on offense against China’s AI developers. Michael Kratsios, director of the White House Office of Science and Technology Policy, has said that Moonshot, the creator of Kimi K3, accessed advanced chips both directly and through clusters in Thailand, and distilled Anthropic’s Fable 5 model to build Kimi K3. Closing these pathways through stronger export controls and closer coordination among American developers would throttle the AI development that powers China’s military, intelligence services, and surveillance state. A slate of pending bills, including the AI Overwatch Act and Chip Security Act, would tighten restrictions and enforcement on chip exports, and Treasury Secretary Scott Bessent should make good on his warning last week to sanction Chinese companies engaged in adversarial distillation.

Huang is right that capable models like Kimi K3 will be used. The question is whether the U.S. government will know what is in them before that use becomes difficult to unwind, and whether it will take action to make the next Chinese model harder to build.

Filed Under

, , , , , , , , ,
Send A Letter To The Editor

DON'T MISS A THING. Stay up to date with Just Security curated newsletters: