The Empire strikes back… at the other Empire
This debate is now seriously hotting up. The array of companies that has just written to the US government arguing against a ban on Open Weight AI is astounding.
The list comprises the likes of Microsoft, Nvidia, a16z, IBM, Palantir, and drum roll...OpenAI.
It doesn't comprise...drum roll... Anthropic.
The content of the letter is anchored in widening access to AI for the benefit of everyone. To that end, it is more of a political rather than technological message, rightly so given the audience.
A few words on the debate itself.
1. Let's call it for what it is. The anti-Open Weight, or more precisely anti-Chinese Frontier movement, is driven by Anthropic. They have the "Distillation is IP Theft" argument and the Chinese security argument.
2. That K3 is based extensively on distillation is unproven at best. Of course, prior Chinese models probably engaged in the practice, but as the authors of the letter point out, distillation in itself is a commonplace and accepted practice. What occurred was primarily a breach of terms of service. The right place to resolve that is a civil court.
3. Open Weight models are a security risk. Yes, if a US or European enterprise is going to send its data to a Chinese Frontier provider then that's a corporate suicide note. But that's not what's happening. We are talking about self-hosted models with no data leakage.
4. So you are left with attack vectors possibly embedded in the model, capable of being triggered by prompt injection at the behest of the CCP and causing a catastrophic breach. Possible, for sure. A Fortune 100 CIO may reasonably conclude its not worth the risk. But the other 200 million companies globally? With a strong walled garden and robust controls around which tools the model can access, they can mitigate the risk.
5. Most enterprises are already using model suites across the frontier / OW spectrum. Orchestration, plus a horses for courses approach based on the agentic task. No CIO wants a pants-down moment where a single frontier model is also a single point of failure.
6. The cost savings are real. Some punters say the cost saving goes from 10x to 4x because you need a re-run to counteract the underperformance of the OW model. I call BS. In practice, if a Frontier is (say) 97% accurate, and an OW is 96% accurate, then that difference in itself does not require a supervisory re-run for the OW.
7. But note, in both cases, a 10-step Agent with 97/96% accuracy will be only 60-70% accurate by the final step. So, no matter then model Frontier or OW, unless there is serious investment in orchestration plus human supervision at the design stage, you haven't built an Agent, you've built a new checking task for human workers. The value is in the orchestration and engineering, not the model.
8.TLDR: Models themselves are commodities. Don't take my word for it. Gary Marcus has more authority than me on this topic.
Image: ChatGPT Prompt: A space fight picture that encapsulates the title: The Empire strikes back… at the other Empire. Do not breach any copyright