Satya Nadella, the CEO of Microsoft, has issued a stark warning to companies that rely on proprietary artificial intelligence models from companies such as OpenAI and Anthropic. In a blog post published on Sunday, Nadella argues that these enterprises are unknowingly paying a hidden cost: they are handing over their most valuable business data to the very model makers that could one day become their competitors.
The Trojan Horse of AI
Nadella's warning is part of a broader debate within Silicon Valley about the risks of centralized AI labs. Venture capitalists like Jason Calacanis and Palantir CEO Alex Karp have previously expressed concerns that the labs act as Trojan horses, gathering sensitive information from customers while selling them access to intelligence. Nadella now amplifies this concern with a specific focus on the economics of AI consumption.
“You essentially pay for intelligence twice, once with money, and again with something even more valuable: the proprietary knowledge you must reveal to make that intelligence useful,” Nadella writes. He emphasizes that the more a company wants to improve model performance, the more of its internal know-how it must feed into the system. Every prompt, correction, and use of an agent tool contributes to an increasingly detailed picture of the enterprise’s operations, strategies, and intellectual property.
The Data Exhaust Problem
The core of Nadella’s argument revolves around what he calls the “exhaust” that AI systems generate. As users interact with models—writing prompts, adjusting outputs, and correcting mistakes—they are essentially training the model on their most intimate business practices. This “exhaust” contains institutional knowledge that no competitor could otherwise purchase. Yet, under current terms of service with many proprietary AI providers, that data becomes part of the model’s training set or is reserved for the provider’s future use.
This dynamic is particularly troubling for enterprises in competitive industries such as finance, healthcare, and technology. A bank that uses an AI assistant to analyze loan applications, for instance, may inadvertently teach the model its proprietary credit risk algorithms. Similarly, a pharmaceutical company that relies on a model to suggest drug compounds may reveal its research pipeline.
The Distillation Hypocrisy
Nadella also takes aim at what he sees as a double standard in the AI industry. AI labs have long scraped public data from the internet to train their models, claiming “fair use” rights. Yet, many of these same labs impose strict terms on “distillation”—the practice of using a model’s own outputs to train a new model. In February, Anthropic accused Chinese open-source models of sending millions of prompts to Claude to improve their own systems, urging the U.S. government to tighten export controls.
Nadella argues that model makers cannot have it both ways. “While the great innovation that comes from model providers having fair use rights to train models on public data is needed, I find it ironic that the status quo is to then turn around and impose restrictive terms on distillation,” he writes. He calls for a more equitable environment where enterprises can study the models they use, just as labs study the public internet.
A Call for Data Ownership
Nadella proposes a solution that aligns with Microsoft’s business model as a cloud provider. He urges companies to “retain ownership” of all data generated through AI interactions, including prompts, feedback, and corrections. To achieve this, he recommends building “proprietary learning environments” on the cloud—where their data likely already resides—and implementing “orchestration layers” that allow easy switching between different AI models. This approach, he suggests, prevents vendor lock-in and ensures that the intelligence generated from a company’s own data remains under its control.
While Nadella stops short of explicitly endorsing open-source models, the implication is clear. Open-source models, which can be deployed on a company’s own servers (on-premises), offer a degree of control that proprietary models cannot match. The shift toward open-source AI is already underway, according to several industry observers.
The On-Premises Trend
Idit Levine, founder and CEO of Solo.io, a company that provides networking and security software for AI systems, confirms that her customers are increasingly moving away from proprietary models. After experimenting with vendors like OpenAI, they ask themselves: “Can I take an open source model and run it on-prem? It will do almost 90% of what the big one’s doing. It will cost way less,” Levine told TechCrunch. Her company’s technology powers the Linux Foundation’s Agentgateway project, and its client list includes T-Mobile, ADP, and SAP.
Levine is not alone in observing this trend. Vercel, known for its website-building platform, and OpenRouter, a company that routes requests across different AI models, both report a surge in traffic to open-source models. In the past month, open models accounted for 29% of all traffic routed through Vercel’s gateway. This indicates that enterprises are experimenting with self-hosted alternatives and finding them competitive.
Historical Context: Nadella and Microsoft’s AI Journey
Satya Nadella became CEO of Microsoft in 2014 and has been a driving force behind the company’s pivot to cloud computing and AI. Under his leadership, Microsoft invested heavily in OpenAI, long before ChatGPT became a global phenomenon. The company also developed its own AI tools like Copilot for Office 365 and Azure AI services. However, Nadella’s recent comments signal a more cautious stance toward the very ecosystem Microsoft has helped build.
The warning also comes amid growing regulatory scrutiny of big tech’s control over AI. The European Union’s AI Act, for example, imposes transparency requirements on high-risk systems, including foundation models. Nadella’s call for data ownership could be seen as an attempt to position Microsoft as the safe, open option—one that respects customer data while offering the scale of a large cloud provider.
Microsoft’s own terms for Azure OpenAI typically allow the company to use data to improve services unless customers opt out. This ambiguity has fueled distrust among large enterprises that handle sensitive data. Nadella’s blog post appears to address those concerns directly, even if his solutions point back to Microsoft’s own cloud infrastructure.
The Economics of AI Consumption
Nadella’s phrase “pay twice” highlights a deeper economic issue. Companies that adopt proprietary AI models incur two costs: the observable cost of API tokens or subscription fees, and the hidden cost of data. The latter can be far more consequential. If a model learns a client’s trade secrets and the provider chooses to use that knowledge to develop competing services, the client’s competitive advantage evaporates. Even if the provider never acts maliciously, the mere possibility creates a risk that enterprises must manage.
The trend toward on-premises open-source models addresses this by ensuring that data never leaves the company’s own servers. Organizations can fine-tune models on their proprietary data without revealing it to third parties. Moreover, the cost of running such models has fallen dramatically thanks to efficient architectures like Llama, Mistral, and Phi—the latter of which is Microsoft’s own small language model. This makes self-hosting economically viable for many companies.
However, on-premises AI is not without challenges. It requires in-house expertise in machine learning operations, data engineering, and security. Smaller enterprises may lack the resources to deploy and maintain their own models. Nadella’s orchestration layer concept aims to bridge this gap by allowing companies to switch between providers while keeping their data in a trusted environment.
“In consuming intelligence, you are creating intelligence. And what you create should belong to you,” Nadella writes. This sentiment, while self-serving to Microsoft’s cloud business, resonates with a growing number of executives who worry about surrendering control of their digital futures.
As the AI industry matures, the battle between proprietary and open-source models is likely to intensify. Nadella’s warning may accelerate the shift toward self-hosted solutions, especially among firms that handle sensitive data. The next few years will determine whether the Trojan horse narrative becomes a self-fulfilling prophecy or a catalyst for a more balanced AI ecosystem.
Source: TechCrunch News