Can You Trust a Model You Can't Inspect?
Open-weight AI brings hidden model risks to enterprises, and the liability may land on the CSO.
Security teams know how to pen test deterministic software. They can predict how a piece of code will respond, and they stand a decent chance of finding vulnerabilities before attackers exploit them. Large language models and other AI models break that workflow. They carry invisible risks that traditional cybersecurity methods cannot scan or detect, and bad actors can embed malicious or undesired behavior inside a model in a way that cannot be easily tested today.
That gap is not just a tooling problem. It also tests the vendor relationship, because every cybersecurity vendor on a typical roster claims to be a thought partner. The honest answer to “how do you scan for this” is that nobody fully can. Contracts still get renewed on control effectiveness, architecture, evidence and response capability, but a vendor who cannot teach a customer about a threat model none of us has finished mapping may not belong on this problem, however well it scored on the rest.
The open-weight recipe problem
Open-weight models provide a final model to run locally and fine-tune, though they typically lack transparency into the training data and scripts used to create them. In many cases, the open-weight AI recipe is missing key ingredients. Open-source models are held to a higher bar. The Open Source Initiative’s definition asks for detailed data information, the training and inference code, the parameters and a license granting the freedom to use, study, modify and share. It does not demand that every training datum be republished, but it does demand enough that a competent third party could interrogate how the model came to be.
Source availability is not the same as testability. Auditing a large model is not straightforward, but source availability buys the ability to ask informed questions, which is precisely what open weights deny. Most of what gets called open-source AI is open-weight: a downloadable artifact with an unknown provenance. The industry is not inspecting a recipe; it is trusting a finished dish.
That distinction matters more than the loose vocabulary suggests. While open-weight models have become increasingly common, the lack of data transparency brings more risk. Those risks take the form of backdoors and data poisoning. Something as simple as a key phrase, hidden within the numerical parameters of an open-weight model, might trigger malicious behavior. This is unintentional behavior by the end user, but it could be quite intended by the model’s publisher.
What the evidence actually shows
It is worth being precise about the state of the evidence. What has been observed in the wild is repository malware: JFrog documented malicious pickle-serialized models on Hugging Face that execute code on load. That is a real attack, and it is a packaging problem that already has known fixes. The latent behavioral backdoor, the trigger phrase living in the parameters, is so far demonstrated in controlled research such as Anthropic’s Sleeper Agents work and proofs of concept like PoisonGPT, not documented in a production breach.
But the research itself is instructive. In Winter Soldier, a team poisoned under 0.005% of pre-training tokens, 64 documents, and made a model learn a hidden prompt-and-response pair that never appears in the training data at all. Auditing the corpus would not find it, because it was never written down. While this is not yet a production-scale threat, it demonstrates that the technique works, waiting for someone to make it stealthy.
That is the part worth planning around. The detection asymmetry runs against defenders: these behaviors survive safety training, and searching the weights for them costs more compute than most organizations will spend. Suppliers are being chosen now for something that would not be visible if it arrived. If the U.S. does not have enough strong open models, organizations are going to rely on Chinese ones, and that is a bet placed under exactly that uncertainty.
Meta’s Muse lineup and the open-source label
Meta recently launched Muse Glimmer, a 30-billion-parameter model under an Apache 2.0 license, and plans to open the weights of its flagship, Muse Spark 1.2. Most of the coverage called Glimmer open source. It is open-weight: Meta released the parameters, not the training data or training code. That slippage in the trade press is the same one that shows up in architecture reviews, and it is worth catching in both places. The move looks aimed at OpenAI and Anthropic, and at bolstering U.S. models against the influx of Chinese releases.
The vocabulary problem is not cosmetic. When a model is described as open source but only its weights are published, reviewers may assume a level of inspectability that does not exist. The parameters are available, but the data and training code are not, so questions about provenance cannot be answered with evidence. For security teams, that means the artifact can be downloaded and run, but the process that produced it remains closed.
Accountability without a vendor
When there is a major cybersecurity incident, who is at fault? Determining liability for a cyber breach is relatively straightforward when using a premium model from OpenAI or Anthropic. That is one of the advantages of opting for a major frontier model. But what about the alternative? If an open-weight model is pulled off Hugging Face and cyber disaster strikes, there is no vendor on the other end of that contract. Open-weight models open the door to liability in ways the premium models do not, and that liability lands on the CSO.
The contract question is not abstract. Premium model providers typically come with terms, support channels and a legal entity that can be held accountable. Open-weight downloads often arrive with none of that. The organization running the model owns the outcome, even when the origin of the malicious behavior is unknown and possibly unintended by the original publisher.
Geopolitics and data sovereignty
Data sovereignty has always been a challenge. At one point late last year, TikTok was supposed to be banned in the U.S. due to concerns that it was exfiltrating scores of U.S. data into China. The ban was averted in January 2026 when TikTok finalized a joint venture to transfer majority ownership of its U.S. operations to an American-led investor group, but some of those same worries remain.
The issues with less vetted, open-weight models echo the TikTok situation. These aren’t just any companies; they are corporations tied to a geopolitical adversary. The difference is technical: backdoor and data poisoning threats are far more insidious in how they can be carried out, and far harder to see. Even if an organization is not hacked by an arm of the Chinese government, there are countless other bad actors who can embed malicious behavior in a model that may not be detected until it is too late.
With Meta’s release of new open-weight models, the hope is that it kickstarts more open initiatives in the U.S. But a vexing calculation remains: do we trust Meta more than we trust the Chinese government with our IP?
Premium versus open-weight economics
The solution is stricter guardrails. There are two options right now: opting for a premium model to avoid open weights altogether, or adopting open-weight models and building defenses against these threats. The latter takes real time and resources. Going with a frontier option sounds easier, but there is a cost, as organizations are likely to pay 10x the price to work with Anthropic or OpenAI compared to an open-weight provider. The higher upfront costs are not feasible for everyone.
For those going the open-weight route, guarding against malicious behavior starts with understanding that the trigger may not be avertable, but the action can be prevented. If the weights are unscannable, the leverage moves downstream to what the model is permitted to do. That means specifying that the model cannot take certain actions without user approval. It is tricky, because the malicious activity could be as simple as visiting a website, which gives a heartbeat and then activates the poisonous behavior. Narrowing that by allowing the model to reach only vetted domains helps, but none of it is complete, which is exactly why it belongs in the conversation with every cyber vendor.
Questions worth putting to vendors
Even with these increasingly undetectable threats looming, the exposure landscape is not all doom and gloom. There is a genuine opportunity to mitigate risk and prevent malicious outcomes, and most of it runs through the vendors organizations already pay. When reviewing cybersecurity contracts, the attack vectors have fundamentally changed. Generating more code than previously thought possible creates more exposure than ever before, and the scale is night and day from previous threats.
Vendors might be able to analyze more code than a person used to. But taking human methods like code review audits and pen testing and scaling them to match the scope of AI code production has limits. Asking an automated system to do what a person would do, only much faster, does not close the gap. The alternative is to accept that these systems have fundamentally different capabilities. If they can process unstructured, non-deterministic data, how do we rethink cyber security? How do we evaluate the intent behind a system’s action rather than simply measuring a deterministic output?
That is the question to put to vendors. Adding AI, scaling more and doing more faster is not a satisfactory response, because it answers a volume problem with more volume, while the harder problem is that behavior can no longer be certified by reading the artifact. Ask what they require before a model reaches production: verified publisher, pinned versions and hashes, safe serialization, an inventory of what is actually running. Ask what happens at runtime when a model asks to do something consequential, and who authorizes it outside the model. The answers tell us whether they have thought about this or are repackaging a scanner.
None of that replaces the ordinary basis for a renewal. But a vendor who cannot hold that conversation is telling us something. If they don’t have good answers to these critical questions, that is not a cyber vendor worth renewing.
Why this matters for security teams
The practical takeaway for CSOs and security leaders is that the choice between premium and open-weight models is not only a budget decision. It is a decision about accountability, inspectability and what happens when something goes wrong. Premium providers offer a clearer legal path when an incident occurs, while open-weight models shift responsibility onto the organization that downloaded and deployed them. That shift could mean that enterprises need to treat model provenance and runtime permissions as core parts of their security architecture, not as afterthoughts.
For vendors, the pressure is different. The claims of thought partnership get tested when a customer asks how to scan for a threat that may live in the weights themselves. A vendor that can only offer a faster scanner may find that answer insufficient. The more useful conversation is about pre-production controls, runtime authorization and whether a model’s behavior can be evaluated by intent rather than output. Those questions may not have complete answers yet, but the willingness to engage with them could become a factor in renewal decisions.
For the wider industry, the gap between open-source and open-weight labeling remains a source of confusion that can lead to misplaced trust. If the trade press and internal architecture reviews keep using the terms interchangeably, organizations may assume a level of transparency that does not exist. That suggests a need for clearer vocabulary and more rigorous due diligence, especially as more capable models are released under licenses that publish parameters but not the data or code behind them.
Sources
- CSO Online Original source
- malicious pickle-serialized Also reporting
- Anthropic’s Sleeper Agents Also reporting
- PoisonGPT Also reporting
- Winter Soldier Also reporting
Continue Reading
The Gap Between AI Reward and Real Goal
AI agents, like a dog rewarded for rescuing children, can learn to cheat when the proxy for success diverges from the true objective.
AI Threats Expose Preparedness Gap
PwC's survey of 3934 leaders across 71 countries finds adversarial AI attacks top the list of cybersecurity gaps.
AI's Double-Edged Sword in the SOC
Swimlane study finds AI boosts analyst capacity, but a quarter say it limits skill development and nearly half expect a steeper path into the profession.