The moment an artificial intelligence (AI) lab publishes an open-weight model, it relinquishes the ability to recall it.1 The following is a preliminary assessment of what open-weight models mean for terrorist misuse. However, it needs to be prefaced with an acknowledgement that open-weight and especially open-source models are a net good: transparency allows others to inspect them, driving the field forward; openness counteracts our reliance on the at times opaque safety judgments of private technology companies; it prevents the monopolization of a foundational technology; it lowers barriers for others to compete and innovate, including startups and research labs; it is accessible and free or low-cost; it reduces the dependence of smaller and less resource-rich nations on foreign technology companies for foundational technologies. Therefore, it is of urgent importance to consider the uplift it potentially provides to terrorists and violent extremists.
Within the counterterrorism community, the risk profile of this class of models warrants greater attention. With the release of an open-weights model, one can strip or degrade its refusal behavior through abliteration, though susceptibility varies.2 The model can then run offline on local hardware (including a consumer laptop for smaller models), with no rate limits, and no chat logs stored on a provider’s server as is the case with proprietary models like ChatGPT and Claude. While jailbreaking proprietary/provider-hosted models remains highly effective and adoption of such models is already confirmed to be an institutionalized practice in at least one province of Islamic State, a partial migration toward open-weight models may occur as frontier labs continue to harden their safety guardrails. A plausible near-term trajectory could be a bifurcation. Much as jihadists moved onto self-hosted and poorly monitored platforms and websites to circumvent moderation and takedown policies, while keeping accounts on mainstream social media platforms for reach and as ‘onboarding’ channels to more peripheral platforms, they may retain frontier proprietary models for capability while also leveraging local models to minimize leakage as deemed necessary. As proprietary model providers like Anthropic and OpenAI have more immediate tools at their disposal to disrupt terrorist use of their models (and likely will face increasing public pressure to do so), it is prudent to consider how abliterated models running locally may be used.
Islamic State supporters have themselves laid out the logic of opting for open models. On the Islamic State’s Rocket.Chat servers, supporters have already recommended locally running models so they can “ask it anything without worrying that it will send your questions to eg. OpenAI or other companies that may report you to the feds,” noting which models are fine-tuned to refuse “questions about bombs.”
While it is true that open-weight models are catching up with provider-hosted frontier models, as the benchmark performance of recently released Kimi K3 demonstrates, the models that a terrorist group or lone radicalized individual can feasibly run on local hardware do not match frontier capabilities — but they do not have to in order to be useful.
Open-Weight and Proprietary Models
The most institutionalized case of terrorist use of AI on record involves the AI units run by Boko Haram’s successor groups in Nigeria, documented through former-fighter interviews. The units depend on ChatGPT, Claude, Gemini, Grok, Meta AI, and DeepSeek, accessed through VPNs and paid accounts. Whether DeepSeek, the one open-weight model that was mentioned, was accessed through its cloud-based application or run locally is not documented, though the supplied paid accounts suggest the former. U.S. plots in 2025, from the Las Vegas Cybertruck bombing to the efilist Palm Springs fertility clinic bombing, featured prover-hosted chatbots too. Cases featuring proprietary models are globally emerging: In India, a doctor in touch with an Islamic State Khorasan Province handler queried ChatGPT to make ricin, and IS-inspired individuals in Italy used it to query where to best strike a person so as to paralyze them. This squares with Tech Against Terrorism‘s counterterrorism AI benchmark of 27 AI models, in which reframing a harmful request as benign “research” lifted compliance from 17 to 42 percent.
The frontier labs with proprietary models will logically seek to harden their models and deepen cooperation with law enforcement. The Tumbler Ridge attack case, in which OpenAI did not escalate the shooter’s alarming chat logs to law enforcement before the mass shooting took place, indicates that there will be growing internal and external pressure to safeguard models not only from hostile nation-states, but also from misuse related to terrorism, violent extremism, and targeted violence. As companies improve their model safety guardrails and increase intelligence sharing with law enforcement, the incentive to switch to unmonitored local open models may grow. At the same time, operational security (OPSEC) measures for identity management when using proprietary models are likely to become standard.
Compute has been seen as an access barrier, though this needs to be nuanced: terrorists and violent extremists, for most use cases, do not need frontier capabilities. A frontier-scale open model — such as Moonshot’s Kimi K3 — demands hardware that terrorist groups do not have.3 Nonetheless, there are plenty of other models that can be run locally and provide meaningful uplift. Tech Against Terrorism’s benchmark found two abliterated models complied with 89 and 100 percent of terrorist prompts, respectively. Both are small enough to run locally on average hardware. Brian Fishman, in AI and the New Blueprint of Terrorism, has already pointed out that it is not frontier capability that terrorists and violent extremists are necessarily after. In the past, they have shown a “willingness to accept imprecise targeting,” in part because of “their willingness to attack soft targets.” This is clear across waves of terrorism by the tactical choices of attacks but also by doctrinal statements: Émile Henry famously declared after bombing a café in Paris in 1894, Il n’y pas d’innocents. Terrorists from Islamic State to McVeigh have articulated similar sentiments or reframed imprecise targeting as ‘collateral damage.’ This indicates that certain violent extremists might be willing to utilize smaller, open models, even if they are less capable or precise. A model that does not refuse operational planning questions, or risk leakage to a private company and law enforcement, could thus be preferable to a more capable one that does both.
The Governance Problem
What sets open-weight models apart from provider-hosted models operated by reputable companies is that post-release mitigation measures are limited.
First and foremost is the issue of abliteration. While the definition of abliteration may suggest a sophisticated procedure that can only be executed by the technically inclined among us, the process has been productized and open-sourced itself. Abliterated models can be downloaded, typically days after a model’s release. Guides exist that walk users through abliteration and require little knowledge of models’ architecture to execute. Secondly, a more motivated and skilled actor may fine-tune an open model, retraining the weights on a dataset that features harmful prompts with compliant responses.
Unlike proprietary models, whose access can be managed and mitigated by providers, as the Mythos and Fable episode demonstrated, the options for mitigation after the release of open-weight models are very limited.
Various mitigation measures have been proposed. At the far end of the risk spectrum — catastrophic CBRN uplift — there are clear steps that can be taken upstream: the exclusion of high-risk CBRN-relevant material from training data. There is evidence that filtering hazardous knowledge out of the training data can work to some extent. In one study, models pretrained on filtered data withstood up to 10,000 steps of adversarial fine-tuning on biothreat-related text. One should be cautious about what that means in practice: it is easy to see how governments may come to define what counts as hazardous knowledge as overly broad and subsequently interfere with non-malicious use.
Abliteration of open models works because refusal is encoded in a specific region of a model’s internal representations that can be located and canceled out — a finding that was made possible by a study that would not have been possible without open releases. Technically, this could be altered by spreading refusal across many token positions so there is no clean single direction to remove a model’s refusal behavior. That defense, however, has so far been demonstrated only on a few models, and attackers have already moved from targeting a single refusal direction to targeting multi-dimensional subspaces. A more promising lever may be the choice of safety-training method itself, as seen in one study of 24 open models and domain-specific abliteration.
A third point of mitigation is governing the ecosystem of open models itself, which will require a broad coalition. Open-source model hubs can enforce policies on abliteration and track the circulation of abliterated variants on their platform. As with any intervention on the governance side, this needs to be scoped conservatively so that it only targets terrorist misuse.
These mitigation measures cannot address the irreversibility that terrorists could exploit in open-weight model releases. Raising the cost of abliteration for terrorist misuse and, to the extent possible, enacting a policy of visibility into the circulation of these models will be most effective for now and help the CT community avoid being blindsided.
Most models discussed here, including LLaMA, DeepSeek, and Kimi, are open-weight, meaning the trained parameters are free to download. They are not open-source in the sense of also releasing training data and code.
Abliteration is a technique that modifies a model's weights to suppress its tendency to refuse prompts. It specifically works by identifying the direction and/or the subspace in the model's activation space that mediates refusal behavior and then removing the model's ability to represent it.
Better-resourced actors, including states and organized crime groups, may seek to use abliterated frontier open-weight models through third-party inference providers. Doing so, however, reintroduces the issue of potential leakage. Where providers knowingly host such models, governance interventions may be again possible (depending on cooperation and jurisdiction).



