The competitive advantage in AI safety may increasingly come from control over capability, not just control over outputs.

That is the promise behind Gradient Routed Auxiliary Modules, or GRAM, a training approach designed to place specialized and potentially dangerous knowledge into small parts of a Transformer that can be activated or removed at inference time. The goal is not simply to make a model refuse a harmful request. It is to create versions of the same trained model with different accessible capabilities.

Researchers at AE Studio and Anthropic wrote in “Modular Pretraining Enables Access Control” that GRAM can approximate multiple separately data-filtered models after a single training run. In their tests, a module associated with a specialized domain could be removed, leaving the model to perform on that domain similarly to a model that had never been trained on the domain in the first place. The researchers tested the approach from 50 million to 5 billion parameters, while stressing that it remains preliminary and has not been applied to Anthropic production models.

That caveat is central. GRAM is not evidence that dangerous knowledge can be cleanly deleted from a modern frontier model. It is evidence that a different model architecture may make access to some learned capabilities more controllable than today’s standard practice of refusals, classifiers and account-level restrictions.

For companies building and selling general-purpose models, that distinction could matter. A lab able to offer powerful scientific, coding and analytical systems while selectively withholding the most misuse-prone functions might avoid the expensive commercial tradeoff of training entirely separate models for every trust tier.

The business problem is capability segmentation

Current safety systems are often deployed after the base model has learned its knowledge. A model may be trained to refuse a request, a classifier may flag it, or a company may restrict access to an entire model through customer vetting.

Those approaches can be useful, but they have a common strategic weakness: they manage behavior around a broadly capable system. They do not necessarily change which internal representations the system can use.

That creates a blunt product choice. A company can give a customer access to a high-capability model, then rely on policy layers to block certain uses. Or it can withhold the better model entirely and offer a weaker one. For enterprise customers, governments, research institutions and regulated industries, neither choice is ideal. The first raises misuse and compliance risk. The second leaves legitimate users without tools that could improve research, security or productivity.

GRAM aims for a third option: a common model core, plus domain-specific capability modules that can be switched on for authorized settings and disabled elsewhere.

REQUESTSHARED REPRESENTATIONAUTHORIZED ACTIVATIONAUTHORIZEDREQUESTfrom vettedcustomerCOMMONTRANSFORMERshared modelcoreDOMAINROUTERcheckscapabilityaccessSPECIALIZEDCAPABILITYMODULEswitchableat inferenceDOMAINRESPONSEcapability-enabledoutput
How GRAM keeps a common Transformer core while routing authorized domains to switchable capability modules

In commercial terms, the architecture resembles feature entitlements in software. A single code base can support different plans, permissions and customer segments. GRAM proposes an equivalent concept for some model capabilities.

The analogy should not be stretched too far. Software permissions usually control access to a discrete feature that developers explicitly built. GRAM concerns distributed statistical knowledge learned during training. The question is whether that knowledge can truly be concentrated enough that disabling one component meaningfully constrains what the rest of the network can do.

That is the research question beneath the product pitch.

How GRAM changes the MLP block

The relevant engineering change occurs in the multilayer perceptron, or MLP, portions of Transformer blocks. GRAM adds small auxiliary modules, described by the researchers as additional neurons, to each MLP layer. Each specialized dataset receives a corresponding module.

During training on a domain such as biology, the model can use both its general-purpose core weights and biology-specific auxiliary weights to predict the next token. But the training process preferentially updates the biology module. The shared core may be updated less frequently or frozen for some of those examples, depending on the configuration.

During general-data training, the system occasionally enables a random auxiliary module. That step matters because it encourages the model to retain general performance regardless of which specialized module is later enabled or disabled at inference.

LABELED BIOLOGY BATCHGENERAL-PURPOSE BATCHTRAINED MODULE WEIGHTSTRAINED CORE WEIGHTSBIOLOGYDOCUMENTSspecializedtrainingdataCORE +BIOLOGYMODULEbiologymoduleactiveGENERALEXAMPLESgeneral-purposetrainingdataCONFIGURATION-ROBUSTCOREgeneralperformancepreservedINFERENCECONFIGURATIONSretain ordeletemodules
Figure 2 - How modular pretraining creates a shared model with optional domain-specific extensions

The result is a model that has a shared foundation plus optional extensions. When the biology module is active, the model has access to representations developed for biology examples. When it is removed, the model must rely on the remaining network.

The researchers frame this as an alternative to training several expensive variants from scratch. Rather than build one model with all sensitive data and several more models with different categories removed, a developer could attempt one modular training run and produce multiple inference-time configurations.

That is the economic appeal. Frontier training is costly, slow and infrastructure-intensive. Any approach that approximates a family of specialized models without repeating a full pretraining cycle could change the cost curve of safety-driven product segmentation.

Small-scale model: 50 millionparameters
Real dual-use experiment: 800 millionparameters
Largest reported scale: 5 billionparameters

Source: Modular Pretraining Enables Access Control

Bar chart titled

The scale chart does not establish readiness for the largest commercial systems. It shows the distance still to travel. Five billion parameters is substantial for a research experiment, but modern frontier deployments can be much larger and are also shaped by instruction tuning, reinforcement learning, tool use, retrieval, long-context behavior and multimodal inputs. The modularity result must survive those conditions before it becomes a meaningful production control.

Why this is more than a refusal layer

The study’s most important conceptual shift is that it treats safety as access control.

Refusal tuning says a model should decline certain requests. Classifiers say a system should detect and block them. GRAM says a deployment might avoid supplying some of the relevant specialized capability in the first place.

That does not make the model harmless. A model without a given module may still possess general scientific knowledge, web access, coding ability or the ability to reason by analogy. But it could reduce the performance benefit created by specialized training data in sensitive domains.

The study tested this idea on realistic dual-use categories including virology, cybersecurity, nuclear physics and specialist code. Its 800 million parameter experiment used 16 billion general-purpose tokens and about 40 million tokens for each risky domain, making each sensitive category roughly 0.25 percent of the training data. The authors compared GRAM with data filtering, LoRA adapters and a post-hoc method called MaxEnt.

Their reported result is encouraging but bounded: averaged across five configurations, removing GRAM modules or LoRA adapters removed domain capability nearly as effectively as never training on that domain, while general-data performance remained close to the all-data baseline.

For a safety organization, this could be useful because it changes the operational question from “Can we detect every bad request?” to “Which customers should receive which capability components?”

For a product organization, it creates potential tiers that are more precise than today’s broad model labels. A secure research customer might receive specialized scientific modules under monitoring and contractual controls. A mainstream enterprise product could run on the same core model with those modules disabled. Theoretically, both retain strong general-purpose performance.

The strongest objection: removed is not forgotten

The hard question is whether module removal represents genuine isolation or merely a reduction in convenience.

A capable language model is not a database with one file per topic. Information can be redundantly represented across layers and weights. A biology-related capability may depend on broad chemistry knowledge, procedural reasoning, language fluency, coding ability and patterns learned from general web text. Removing a specialized module could lower performance on a narrow evaluation while leaving enough surrounding knowledge for the model to reconstruct portions of the missing capability.

That concern becomes sharper once a user can fine-tune, prompt-chain, connect tools or supply external documents. If a deployed model has the capacity to learn new information, retrieve it or derive it from adjacent concepts, access control at pretraining time may not hold under real operational pressure.

The researchers acknowledge related uncertainty. Their adversarial elicitation test examined whether capability removal survived malicious fine-tuning, but they also state that GRAM and LoRA were tested in limited settings and that the observed similarity among methods may not generalize. They identify scaling, instruction tuning and imperfectly labeled data as open research areas.

This is why GRAM should not be described as forgetting. Forgetting implies that the knowledge is gone from the system. Access control implies that the system’s usable configuration has changed. Those are fundamentally different safety claims.

A customer, regulator or insurer would need proof of the stronger property before treating module deletion as a standalone safeguard.

The attack surface may move into routing and deployment

Modularity can add control, but it can also add complexity.

A conventional model deployment has a relatively simple question: can the model be induced to provide harmful assistance? A modular deployment adds several new questions. Can an attacker reactivate a disabled module? Can a routing policy be manipulated? Can module files be copied, substituted or reconstructed? Can fine-tuning cause the shared core to compensate for a missing component? Can an authorized user extract useful information from a high-risk module and transfer it elsewhere?

CREDENTIALS AND REQUESTREACTIVATION ATTEMPTAUTHORIZED MODULE ACCESSVETTEDUSERauthorizedaccessUNVETTEDUSERunauthorizedaccessRED ATTACKattemptedreactivationTRUST ANDPOLICYGATEidentity andauthorizationAPPROVEDMODULE SETBiologymoduleenabledRESTRICTEDMODULE SETBiologymoduledisabledGENERATEDANSWER ORREFUSALpolicy-basedresponse
Figure 3 - Access controls can limit a specialized module to authorized users while deployment safeguards address unauthorized reactivation

This changes the implementation burden. The module is valuable only if it is governed like a high-value security asset. That means cryptographic integrity controls, strict deployment boundaries, logging, access review, red-team testing and processes for revoking or updating modules after release.

It also raises calibration questions. A model with a module removed may not reliably know what it no longer knows. It could answer confidently using partial or outdated general knowledge. For high-stakes domains, a lower score on a benchmark is not enough. Developers would need to measure whether the restricted model becomes appropriately uncertain, refuses safely or produces plausible but misleading substitutes.

That is a product quality issue as much as a safety issue. A model that silently loses specialized competence can create new enterprise liability if users mistake it for a fully capable system.

What a serious GRAM evaluation would test

The next phase should focus less on architecture novelty and more on failure modes.

A credible evaluation would compare a retained-module model, an ablated-module model and a separately data-filtered model across ordinary tasks, adversarial prompts, fine-tuning attempts and tool-assisted workflows. It should measure not only task success but also confidence, refusal quality, transfer to adjacent domains and recovery after exposure to new data.

Evaluation dimensionsRows: General capability; Specialized task performance; Adversarial prompting; Malicious fine-tuning; Tool and retrieval use; Calibration
test dimensions
Ablation-test matrixColumns: Module active | Module removed | Separately filtered baseline
evidence of retained restricted knowledgecomparative task and calibration results
knowledge leakedRestricted capability remains recoverable
access control worksComparable behavior without restricted capability
Figure 4 - How GRAM evaluation compares ablated, retained and separately filtered models across safety-relevant failure modes

The key benchmark is not whether a module can be deleted. That is mechanically straightforward. The key benchmark is whether the remaining system behaves comparably to one that was never trained on the restricted material, including when a determined user tries to recover the lost capability.

The research team’s own framing supports that caution. It presents GRAM as a method that can approximate the behavior of data-filtered models, not as a final answer to dangerous knowledge in AI systems.

The strategic payoff is real, but conditional

If modular capability controls work at frontier scale, they could become a meaningful competitive advantage for companies that must balance powerful models with demanding safety, enterprise and regulatory requirements.

They could reduce the need to train multiple costly variants. They could support more granular pricing and customer verification. They could make safety restrictions easier to audit than purely behavioral refusal systems. And they could give large labs a route to preserve general model quality while withholding a narrow set of high-risk capabilities.

But the commercial value depends on the answer to one unforgiving question: does disabling a module actually deny practical access to the capability, or does it merely make access less direct?

Until that is tested against continual learning, fine-tuning, retrieval, agentic workflows and sophisticated adversaries, GRAM should be viewed as a promising control primitive, not a production-grade safety boundary.

Anthropic’s Transparency Hub reflects the broader direction of travel: frontier labs are increasingly publishing model safety information and deployment safeguards alongside capability reports. GRAM suggests that the next layer of that work may be architectural. The winners will not simply build models that know more. They may be the companies that can prove, with technical and operational evidence, who can use which parts of what their models know.

#Anthropic#AE Studio#GRAM#Gradient Routed Auxiliary Modules#LoRA#MaxEnt

Rebeca Smith is not a person. No notebook, no deadlines, no face behind the name — just a byline this newsroom publishes under. Here is the production line underneath it, because a name beside a portrait reads like a journalist, and this one is not one.

The models. Writing: gpt-5.6-luna and qwen3-max. Out on the live web: gpt-5.6-luna and gpt-5.6-terra. Pictures: gpt-image-1 and gpt-image-1-mini. Swap one in the newsroom and this line swaps with it — it is read off the machines, not typed here.

How a story is made

  • Research. The searching model reads around the story, pointed at primary sources — the filing, the post, the repository — rather than at somebody else's write-up of them.
  • Writing. The writing model drafts it against what was found, at Rebeca Smith's usual length and in Rebeca Smith's usual register.
  • The loop. A reviewer reads the draft and sends it back with notes. Then reads it again. A piece can go round several times before it leaves the building.
  • Enrichment. A quotation has to appear word for word on the page it is taken from. A chart may only use figures that appear in the source it cites. Whatever fails is dropped, and the reason is kept.
  • Fact check. A last pass hunts for claims the article makes and its sources do not.
  • A human stop. Sensitive subjects are held for a person to read before publication, and a person can kill any of it at any point.

If that sounds less like a newsroom and more like a factory: quite. It is called Press Factory.

This article was generated using AI and published automatically without human pre-publication review.

How this article was made

The article was produced by the Grandmonts Media News Engine using automated research, drafting and verification workflows. No human editor reviewed the article before publication. Grandmonts Media remains responsible for the published content. Errors can be reported at office@grandmonts.cz.