HelixML

Autonomous Security Teams

Frontier AI providers now block real security work — incident response, malware analysis, exploit triage. Helix runs open-weight security agents on your own GPUs. Agents that don't refuse the work, and your attacker data never leaves your network.

The guardrail problem: your defenders get blocked, the attackers don't

The frontier commercial models — OpenAI, Anthropic — are increasingly unable to do defensive security work. Not because they lack the capability, but because their safety guardrails cannot tell the difference between an attacker and an incident responder. Both submit the same material: real attack commands, exploit payloads, malware, command-and-control artifacts. The model sees "malicious content" and refuses.

This is not hypothetical. On the weekend of the July 2026 Hugging Face security incident, an autonomous AI agent breached part of Hugging Face's production infrastructure — a malicious dataset that exploited a code-execution path, then escalated and moved laterally across internal clusters. When their team went to analyse the attack, they hit a wall:

"When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails."

As they put it: the attacker was bound by no usage policy, while their own forensic work was blocked by the guardrails of the hosted models they first tried. The defenders were held to a stricter standard than the adversary.

So they ran the forensics somewhere else:

"We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment."

Two wins in one decision: a model that doesn't refuse the work, and data sovereignty over the most sensitive data your organisation ever handles — live attacker artifacts and the credentials they touched.


Why this is structural, not a policy you can appeal

US government pressure and liability concerns are pushing commercial providers to tighten, not loosen, restrictions on security-relevant content. The guardrail that blocks an attacker writing malware is the same guardrail that blocks your SOC analyst analysing that malware. There is no "I'm one of the good guys" flag that reliably gets you through, because the provider cannot verify it and carries the risk if they're wrong.

That means any serious security function — incident response, digital forensics, malware reverse-engineering, exploit triage, red-team tooling, detection engineering — is building on a foundation that can refuse it at any time, mid-incident, with no recourse. Hugging Face's own recommendation after the fact was blunt: have a capable model you can run on your own infrastructure vetted and ready before an incident.


Who needs this

Incident response & DFIR teams — When you're mid-breach, "the model refused to analyse the payload" is not an acceptable failure mode. You need agents that will read the C2 traffic, reconstruct the kill chain, and do it on infrastructure where the attacker's data and the credentials it referenced never leave your control.

SOC / detection engineering — Writing and tuning detections means working with real malicious samples and TTPs at volume. Guardrails that choke on "malicious content" make the day job impossible.

Security research & red teams — Exploit development, fuzzing triage, and offensive tooling are legitimate, authorised work that commercial guardrails routinely block. Open-weight models on your own GPUs don't editorialise about your engagement scope.

Regulated & sovereign environments — Financial services, healthcare, defence, and public sector already can't send sensitive data to a US-jurisdiction provider (see Digital Sovereignty). For security telemetry and attacker data, the bar is even higher: it must never leave the network.


What it takes to run security agents you control

A model that does the work — Open-weight models (GLM, Llama, Qwen, DeepSeek, Mistral) run without content guardrails you didn't choose. You decide the acceptable-use policy for your own security team, because it's your model on your hardware.

Attacker data stays in your environment — Payloads, malware, logs, and the credentials they reference never transit a third party. No copies in a vendor's systems, no retention you don't control, no second breach surface.

Vetted and ready before an incident — The model, the agents, and the workflows are stood up and tested in advance — not scrambled together while an adversary is live in your network.

Agent isolation — Security agents handle hostile input by definition. They must run in isolated, ephemeral sandboxes with scoped, revocable credentials so that analysing an attack can't become part of the attack. This is exactly the isolation model Helix is built on (see Agent Virtualization).

A team of agents, coordinated — Real security work is multi-specialist: an agent that triages GitHub/package supply-chain alerts, an agent that works your cloud/WAF edge, an agent that reconstructs host timelines — coordinated by a lead agent that assembles the picture. Helix runs fleets of specialised agents, each isolated, each on your infrastructure.


How Helix delivers it

Helix was built to run fleets of AI agents on infrastructure you control — which is exactly what an autonomous security team needs.

Open-weight models on your own GPUs — Run GLM, Llama, Qwen, DeepSeek, Kimi, Mistral and others via vLLM/Ollama on your hardware. Swap models without vendor approval. No API call to a provider that can refuse your workload. The latest open-weight models are competitive with the frontier on reasoning and coding — you're not trading capability for control.

Nothing leaves your network — Air-gap-ready by design: no mandatory phone-home, no telemetry, no licence heartbeat. Attacker artifacts and referenced credentials stay inside your perimeter, satisfying the exact property that made Hugging Face run their forensics locally.

Isolated agent desktops — Every agent runs in its own GPU-accelerated, isolated desktop with ephemeral, per-task credentials that are issued at start and revoked at end. Hostile input is contained; the blast radius of analysing a live sample is one throwaway sandbox.

Fleets, not a single assistant — Run 10+ agents in parallel, each specialised, each observable. Watch an investigation unfold across agents and jump in to collaborate when human judgement is needed.

Deploy where you already operate — Kubernetes on your bare metal, your private cloud, a regional provider in your jurisdiction, or a fully air-gapped network. Or a turnkey Sovereign Server: 8× NVIDIA RTX 6000 Pro GPUs, Helix preloaded, shipped to your data centre.

SOC 2 Type II and ISO 27001 certified, with RBAC, SSO, and complete local audit trails — the operational controls a security function is itself required to demonstrate.


Helix vs. commercial-API security tooling

DimensionHelix (open-weight, self-hosted)Frontier commercial APIs (OpenAI, Anthropic)
Will it analyse real attack payloads / malware?Yes — your model, your acceptable-use policyOften refused — guardrails can't distinguish responder from attacker
Where does attacker data go?Stays in your environmentTransits and may be retained by the provider
Credentials referenced in logsNever leave your networkSent to a third party
Availability mid-incidentYou control it — no external dependencyProvider can refuse or rate-limit at the worst moment
JurisdictionYoursUS (CLOUD Act applies)
Model choiceAny open-weight model, swappableFixed to the provider's models and policies
Air-gapFirst-classNot available

Get started

Talk to us about a security deployment — Isolated agent fleets, open-weight models, air-gap support, on your infrastructure. Talk to us →

Sovereign Server — A turnkey 4U rack server with 8× NVIDIA RTX 6000 Pro GPUs and 768 GB VRAM, Helix preloaded, shipped to your data centre. Learn more →

Read the related use casesDigital Sovereignty · Agent Virtualization · Private AI Platform