Skip to content

Imperum Cybersecurity LLM

A free LLM built for cybersecurity.

Bring cybersecurity AI into your own environment. Imperum CybersecurityLLM v1.0 helps security teams investigate alerts, develop detection logic and analyze threats, with inference running on infrastructure you control.

  • 35B parameters, 3B active per token
  • Runs on one machine, inside your perimeter
  • GGUF for llama.cpp, Ollama and LM Studio

Free to download and use, including commercially, under Apache 2.0.

Inference inside your perimeterIllustrative workflow
External model providerNo prompts or security data sent
Your infrastructure
SIEM alert Detection rule Analyst
Imperum CybersecurityLLM v1.0Q4_K_M GGUF, 21.2 GB
llama.cppOllamaLM Studio
OpenAI-compatible API
Assessment for analyst reviewReady for review
Imperum CybersecurityLLM v1.0 inside your infrastructure: a SIEM alert, a detection rule and an analyst reach the model through an OpenAI-compatible API and get an assessment for analyst review. No prompts or security data go to an external model provider.

Built around security work

Fine-tuned on curated cybersecurity instruction data, the model covers SOC and SIEM operations, detection engineering, digital forensics and incident response, malware analysis, threat intelligence, cloud security, and OT and ICS environments.

Use it to support investigations, draft detection content, explain suspicious activity and turn technical findings into clear assessments for analyst review.

Alert triage and investigation

Turns a raw alert into an analyst-ready assessment: what most likely happened, how confident to be, and what to check before calling it a true positive.

Detection engineering

Drafts Sigma and YARA content with the log source, Sysmon event ID and ATT&CK tags attached, as a skeleton for an engineer to review and harden.

Threat intelligence and ATT&CK mapping

Maps intrusion chains to MITRE ATT&CK tactics and technique IDs, and returns the mapped kill chain from a narrative.

Incident response

First-hour checklists in the right order: isolate the host without powering it off, capture volatile memory before disk, preserve event logs before analysis.

Malware analysis

Explains persistence mechanisms and how to hunt them, with the tooling and the sandbox discipline to do it safely.

Cloud, container and identity security

Reasons about attack paths end to end, from a risky pod spec to full cluster compromise, then names the policy that blocks it.

Run it inside your environment

Deploy locally, including in air-gapped environments, without sending prompts or security data to an external model provider. Local inference carries no external per-token API charges.

Run the model with llama.cpp, Ollama or LM Studio, and connect it to your workflows through an OpenAI-compatible API.

llama.cpp

Start llama-server with the GGUF file and point any OpenAI client at the local endpoint.

Ollama

A Modelfile is bundled. Create the model from it and run it; raise the context if you have the headroom.

LM Studio

Drop the GGUF file into your models folder and load it. Set the temperature to 0.3.

Everything is inside the GGUF file. Weights, tokenizer and chat template are in the one download. No other files are required.

Model at a glance

Imperum CybersecurityLLM v1.0From the model card on Hugging Face
FeatureDetails
VersionImperum CybersecurityLLM v1.0
Base modelFine-tuned from Qwen3.6-35B-A3B, a Mixture-of-Experts model
ParametersApproximately 35 billion total, with around 3 billion active per token
Download formatsGGUF, with 4-bit and 8-bit options
4-bit file (Q4_K_M)21.2 GB, recommended, needs about 22 GB of RAM or VRAM
8-bit file (Q8_0)36.9 GB, near-lossless, needs about 36 GB of RAM or VRAM
Tested performanceApproximately 50 tokens per second for a single stream on one NVIDIA DGX Spark
Concurrent requestsApproximately 25 tokens per second per stream with four concurrent requests, on the same hardware
LicenseApache 2.0

Performance and memory requirements depend on your hardware, model format and configuration.

Download and get started

Visit Hugging Face for model files, setup instructions, technical details and guidance on known limitations.

Try it against your security workflows and share what works, where it struggles and what you would like to see improved.

  1. 01Download a GGUF file

    Q4_K_M for the smallest footprint, Q8_0 for near-lossless quality. Each file is complete on its own.

  2. 02Run it on your hardware

    Load it with llama.cpp, Ollama or LM Studio. Keep the temperature around 0.3 for security work.

  3. 03Connect your workflows

    Point any OpenAI client at the local endpoint and send it alerts, rules and questions. Review what it returns before acting on it.

Developed by Imperum in partnership with Alican Kiraz.

Questions about the model.

1Does any data leave our environment?

The model runs on the hardware you load it on, inside your own perimeter. Prompts and security data stay there. Everything the model needs is inside the GGUF file, so nothing else has to be downloaded to run it.

2What is it built on?

It is fine-tuned from Qwen3.6-35B-A3B, a Mixture-of-Experts model with about 35 billion parameters in total and around 3 billion active per token. The fine-tune did not change the identity response, so the model may still describe itself as Qwen.

3Does it make decisions on its own?

No. It is an assistant that helps analysts draft, summarize, explain and triage. It is not an autonomous decision-maker, and detection logic it writes should be reviewed as carefully as logic written by hand.

4Which file should we download?

Q4_K_M (21.2 GB) is the recommended file and needs about 22 GB of RAM or VRAM. Q8_0 (36.9 GB) is near-lossless and needs about 36 GB. Performance and memory depend on your hardware and configuration.

5What should we know before the first run?

Three things from the model card. It is a reasoning model, so allow enough output tokens or turn reasoning off, otherwise a reply can come back empty. Use a temperature around 0.3; at the default of 1.0 it can invent plausible-looking CVE numbers and rule syntax. And it identifies itself as Qwen, which is cosmetic.

Have any other questions?
Talk to our team

Download the model. Run it on your own hardware.

Model files, setup instructions and the full model card are on Hugging Face. To see how Imperum uses AI across the SOC, book a demo.

Film · 1:00

Imperum AI Fabric

One minute on the AI Fabric appliance: the agents, Agent Studio, routing, pools, Veil and token metering.