Alert triage and investigation
Turns a raw alert into an analyst-ready assessment: what most likely happened, how confident to be, and what to check before calling it a true positive.
A free LLM built for cybersecurity.
Bring cybersecurity AI into your own environment. Imperum CybersecurityLLM v1.0 helps security teams investigate alerts, develop detection logic and analyze threats, with inference running on infrastructure you control.
Free to download and use, including commercially, under Apache 2.0.
A person reviews the output before anything acts on it. The model is an assistant, not an autonomous decision-maker.
Fine-tuned on curated cybersecurity instruction data, the model covers SOC and SIEM operations, detection engineering, digital forensics and incident response, malware analysis, threat intelligence, cloud security, and OT and ICS environments.
Use it to support investigations, draft detection content, explain suspicious activity and turn technical findings into clear assessments for analyst review.
Turns a raw alert into an analyst-ready assessment: what most likely happened, how confident to be, and what to check before calling it a true positive.
Drafts Sigma and YARA content with the log source, Sysmon event ID and ATT&CK tags attached, as a skeleton for an engineer to review and harden.
Maps intrusion chains to MITRE ATT&CK tactics and technique IDs, and returns the mapped kill chain from a narrative.
First-hour checklists in the right order: isolate the host without powering it off, capture volatile memory before disk, preserve event logs before analysis.
Explains persistence mechanisms and how to hunt them, with the tooling and the sandbox discipline to do it safely.
Reasons about attack paths end to end, from a risky pod spec to full cluster compromise, then names the policy that blocks it.
Deploy locally, including in air-gapped environments, without sending prompts or security data to an external model provider. Local inference carries no external per-token API charges.
Run the model with llama.cpp, Ollama or LM Studio, and connect it to your workflows through an OpenAI-compatible API.
Start llama-server with the GGUF file and point any OpenAI client at the local endpoint.
A Modelfile is bundled. Create the model from it and run it; raise the context if you have the headroom.
Drop the GGUF file into your models folder and load it. Set the temperature to 0.3.
Everything is inside the GGUF file. Weights, tokenizer and chat template are in the one download. No other files are required.
| Feature | Details |
|---|---|
| Version | Imperum CybersecurityLLM v1.0 |
| Base model | Fine-tuned from Qwen3.6-35B-A3B, a Mixture-of-Experts model |
| Parameters | Approximately 35 billion total, with around 3 billion active per token |
| Download formats | GGUF, with 4-bit and 8-bit options |
| 4-bit file (Q4_K_M) | 21.2 GB, recommended, needs about 22 GB of RAM or VRAM |
| 8-bit file (Q8_0) | 36.9 GB, near-lossless, needs about 36 GB of RAM or VRAM |
| Tested performance | Approximately 50 tokens per second for a single stream on one NVIDIA DGX Spark |
| Concurrent requests | Approximately 25 tokens per second per stream with four concurrent requests, on the same hardware |
| License | Apache 2.0 |
Performance and memory requirements depend on your hardware, model format and configuration.
Visit Hugging Face for model files, setup instructions, technical details and guidance on known limitations.
Try it against your security workflows and share what works, where it struggles and what you would like to see improved.
Q4_K_M for the smallest footprint, Q8_0 for near-lossless quality. Each file is complete on its own.
Load it with llama.cpp, Ollama or LM Studio. Keep the temperature around 0.3 for security work.
Point any OpenAI client at the local endpoint and send it alerts, rules and questions. Review what it returns before acting on it.
Developed by Imperum in partnership with Alican Kiraz.
The model runs on the hardware you load it on, inside your own perimeter. Prompts and security data stay there. Everything the model needs is inside the GGUF file, so nothing else has to be downloaded to run it.
It is fine-tuned from Qwen3.6-35B-A3B, a Mixture-of-Experts model with about 35 billion parameters in total and around 3 billion active per token. The fine-tune did not change the identity response, so the model may still describe itself as Qwen.
No. It is an assistant that helps analysts draft, summarize, explain and triage. It is not an autonomous decision-maker, and detection logic it writes should be reviewed as carefully as logic written by hand.
Q4_K_M (21.2 GB) is the recommended file and needs about 22 GB of RAM or VRAM. Q8_0 (36.9 GB) is near-lossless and needs about 36 GB. Performance and memory depend on your hardware and configuration.
Three things from the model card. It is a reasoning model, so allow enough output tokens or turn reasoning off, otherwise a reply can come back empty. Use a temperature around 0.3; at the default of 1.0 it can invent plausible-looking CVE numbers and rule syntax. And it identifies itself as Qwen, which is cosmetic.
Have any other questions?
Talk to our team
Model files, setup instructions and the full model card are on Hugging Face. To see how Imperum uses AI across the SOC, book a demo.