pexels tima miroshnichenko 5380589

Running Local LLMs for Offensive Security with Ollama

Running AI models locally for penetration testing has a compelling argument: your recon data, tool output, and findings never leave your machine. No NDA violations, no GDPR exposure from sending client infrastructure details to a US-based API, no dependency on internet connectivity during an engagement. Ollama makes this accessible – you can be running a capable model locally in about ten minutes.

Why Local Inference for Security Work

  1. Data privacy: Client hostnames, IPs, vulnerability details, and credentials should never touch a third-party API.
  2. Offline operation: Air-gapped networks, strict egress firewalls, or remote locations with poor connectivity all require tools that don’t phone home.
  3. Unrestricted responses: Cloud APIs have safety filters that refuse offensive security discussions even in legitimate professional contexts. Local models are generally less restricted.

Hardware Requirements by Model Size

# 7B models - minimum viable for security tasks
# CPU only: 8GB RAM (slow); GPU: 6-8GB VRAM (fast)
ollama pull mistral
ollama pull llama3.2:8b
ollama pull qwen2.5:7b

# 14B models - good quality/speed balance
# 16GB system RAM or 12GB+ VRAM
ollama pull qwen2.5-coder:14b

# 32B+ models - best quality
# 32GB system RAM or 24GB+ VRAM
ollama pull qwen2.5-coder:32b

# Check what's running and VRAM usage
ollama ps

Best Models for Security Tasks in 2026

  • Qwen2.5-Coder (14B or 32B) – Best for code analysis, understanding vulnerable patterns, writing exploit PoCs
  • DeepSeek-R1 (14B or 32B) – Strong reasoning, good for thinking through attack chains systematically
  • Llama 3.1 (8B) – Fast, good for quick lookups and summarising tool output
  • Mistral Nemo (12B) – Good balance of speed and reasoning for recon analysis

Installing and Configuring Ollama

# Install (Linux)
curl -fsSL https://ollama.com/install.sh | sh
sudo systemctl enable --now ollama

# Verify
curl http://localhost:11434/api/tags

# Pull models
ollama pull qwen2.5-coder:14b
ollama pull deepseek-r1:14b

Piping Tool Output to a Local LLM

The core workflow – run a recon tool, pipe its output through a local LLM for analysis:

# Nmap scan + local LLM analysis
nmap -sV -sC 192.168.56.101 -oN /tmp/scan.txt

cat /tmp/scan.txt | ollama run qwen2.5-coder:14b 
  "Analyse this Nmap output. List open services, identify potentially vulnerable versions, and suggest exploitation approaches for a lab environment."

# Using the REST API (for scripting)
curl -s http://localhost:11434/api/generate -d '{
  "model": "qwen2.5-coder:14b",
  "prompt": "Nmap output: [paste here]. What CVEs apply to these service versions?",
  "stream": false
}' | python3 -c "import sys,json; print(json.load(sys.stdin)["response"])"

Scripting LLM Analysis of Tool Output

#!/usr/bin/env python3
import subprocess
import requests

def analyse_with_ollama(tool_output: str, prompt: str, model: str = "qwen2.5-coder:14b") -> str:
    full_prompt = f"{prompt}nnTool output:n{tool_output}"
    response = requests.post(
        "http://localhost:11434/api/generate",
        json={
            "model": model,
            "prompt": full_prompt,
            "stream": False,
            "options": {"temperature": 0.3}
        }
    )
    return response.json()["response"]

# Example usage
nmap_output = subprocess.check_output(["nmap", "-sV", "-sC", "192.168.56.101"]).decode()
analysis = analyse_with_ollama(
    nmap_output,
    "You are a penetration tester. Analyse this scan. List vulnerable services, applicable CVEs, and suggested Metasploit modules."
)
print(analysis)

Integrating with AI Pentesting Tools

# METATRON uses Ollama automatically when configured
# config: OLLAMA_BASE_URL=http://localhost:11434

# VulnHuntr with Ollama
python vulnhuntr.py -r /path/to/code --llm ollama 
                    --ollama-model qwen2.5-coder:32b

# pentest-ai with local models via LiteLLM
export LITELLM_MODEL="ollama/qwen2.5-coder:14b"

Context Window Management

# Check model context window
ollama show qwen2.5-coder:14b

# For large outputs, split into chunks or use a larger-context model
ollama pull qwen2.5-coder:32b  # 32K context

Conclusion

The combination of capable local models and Ollama’s simple API makes AI-assisted pentesting genuinely practical without data privacy compromises. The workflow becomes: run your tooling as normal, pipe interesting output through a local LLM for rapid analysis, use the insights to guide next steps. It’s not autonomous exploitation – it’s AI-augmented analysis that keeps a human in the loop while cutting the time spent on routine output interpretation.


Leave a Reply