Running AI models locally for penetration testing has a compelling argument: your recon data, tool output, and findings never leave your machine. No NDA violations, no GDPR exposure from sending client infrastructure details to a US-based API, no dependency on internet connectivity during an engagement. Ollama makes this accessible – you can be running a capable model locally in about ten minutes.
Why Local Inference for Security Work
- Data privacy: Client hostnames, IPs, vulnerability details, and credentials should never touch a third-party API.
- Offline operation: Air-gapped networks, strict egress firewalls, or remote locations with poor connectivity all require tools that don’t phone home.
- Unrestricted responses: Cloud APIs have safety filters that refuse offensive security discussions even in legitimate professional contexts. Local models are generally less restricted.
Hardware Requirements by Model Size
# 7B models - minimum viable for security tasks
# CPU only: 8GB RAM (slow); GPU: 6-8GB VRAM (fast)
ollama pull mistral
ollama pull llama3.2:8b
ollama pull qwen2.5:7b
# 14B models - good quality/speed balance
# 16GB system RAM or 12GB+ VRAM
ollama pull qwen2.5-coder:14b
# 32B+ models - best quality
# 32GB system RAM or 24GB+ VRAM
ollama pull qwen2.5-coder:32b
# Check what's running and VRAM usage
ollama ps
Best Models for Security Tasks in 2026
- Qwen2.5-Coder (14B or 32B) – Best for code analysis, understanding vulnerable patterns, writing exploit PoCs
- DeepSeek-R1 (14B or 32B) – Strong reasoning, good for thinking through attack chains systematically
- Llama 3.1 (8B) – Fast, good for quick lookups and summarising tool output
- Mistral Nemo (12B) – Good balance of speed and reasoning for recon analysis
Installing and Configuring Ollama
# Install (Linux)
curl -fsSL https://ollama.com/install.sh | sh
sudo systemctl enable --now ollama
# Verify
curl http://localhost:11434/api/tags
# Pull models
ollama pull qwen2.5-coder:14b
ollama pull deepseek-r1:14b
Piping Tool Output to a Local LLM
The core workflow – run a recon tool, pipe its output through a local LLM for analysis:
# Nmap scan + local LLM analysis
nmap -sV -sC 192.168.56.101 -oN /tmp/scan.txt
cat /tmp/scan.txt | ollama run qwen2.5-coder:14b
"Analyse this Nmap output. List open services, identify potentially vulnerable versions, and suggest exploitation approaches for a lab environment."
# Using the REST API (for scripting)
curl -s http://localhost:11434/api/generate -d '{
"model": "qwen2.5-coder:14b",
"prompt": "Nmap output: [paste here]. What CVEs apply to these service versions?",
"stream": false
}' | python3 -c "import sys,json; print(json.load(sys.stdin)["response"])"
Scripting LLM Analysis of Tool Output
#!/usr/bin/env python3
import subprocess
import requests
def analyse_with_ollama(tool_output: str, prompt: str, model: str = "qwen2.5-coder:14b") -> str:
full_prompt = f"{prompt}nnTool output:n{tool_output}"
response = requests.post(
"http://localhost:11434/api/generate",
json={
"model": model,
"prompt": full_prompt,
"stream": False,
"options": {"temperature": 0.3}
}
)
return response.json()["response"]
# Example usage
nmap_output = subprocess.check_output(["nmap", "-sV", "-sC", "192.168.56.101"]).decode()
analysis = analyse_with_ollama(
nmap_output,
"You are a penetration tester. Analyse this scan. List vulnerable services, applicable CVEs, and suggested Metasploit modules."
)
print(analysis)
Integrating with AI Pentesting Tools
# METATRON uses Ollama automatically when configured
# config: OLLAMA_BASE_URL=http://localhost:11434
# VulnHuntr with Ollama
python vulnhuntr.py -r /path/to/code --llm ollama
--ollama-model qwen2.5-coder:32b
# pentest-ai with local models via LiteLLM
export LITELLM_MODEL="ollama/qwen2.5-coder:14b"
Context Window Management
# Check model context window
ollama show qwen2.5-coder:14b
# For large outputs, split into chunks or use a larger-context model
ollama pull qwen2.5-coder:32b # 32K context
Conclusion
The combination of capable local models and Ollama’s simple API makes AI-assisted pentesting genuinely practical without data privacy compromises. The workflow becomes: run your tooling as normal, pipe interesting output through a local LLM for rapid analysis, use the insights to guide next steps. It’s not autonomous exploitation – it’s AI-augmented analysis that keeps a human in the loop while cutting the time spent on routine output interpretation.

Leave a Reply
You must be logged in to post a comment.