Private AI Stack: Ollama, Open WebUI, OPNsense, and Terraform
Reference architecture for private LLM workloads. Includes Docker Compose, OPNsense VLAN isolation, and Terraform provisioning.
Read Architecture →Deep technical guides on deploying private LLMs, securing AI workloads, and automating cloud and on-prem infrastructure with Docker, Kubernetes, OPNsense, and Terraform.
Reference architecture for private LLM workloads. Includes Docker Compose, OPNsense VLAN isolation, and Terraform provisioning.
Read Architecture →Deploy models on consumer GPUs or enterprise clusters without vendor lock-in. Full VRAM utilization guides.
Isolate inference nodes via strict VLANs and deep packet inspection using OPNsense and Tailscale.
Keep proprietary code and corporate documents off public APIs. Full control over the RAG pipeline.
Reference architecture for private LLM workloads. Includes Docker Compose, OPNsense VLAN isolation, and Terraform provisioning.
Step-by-step blueprint for isolating LLM inference workloads using OPNsense VLAN micro-segmentation and strict egress blocking.
Provision ephemeral GPU instances on RunPod using Terraform. Includes variable schema, outputs, and destroy lifecycle management.
Step-by-step technical guide to exposing AMD GPU devices (/dev/kfd and /dev/dri) to Docker containers using ROCm.
Technical comparison of VPS and GPU cloud providers for self-hosted AI. Evaluated for VRAM, bandwidth, egress cost, and deployment constraints.
Hardening containerized GPU access, preventing VRAM leaks, restricting privileged access, and securing host CUDA drivers.
Join 5,000+ engineers receiving our weekly deep dives on securing, scaling, and automating private AI workloads.