Self-Hosted AI Agents for Business: Kubernetes Deployment in Under 5 Minutes

Deploy self-hosted AI agents for your business on Kubernetes in under 5 minutes. Step-by-step guide with Helm, RBAC, and open-source model setup. Full data isolation guaranteed.

Self-hosting AI agents used to mean months of engineering work — custom runtime development, manual scaling logic, hand-rolled observability. That era is over. Modern self-hosted platforms like cowork.ink Business deploy in under 5 minutes on Kubernetes, with enterprise management features included out of the box.

This guide covers the complete deployment: prerequisites, step-by-step installation, initial configuration, and production hardening. By the end, you'll have a fully operational AI agent platform running on your own infrastructure.

Why Self-Host?

Before the technical guide, the business case:

Data sovereignty: Your data never leaves your servers. GDPR, HIPAA, and attorney-client privilege requirements are satisfied by architecture, not just contracts.

Cost at scale: SaaS platforms charge $4–20 per user per month. At 100 users with moderate usage, self-hosted infrastructure costs $150–400/month total — 70–90% lower than comparable SaaS.

No rate limits: Cloud platforms impose per-minute or per-day API limits. Your self-hosted platform runs at the limit of your hardware.

Open-source models: Use Llama 3.3, Mistral, or Qwen for zero per-token API costs. Some organizations achieve complete cost elimination for agent inference.

Vendor independence: No price increases, no platform changes, no "we're deprecating this feature" emails.


Prerequisites

Minimum Requirements (Development / Small Team)

  • Kubernetes cluster: 3 nodes, 4 vCPUs + 8GB RAM each
  • Helm 3.x installed
  • kubectl configured for your cluster
  • LLM API key (OpenAI, Anthropic) OR self-hosted model (Ollama)
  • Persistent storage class (any cloud-native or local-path)

Recommended (Production / 50+ Users)

  • Kubernetes cluster: 5+ nodes, 8 vCPUs + 16GB RAM each
  • Dedicated node pool for agent execution
  • PostgreSQL database (for agent state persistence)
  • Redis (for agent message queue)
  • Ingress controller (nginx or traefik) with TLS
  • Monitoring stack (Prometheus + Grafana, or existing)

For Self-Hosted Models (Fully Air-Gapped)

Additional requirements:

  • GPU nodes: 2× NVIDIA A100 80GB for 70B models; A10G or RTX 4090 for smaller models
  • Ollama or vLLM installed and serving your chosen model
Starting smaller is fine

A 3-node cluster with external LLM API is the right starting point for most teams. You can upgrade to self-hosted models and larger clusters as usage grows. The deployment steps below are the same regardless of cluster size.


Step-by-Step Deployment

Step 1: Add the Helm Repository

helm repo add cowork https://charts.cowork.ink
helm repo update

Verify the repository is available:

helm search repo cowork
# Should show: cowork/business    1.x.x    AI Agent Platform for Business

Step 2: Create the Namespace

kubectl create namespace cowork

Step 3: Create the Configuration Values File

Create a values.yaml file with your deployment configuration:

# values.yaml — cowork.ink Business configuration

# LLM Provider Configuration
llm:
  provider: openai        # openai | anthropic | ollama | custom
  apiKey: "YOUR_API_KEY"  # Set via secret in production (see Step 4b)
  model: "gpt-4o-mini"   # Default model for agents

# Storage (use your cluster's storage class)
persistence:
  storageClass: "standard"  # or "gp3", "fast", "nfs", etc.
  size: "50Gi"

# Ingress (optional but recommended for production)
ingress:
  enabled: true
  host: "agents.yourdomain.com"
  tls: true
  certManager: true         # Uses cert-manager for auto TLS

# Admin configuration
admin:
  email: "admin@yourdomain.com"
  # Password set on first login

# Scaling
replicaCount: 1             # Increase for HA
agentsPerNode: 200          # Max concurrent agents per node

# SSO (optional — configure after initial deployment)
sso:
  enabled: false
  provider: ""              # okta | google | azure | saml

Step 4a: Deploy (Development/Quick Start)

helm install cowork-business cowork/business \
  --namespace cowork \
  --values values.yaml \
  --wait

Wait for all pods to be ready (usually 60–90 seconds):

kubectl get pods -n cowork --watch

Step 4b: Deploy (Production — API Keys as Secrets)

In production, don't put API keys in values.yaml. Use Kubernetes Secrets:

# Create the API key secret
kubectl create secret generic cowork-llm-secret \
  --from-literal=apiKey="YOUR_ACTUAL_API_KEY" \
  --namespace cowork

# Update values.yaml to reference the secret
# llm:
#   apiKeySecretName: cowork-llm-secret

helm install cowork-business cowork/business \
  --namespace cowork \
  --values values.yaml

Step 5: Verify the Deployment

# Check all pods are running
kubectl get pods -n cowork

# Expected output:
# NAME                              READY   STATUS    RESTARTS
# cowork-api-7d8f9b-xxx             1/1     Running   0
# cowork-worker-5c6d7e-xxx          1/1     Running   0
# cowork-worker-5c6d7e-yyy          1/1     Running   0
# cowork-redis-0                    1/1     Running   0
# cowork-postgres-0                 1/1     Running   0

Step 6: Access the Admin Panel

# If using ingress: https://agents.yourdomain.com
# If not: port-forward for initial access
kubectl port-forward -n cowork svc/cowork-api 8080:80

# Visit: http://localhost:8080

Complete the first-login setup wizard: set admin password, verify LLM connection, create your first team.


Initial Configuration (10 Minutes)

After deployment, run through this setup checklist:

Admin Panel Setup

  1. Set admin password — prompted on first login
  2. Create your first team — Admin → Teams → New Team
  3. Invite team members — Admin → Users → Invite
  4. Create roles — Admin → Roles → Define role permissions

Create Your First Agent

  1. Go to Agents → New Agent
  2. Set the agent name, description, and system prompt
  3. Configure tools (what actions can this agent take?)
  4. Assign to a team
  5. Run a test query

Configure Integrations

Connect your business systems:

  • Email (Gmail/Outlook): Settings → Integrations → Email
  • Slack: Settings → Integrations → Slack
  • CRM (HubSpot/Salesforce): Settings → Integrations → CRM
  • Custom APIs: Settings → Tools → Custom API

Production Hardening

TLS and Ingress

If not using cert-manager:

# Create TLS secret from your certificate
kubectl create secret tls cowork-tls \
  --cert=path/to/tls.crt \
  --key=path/to/tls.key \
  --namespace cowork

Network Policies

Restrict agent network access to only required endpoints:

# network-policy.yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: cowork-agent-policy
  namespace: cowork
spec:
  podSelector:
    matchLabels:
      app: cowork-worker
  egress:
  - to:
    - ipBlock:
        cidr: 0.0.0.0/0
        except:
          - 169.254.0.0/16  # Block cloud metadata endpoints
    ports:
    - protocol: TCP
      port: 443  # HTTPS only

Resource Limits

Set resource limits to prevent runaway agent costs:

# Add to values.yaml
resources:
  worker:
    limits:
      cpu: "2"
      memory: "4Gi"
    requests:
      cpu: "500m"
      memory: "1Gi"

Backup Configuration

# PostgreSQL backup (add to your cron scheduler)
kubectl exec -n cowork cowork-postgres-0 -- \
  pg_dump -U cowork cowork_db > backup-$(date +%Y%m%d).sql

Adding Self-Hosted Models (Optional)

To eliminate LLM API costs entirely, add Ollama to the cluster:

# Deploy Ollama (GPU node required for 70B models)
kubectl apply -f - <<EOF
apiVersion: apps/v1
kind: Deployment
metadata:
  name: ollama
  namespace: cowork
spec:
  replicas: 1
  selector:
    matchLabels:
      app: ollama
  template:
    spec:
      containers:
      - name: ollama
        image: ollama/ollama:latest
        resources:
          limits:
            nvidia.com/gpu: "1"
        ports:
        - containerPort: 11434
---
apiVersion: v1
kind: Service
metadata:
  name: ollama-service
  namespace: cowork
spec:
  selector:
    app: ollama
  ports:
  - port: 11434
EOF

# Pull the model
kubectl exec -n cowork ollama-pod-name -- ollama pull llama3.3

# Update cowork config to use Ollama
helm upgrade cowork-business cowork/business \
  --namespace cowork \
  --set llm.provider=ollama \
  --set llm.endpoint=http://ollama-service:11434 \
  --set llm.model=llama3.3

Monitoring and Observability

cowork.ink Business exposes Prometheus metrics by default:

# Add Prometheus scrape config
kubectl apply -f - <<EOF
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: cowork-monitor
  namespace: cowork
spec:
  selector:
    matchLabels:
      app: cowork-api
  endpoints:
  - port: metrics
    interval: 30s
EOF

Key metrics to watch:

  • cowork_agent_tasks_total — total tasks executed
  • cowork_agent_task_duration_seconds — latency histogram
  • cowork_agent_errors_total — error rate
  • cowork_token_usage_total — LLM token consumption
  • cowork_agent_active — currently running agents

Troubleshooting Common Issues

IssueDiagnosisFix
Pods in CrashLoopBackOffkubectl logs -n cowork pod-nameUsually config error — check values.yaml
LLM API errorsCheck API key, model nameVerify in Settings → LLM Configuration
Slow agent responsesCPU/memory pressureCheck kubectl top pods -n cowork
Agents not scalingResource limits hitIncrease limits or add nodes
Database connection errorsPostgreSQL pod healthCheck kubectl describe pod cowork-postgres-0
Need help with your deployment?

The cowork.ink Business documentation includes a troubleshooting guide, and the community forum has active responses for deployment questions. For enterprise support contracts, contact the team via the business page.


Upgrade Path

Upgrading to new versions:

# Check available versions
helm search repo cowork --versions

# Update Helm repo
helm repo update

# Upgrade (rolling update, zero downtime)
helm upgrade cowork-business cowork/business \
  --namespace cowork \
  --values values.yaml \
  --reuse-values  # Keep your existing configuration

Review the changelog before upgrading major versions — configuration keys may change.


You now have a production-grade self-hosted AI agent platform. The GoGogot runtime underneath handles scaling and reliability; the cowork.ink Business management layer handles governance and visibility. Your data stays on your infrastructure, your costs are predictable, and you have no vendor lock-in.

For expanding your deployment with more use cases, see the AI agents for business automation guide. For security hardening beyond what's covered here, see the AI agent security guide.

Frequently Asked Questions

What does it mean to self-host AI agents?
Self-hosting AI agents means running the agent software on infrastructure you control — your own servers, a private cloud VPC, or on-premise hardware — rather than using a vendor's cloud platform. Your data, agent configurations, and conversation history stay entirely within your infrastructure.
Do I need Kubernetes to self-host AI agents?
Kubernetes is the recommended production deployment method for scale and reliability, but it's not required. cowork.ink Business also supports Docker Compose for smaller deployments (1–20 agents). Kubernetes becomes valuable when you need horizontal scaling, automatic failover, and enterprise-grade reliability.
How many AI agents can I run on self-hosted infrastructure?
cowork.ink Business supports 200 agents per Kubernetes node. A 3-node cluster (modest hardware) handles 600 concurrent agents. Vertical scaling (adding nodes) is linear — adding a node adds 200 agent capacity.
What hardware do I need to self-host AI agents?
For API-based models (OpenAI, Anthropic): any 3 VMs with 4 vCPUs and 8GB RAM each runs comfortably. For self-hosted open-source models (Llama 3.3 70B): 2× NVIDIA A100 80GB or equivalent. For smaller models (Mistral 7B, Llama 8B): a single gaming GPU (RTX 4090, A10G) handles moderate loads.
Home Blog Company