Self-hosting AI agents used to mean months of engineering work — custom runtime development, manual scaling logic, hand-rolled observability. That era is over. Modern self-hosted platforms like cowork.ink Business deploy in under 5 minutes on Kubernetes, with enterprise management features included out of the box.
This guide covers the complete deployment: prerequisites, step-by-step installation, initial configuration, and production hardening. By the end, you'll have a fully operational AI agent platform running on your own infrastructure.
Why Self-Host?
Before the technical guide, the business case:
Data sovereignty: Your data never leaves your servers. GDPR, HIPAA, and attorney-client privilege requirements are satisfied by architecture, not just contracts.
Cost at scale: SaaS platforms charge $4–20 per user per month. At 100 users with moderate usage, self-hosted infrastructure costs $150–400/month total — 70–90% lower than comparable SaaS.
No rate limits: Cloud platforms impose per-minute or per-day API limits. Your self-hosted platform runs at the limit of your hardware.
Open-source models: Use Llama 3.3, Mistral, or Qwen for zero per-token API costs. Some organizations achieve complete cost elimination for agent inference.
Vendor independence: No price increases, no platform changes, no "we're deprecating this feature" emails.
Prerequisites
Minimum Requirements (Development / Small Team)
- Kubernetes cluster: 3 nodes, 4 vCPUs + 8GB RAM each
- Helm 3.x installed
kubectlconfigured for your cluster- LLM API key (OpenAI, Anthropic) OR self-hosted model (Ollama)
- Persistent storage class (any cloud-native or local-path)
Recommended (Production / 50+ Users)
- Kubernetes cluster: 5+ nodes, 8 vCPUs + 16GB RAM each
- Dedicated node pool for agent execution
- PostgreSQL database (for agent state persistence)
- Redis (for agent message queue)
- Ingress controller (nginx or traefik) with TLS
- Monitoring stack (Prometheus + Grafana, or existing)
For Self-Hosted Models (Fully Air-Gapped)
Additional requirements:
- GPU nodes: 2× NVIDIA A100 80GB for 70B models; A10G or RTX 4090 for smaller models
- Ollama or vLLM installed and serving your chosen model
A 3-node cluster with external LLM API is the right starting point for most teams. You can upgrade to self-hosted models and larger clusters as usage grows. The deployment steps below are the same regardless of cluster size.
Step-by-Step Deployment
Step 1: Add the Helm Repository
helm repo add cowork https://charts.cowork.ink
helm repo update
Verify the repository is available:
helm search repo cowork
# Should show: cowork/business 1.x.x AI Agent Platform for Business
Step 2: Create the Namespace
kubectl create namespace cowork
Step 3: Create the Configuration Values File
Create a values.yaml file with your deployment configuration:
# values.yaml — cowork.ink Business configuration
# LLM Provider Configuration
llm:
provider: openai # openai | anthropic | ollama | custom
apiKey: "YOUR_API_KEY" # Set via secret in production (see Step 4b)
model: "gpt-4o-mini" # Default model for agents
# Storage (use your cluster's storage class)
persistence:
storageClass: "standard" # or "gp3", "fast", "nfs", etc.
size: "50Gi"
# Ingress (optional but recommended for production)
ingress:
enabled: true
host: "agents.yourdomain.com"
tls: true
certManager: true # Uses cert-manager for auto TLS
# Admin configuration
admin:
email: "admin@yourdomain.com"
# Password set on first login
# Scaling
replicaCount: 1 # Increase for HA
agentsPerNode: 200 # Max concurrent agents per node
# SSO (optional — configure after initial deployment)
sso:
enabled: false
provider: "" # okta | google | azure | saml
Step 4a: Deploy (Development/Quick Start)
helm install cowork-business cowork/business \
--namespace cowork \
--values values.yaml \
--wait
Wait for all pods to be ready (usually 60–90 seconds):
kubectl get pods -n cowork --watch
Step 4b: Deploy (Production — API Keys as Secrets)
In production, don't put API keys in values.yaml. Use Kubernetes Secrets:
# Create the API key secret
kubectl create secret generic cowork-llm-secret \
--from-literal=apiKey="YOUR_ACTUAL_API_KEY" \
--namespace cowork
# Update values.yaml to reference the secret
# llm:
# apiKeySecretName: cowork-llm-secret
helm install cowork-business cowork/business \
--namespace cowork \
--values values.yaml
Step 5: Verify the Deployment
# Check all pods are running
kubectl get pods -n cowork
# Expected output:
# NAME READY STATUS RESTARTS
# cowork-api-7d8f9b-xxx 1/1 Running 0
# cowork-worker-5c6d7e-xxx 1/1 Running 0
# cowork-worker-5c6d7e-yyy 1/1 Running 0
# cowork-redis-0 1/1 Running 0
# cowork-postgres-0 1/1 Running 0
Step 6: Access the Admin Panel
# If using ingress: https://agents.yourdomain.com
# If not: port-forward for initial access
kubectl port-forward -n cowork svc/cowork-api 8080:80
# Visit: http://localhost:8080
Complete the first-login setup wizard: set admin password, verify LLM connection, create your first team.
Initial Configuration (10 Minutes)
After deployment, run through this setup checklist:
Admin Panel Setup
- Set admin password — prompted on first login
- Create your first team — Admin → Teams → New Team
- Invite team members — Admin → Users → Invite
- Create roles — Admin → Roles → Define role permissions
Create Your First Agent
- Go to Agents → New Agent
- Set the agent name, description, and system prompt
- Configure tools (what actions can this agent take?)
- Assign to a team
- Run a test query
Configure Integrations
Connect your business systems:
- Email (Gmail/Outlook): Settings → Integrations → Email
- Slack: Settings → Integrations → Slack
- CRM (HubSpot/Salesforce): Settings → Integrations → CRM
- Custom APIs: Settings → Tools → Custom API
Production Hardening
TLS and Ingress
If not using cert-manager:
# Create TLS secret from your certificate
kubectl create secret tls cowork-tls \
--cert=path/to/tls.crt \
--key=path/to/tls.key \
--namespace cowork
Network Policies
Restrict agent network access to only required endpoints:
# network-policy.yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: cowork-agent-policy
namespace: cowork
spec:
podSelector:
matchLabels:
app: cowork-worker
egress:
- to:
- ipBlock:
cidr: 0.0.0.0/0
except:
- 169.254.0.0/16 # Block cloud metadata endpoints
ports:
- protocol: TCP
port: 443 # HTTPS only
Resource Limits
Set resource limits to prevent runaway agent costs:
# Add to values.yaml
resources:
worker:
limits:
cpu: "2"
memory: "4Gi"
requests:
cpu: "500m"
memory: "1Gi"
Backup Configuration
# PostgreSQL backup (add to your cron scheduler)
kubectl exec -n cowork cowork-postgres-0 -- \
pg_dump -U cowork cowork_db > backup-$(date +%Y%m%d).sql
Adding Self-Hosted Models (Optional)
To eliminate LLM API costs entirely, add Ollama to the cluster:
# Deploy Ollama (GPU node required for 70B models)
kubectl apply -f - <<EOF
apiVersion: apps/v1
kind: Deployment
metadata:
name: ollama
namespace: cowork
spec:
replicas: 1
selector:
matchLabels:
app: ollama
template:
spec:
containers:
- name: ollama
image: ollama/ollama:latest
resources:
limits:
nvidia.com/gpu: "1"
ports:
- containerPort: 11434
---
apiVersion: v1
kind: Service
metadata:
name: ollama-service
namespace: cowork
spec:
selector:
app: ollama
ports:
- port: 11434
EOF
# Pull the model
kubectl exec -n cowork ollama-pod-name -- ollama pull llama3.3
# Update cowork config to use Ollama
helm upgrade cowork-business cowork/business \
--namespace cowork \
--set llm.provider=ollama \
--set llm.endpoint=http://ollama-service:11434 \
--set llm.model=llama3.3
Monitoring and Observability
cowork.ink Business exposes Prometheus metrics by default:
# Add Prometheus scrape config
kubectl apply -f - <<EOF
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: cowork-monitor
namespace: cowork
spec:
selector:
matchLabels:
app: cowork-api
endpoints:
- port: metrics
interval: 30s
EOF
Key metrics to watch:
cowork_agent_tasks_total— total tasks executedcowork_agent_task_duration_seconds— latency histogramcowork_agent_errors_total— error ratecowork_token_usage_total— LLM token consumptioncowork_agent_active— currently running agents
Troubleshooting Common Issues
| Issue | Diagnosis | Fix |
|---|---|---|
| Pods in CrashLoopBackOff | kubectl logs -n cowork pod-name | Usually config error — check values.yaml |
| LLM API errors | Check API key, model name | Verify in Settings → LLM Configuration |
| Slow agent responses | CPU/memory pressure | Check kubectl top pods -n cowork |
| Agents not scaling | Resource limits hit | Increase limits or add nodes |
| Database connection errors | PostgreSQL pod health | Check kubectl describe pod cowork-postgres-0 |
The cowork.ink Business documentation includes a troubleshooting guide, and the community forum has active responses for deployment questions. For enterprise support contracts, contact the team via the business page.
Upgrade Path
Upgrading to new versions:
# Check available versions
helm search repo cowork --versions
# Update Helm repo
helm repo update
# Upgrade (rolling update, zero downtime)
helm upgrade cowork-business cowork/business \
--namespace cowork \
--values values.yaml \
--reuse-values # Keep your existing configuration
Review the changelog before upgrading major versions — configuration keys may change.
You now have a production-grade self-hosted AI agent platform. The GoGogot runtime underneath handles scaling and reliability; the cowork.ink Business management layer handles governance and visibility. Your data stays on your infrastructure, your costs are predictable, and you have no vendor lock-in.
For expanding your deployment with more use cases, see the AI agents for business automation guide. For security hardening beyond what's covered here, see the AI agent security guide.