Quick Answer: An AI agent knowledge base gives your agent access to private, up-to-date facts through a RAG pipeline: chunk your documents → embed them → store in a vector database → retrieve at query time. The quality of that pipeline determines whether your agent reasons correctly or hallucinates.
An agent without a knowledge base is a brilliant amnesiac. It may understand your question perfectly — and still answer it wrong because it lacks the facts. Building an AI agent knowledge base is the engineering discipline that closes that gap.
This guide covers every decision in the pipeline: which document formats to use, how to chunk them for maximum retrieval accuracy, which indexing strategy matches your query patterns, and how to keep everything fresh. Teams at cowork.ink connect shared knowledge sources to agents that every engineer on the team can query — no duplicated RAG plumbing per project.
What Is an AI Agent Knowledge Base?
An AI agent knowledge base is a curated, indexed repository of information that an agent retrieves at runtime — rather than relying on what was baked into its training weights.
The mechanism is Retrieval-Augmented Generation (RAG): when a user asks a question, the agent embeds the query, searches the knowledge base for semantically similar chunks, and injects those chunks as context before generating a response. The result: the agent answers from your data, not from hallucination.
This is distinct from the agent's working memory and context window — the knowledge base is persistent storage that outlives any single conversation. For a deeper look at memory types, see our guide to AI agent memory architectures.
Knowledge Base vs. Fine-Tuning
The first architectural decision: should you store knowledge in a database or bake it into the model via fine-tuning?
| Criterion | Knowledge Base (RAG) | Fine-Tuning |
|---|---|---|
| Updates | Real-time, incremental | Retraining cycle (days/weeks) |
| Fact precision | High — cites sources | Moderate — weights blend facts |
| Cost | Storage + inference | Training GPU costs ($100–$10K+) |
| Hallucination risk | Low with good retrieval | Higher — model may confabulate |
| Private data | Strong — data stays in your DB | Requires sharing data with trainer |
| Best for | Evolving docs, product knowledge | Style, tone, task specialization |
For most teams, the answer is RAG first. Fine-tune only to teach the agent how to behave — not what to know. See our deep-dive on RAG vs. fine-tuning for the full comparison.
Step 1: Choose Your Document Formats
The format your documents arrive in directly affects retrieval quality. Some formats chunk cleanly; others require heavy preprocessing.
| Format | Best For | Retrieval Quality | Preparation Needed |
|---|---|---|---|
| Markdown (.md) | Docs, wikis, READMEs | Highest | None — use heading-aware chunker |
| Plain text (.txt) | Logs, simple Q&A | Good | Minimal |
| JSON / JSONL | Product catalogs, configs | Excellent for structured facts | Flatten nested objects |
| CSV | FAQs, tabular data | Good for rows | Convert rows to text records |
| Manuals, compliance docs | Variable | Requires Docling or PyMuPDF to extract clean text | |
| HTML | Web content, help pages | Moderate | Strip nav/footer boilerplate |
| DOCX | Corporate docs | Moderate | Extract with python-docx |
The practical recommendation: standardize on Markdown for prose, JSON for structured facts. Convert everything else to one of these two formats during your ingestion pipeline. Markdown's heading hierarchy is especially valuable — it gives your chunker natural split points that align with semantic meaning.
PDFs embed text in a layout grid — the extracted order is often wrong. Multi-column PDFs and scanned documents require OCR. Always validate extracted text before indexing. IBM's open-source Docling library handles PDFs, DOCX, and images into clean Markdown — worth adding to your ingestion stack.
Step 2: Design Your Chunking Strategy
Chunking is the most consequential step in the pipeline. Chunk too large and retrieval becomes imprecise; chunk too small and you lose context. The NVIDIA developer blog benchmark tested every major strategy across real QA tasks — here's what they found.
Chunking strategies compared
| Strategy | How It Works | Best For | Avg Accuracy |
|---|---|---|---|
| Fixed-size | Split every N tokens with overlap | Quick baselines | 0.603–0.645 |
| Recursive | Split on paragraph → sentence → char | Unstructured articles | 0.630 |
| Structural / doc-based | Split on Markdown headers, HTML tags | Markdown, code, HTML | 0.638 |
| Semantic | Split where sentence similarity drops | Academic, legal, technical | 0.641 |
| Page-level | One chunk per page | Dense reports, books | 0.648 |
| Hierarchical | Multi-level: summary → section → para | Long documents, complex queries | High — flexible |
| LLM-based | LLM defines chunk boundaries | High-value, varied documents | Highest, expensive |
512 tokens with 50–100 token overlap is the empirically validated starting point. It achieves 0.640+ average accuracy across QA, summarization, and retrieval tasks. For factoid queries, go down to 256 tokens. For analytical or synthesis queries, go up to 1,024 or page-level.
Overlap is not optional
Every chunking strategy should include token overlap between adjacent chunks. Without overlap, a sentence that crosses a chunk boundary becomes invisible to both chunks. A 10–20% overlap (50–100 tokens for 512-token chunks) recovers this signal with minimal storage cost.
Structural chunking for Markdown
If your knowledge base is primarily documentation or wikis, use a Markdown-aware splitter that respects heading levels. Splitting at ## boundaries preserves semantic coherence far better than arbitrary token counts. LlamaIndex's MarkdownNodeParser and LangChain's MarkdownHeaderTextSplitter both do this correctly.
Step 3: Choose Your Indexing Method
Once chunked and embedded, how you index determines retrieval precision. Three methods dominate production systems:
Vector indexing (dense retrieval)
Chunk text is converted to a high-dimensional vector via an embedding model. Queries are also embedded; the database returns the top-k chunks by cosine or dot-product similarity via Approximate Nearest Neighbor (ANN) search.
Best for: Conceptual, semantic queries — "explain our refund process," "how does the rate limiter work."
Fails on: Exact matches for product codes, SKUs, brand names, proprietary terminology — anything where semantics doesn't help.
Keyword indexing (BM25, sparse retrieval)
Traditional inverted index. Scores chunks by term frequency and inverse document frequency. Fast, deterministic, no embedding cost.
Best for: Exact term lookup, product codes, API endpoint names, technical identifiers.
Fails on: Synonyms and conceptual queries — "shutting it down" won't find chunks containing "power off."
Hybrid search (recommended for production)
Combines dense vector scores and BM25 sparse scores using Reciprocal Rank Fusion (RRF) or Relative Score Fusion. Covers semantic queries AND exact-match queries in one pass.
According to Weaviate's chunking and retrieval research, hybrid search consistently outperforms pure vector search on mixed-intent enterprise corpora. Most modern vector databases — Weaviate, Qdrant, Turbopuffer, Azure AI Search, and Elastic — support hybrid natively.
Step 4: Pick Your Vector Database
The vector database is where your chunks live and where retrieval happens. Your choice affects latency, cost, and the types of queries you can run.
| Database | Best For | Hybrid Search | Self-Hosted | Managed |
|---|---|---|---|---|
| Chroma | Local dev, prototypes | No | Yes | No |
| Pinecone | Serverless production | Yes | No | Yes |
| Weaviate | Hybrid + graph-compatible | Yes | Yes | Yes |
| Qdrant | High-speed production | Yes | Yes | Yes |
| pgvector | Existing Postgres users | Partial | Yes | Yes |
| Turbopuffer | Production teams (Cursor, Notion, Linear) | Yes | No | Yes |
| FAISS | Research, local high-performance | No | Yes | No |
| Azure AI Search | Enterprise + Microsoft stack | Yes | No | Yes |
Decision heuristic:
- Prototyping: Chroma — zero setup, in-memory or local file
- Production, small team: Qdrant self-hosted or pgvector (if you already run Postgres)
- Production, enterprise: Pinecone serverless or Weaviate Cloud
- Microsoft-native stack: Azure AI Search with agentic retrieval pipeline
For team agents on cowork.ink, knowledge sources are connected at the workspace level — every agent in the team's workspace automatically retrieves from the same indexed corpus.
Step 5: Select an Embedding Model
The embedding model converts text chunks to vectors. Better embeddings → better semantic similarity → better retrieval.
| Model | Accuracy | Cost | Best For |
|---|---|---|---|
| text-embedding-3-large (OpenAI) | High | Moderate | General-purpose production |
| voyage-3.5-lite (Anthropic/Voyage) | 66.1% | Low | Best cost/accuracy ratio |
| mistral-embed | 77.8% | Moderate | Precision-critical use cases |
| bge-large-en-1.5 | High | Free (open-source) | On-premise, privacy-first |
| all-MiniLM-L6-v2 | Moderate | Free | Lightweight, fast, local |
One critical constraint: use the same embedding model for ingestion and query time. Embedding models are not interchangeable — switching models requires re-embedding your entire corpus.
Step 6: Keep the Knowledge Base Fresh
A knowledge base built once and left to age becomes a liability. Documents update; the index doesn't; the agent serves stale answers.
Build an incremental ingestion pipeline from day one:
- Detect changes — webhooks on document updates, git hooks for docs-as-code repos, or polling on a schedule
- Re-ingest changed documents only — full re-embedding is expensive; track document hashes to skip unchanged files
- Store metadata with every chunk — source URL, timestamp, document version, author — this enables filtered retrieval and stale-chunk identification
- Invalidate and replace stale vectors — when a document changes, delete its old chunks from the vector DB before inserting new ones
Both LlamaIndex and Amazon Bedrock Knowledge Bases support incremental refresh pipelines out of the box. For teams, connecting your knowledge base to agents via Model Context Protocol (MCP) lets you update sources once and have all connected agents retrieve from the latest version automatically.
Step 7: Evaluate Retrieval Quality
You can't improve what you don't measure. Use RAGAS to evaluate your pipeline on four metrics:
- Faithfulness — does the answer match the retrieved context? (target: >0.8)
- Context precision — are the retrieved chunks relevant to the question? (target: >0.75)
- Context recall — did retrieval find all the necessary information?
- Answer relevancy — is the final answer on-topic?
Run RAGAS on a golden question set of 50–100 Q&A pairs that cover your knowledge domain. Use it to compare chunking strategies, embedding models, and retrieval methods before committing to production.
The Build Checklist
Don't optimize everything at once. Get a working pipeline first, then improve each stage based on RAGAS metrics.
- Standardize formats — convert all sources to Markdown or JSON
- Baseline chunking — 512 tokens, 10% overlap, structural splitting for Markdown
- Choose embedding model —
text-embedding-3-small(fast + cheap) for prototype;voyage-3.5-liteormistral-embedfor production - Pick a vector DB — Chroma for local; Qdrant or Weaviate for production
- Enable hybrid search — add BM25 index alongside vector index
- Add metadata to every chunk — source, timestamp, section title
- Build the ingestion pipeline — incremental, hash-based, with soft deletes for old vectors
- Build the retrieval layer — top-k with metadata filtering, re-ranking for high-stakes queries
- Evaluate with RAGAS — run against your golden Q&A set
- Connect to your agents — expose via MCP, LangChain retriever, or LlamaIndex query engine
Get Started
A well-built AI agent knowledge base is the foundation that makes everything else — multi-agent collaboration, agentic RAG pipelines, automated code review — actually work at quality.
Try cowork.ink free — connect your team's knowledge sources to shared AI agents in minutes. Every engineer on the team queries the same fresh, indexed corpus — no duplicated RAG plumbing per person.