AI Agent Knowledge Base: Formats, Chunking & Indexing

Build a BETTER AI agent knowledge base. NVIDIA benchmarks on chunk sizes, format tradeoffs, hybrid search, and RAG pipelines. Step-by-step guide for engineers.

Quick Answer: An AI agent knowledge base gives your agent access to private, up-to-date facts through a RAG pipeline: chunk your documents → embed them → store in a vector database → retrieve at query time. The quality of that pipeline determines whether your agent reasons correctly or hallucinates.


An agent without a knowledge base is a brilliant amnesiac. It may understand your question perfectly — and still answer it wrong because it lacks the facts. Building an AI agent knowledge base is the engineering discipline that closes that gap.

This guide covers every decision in the pipeline: which document formats to use, how to chunk them for maximum retrieval accuracy, which indexing strategy matches your query patterns, and how to keep everything fresh. Teams at cowork.ink connect shared knowledge sources to agents that every engineer on the team can query — no duplicated RAG plumbing per project.


What Is an AI Agent Knowledge Base?

An AI agent knowledge base is a curated, indexed repository of information that an agent retrieves at runtime — rather than relying on what was baked into its training weights.

The mechanism is Retrieval-Augmented Generation (RAG): when a user asks a question, the agent embeds the query, searches the knowledge base for semantically similar chunks, and injects those chunks as context before generating a response. The result: the agent answers from your data, not from hallucination.

This is distinct from the agent's working memory and context window — the knowledge base is persistent storage that outlives any single conversation. For a deeper look at memory types, see our guide to AI agent memory architectures.


Knowledge Base vs. Fine-Tuning

The first architectural decision: should you store knowledge in a database or bake it into the model via fine-tuning?

CriterionKnowledge Base (RAG)Fine-Tuning
UpdatesReal-time, incrementalRetraining cycle (days/weeks)
Fact precisionHigh — cites sourcesModerate — weights blend facts
CostStorage + inferenceTraining GPU costs ($100–$10K+)
Hallucination riskLow with good retrievalHigher — model may confabulate
Private dataStrong — data stays in your DBRequires sharing data with trainer
Best forEvolving docs, product knowledgeStyle, tone, task specialization

For most teams, the answer is RAG first. Fine-tune only to teach the agent how to behave — not what to know. See our deep-dive on RAG vs. fine-tuning for the full comparison.


Step 1: Choose Your Document Formats

The format your documents arrive in directly affects retrieval quality. Some formats chunk cleanly; others require heavy preprocessing.

FormatBest ForRetrieval QualityPreparation Needed
Markdown (.md)Docs, wikis, READMEsHighestNone — use heading-aware chunker
Plain text (.txt)Logs, simple Q&AGoodMinimal
JSON / JSONLProduct catalogs, configsExcellent for structured factsFlatten nested objects
CSVFAQs, tabular dataGood for rowsConvert rows to text records
PDFManuals, compliance docsVariableRequires Docling or PyMuPDF to extract clean text
HTMLWeb content, help pagesModerateStrip nav/footer boilerplate
DOCXCorporate docsModerateExtract with python-docx

The practical recommendation: standardize on Markdown for prose, JSON for structured facts. Convert everything else to one of these two formats during your ingestion pipeline. Markdown's heading hierarchy is especially valuable — it gives your chunker natural split points that align with semantic meaning.

PDF Warning

PDFs embed text in a layout grid — the extracted order is often wrong. Multi-column PDFs and scanned documents require OCR. Always validate extracted text before indexing. IBM's open-source Docling library handles PDFs, DOCX, and images into clean Markdown — worth adding to your ingestion stack.


Step 2: Design Your Chunking Strategy

Chunking is the most consequential step in the pipeline. Chunk too large and retrieval becomes imprecise; chunk too small and you lose context. The NVIDIA developer blog benchmark tested every major strategy across real QA tasks — here's what they found.

Chunking strategies compared

StrategyHow It WorksBest ForAvg Accuracy
Fixed-sizeSplit every N tokens with overlapQuick baselines0.603–0.645
RecursiveSplit on paragraph → sentence → charUnstructured articles0.630
Structural / doc-basedSplit on Markdown headers, HTML tagsMarkdown, code, HTML0.638
SemanticSplit where sentence similarity dropsAcademic, legal, technical0.641
Page-levelOne chunk per pageDense reports, books0.648
HierarchicalMulti-level: summary → section → paraLong documents, complex queriesHigh — flexible
LLM-basedLLM defines chunk boundariesHigh-value, varied documentsHighest, expensive
Starting Baseline

512 tokens with 50–100 token overlap is the empirically validated starting point. It achieves 0.640+ average accuracy across QA, summarization, and retrieval tasks. For factoid queries, go down to 256 tokens. For analytical or synthesis queries, go up to 1,024 or page-level.

Overlap is not optional

Every chunking strategy should include token overlap between adjacent chunks. Without overlap, a sentence that crosses a chunk boundary becomes invisible to both chunks. A 10–20% overlap (50–100 tokens for 512-token chunks) recovers this signal with minimal storage cost.

Structural chunking for Markdown

If your knowledge base is primarily documentation or wikis, use a Markdown-aware splitter that respects heading levels. Splitting at ## boundaries preserves semantic coherence far better than arbitrary token counts. LlamaIndex's MarkdownNodeParser and LangChain's MarkdownHeaderTextSplitter both do this correctly.


Step 3: Choose Your Indexing Method

Once chunked and embedded, how you index determines retrieval precision. Three methods dominate production systems:

Vector indexing (dense retrieval)

Chunk text is converted to a high-dimensional vector via an embedding model. Queries are also embedded; the database returns the top-k chunks by cosine or dot-product similarity via Approximate Nearest Neighbor (ANN) search.

Best for: Conceptual, semantic queries — "explain our refund process," "how does the rate limiter work."

Fails on: Exact matches for product codes, SKUs, brand names, proprietary terminology — anything where semantics doesn't help.

Keyword indexing (BM25, sparse retrieval)

Traditional inverted index. Scores chunks by term frequency and inverse document frequency. Fast, deterministic, no embedding cost.

Best for: Exact term lookup, product codes, API endpoint names, technical identifiers.

Fails on: Synonyms and conceptual queries — "shutting it down" won't find chunks containing "power off."

Hybrid search (recommended for production)

Combines dense vector scores and BM25 sparse scores using Reciprocal Rank Fusion (RRF) or Relative Score Fusion. Covers semantic queries AND exact-match queries in one pass.

According to Weaviate's chunking and retrieval research, hybrid search consistently outperforms pure vector search on mixed-intent enterprise corpora. Most modern vector databases — Weaviate, Qdrant, Turbopuffer, Azure AI Search, and Elastic — support hybrid natively.


Step 4: Pick Your Vector Database

The vector database is where your chunks live and where retrieval happens. Your choice affects latency, cost, and the types of queries you can run.

DatabaseBest ForHybrid SearchSelf-HostedManaged
ChromaLocal dev, prototypesNoYesNo
PineconeServerless productionYesNoYes
WeaviateHybrid + graph-compatibleYesYesYes
QdrantHigh-speed productionYesYesYes
pgvectorExisting Postgres usersPartialYesYes
TurbopufferProduction teams (Cursor, Notion, Linear)YesNoYes
FAISSResearch, local high-performanceNoYesNo
Azure AI SearchEnterprise + Microsoft stackYesNoYes

Decision heuristic:

  • Prototyping: Chroma — zero setup, in-memory or local file
  • Production, small team: Qdrant self-hosted or pgvector (if you already run Postgres)
  • Production, enterprise: Pinecone serverless or Weaviate Cloud
  • Microsoft-native stack: Azure AI Search with agentic retrieval pipeline

For team agents on cowork.ink, knowledge sources are connected at the workspace level — every agent in the team's workspace automatically retrieves from the same indexed corpus.


Step 5: Select an Embedding Model

The embedding model converts text chunks to vectors. Better embeddings → better semantic similarity → better retrieval.

ModelAccuracyCostBest For
text-embedding-3-large (OpenAI)HighModerateGeneral-purpose production
voyage-3.5-lite (Anthropic/Voyage)66.1%LowBest cost/accuracy ratio
mistral-embed77.8%ModeratePrecision-critical use cases
bge-large-en-1.5HighFree (open-source)On-premise, privacy-first
all-MiniLM-L6-v2ModerateFreeLightweight, fast, local

One critical constraint: use the same embedding model for ingestion and query time. Embedding models are not interchangeable — switching models requires re-embedding your entire corpus.


Step 6: Keep the Knowledge Base Fresh

A knowledge base built once and left to age becomes a liability. Documents update; the index doesn't; the agent serves stale answers.

Build an incremental ingestion pipeline from day one:

  1. Detect changes — webhooks on document updates, git hooks for docs-as-code repos, or polling on a schedule
  2. Re-ingest changed documents only — full re-embedding is expensive; track document hashes to skip unchanged files
  3. Store metadata with every chunk — source URL, timestamp, document version, author — this enables filtered retrieval and stale-chunk identification
  4. Invalidate and replace stale vectors — when a document changes, delete its old chunks from the vector DB before inserting new ones

Both LlamaIndex and Amazon Bedrock Knowledge Bases support incremental refresh pipelines out of the box. For teams, connecting your knowledge base to agents via Model Context Protocol (MCP) lets you update sources once and have all connected agents retrieve from the latest version automatically.


Step 7: Evaluate Retrieval Quality

You can't improve what you don't measure. Use RAGAS to evaluate your pipeline on four metrics:

  • Faithfulness — does the answer match the retrieved context? (target: >0.8)
  • Context precision — are the retrieved chunks relevant to the question? (target: >0.75)
  • Context recall — did retrieval find all the necessary information?
  • Answer relevancy — is the final answer on-topic?

Run RAGAS on a golden question set of 50–100 Q&A pairs that cover your knowledge domain. Use it to compare chunking strategies, embedding models, and retrieval methods before committing to production.


The Build Checklist

Ship in this order

Don't optimize everything at once. Get a working pipeline first, then improve each stage based on RAGAS metrics.

  1. Standardize formats — convert all sources to Markdown or JSON
  2. Baseline chunking — 512 tokens, 10% overlap, structural splitting for Markdown
  3. Choose embedding model — text-embedding-3-small (fast + cheap) for prototype; voyage-3.5-lite or mistral-embed for production
  4. Pick a vector DB — Chroma for local; Qdrant or Weaviate for production
  5. Enable hybrid search — add BM25 index alongside vector index
  6. Add metadata to every chunk — source, timestamp, section title
  7. Build the ingestion pipeline — incremental, hash-based, with soft deletes for old vectors
  8. Build the retrieval layer — top-k with metadata filtering, re-ranking for high-stakes queries
  9. Evaluate with RAGAS — run against your golden Q&A set
  10. Connect to your agents — expose via MCP, LangChain retriever, or LlamaIndex query engine

Get Started

A well-built AI agent knowledge base is the foundation that makes everything else — multi-agent collaboration, agentic RAG pipelines, automated code review — actually work at quality.

Try cowork.ink free — connect your team's knowledge sources to shared AI agents in minutes. Every engineer on the team queries the same fresh, indexed corpus — no duplicated RAG plumbing per person.

Frequently Asked Questions

What is an AI agent knowledge base?
An AI agent knowledge base is a structured collection of documents, facts, and data that an agent retrieves at runtime to answer questions accurately. It works through a RAG pipeline — documents are chunked, embedded into vectors, and stored in a vector database for semantic retrieval. This is how agents stay factual without hallucinating.
What file formats work best for an AI knowledge base?
Markdown (.md) produces the highest retrieval quality because its heading structure maps directly to meaningful chunks. JSON and CSV excel for structured data like product catalogs. PDFs work but require preprocessing with tools like Docling before ingestion. For most teams, a Markdown-first strategy with JSON for structured facts is optimal.
What is the best chunk size for RAG?
NVIDIA research shows 512 tokens with 50–100 token overlap is the best starting point for most use cases, achieving 0.640+ average retrieval accuracy. For precise factoid queries, 256–512 tokens works best; for complex analytical questions, use 1,024-token or page-level chunks.
What is the difference between vector, keyword, and hybrid search?
Vector search finds semantically similar content via embeddings — great for conceptual queries. Keyword (BM25) search matches exact terms — reliable for product codes and proper nouns. Hybrid search combines both using Reciprocal Rank Fusion for the best coverage. Most production RAG systems use [hybrid search](/blog/agentic-rag/) because pure semantic search fails on proprietary terminology.
How do I keep an AI knowledge base from going stale?
Implement an incremental ingestion pipeline triggered by source changes (webhooks, git hooks, or scheduled polling). Each update re-chunks and re-embeds only the changed documents. Use timestamp metadata to invalidate stale vectors. Teams using cowork.ink can connect knowledge sources directly to shared agents so the whole team always queries fresh data.
Home Blog Company