LlamaIndex Launches legal-kb Agentic Retrieval on Index v2

The new Retrieval Harness pattern lets AI agents crawl, search, and verify legal documents using four filesystem-style tools.

By Central
LlamaIndex's legal-kb demonstrates agentic retrieval with multi-step reasoning and visual citations on Index v2.
Highlights
  • legal-kb uses a Retrieval Harness pattern that equips language models with filesystem-style tools for document reasoning.
  • The agent within legal-kb enforces a strict tool order: findFiles, retrieve, readFile, and grepFile before citing sources.
  • Uploaded legal documents are automatically parsed, versioned, and indexed in LlamaCloud Index v2 for live agent queries.

LlamaIndex has released legal-kb, a public reference application on GitHub that reimagines how AI agents interact with legal document collections. Rather than relying on a single embedding search per query, the project introduces a pattern the team calls a Retrieval Harness — an agentic approach that equips language models with filesystem-style tools to crawl, search, and verify information across large, versioned knowledge bases. Powered by LlamaIndex Index v2 (the LlamaParse Platform), legal-kb demonstrates a shift from single-shot retrieval toward multi-step, tool-driven document reasoning that could change how developers build AI systems for regulated industries.

What Is legal-kb and How Does It Work

legal-kb is a working TanStack Start web application, not a library or a framework. Users sign in, create a project, upload legal documents, and then interact with an AI agent through a chat interface. Each project is mirrored as a managed LlamaCloud Index v2. Uploaded files are automatically parsed and indexed in the background, and the chat agent queries that index live during each turn of the conversation. The full source code is available on GitHub under the run-llama organization.

The Retrieval Harness in Plain Terms

The harness provides a persistent data pipeline over your document collection. It connects to a data source, indexes the content, and keeps the index updated as files change. On top of that pipeline, it exposes a set of tools to the agent that are deliberately modeled on filesystem operations: list files, read a file, grep inside a file, or run a hybrid semantic and keyword search. Because the tools are generic, developers can plug the same harness into their own agents without reinventing the retrieval layer.

The Four Agent Tools and How They Map to Index v2

The agent defined in src/lib/agent.tscodecodecodecode receives exactly four tools, each backed by a corresponding Index v2 retrieval API. The system prompt enforces a strict order: the agent must call findFilescodecodecodecode first to establish the document inventory, then narrow the search with retrievecodecodecodecode, and finally confirm exact wording with readFilecodecodecodecode or grepFilecodecodecodecode before citing any source.

  • retrieve — backed by beta.retrieval.retrievecodecodecodecode. Runs hybrid semantic search with optional reranking. Key parameters include querycodecodecodecode, top_kcodecodecodecode, score_thresholdcodecodecodecode, rerank_top_ncodecodecodecode, and metadata filters for file_namecodecodecodecode and file_versioncodecodecodecode.
  • findFiles — backed by beta.retrieval.findcodecodecodecode. Searches files by exact name or substring, with automatic pagination.
  • readFile — backed by beta.retrieval.readcodecodecodecode. Reads raw file content with offset and length windows.
  • grepFile — backed by beta.retrieval.grepcodecodecodecode. Matches a regex pattern inside a single file and returns character positions.

Every retrieved chunk receives a short citation identifier such as cite:c7f2qacodecodecodecode. The agent references that identifier inline, and the UI renders a clickable chip that opens a page screenshot with bounding-box rectangles over the cited text.

How It Works Under the Hood: Uploads, Versioning, and Agent Loop

The upload pipeline is clearly structured in src/lib/files.tscodecodecodecode. File bytes are pushed to the project’s LlamaCloud source directory, and corresponding Filecodecodecodecode and ProjectFilecodecodecodecode records are written to PostgreSQL via Prisma. An index sync is triggered but not awaited — the UI polls the status until the index is ready. Versioning is scoped to the (project, filename) pair. Re-uploading the same file, such as nda.pdfcodecodecodecode, to the same project produces versions v1, v2, and v3 side by side. The retrieval layer filters on the versioncodecodecodecode metadata field, giving the knowledge base built-in version control.

The agent uses the ToolLoopAgentcodecodecodecode from the Vercel AI SDK 6. Developers can choose OpenAI or Anthropic models per turn and bring their own API keys. Reasoning is streamed: Claude models use extended thinking, and OpenAI reasoning models use a medium reasoning effort. The tool closure wraps the Index v2 retrieval APIs, passing results back to the agent as structured output with both formatted text and citation objects.

Naive RAG vs the Agentic Retrieval Harness

The fundamental difference between single-shot RAG and this agentic approach is in execution behavior. Single-shot RAG runs one vector search per query and returns fixed top-k chunks. The agentic Retrieval Harness, by contrast, runs a multi-step tool loop: find files, retrieve semantically, then read or grep to confirm exact wording. It supports hybrid semantic search, keyword search, and regex grep. The agent can read full files or windows on demand rather than relying on a fixed chunk size. The persistent pipeline with sync and versioning ensures freshness, and parameters like top_kcodecodecodecode, score_thresholdcodecodecodecode, and rerank_top_ncodecodecodecode are fully exposed for precision control. The citations include visual page screenshots with bounding boxes, not just chunk identifiers. This design is best suited for long-horizon document tasks — contract analysis, due diligence, and versioned policy review — rather than short question answering.

Use Cases with Examples

The design targets domains where agents must navigate large, evolving document sets. Legal and fintech are the stated examples. Consider a contract question: “What notice is needed to terminate the MSA?” The agent lists files, runs a semantic retrieve, then greps the exact clause and returns an answer with a citation to the specific page. In due diligence across a data room, an agent can find files by name and read each candidate, cross-checking clauses without a human opening every PDF. For a versioned policy base, the file_versioncodecodecodecode filter allows the agent to query a specific version, supporting change tracking over time.

What This Means for Developers

legal-kb is not a polished product aimed at end users — it is a reference implementation that demonstrates a pattern. Developers working on document-heavy AI applications should examine the source code to understand how the Retrieval Harness integrates with Index v2, how the tool loop enforces a disciplined search order, and how versioning and visual citations can be wired into a real application. The repository at github.com/run-llama/legal-kb is the best starting point. Clone it, run the TanStack Start app with your own API keys, and experiment with how the four-tool loop changes the quality and verifiability of agentic document retrieval. The pattern has clear potential for any organization that needs AI agents to reason over large, version-controlled document collections with verifiable citations.

Share This Article