{"id":62177,"date":"2026-07-05T13:29:32","date_gmt":"2026-07-05T17:29:32","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=62177"},"modified":"2026-07-05T13:29:32","modified_gmt":"2026-07-05T17:29:32","slug":"llamaindex-legal-kb-agentic-retrieval","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/llamaindex-legal-kb-agentic-retrieval\/","title":{"rendered":"LlamaIndex Launches legal-kb Agentic Retrieval on Index v2"},"content":{"rendered":"<p><a href=\"https:\/\/www.llamaindex.ai\/\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">LlamaIndex<\/a> has released <strong>legal-kb<\/strong>, a public reference application on GitHub that reimagines how <a href=\"https:\/\/overcentral.com\/en\/patronus-ai-50m-stress-test-ai-agents\/\" title=\"Patronus AI lands $50M to build digital worlds that stress-test AI agents\" data-iacss-internal=\"1\">AI agents<\/a> interact with legal document collections. Rather than relying on a single embedding search per query, the project introduces a pattern the team calls a Retrieval Harness \u2014 an agentic approach that equips language models with filesystem-style tools to crawl, search, and verify information across large, versioned knowledge bases. Powered by LlamaIndex Index v2 (the LlamaParse Platform), legal-kb demonstrates a shift from single-shot retrieval toward multi-step, tool-driven document reasoning that could change how developers build AI systems for regulated industries.<\/p>\n<h2 id=\"h-what-is-legal-kb\">What Is legal-kb and How Does It Work<\/h2>\n<p><strong>legal-kb<\/strong> is a working TanStack Start web application, not a library or a framework. Users sign in, create a project, upload legal documents, and then interact with an <a href=\"https:\/\/overcentral.com\/en\/ai-agent-development-sluggish\/\" title=\"Zuckerberg Confirms AI Agents Development Slower Than Hoped\" data-iacss-internal=\"1\">AI agent<\/a> through a chat interface. Each project is mirrored as a managed LlamaCloud Index v2. Uploaded files are automatically parsed and indexed in the background, and the chat agent queries that index live during each turn of the conversation. The full source code is available on GitHub under the run-llama organization.<\/p>\n<h2 id=\"h-the-retrieval-harness-in-plain-terms\">The Retrieval Harness in Plain Terms<\/h2>\n<p>The harness provides a persistent data pipeline over your document collection. It connects to a data source, indexes the content, and keeps the index updated as files change. On top of that pipeline, it exposes a set of tools to the agent that are deliberately modeled on filesystem operations: list files, read a file, grep inside a file, or run a hybrid semantic and keyword search. Because the tools are generic, developers can plug the same harness into their own agents without reinventing the retrieval layer.<\/p>\n<h2 id=\"h-the-four-agent-tools-and-how-they-map-to-index-v2\">The Four Agent Tools and How They Map to Index v2<\/h2>\n<p>The agent defined in <code>src\/lib\/agent.ts<\/code>codecodecodecode receives exactly four tools, each backed by a corresponding Index v2 retrieval API. The system prompt enforces a strict order: the agent must call <code>findFiles<\/code>codecodecodecode first to establish the document inventory, then narrow the search with <code>retrieve<\/code>codecodecodecode, and finally confirm exact wording with <code>readFile<\/code>codecodecodecode or <code>grepFile<\/code>codecodecodecode before citing any source.<\/p>\n<ul>\n<li><strong>retrieve<\/strong> \u2014 backed by <code>beta.retrieval.retrieve<\/code>codecodecodecode. Runs hybrid semantic search with optional reranking. Key parameters include <code>query<\/code>codecodecodecode, <code>top_k<\/code>codecodecodecode, <code>score_threshold<\/code>codecodecodecode, <code>rerank_top_n<\/code>codecodecodecode, and metadata filters for <code>file_name<\/code>codecodecodecode and <code>file_version<\/code>codecodecodecode.<\/li>\n<li><strong>findFiles<\/strong> \u2014 backed by <code>beta.retrieval.find<\/code>codecodecodecode. Searches files by exact name or substring, with automatic pagination.<\/li>\n<li><strong>readFile<\/strong> \u2014 backed by <code>beta.retrieval.read<\/code>codecodecodecode. Reads raw file content with offset and length windows.<\/li>\n<li><strong>grepFile<\/strong> \u2014 backed by <code>beta.retrieval.grep<\/code>codecodecodecode. Matches a regex pattern inside a single file and returns character positions.<\/li>\n<\/ul>\n<p>Every retrieved chunk receives a short citation identifier such as <code>cite:c7f2qa<\/code>codecodecodecode. The agent references that identifier inline, and the UI renders a clickable chip that opens a page screenshot with bounding-box rectangles over the cited text.<\/p>\n<h2 id=\"h-how-it-works-under-the-hood-uploads-versioning-and-agent-loop\">How It Works Under the Hood: Uploads, Versioning, and Agent Loop<\/h2>\n<p>The upload pipeline is clearly structured in <code>src\/lib\/files.ts<\/code>codecodecodecode. File bytes are pushed to the project&#8217;s LlamaCloud source directory, and corresponding <code>File<\/code>codecodecodecode and <code>ProjectFile<\/code>codecodecodecode records are written to PostgreSQL via Prisma. An index sync is triggered but not awaited \u2014 the UI polls the status until the index is ready. Versioning is scoped to the (project, filename) pair. Re-uploading the same file, such as <code>nda.pdf<\/code>codecodecodecode, to the same project produces versions v1, v2, and v3 side by side. The retrieval layer filters on the <code>version<\/code>codecodecodecode metadata field, giving the knowledge base built-in version control.<\/p>\n<p>The agent uses the <code>ToolLoopAgent<\/code>codecodecodecode from the Vercel AI SDK 6. Developers can choose OpenAI or <a href=\"https:\/\/overcentral.com\/en\/amazon-ceo-anthropic-ai-export-ban\/\" title=\"Amazon CEO Drives Government Crackdown on Anthropic Models\" data-iacss-internal=\"1\">Anthropic models<\/a> per turn and bring their own API keys. Reasoning is streamed: Claude models use extended thinking, and OpenAI reasoning models use a medium reasoning effort. The tool closure wraps the Index v2 retrieval APIs, passing results back to the agent as structured output with both formatted text and citation objects.<\/p>\n<h2 id=\"h-naive-rag-vs-the-agentic-retrieval-harness\">Naive RAG vs the Agentic Retrieval Harness<\/h2>\n<p>The fundamental difference between single-shot RAG and this agentic approach is in execution behavior. Single-shot RAG runs one vector search per query and returns fixed top-k chunks. The agentic Retrieval Harness, by contrast, runs a multi-step tool loop: find files, retrieve semantically, then read or grep to confirm exact wording. It supports hybrid semantic search, keyword search, and regex grep. The agent can read full files or windows on demand rather than relying on a fixed chunk size. The persistent pipeline with sync and versioning ensures freshness, and parameters like <code>top_k<\/code>codecodecodecode, <code>score_threshold<\/code>codecodecodecode, and <code>rerank_top_n<\/code>codecodecodecode are fully exposed for precision control. The citations include visual page screenshots with bounding boxes, not just chunk identifiers. This design is best suited for long-horizon document tasks \u2014 contract analysis, due diligence, and versioned policy review \u2014 rather than short question answering.<\/p>\n<h2 id=\"h-use-cases-with-examples\">Use Cases with Examples<\/h2>\n<p>The design targets domains where agents must navigate large, evolving document sets. Legal and fintech are the stated examples. Consider a contract question: &#8220;What notice is needed to terminate the MSA?&#8221; The agent lists files, runs a semantic retrieve, then greps the exact clause and returns an answer with a citation to the specific page. In due diligence across a data room, an agent can find files by name and read each candidate, cross-checking clauses without a human opening every PDF. For a versioned policy base, the <code>file_version<\/code>codecodecodecode filter allows the agent to query a specific version, supporting change tracking over time.<\/p>\n<h2 id=\"h-what-this-means-for-developers\">What This Means for Developers<\/h2>\n<p>legal-kb is not a polished product aimed at end users \u2014 it is a reference implementation that demonstrates a pattern. Developers working on document-heavy AI applications should examine the source code to understand how the Retrieval Harness integrates with Index v2, how the tool loop enforces a disciplined search order, and how versioning and visual citations can be wired into a real application. The repository at <a href=\"https:\/\/github.com\/run-llama\/legal-kb\" target=\"_blank\" rel=\"noopener\">github.com\/run-llama\/legal-kb<\/a> is the best starting point. Clone it, run the TanStack Start app with your own API keys, and experiment with how the four-tool loop changes the quality and verifiability of agentic document retrieval. The pattern has clear potential for any organization that needs AI agents to reason over large, version-controlled document collections with verifiable citations.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>LlamaIndex has released legal-kb, a public reference application on GitHub that reimagines how AI agents interact with legal document collections. Rather than relying on a single embedding search per query, the project introduces a pattern the team calls a Retrieval Harness \u2014 an agentic approach that equips language models with filesystem-style tools to crawl, search, [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":84427,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/62177.png","fifu_image_alt":"LlamaIndex Launches legal-kb Agentic Retrieval on Index v2","footnotes":""},"categories":[349],"tags":[],"class_list":["post-62177","post","type-post","status-publish","format-standard","has-post-thumbnail","category-articles"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/62177.png","fifu_image_alt":"LlamaIndex Launches legal-kb Agentic Retrieval on Index v2","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/62177","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=62177"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/62177\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/84427"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=62177"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=62177"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=62177"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}