D-RAC is a research effort from the team at Yellow.ai, and we hope it gives the field a foundation for cost-efficient, retrieval-aware document ingestion in production AI systems.
Every enterprise AI agent is only as good as the knowledge it can find. When a customer asks about a policy term buried in a product brochure, or an employee needs a fee from a table on page 40 of a banking document, the answer depends on a step most people never see: how that document was turned into searchable pieces before it was ever indexed.
Our new paper, Document Retrieval-Aware Chunking (D-RAC), shows that enterprises no longer have to choose between ingestion that’s cheap and ingestion that works.
The problem: The hidden bottleneck in enterprise RAG
Retrieval-Augmented Generation (RAG) systems have to ingest whatever an enterprise knowledge base contains: PDFs, Word files, slide decks, spreadsheets, and scans. The most common of these, PDF, is a presentation format, not a semantic one. It describes where ink goes on a page, not what the content means.
That creates real problems for the usual ingestion approaches.
Rule-based extraction and fixed-size chunking pull text out by position and slice it into equal pieces. It’s fast and cheap, but multi-column pages get interleaved, headers and footers bleed into body text, heading hierarchy is lost, and tables collapse into strings of disconnected values. A cell that says “8” means nothing once it’s separated from its row and column.
Layout-analysis models recover more structure, but they need dedicated deployments, struggle with the marketing-style layouts common in enterprise documents, and still output tables as grids that embed poorly.
Agentic chunking hands the text to a frontier LLM and asks it to rewrite the document into coherent chunks. Retrieval quality improves, but the model must regenerate the entire document as output. Output tokens are 4 to 8 times more expensive than input tokens, so costs climb fast, and every rewrite is a chance for the model to silently alter content you need to trust.
Building on W-RAC
Earlier this year we introduced Web Retrieval-Aware Chunking (W-RAC), which reframed chunking as a planning problem rather than a writing problem. Web pages are parsed into small, ID-tagged units, and an LLM decides how to group them by returning lists of IDs instead of rewriting text. On web content, that cut chunking output tokens by 84.6% and total LLM cost by roughly half.
But W-RAC relied on HTML, which has structure you can parse deterministically. PDFs don’t. D-RAC asks a simple question: can a single multimodal pass recover enough structure from rendered pages that the entire W-RAC approach works unchanged? The answer is yes.
How D-RAC works
A gloD-RAC is a four-stage pipeline. Only two stages involve an LLM, and the model reads document content exactly once.
Stage 1: Normalize everything to PDF. Nearly every format has a faithful, deterministic PDF rendering. Word, PowerPoint, and Excel files go through standard tooling like headless LibreOffice, HTML is printed to PDF, and scans are wrapped as PDF. No LLM, no per-format parsers. Each page is then rendered as an image.
Stage 2: One retrieval-aware multimodal conversion. A multimodal model (we used Gemma-3 via AWS Bedrock) reads the page images and writes Markdown. This isn’t generic OCR; the output is shaped for retrieval:
- All text is preserved verbatim, never summarized.
- Every table row becomes its own self-contained sentence, written with the column headers as context.
- Rows are never merged. A table with one option at a 16-year policy term and another at 20 years becomes two separate sentences, never “a policy term of 16 or 20 years.” That kind of merge quietly wrecks precision, because a question about one option retrieves a sentence asserting both.
- Logos, charts, and decorative images are skipped rather than described, so no invented captions end up in the index.
- Heading hierarchy is rebuilt explicitly, and each page is tagged so every chunk stays traceable to its source page.
Stage 3: Deterministic parsing and sectioning. The Markdown is parsed into ID-addressable headers and content blocks. Long documents are split at heading boundaries into manageable sections, and each section carries a short chain of its parent headings, so the planner knows where it sits in the document without re-reading anything.
Stage 4: Chunk planning over IDs. The planner sees only element IDs, short text previews, and hierarchy, and returns chunk plans like [[“h1″,”h2″,”p1″,”p2”], [“h1″,”h3″,”p3″,”p4″,”p5”]]. Coverage is checked programmatically, so every piece of content lands in exactly one chunk. Final chunks are rebuilt locally from the verbatim converted text and prefixed with their heading breadcrumb (for example, Plan Overview > Eligibility > Age Limits) before embedding.
The key idea is that the expensive understanding step happens once, while the chunking step is cheap and repeatable.
The Results
We evaluated D-RAC on the 236-document, 795-page PDF subset of the RAG-Multi-Corpus benchmark, spanning five enterprise domains: automotive, academia, cloud services, enterprise technology, and banking. Retrieval was tested with 762 annotated queries covering procedural, comparative, descriptive, analytical, boolean, temporal, and open-ended questions. All three systems used the same embedding model and the same relevance judge.
Retrieval quality matches agentic chunking
| System | Recall@6 | Recall@3 | MRR | NDCG@6 |
| Fixed-size (rule-based extraction) | 0.717 | 0.666 | 0.602 | 0.764 |
| Agentic chunking | 0.795 | 0.726 | 0.682 | 0.793 |
| D-RAC | 0.798 | 0.743 | 0.690 | 0.801 |
Compared with fixed-size chunking, D-RAC improves Recall@6 by 11.3% and MRR by 14.6%. That’s strong evidence that the bottleneck in traditional PDF ingestion is structure-destroying extraction, not the embedding model.
Against agentic chunking, D-RAC matches or exceeds it on every overall metric. And it did so the hard way: the agentic reference chunks were produced from the corpus’s clean structured sources, while D-RAC worked from rendered PDF pages.
The biggest gains showed up where chunk boundaries and table handling matter most. On temporal questions, Recall@6 rose from 0.73 with fixed-size chunking to 0.85 with D-RAC; comparative and analytical queries improved as well. Boolean yes/no questions were the one category where agentic chunking kept an edge. Performance was also consistent across industries, with Recall@6 ranging only from 0.790 to 0.812 across the four evaluated organizations.
Chunking cost drops by 77.8% to 85.6%
Because the planner outputs only ID arrays, D-RAC’s chunking stage produced 11,714 output tokens across the whole corpus, compared with 270,454 for agentic chunking: a 95.7% reduction. Token counts were measured from the actual prompts and responses, not estimated.
At published model pricing, chunking the full corpus cost:
- GPT-4.1: $2.82 with agentic chunking vs. $0.62 with D-RAC (77.8% lower)
- Gemini 2.5 Pro: $3.11 with agentic chunking vs. $0.45 with D-RAC (85.6% lower)
Chunking time fell by 75%, from 2,167.5 seconds to 541.8 seconds.
Robust and scalable
The entire corpus was converted and chunked in about 72 minutes with zero conversion or chunking errors, producing 1,748 retrieval-ready chunks averaging around 700 characters, right in the range dense retrievers favor. In a separate stress test on a 503-page financial prospectus, conversion took 21.6 minutes with the 27B model (13.4 minutes with 12B), and planning a 5,060-element document took just over a minute. Cost grows linearly with page count, because pages are processed independently.
Why this matters for enterprises
Re-indexing becomes cheap. Knowledge bases change constantly. With D-RAC, the converted Markdown is stored, so changing your chunking strategy, whether chunk size, entity-aware grouping, or per-tenant policies, only requires re-planning, which takes seconds. Agentic chunking pays for full regeneration every time. Extrapolated to a one-million-page corpus, each full re-chunk drops from roughly $3,540 to roughly $785 at GPT-4.1 pricing.
What you index is what the document says. D-RAC’s chunking output contains no prose at all, only IDs validated against the source. No chunk text can be silently rewritten during chunking.
One pipeline for every format. Anything that renders to PDF becomes a first-class input, with no separate ingestion paths to maintain for decks, spreadsheets, Word files, and scans.
Tables finally work. Row-level prose means pricing tables, fee schedules, and eligibility grids become individually retrievable facts rather than noise.
What’s next
Together, W-RAC and D-RAC form a unified ingestion foundation: W-RAC for structured web content, and D-RAC for everything else. Because the converted documents and their IDs are durable, the approach extends naturally to entity-aware chunking, graph-based retrieval, and policy-driven chunk recomposition, all areas we’re actively exploring.
Read the full paper, including the complete methodology, prompts, and benchmark tables, on arXiv: arXiv:2609.24220. The benchmark is available on GitHub.
The paper was authored by Uday Allu, Abhivanth Sivaprakash, Pratik Singh, and Aman Manocha of the Yellow.ai AI Research Team.